Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces

arXiv cs.LG 论文

摘要

This paper extends Topological DeepONets to handle functional measurements on Hausdorff locally convex spaces, replacing point samples with continuous linear functionals and introducing fixed and adaptive measurement systems. The framework is validated on several benchmarks including a non-normable input space, and demonstrates compact, discretization-portable coordinates for operator learning.

arXiv:2608.06428v1 Announce Type: new Abstract: Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode an input function through point values on a fixed discretization. Building on the Topological DeepONet framework of Ismailov (arXiv:2603.11972), we replace point samples by continuous linear functionals drawn from the continuous dual of a Hausdorff locally convex space $({V},\{p_\alpha\}_{\alpha\in A})$, whose topology is generated by a point-separating family of seminorms rather than a single norm, and develop fixed and adaptive functional measurement systems. Measurements are combined with the coefficient-space Two-Step procedure of Lee and Shin (arXiv:2309.01020), while a training-only decoder and regularization stabilize the adaptive coordinates. We derive a discrete error decomposition separating measurement, output-basis, and neural-approximation errors, together with a Barron-rate refinement. The framework is evaluated on the antiderivative operator, a non-normable locally convex input space, heterogeneous Darcy flow, a controlled operator, and fixed-time and time-evolving Navier-Stokes vorticity operators. In the heterogeneous Darcy problem, the functional models retain nearly resolution-independent errors of 5.5-5.6% on unseen grids, while in the controlled problem adaptive measurements reduce the mean error below 1.2%. For the fixed-time Navier-Stokes problem, the Adaptive Topological DeepONet is the most accurate DeepONet-based model, attaining a mean relative $L^2$ error of 1.685% +/- 0.017% using 128 functional coordinates. A comparably sized Fourier neural operator (FNO; arXiv:2010.08895) achieves the lower error 0.832% +/- 0.172%, but requires the full 64x64 input field, twice the training time, and 10.7x greater peak GPU memory. The formulation provides compact, interpretable, and discretization-portable coordinates in the continuous dual $V'$, including for non-normable input spaces.
查看原文
查看缓存全文

缓存时间: 2026/08/10 08:01

# Functional Measurements on Hausdorff Locally Convex Spaces
Source: [https://arxiv.org/html/2608.06428](https://arxiv.org/html/2608.06428)
## Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces

Khemraj Shukla Division of Applied Mathematics, Brown University Providence, Rhode Island 02912, USAGeorge Em Karniadakis Division of Applied Mathematics, Brown University Providence, Rhode Island 02912, USA

###### Abstract

Deep Operator Networks \(DeepONets\)\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]typically encode an input function through point values on a fixed discretization\. Building on the Topological DeepONet framework of Ismailov\[[6](https://arxiv.org/html/2608.06428#bib.bib1)\], we replace point samples by continuous linear functionals drawn from the continuous dual\(𝒱,\{pα\}α∈A\)\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\)111Here𝒱\\mathcal\{V\}is a vector space and\{pα\}α∈A\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}is a point\-separating family of seminorms indexed by a \(possibly infinite\) index setAA; the seminorms generate the locally convex topology on𝒱\\mathcal\{V\}, and𝒱′\\mathcal\{V\}^\{\\prime\}denotes its continuous dual, the space of continuous linear functionals on𝒱\\mathcal\{V\}\., whose topology is generated by a point\-separating family of seminorms rather than a single norm, and develop fixed and adaptive functional measurement systems\. These measurements are combined with a coefficient\-space Two\-Step procedure\[[12](https://arxiv.org/html/2608.06428#bib.bib2)\], while a training\-only decoder and regularization stabilize the adaptive coordinates\. We also derive a discrete error decomposition separating measurement, output\-basis, and neural\-approximation errors, together with a Barron\-rate refinement\. The framework is evaluated on the antiderivative operator, a genuinely non\-normable locally convex input space, heterogeneous Darcy flow, a controlled operator, and fixed\-time and time\-evolving Navier–Stokes vorticity operators\. The non\-normable benchmark directly demonstrates that the method can learn operators when the input topology is generated by a family of seminorms and cannot be represented by a single norm\. In the heterogeneous Darcy problem, the functional models retain nearly resolution\-independent errors of approximately5\.5%5\.5\\%–5\.6%5\.6\\%on unseen grids, while in the controlled problem adaptive measurements reduce the mean error to below1\.2%1\.2\\%\. For the fixed\-time Navier–Stokes problem, the Adaptive Topological DeepONet is the most accurate DeepONet\-based model, attaining a three\-seed mean relativeL2L^\{2\}error of1\.685%±0\.017%1\.685\\%\\pm 0\.017\\%using only128128functional coordinates\. A comparably sized Fourier neural operator \(FNO\)\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\]achieves the lower error0\.832%±0\.172%0\.832\\%\\pm 0\.172\\%on the uniform periodic grid, but requires the full64×6464\\times 64input field, approximately twice the training time, and about10\.7×10\.7\\timesgreater peak GPU memory\. The proposed formulation therefore provides compact, interpretable, and discretization\-portable coordinates in the continuous dual𝒱′\\mathcal\{V\}^\{\\prime\}, including for input spaces that are not normable\.

## 1Introduction

Many problems in computational science and engineering require the repeated evaluation of a possibly nonlinear operator

𝒢:𝒱⟶𝒰,v⟼𝒢​\(v\),\\mathcal\{G\}:\\mathcal\{V\}\\longrightarrow\\mathcal\{U\},\\qquad v\\longmapsto\\mathcal\{G\}\(v\),\(1\)where𝒱\\mathcal\{V\}and𝒰\\mathcal\{U\}are spaces of input and output functions, respectively\[[25](https://arxiv.org/html/2608.06428#bib.bib12),[18](https://arxiv.org/html/2608.06428#bib.bib13),[17](https://arxiv.org/html/2608.06428#bib.bib14),[3](https://arxiv.org/html/2608.06428#bib.bib15),[11](https://arxiv.org/html/2608.06428#bib.bib16),[26](https://arxiv.org/html/2608.06428#bib.bib17),[22](https://arxiv.org/html/2608.06428#bib.bib31),[8](https://arxiv.org/html/2608.06428#bib.bib32),[21](https://arxiv.org/html/2608.06428#bib.bib34),[2](https://arxiv.org/html/2608.06428#bib.bib46)\]\. Examples include parameter\-to\-solution maps for partial differential equations, integral operators, constitutive relations, and evolution operators that advance an initial state to its future configuration\. Within this broader area, neural operators learn maps between function spaces and have been developed using branch–trunk, spectral, graph\-based, wavelet, transformer, and neural\-field representations\[[16](https://arxiv.org/html/2608.06428#bib.bib3),[9](https://arxiv.org/html/2608.06428#bib.bib7),[13](https://arxiv.org/html/2608.06428#bib.bib4),[4](https://arxiv.org/html/2608.06428#bib.bib38),[5](https://arxiv.org/html/2608.06428#bib.bib40),[19](https://arxiv.org/html/2608.06428#bib.bib41),[24](https://arxiv.org/html/2608.06428#bib.bib49)\]\. Learning such maps directly from data yields efficient surrogates for simulation, optimization, inverse problems, design, and uncertainty quantification\[[9](https://arxiv.org/html/2608.06428#bib.bib7),[27](https://arxiv.org/html/2608.06428#bib.bib29)\]\. Neural operators are designed to learn mappings between infinite\-dimensional function spaces rather than between finite\-dimensional vectors tied to a single discretization\[[9](https://arxiv.org/html/2608.06428#bib.bib7),[13](https://arxiv.org/html/2608.06428#bib.bib4)\]\. Prominent architectures include graph\-based neural operators, Fourier Neural Operators \(FNOs\), and Deep Operator Networks \(DeepONets\), which have been applied to parametric PDEs, fluid mechanics, porous\-media flow, solid mechanics, and multiscale systems\[[13](https://arxiv.org/html/2608.06428#bib.bib4),[23](https://arxiv.org/html/2608.06428#bib.bib26),[29](https://arxiv.org/html/2608.06428#bib.bib9),[25](https://arxiv.org/html/2608.06428#bib.bib12),[18](https://arxiv.org/html/2608.06428#bib.bib13)\]\. Deep Operator Networks approximate nonlinear operators through a branch–trunk architecture motivated by universal approximation results for operators\[[1](https://arxiv.org/html/2608.06428#bib.bib5),[16](https://arxiv.org/html/2608.06428#bib.bib3)\]\. Given an input functionvvsampled atmmprescribed sensor locations, the branch network maps the resulting vector of pointwise values to a set ofppexpansion coefficients, while the trunk network evaluates a learned output basis at the query coordinateyy:

𝒢^​\(v\)​\(y\)=∑k=1pbk​\(v​\(x1\),…,v​\(xm\)\)​tk​\(y\)\.\\widehat\{\\mathcal\{G\}\}\(v\)\(y\)=\\sum\_\{k=1\}^\{p\}b\_\{k\}\\\!\\bigl\(v\(x\_\{1\}\),\\ldots,v\(x\_\{m\}\)\\bigr\)\\,t\_\{k\}\(y\)\.\(2\)Several extensions have improved DeepONets’ flexibility and physical consistency\. Physics\-informed DeepONets incorporate governing equations into the training objective\[[29](https://arxiv.org/html/2608.06428#bib.bib9)\]\. MIONet handles operators with multiple input functions\[[7](https://arxiv.org/html/2608.06428#bib.bib10)\]\. SVD\-DeepONet links the trunk representation to singular\-value decomposition and reduced\-order modeling\[[28](https://arxiv.org/html/2608.06428#bib.bib6)\], while Two\-Step DeepONet separates output\-basis learning from the branch coefficient map to improve stability and generalization\[[12](https://arxiv.org/html/2608.06428#bib.bib2)\]\. Latent\-space and nonlinear\-decoder formulations address solution sets that resist low\-dimensional linear description\[[23](https://arxiv.org/html/2608.06428#bib.bib26)\], and multiscale architectures target oscillatory and high\-frequency phenomena\[[15](https://arxiv.org/html/2608.06428#bib.bib11)\]\. Despite these advances, a fundamental limitation of the original DeepONet branch network is its reliance on a fixed\-length vector of pointwise function values\. As a consequence, inputs observed on different meshes, resolutions, or experimental sampling patterns must be interpolated, resampled, or padded to a common representation before training\. Furthermore, pointwise values can be an inefficient coordinate system when the target operator depends primarily on global structures, moments, averages, or spatial correlations of the input function\. Variable\-Input Deep Operator Networks \(VIDONs\)\[[20](https://arxiv.org/html/2608.06428#bib.bib22)\], BelNet\[[31](https://arxiv.org/html/2608.06428#bib.bib23)\], and discretization\-invariant extensions\[[32](https://arxiv.org/html/2608.06428#bib.bib24)\]have relaxed the fixed\-grid constraint by allowing sensor locations to vary across samples\. These developments show that the fixed\-grid restriction is not intrinsic to operator learning and motivate a more fundamental question: can the branch representation be formulated directly in terms of the topology and continuous dual of the underlying input space? Topological DeepONets address this question by extending the classical framework from Banach spaces of continuous functions to general Hausdorff locally convex spaces \(HLCSs\), whose topology is generated by a family of seminorms\{pα\}α∈A\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\[[6](https://arxiv.org/html/2608.06428#bib.bib1)\]\. Rather than samplingvvat prescribed points, Topological DeepONets encode the input through a finite collection of continuous linear functionalsℓ1,…,ℓq\\ell\_\{1\},\\ldots,\\ell\_\{q\}drawn from the continuous dual𝒱′\\mathcal\{V\}^\{\\prime\}\. The resulting coordinate vector\(ℓ1​\(v\),…,ℓq​\(v\)\)∈ℝq\(\\ell\_\{1\}\(v\),\\ldots,\\ell\_\{q\}\(v\)\)\\in\\mathbb\{R\}^\{q\}is then processed by the branch network in exactly the same way as before\. Classical point\-sensor DeepONets are recovered as a special case when each functional is a continuous point\-evaluation map\. This functional perspective separates the infinite\-dimensional mathematical input from the finite\-dimensional branch encoding, and it provides a natural mechanism for handling variable\-resolution observations: the sameqqfunctionals, approximated by quadrature rules adapted to whatever grid is available, convert heterogeneous input representations into a common coordinate space\. In this work we develop and evaluate two concrete instantiations of the topological framework\. The fixed Topological DeepONet employs prescribed global functionals \(e\.g\. inner products against a Lagrange or spectral basis\), while the adaptive Topological DeepONet learns the measurement functionals directly from operator data, with a training\-only reconstruction decoder and soft regularization to prevent coordinate collapse\. We integrate both variants with an enhanced Two\-Step DeepONet strategy\[[12](https://arxiv.org/html/2608.06428#bib.bib2)\], which first constructs a stable low\-dimensional output basis via weighted singular\-value decomposition and then trains the branch network to predict the corresponding coefficients\. This separation allows point\-sensor, fixed\-functional, and adaptive\-functional branch inputs to be compared under a uniform output representation and on equal parameter budgets\. Full mathematical details of all three input encodings, the Two\-Step construction, and the training objectives are developed in Section[3](https://arxiv.org/html/2608.06428#S3)\. We evaluate the proposed framework on three benchmark operator\-learning problems: the antiderivative operator, the heterogeneous Darcy\-flow equation, and the two\-dimensional Navier–Stokes vorticity problem, the latter in both fixed\-time and time\-evolving settings\. These benchmarks span integral, elliptic, and nonlinear time\-dependent operators, and the Darcy and Navier–Stokes datasets follow standard neural\-operator evaluation protocols from the FNO literature\[[13](https://arxiv.org/html/2608.06428#bib.bib4),[9](https://arxiv.org/html/2608.06428#bib.bib7)\]\.

##### Contributions\.

The principal contributions of this work are as follows\.

1. 1\.We provide a computational realization of Topological DeepONets in which finite\-dimensional branch coordinates are constructed from continuous linear functionals in the dual space𝒱′\\mathcal\{V\}^\{\\prime\}of a Hausdorff locally convex input space\(𝒱,\{pα\}α∈A\)\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\)\. Unlike canonical point\-sensor DeepONets, the resulting coordinates may represent weak, local, spectral, finite\-element, or distributional measurements compatible with the topology of𝒱\\mathcal\{V\}\.
2. 2\.We introduce fixed and supervised\-adaptive functional measurement systems\. The adaptive coordinates remain continuous linear functionals because they are learned within the span of an admissible dual dictionary\. A training\-only decoder and soft regularization are used to control information loss, feature collapse, and excessive drift from the structured initialization\.
3. 3\.We integrate these measurements with a coefficient\-space Two\-Step DeepONet construction\. A weighted singular value decomposition produces a rank\-stable output basis, and the branch model is trained directly to predict the corresponding reduced coefficients\. This separates input representation learning from output\-subspace approximation\.
4. 4\.We derive a discrete error decomposition that separates measurement reconstruction, output\-basis truncation, and neural approximation errors, together with a Barron\-rate refinement of the neural term\. For the antiderivative operator, we evaluate the decomposition overq∈\{8,16,32,64,128\}q\\in\\\{8,16,32,64,128\\\}and verify that the theorem\-consistent bound holds for every test sample while the measurement and reconstruction errors decrease as the number of functionals increases\.
5. 5\.We demonstrate discretization transfer in a heterogeneous\-resolution Darcy benchmark\. The models are trained using coefficient fields observed on33×3333\\times 33,49×4949\\times 49,65×6565\\times 65, and85×8585\\times 85grids and are evaluated on previously unseen57×5757\\times 57,73×7373\\times 73, and97×9797\\times 97grids\. The same functionals are evaluated directly on each native grid using grid\-dependent quadrature, without first interpolating every input to a common mesh\.
6. 6\.We isolate the role of functional representation through a controlled operator with a known task\-relevant input subspace\. Under a common latent dimension, output basis, branch architecture, and test set, we compare the proposed measurements with PCA, random projections, fixed and optimized sensors, a learned dense bottleneck, decoder variants, and a full\-field multilayer perceptron\.
7. 7\.We provide matched computational comparisons on the antiderivative, Darcy, and two\-dimensional Navier–Stokes benchmarks\. For the fixed\-time Navier–Stokes problem, the Adaptive Topological DeepONet is the most accurate DeepONet\-based model, while a parameter\-matched Fourier neural operator achieves lower error on the uniform periodic grid\. This comparison clarifies that the principal contribution of the topological formulation is compact, interpretable, operator\-adapted, and discretization\-portable functional coordinates rather than universal superiority over grid\-specific spectral architectures\.

##### Paper organization\.

Section[2](https://arxiv.org/html/2608.06428#S2)reviews the required function\-space background and fixes notation\. Section[3](https://arxiv.org/html/2608.06428#S3)presents the fixed and adaptive functional measurements, the enhanced Two\-Step construction, and the training objectives, together with the complete algorithm and architecture \(Section[3\.3](https://arxiv.org/html/2608.06428#S3.SS3)\)\. Section[4](https://arxiv.org/html/2608.06428#S4)establishes a discrete approximation error bound for Topological DeepONets\. Section[5](https://arxiv.org/html/2608.06428#S5)describes the benchmark problems, numerical settings, and results for the antiderivative, Darcy\-flow, and fixed\-time and time\-evolving Navier–Stokes problems, and Section[6](https://arxiv.org/html/2608.06428#S6)summarizes the conclusions and directions for future work\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x1.png)Figure 1:Topological DeepONet: operator learning framework𝒢:𝒱→𝒰\\mathcal\{G\}:\\mathcal\{V\}\\to\\mathcal\{U\}where𝒱\\mathcal\{V\}is a Hausdorff locally convex input space and𝒰\\mathcal\{U\}is the Hilbert output space\.*Left*: space hierarchyBanach\(𝒱,∥⋅∥\)⊂Fre´chet⊂HLCS\(𝒱,\{pα\}\)\\mathrm\{Banach\}\(\\mathcal\{V\},\\\|\\cdot\\\|\)\\subset\\mathrm\{Fr\\acute\{e\}chet\}\\subset\\mathrm\{HLCS\}\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\)\.*Right*: branch–trunk architecture; the input is encoded by functional measurementsℓj∈𝒱′\\ell\_\{j\}\\in\\mathcal\{V\}^\{\\prime\}\(point evaluations being the canonical special case\), and continuity ofℬθ\\mathcal\{B\}\_\{\\theta\}is enforced with respect to the seminorm topology\{pα\}\\\{p\_\{\\alpha\}\\\}on𝒱\\mathcal\{V\}rather than a single norm\.

## 2Mathematical Preliminaries

This section collects the functional\-analytic background required for the topological formulation: seminorms and locally convex topologies \([Section 2\.1](https://arxiv.org/html/2608.06428#S2.SS1)\), the Hausdorff separation property \([Section 2\.2](https://arxiv.org/html/2608.06428#S2.SS2)\), and continuous linear functionals on such spaces \([Section 2\.3](https://arxiv.org/html/2608.06428#S2.SS3)\)\. Throughout the paper, generic operator inputs are denoted byv∈𝒱v\\in\\mathcal\{V\}and outputs byu∈𝒰u\\in\\mathcal\{U\}; in the PDE benchmarks of Section[5](https://arxiv.org/html/2608.06428#S5), where the input is a coefficient or initial field, we follow the neural\-operator literature and writeaafor the input\.

### 2\.1Seminorms and Locally Convex Spaces

A seminorm on a vector space𝒱\\mathcal\{V\}is a mapp:𝒱→\[0,∞\)p:\\mathcal\{V\}\\to\[0,\\infty\)satisfying

1. 1\.p​\(λ​v\)=\|λ\|​p​\(v\)p\(\\lambda v\)=\|\\lambda\|\\,p\(v\)for allv∈𝒱v\\in\\mathcal\{V\},λ∈ℝ\\lambda\\in\\mathbb\{R\},
2. 2\.p​\(u\+v\)≤p​\(u\)\+p​\(v\)p\(u\+v\)\\leq p\(u\)\+p\(v\)for allu,v∈𝒱u,v\\in\\mathcal\{V\}\.

Unlike a norm, a seminorm permitsp​\(v\)=0p\(v\)=0forv≠0v\\neq 0\. A locally convex space is a vector space𝒱\\mathcal\{V\}whose topology is generated by a family of seminorms\{pα\}α∈A\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}; a netvλ→vv\_\{\\lambda\}\\to vin this topology if and only ifpα​\(vλ−v\)→0p\_\{\\alpha\}\(v\_\{\\lambda\}\-v\)\\to 0for everyα∈A\\alpha\\in A\.

### 2\.2Hausdorff Locally Convex Spaces \(HLCS\)

A locally convex space\(𝒱,\{pα\}α∈A\)\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\)is called Hausdorff \(or separated\) if the seminorm family separates points: for everyu≠vu\\neq vin𝒱\\mathcal\{V\}there existsα∈A\\alpha\\in Asuch thatpα​\(u−v\)\>0p\_\{\\alpha\}\(u\-v\)\>0\. This condition ensures that limits of convergent nets are unique\. The class of Hausdorff locally convex spaces \(HLCSs\) encompasses Banach spaces \(a single norm, complete\), Fréchet spaces \(countably many seminorms, metrizable, complete\), and spaces of distributions or smooth functions such asC∞​\(Ω\)C^\{\\infty\}\(\\Omega\),𝒮​\(ℝn\)\\mathcal\{S\}\(\\mathbb\{R\}^\{n\}\), and𝒟′​\(Ω\)\\mathcal\{D\}^\{\\prime\}\(\\Omega\)\.

### 2\.3Continuous Dual and Continuous Functionals

The continuous dual𝒱′\\mathcal\{V\}^\{\\prime\}of a locally convex space𝒱\\mathcal\{V\}consists of all continuous linear functionalsℓ:𝒱→ℝ\\ell:\\mathcal\{V\}\\to\\mathbb\{R\}\. A linear functionalℓ\\ellis continuous in the seminorm topology if and only if there exist finitely many seminormspα1,…,pαkp\_\{\\alpha\_\{1\}\},\\ldots,p\_\{\\alpha\_\{k\}\}and a constantC\>0C\>0such that

\|ℓ​\(v\)\|≤C​maxi⁡pαi​\(v\)for all​v∈𝒱\.\|\\ell\(v\)\|\\leq C\\max\_\{i\}\\,p\_\{\\alpha\_\{i\}\}\(v\)\\quad\\text\{for all \}v\\in\\mathcal\{V\}\.Point evaluationsℓ​\(v\)=v​\(x0\)\\ell\(v\)=v\(x\_\{0\}\)belong to𝒱′\\mathcal\{V\}^\{\\prime\}whenever pointwise evaluation is continuous in the topology of𝒱\\mathcal\{V\}, as inC​\(Ω\)C\(\\Omega\)with the uniform norm\. More general examples include inner productsℓj​\(v\)=∫Ωv​\(x\)​φj​\(x\)​dx\\ell\_\{j\}\(v\)=\\int\_\{\\Omega\}v\(x\)\\varphi\_\{j\}\(x\)\\,\\mathrm\{d\}xagainst fixed test functionsφj∈L2​\(Ω\)\\varphi\_\{j\}\\in L^\{2\}\(\\Omega\)\. The Banach\-space setting does not, by itself, imply that point sensors are admissible\. The allowable measurements are arbitrary elements of the continuous dual𝒱′\\mathcal\{V\}^\{\\prime\}\. Point evaluation is continuous onC​\(Ω\)C\(\\Omega\)equipped with the supremum norm, but it is neither well defined on equivalence classes inL2​\(Ω\)L^\{2\}\(\\Omega\)nor continuous in theL2L^\{2\}topology\. Thus, canonical point\-sensor DeepONet is a narrower special case obtained only when point evaluations belong to the chosen continuous dual\.

## 3Methodology

In this section, we first present a detailed comparison between the conventional DeepONet and the Topological DeepONet \([Section 3\.1](https://arxiv.org/html/2608.06428#S3.SS1)\)\. We then introduce the proposed Two\-Step Topological DeepONet \([Section 3\.2](https://arxiv.org/html/2608.06428#S3.SS2)\) and summarize the complete algorithm and architecture \([Section 3\.3](https://arxiv.org/html/2608.06428#S3.SS3)\)\.

### 3\.1Canonical versus Topological DeepONet

Figure[1](https://arxiv.org/html/2608.06428#S1.F1)presents the overall architecture of the proposed Topological DeepONet framework\. The left panel situates the framework within a strict hierarchy of topological vector spaces\. At the outermost level sits the Hausdorff locally convex space\(𝒱,\{pα\}α∈A\)\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\), whose topology is generated by an arbitrary family of seminorms\{pα\}\\\{p\_\{\\alpha\}\\\}that jointly separate points — the defining property that replaces the single norm of classical functional analysis\. Nested within it is the class of Fréchet spaces, which impose the additional regularity that the seminorm family be countable and the space be complete with respect to the induced metric\. The innermost class, the Banach space\(𝒱,∥⋅∥\)\(\\mathcal\{V\},\\\|\\cdot\\\|\), corresponds to the degenerate case\|A\|=1\|A\|=1in which a single norm suffices; this is precisely the setting assumed by the original DeepONet of Lu et al\.\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]\. The right panel depicts the operator learning pipeline: an input functionv∈\(𝒱,\{pα\}\)v\\in\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\)is first encoded bymmcontinuous functional measurements into the vector𝐳=\[ℓ1​\(v\),…,ℓm​\(v\)\]∈ℝm\\mathbf\{z\}=\[\\ell\_\{1\}\(v\),\\ldots,\\ell\_\{m\}\(v\)\]\\in\\mathbb\{R\}^\{m\}, which is processed by the Branch Netℬθ:ℝm→ℝp\\mathcal\{B\}\_\{\\theta\}:\\mathbb\{R\}^\{m\}\\to\\mathbb\{R\}^\{p\}, while a query pointy∈ℝdy\\in\\mathbb\{R\}^\{d\}is processed independently by the Trunk Net𝒯ϕ:ℝd→ℝp\\mathcal\{T\}\_\{\\phi\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{p\}\. The operator output is reconstructed as

𝒢^​\(v\)​\(y\)=∑k=1pbk​\(v\)​tk​\(y\)\+b0,\\hat\{\\mathcal\{G\}\}\(v\)\(y\)\\;=\\;\\sum\_\{k=1\}^\{p\}b\_\{k\}\(v\)\\,t\_\{k\}\(y\)\\;\+\\;b\_\{0\},\(3\)an inner product between the branch and trunk embeddings\. The dashed arrow connecting the Banach sub\-box to the right panel makes explicit that all existing DeepONet theory is recovered as a strict special case of this more general construction\. The critical topological subtlety introduced by this framework is indicated by the brace annotation alongside the Branch Net in Figure[1](https://arxiv.org/html/2608.06428#S1.F1): the composite branch mapv↦ℬθ​\(𝐳​\(v\)\)v\\mapsto\\mathcal\{B\}\_\{\\theta\}\(\\mathbf\{z\}\(v\)\)is required to be continuous with respect to the seminorm topology on𝒱\\mathcal\{V\}, meaning thatpα​\(v−v′\)→0p\_\{\\alpha\}\(v\-v^\{\\prime\}\)\\to 0for allα∈A\\alpha\\in Amust imply‖ℬθ​\(𝐳​\(v\)\)−ℬθ​\(𝐳​\(v′\)\)‖ℝp→0\\\|\\mathcal\{B\}\_\{\\theta\}\(\\mathbf\{z\}\(v\)\)\-\\mathcal\{B\}\_\{\\theta\}\(\\mathbf\{z\}\(v^\{\\prime\}\)\)\\\|\_\{\\mathbb\{R\}^\{p\}\}\\to 0\. In a Banach space this reduces to ordinary norm continuity, a condition already implicit in classical DeepONet\. The class of admissible continuous linear functionals depends on the chosen topology\. Ifτ1⊂τ2\\tau\_\{1\}\\subset\\tau\_\{2\}are two locally convex topologies on the same vector space𝒱\\mathcal\{V\}, thenτ1\\tau\_\{1\}is coarser thanτ2\\tau\_\{2\}and

\(𝒱,τ1\)′⊆\(𝒱,τ2\)′\.\(\\mathcal\{V\},\\tau\_\{1\}\)^\{\\prime\}\\subseteq\(\\mathcal\{V\},\\tau\_\{2\}\)^\{\\prime\}\.Thus, coarsening the domain topology generally imposes a stronger continuity requirement and may reduce the continuous dual, whereas refining the topology may enlarge the class of continuous linear functionals\. Moreover, nonnormability does not mean that the topology is simply coarser than every norm topology; rather, it means that the topology cannot be generated by a single norm\. This distinction is relevant when𝒱\\mathcal\{V\}contains spaces such as distributions𝒟′​\(Ω\)\\mathcal\{D\}^\{\\prime\}\(\\Omega\), Schwartz functions𝒮​\(ℝn\)\\mathcal\{S\}\(\\mathbb\{R\}^\{n\}\), or smooth functionsC∞​\(Ω\)C^\{\\infty\}\(\\Omega\), whose natural locally convex topologies are generated by families of seminorms\. By grounding the operator𝒢:𝒱→𝒰\\mathcal\{G\}:\\mathcal\{V\}\\to\\mathcal\{U\}in theτ𝒱\\tau\_\{\\mathcal\{V\}\}\-to\-τ𝒰\\tau\_\{\\mathcal\{U\}\}continuity of thehlcstopology rather than in any norm, the Topological DeepONet extends universal approximation guarantees to these infinite\-dimensional, non\-normable spaces\. Figure[2](https://arxiv.org/html/2608.06428#S3.F2)isolates the central architectural departure from canonical DeepONet by placing the two branch input paradigms side by side\. In the left panel, the canonical approach is depicted: the underlying functionv∈C​\(Ω\)⊂𝒱v\\in C\(\\Omega\)\\subset\\mathcal\{V\}, shown as a smooth curve, is sampled atmmpre\-fixed sensor locations\{xi\}i=1m\\\{x\_\{i\}\\\}\_\{i=1\}^\{m\}\(marked as discrete dots with dashed projection lines to the domain axis\), reducing it to a finite\-dimensional vector𝐯=\[v​\(x1\),…,v​\(xm\)\]∈ℝm\\mathbf\{v\}=\[v\(x\_\{1\}\),\\ldots,v\(x\_\{m\}\)\]\\in\\mathbb\{R\}^\{m\}before any learning takes place\. This discretisation step is inherent to the standard formulation — the Branch Netℬθ\\mathcal\{B\}\_\{\\theta\}sees onlyℝm\\mathbb\{R\}^\{m\}, and its continuity is measured in the Euclidean norm‖𝐯−𝐯′‖ℝm\\\|\\mathbf\{v\}\-\\mathbf\{v\}^\{\\prime\}\\\|\_\{\\mathbb\{R\}^\{m\}\}\. The right panel shows the topological alternative: the same functionvvis treated as an element of the Hausdorff locally convex space\(𝒱,\{pα\}α∈A\)\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\), visualised as a shaded continuous curve to emphasise that the branch coordinates draw on the entire functional structure ofvvrather than onmmprescribed pointwise values\. The branch input is the coordinate vector𝐳m​\(v\)=\[ℓ1​\(v\),…,ℓm​\(v\)\]\\mathbf\{z\}\_\{m\}\(v\)=\[\\ell\_\{1\}\(v\),\\ldots,\\ell\_\{m\}\(v\)\], where eachℓj∈𝒱′\\ell\_\{j\}\\in\\mathcal\{V\}^\{\\prime\}is continuous with respect to the seminorm family:pα​\(v−v′\)→0p\_\{\\alpha\}\(v\-v^\{\\prime\}\)\\to 0for allα∈A\\alpha\\in Aforces𝐳m​\(v\)→𝐳m​\(v′\)\\mathbf\{z\}\_\{m\}\(v\)\\to\\mathbf\{z\}\_\{m\}\(v^\{\\prime\}\)and hence drives the branch outputs together\. Consequently, the Topological DeepONet requires no prescribed sensor grid: the same functionals can be approximated by quadrature on whatever discretization ofvvis available, and the finite\-dimensional encoding is adapted to the topology of𝒱\\mathcal\{V\}rather than tied to a particular mesh\. Taken together, the two figures articulate both the theoretical scope and the practical motivation of the proposed framework\. Figure[1](https://arxiv.org/html/2608.06428#S1.F1)establishes that Banach\-space DeepONet occupies the innermost level of a well\-ordered hierarchy of function spaces, and that the branch–trunk architecture carries over to every level of that hierarchy upon replacing norm continuity with seminorm continuity\. Figure[2](https://arxiv.org/html/2608.06428#S3.F2)then makes the departure operationally concrete: the key change is not in the network architecture itself — the branch–trunk decomposition and the inner\-product reconstruction formula \([3](https://arxiv.org/html/2608.06428#S3.E3)\) remain intact — but in how the branch coordinates are generated: point evaluations tied to a fixed grid are replaced by continuous functionals compatible with the intrinsic topology of𝒱\\mathcal\{V\}\. This change is what enables the Topological DeepONet to handle function classes inaccessible to norm\-based formulations, including spaces of smooth functions, distributions, and other non\-normable locally convex spaces that arise naturally in the analysis of partial differential equations and quantum field theory\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x2.png)Figure 2:Construction of the branch\-network input\.*Left*: canonical DeepONet\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]evaluatesvvatmmfixed sensors and forms𝐯m=\[v​\(x1\),…,v​\(xm\)\]∈ℝm\\mathbf\{v\}\_\{m\}=\[v\(x\_\{1\}\),\\ldots,v\(x\_\{m\}\)\]\\in\\mathbb\{R\}^\{m\}\.*Right*: Topological DeepONet applies continuous functionalsℓj:𝒱→ℝ\\ell\_\{j\}:\\mathcal\{V\}\\to\\mathbb\{R\}and forms the topology\-aware coordinate vector𝐳m​\(v\)=\[ℓ1​\(v\),…,ℓm​\(v\)\]∈ℝm\\mathbf\{z\}\_\{m\}\(v\)=\[\\ell\_\{1\}\(v\),\\ldots,\\ell\_\{m\}\(v\)\]\\in\\mathbb\{R\}^\{m\}, for example withℓj​\(v\)=∫Ωv​\(x\)​Lj​\(x\)​dx\\ell\_\{j\}\(v\)=\\int\_\{\\Omega\}v\(x\)L\_\{j\}\(x\)\\,\\mathrm\{d\}xandLjL\_\{j\}a Lagrange basis function\. Both branch networks receive a finite\-dimensional vector, but the topological construction is not restricted to point evaluations and can use functionals compatible with the seminorm topology\{pα\}α∈A\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}of𝒱\\mathcal\{V\}\. Here the number of functionals is denotedmmfor direct comparison with themmpoint sensors; in Section[3\.2](https://arxiv.org/html/2608.06428#S3.SS2)the \(possibly smaller\) number of learned functional coordinates is denotedq≤mq\\leq m\.
### 3\.2Two\-Step Adaptive Topological DeepONet

#### 3\.2\.1Operator and function spaces

Let𝒱\\mathcal\{V\}be a Hausdorff locally convex topological vector space and let𝒰\\mathcal\{U\}be a separable Hilbert space with inner product⟨⋅,⋅⟩𝒰\\langle\\cdot,\\cdot\\rangle\_\{\\mathcal\{U\}\}and induced norm∥⋅∥𝒰\\\|\\cdot\\\|\_\{\\mathcal\{U\}\}\. We seek to approximate a continuous nonlinear operator

𝒢:𝒱→𝒰,\\mathcal\{G\}:\\mathcal\{V\}\\rightarrow\\mathcal\{U\},\(4\)from a finite training set𝒟N=\{\(vi,ui\)\}i=1N\\mathcal\{D\}\_\{N\}=\\\{\(v\_\{i\},u\_\{i\}\)\\\}\_\{i=1\}^\{N\}, whereui=𝒢​\(vi\)u\_\{i\}=\\mathcal\{G\}\(v\_\{i\}\)\. No norm is assumed on𝒱\\mathcal\{V\}; only a finite family of continuous observations is required for the computational realisation\.

#### 3\.2\.2Base observation map

Letλ1,…,λm∈𝒱′\\lambda\_\{1\},\\ldots,\\lambda\_\{m\}\\in\\mathcal\{V\}^\{\\prime\}be continuous linear functionals on𝒱\\mathcal\{V\}\. The*base observation map*is

Λm:𝒱→ℝm,Λm​\(v\)=\[λ1​\(v\),…,λm​\(v\)\]⊤\.\\Lambda\_\{m\}:\\mathcal\{V\}\\rightarrow\\mathbb\{R\}^\{m\},\\qquad\\Lambda\_\{m\}\(v\)=\\bigl\[\\lambda\_\{1\}\(v\),\\,\\ldots,\\,\\lambda\_\{m\}\(v\)\\bigr\]^\{\\top\}\.\(5\)For function spaces on which point evaluation is continuous, such asC​\(Ω¯\)C\(\\overline\{\\Omega\}\)with the supremum norm, one may takeλj​\(v\)=v​\(xj\)\\lambda\_\{j\}\(v\)=v\(x\_\{j\}\)\. When point evaluation is not well\-defined or not continuous, eachλj\\lambda\_\{j\}may instead be a weak measurementλj​\(v\)=⟨v,φj⟩\\lambda\_\{j\}\(v\)=\\langle v,\\varphi\_\{j\}\\ranglefor a test functionφj\\varphi\_\{j\}, a local average, or a finite\-element degree of freedom\. In the Darcy and Navier–Stokes experiments, point\-sensor models are used only as discrete numerical baselines after spatial discretization\. They are not interpreted as continuum elements ofL2​\(Ω\)′L^\{2\}\(\\Omega\)^\{\\prime\}\. The training observation matrix is

S=\[Λm​\(v1\),…,Λm​\(vN\)\]⊤∈ℝN×m\.S=\\bigl\[\\Lambda\_\{m\}\(v\_\{1\}\),\\,\\ldots,\\,\\Lambda\_\{m\}\(v\_\{N\}\)\\bigr\]^\{\\top\}\\in\\mathbb\{R\}^\{N\\times m\}\.\(6\)
Algorithm 1 Fixed and Adaptive Topological DeepONetInput:Training data𝒟N=\{\(vi,ui\)\}i=1N\\mathcal\{D\}\_\{N\}=\\\{\(v\_\{i\},u\_\{i\}\)\\\}\_\{i=1\}^\{N\}; base observationsS∈ℝN×mS\\in\\mathbb\{R\}^\{N\\times m\}; outputsU∈ℝny×NU\\in\\mathbb\{R\}^\{n\_\{y\}\\times N\}; initial measurement matrixM0∈ℝm×qM\_\{0\}\\in\\mathbb\{R\}^\{m\\times q\}; mode∈\{Fixed,Adaptive\}\\in\\\{\\textsc\{Fixed\},\\textsc\{Adaptive\}\\\}Output:Operator surrogate𝒢^\\widehat\{\\mathcal\{G\}\}Stage I: output basis1\.Initialize trunkϕμ\\bm\{\\phi\}\_\{\\mu\}and free coefficientsA∈ℝp0×NA\\in\\mathbb\{R\}^\{p\_\{0\}\\times N\}\.2\.MinimizeℒI=‖Wy1/2​\(Φμ​A−U\)‖F2N​tr⁡\(Wy\)\\mathcal\{L\}\_\{\\mathrm\{I\}\}=\\frac\{\\\|W\_\{y\}^\{1/2\}\(\\Phi\_\{\\mu\}A\-U\)\\\|\_\{F\}^\{2\}\}\{N\\,\\operatorname\{tr\}\(W\_\{y\}\)\}with respect to\(μ,A\)\(\\mu,A\)\.3\.ComputeWy1/2​Φμ⋆=P​Σ​V⊤,r=\#​\{σk\>τrank​σ1\}\.W\_\{y\}^\{1/2\}\\Phi\_\{\\mu^\{\\star\}\}=P\\Sigma V^\{\\top\},\\qquad r=\\\#\\\{\\sigma\_\{k\}\>\\tau\_\{\\mathrm\{rank\}\}\\sigma\_\{1\}\\\}\.4\.SetQ=Wy−1/2​Pr,C⋆=Σr​Vr⊤​A⋆,𝝍​\(y\)=Σr−1​Vr⊤​ϕμ⋆​\(y\),Q=W\_\{y\}^\{\-1/2\}P\_\{r\},\\qquad C^\{\\star\}=\\Sigma\_\{r\}V\_\{r\}^\{\\top\}A^\{\\star\},\\qquad\\bm\{\\psi\}\(y\)=\\Sigma\_\{r\}^\{\-1\}V\_\{r\}^\{\\top\}\\bm\{\\phi\}\_\{\\mu^\{\\star\}\}\(y\),and freezeμ⋆\\mu^\{\\star\}\.Stage II: branch and measurement training5\.SetM=\{M0,Fixed,trainable​M0,Adaptive\.M=\\begin\{cases\}M\_\{0\},&\\textsc\{Fixed\},\\\\ \\text\{trainable \}M\_\{0\},&\\textsc\{Adaptive\}\.\\end\{cases\}Initializeℬθ:ℝq→ℝr\\mathcal\{B\}\_\{\\theta\}:\\mathbb\{R\}^\{q\}\\to\\mathbb\{R\}^\{r\}\. For the adaptive model, also initialize the training\-only decoderDω:ℝq→ℝmD\_\{\\omega\}:\\mathbb\{R\}^\{q\}\\to\\mathbb\{R\}^\{m\}\.6\.Compute and store normalization statistics fromS​MSM\.7\.foreach epochdo7\.1\.Draw a mini\-batch, optionally using hard\-example probabilities\.7\.2\.FormZ~ℬ=standardize⁡\(Sℬ​M\),C^ℬ=ℬθ​\(Z~ℬ\)\.\\widetilde\{Z\}\_\{\\mathcal\{B\}\}=\\operatorname\{standardize\}\(S\_\{\\mathcal\{B\}\}M\),\\qquad\\widehat\{C\}\_\{\\mathcal\{B\}\}=\\mathcal\{B\}\_\{\\theta\}\(\\widetilde\{Z\}\_\{\\mathcal\{B\}\}\)\.7\.3\.Minimizeℒ=\{ℒc,Fixed,ℒc\+λrec​ℒrec\+λorth​ℒorth\+λdrift​ℒdrift,Adaptive,\\mathcal\{L\}=\\begin\{cases\}\\mathcal\{L\}\_\{c\},&\\textsc\{Fixed\},\\\\\[2\.84526pt\] \\mathcal\{L\}\_\{c\}\+\\lambda\_\{\\mathrm\{rec\}\}\\mathcal\{L\}\_\{\\mathrm\{rec\}\}\+\\lambda\_\{\\mathrm\{orth\}\}\\mathcal\{L\}\_\{\\mathrm\{orth\}\}\+\\lambda\_\{\\mathrm\{drift\}\}\\mathcal\{L\}\_\{\\mathrm\{drift\}\},&\\textsc\{Adaptive\},\\end\{cases\}with losses defined in[equations14](https://arxiv.org/html/2608.06428#S3.E14),[15](https://arxiv.org/html/2608.06428#S3.E15)and[16](https://arxiv.org/html/2608.06428#S3.E16)\.7\.4\.Updateθ\\theta; for the adaptive model also updateMMandω\\omega\. Retain the best validation checkpoint\.8\.end for9\.DiscardDωD\_\{\\omega\}after adaptive training\.Inference9\.For a new inputvv, compute𝐬=Λm​\(v\),𝐳~=standardize⁡\(M⋆⊤​𝐬\),𝐜^=ℬθ⋆​\(𝐳~\)\.\\mathbf\{s\}=\\Lambda\_\{m\}\(v\),\\qquad\\widetilde\{\\mathbf\{z\}\}=\\operatorname\{standardize\}\(M^\{\\star\\top\}\\mathbf\{s\}\),\\qquad\\widehat\{\\mathbf\{c\}\}=\\mathcal\{B\}\_\{\\theta^\{\\star\}\}\(\\widetilde\{\\mathbf\{z\}\}\)\.10\.return𝒢^​\(v\)​\(y\)=𝝍​\(y\)⊤​𝐜^\.\\widehat\{\\mathcal\{G\}\}\(v\)\(y\)=\\bm\{\\psi\}\(y\)^\{\\top\}\\widehat\{\\mathbf\{c\}\}\.

#### 3\.2\.3Learned measurement map

A conventional DeepONet passesΛm​\(v\)\\Lambda\_\{m\}\(v\)directly to the branch network\. Instead, we learn a compact linear mapM∈ℝm×qM\\in\\mathbb\{R\}^\{m\\times q\}\(q≤mq\\leq m\) that produces*operator\-adapted*features:

M⊤​Λm​\(v\)∈ℝq,ℓkM​\(v\)=∑j=1mMj​k​λj​\(v\),k=1,…,q\.M^\{\\top\}\\Lambda\_\{m\}\(v\)\\in\\mathbb\{R\}^\{q\},\\qquad\\ell\_\{k\}^\{M\}\(v\)=\\sum\_\{j=1\}^\{m\}M\_\{jk\}\\,\\lambda\_\{j\}\(v\),\\quad k=1,\\ldots,q\.\(7\)EachℓkM\\ell\_\{k\}^\{M\}is a finite linear combination of elements of𝒱′\\mathcal\{V\}^\{\\prime\}, hence continuous on𝒱\\mathcal\{V\}\.MMis initialised from a structured weak\-measurement basis \(e\.g\. cosine or polynomial test functions\) and adapted during training\.

#### 3\.2\.4Output basis: Stage I

Lety1,…,ynyy\_\{1\},\\ldots,y\_\{n\_\{y\}\}be aligned output locations in the output domainΩy⊂ℝd\\Omega\_\{y\}\\subset\\mathbb\{R\}^\{d\},Wy=diag​\(w1y,…,wnyy\)≻0W\_\{y\}=\\mathrm\{diag\}\(w\_\{1\}^\{y\},\\ldots,w\_\{n\_\{y\}\}^\{y\}\)\\succ 0a quadrature weight matrix, andU∈ℝny×NU\\in\\mathbb\{R\}^\{n\_\{y\}\\times N\}the snapshot matrix whoseii\-th column storesuiu\_\{i\}evaluated on the grid\. We learn a trunk networkϕμ:Ωy→ℝp0\\bm\{\\phi\}\_\{\\mu\}:\\Omega\_\{y\}\\rightarrow\\mathbb\{R\}^\{p\_\{0\}\}, with output widthp0∈ℕp\_\{0\}\\in\\mathbb\{N\}, by solving

\(μ⋆,A⋆\)∈arg​minμ,A∈ℝp0×N⁡1N​tr​\(Wy\)​‖Wy1/2​\(Φμ​A−U\)‖F2,\(\\mu^\{\\star\},A^\{\\star\}\)\\in\\operatorname\*\{arg\\,min\}\_\{\\mu,\\,A\\in\\mathbb\{R\}^\{p\_\{0\}\\times N\}\}\\frac\{1\}\{N\\,\\mathrm\{tr\}\(W\_\{y\}\)\}\\left\\\|W\_\{y\}^\{1/2\}\\\!\\left\(\\Phi\_\{\\mu\}A\-U\\right\)\\right\\\|\_\{F\}^\{2\},\(8\)whereΦμ∈ℝny×p0\\Phi\_\{\\mu\}\\in\\mathbb\{R\}^\{n\_\{y\}\\times p\_\{0\}\}stacks the trunk evaluations on the grid\. The free coefficient matrixAAdecouples output\-subspace learning from input regression\. After training, a rank\-revealing weighted SVD

Wy1/2​Φμ⋆=P​Σ​V⊤W\_\{y\}^\{1/2\}\\Phi\_\{\\mu^\{\\star\}\}=P\\Sigma V^\{\\top\}\(9\)yields the*weighted\-orthonormal basis*

Q=Wy−1/2​Pr∈ℝny×r,Q⊤​Wy​Q=Ir,Q=W\_\{y\}^\{\-1/2\}P\_\{r\}\\in\\mathbb\{R\}^\{n\_\{y\}\\times r\},\\qquad Q^\{\\top\}W\_\{y\}Q=I\_\{r\},\(10\)wherer=\#​\{k:σk\>τrank​σ1\}r=\\\#\\\{k:\\sigma\_\{k\}\>\\tau\_\{\\mathrm\{rank\}\}\\,\\sigma\_\{1\}\\\}is the numerical rank determined by a prescribed toleranceτrank∈\(0,1\)\\tau\_\{\\mathrm\{rank\}\}\\in\(0,1\)\. The Stage II coefficient targets areC⋆=Σr​Vr⊤​A⋆∈ℝr×NC^\{\\star\}=\\Sigma\_\{r\}V\_\{r\}^\{\\top\}A^\{\\star\}\\in\\mathbb\{R\}^\{r\\times N\}\. BecauseQQis weighted\-orthonormal, output error and coefficient error are isometric: for anya,b∈ℝra,b\\in\\mathbb\{R\}^\{r\},

‖Q​\(a−b\)‖Wy=‖a−b‖2\.\\\|Q\(a\-b\)\\\|\_\{W\_\{y\}\}=\\\|a\-b\\\|\_\{2\}\.\(11\)

#### 3\.2\.5Coefficient prediction: Stage II

WithQQandC⋆C^\{\\star\}fixed, the Branch Netℬθ:ℝq→ℝr\\mathcal\{B\}\_\{\\theta\}:\\mathbb\{R\}^\{q\}\\rightarrow\\mathbb\{R\}^\{r\}and the measurement matrixMMare trained jointly\. Standardised measurements𝐳~​\(v\)=diag​\(𝝈z\)−1​\(M⊤​Λm​\(v\)−𝝁z\)\\tilde\{\\mathbf\{z\}\}\(v\)=\\mathrm\{diag\}\(\\bm\{\\sigma\}\_\{z\}\)^\{\-1\}\(M^\{\\top\}\\Lambda\_\{m\}\(v\)\-\\bm\{\\mu\}\_\{z\}\)serve as input, where𝝁z∈ℝq\\bm\{\\mu\}\_\{z\}\\in\\mathbb\{R\}^\{q\}and𝝈z∈ℝq\\bm\{\\sigma\}\_\{z\}\\in\\mathbb\{R\}^\{q\}denote the componentwise mean and standard deviation of the training features, and the prediction is

𝒢^​\(v\)​\(y\)=𝝍​\(y\)⊤​ℬθ​\(𝐳~​\(v\)\),\\hat\{\\mathcal\{G\}\}\(v\)\(y\)=\\bm\{\\psi\}\(y\)^\{\\top\}\\mathcal\{B\}\_\{\\theta\}\\\!\\left\(\\tilde\{\\mathbf\{z\}\}\(v\)\\right\),\(12\)where𝝍​\(y\)⊤=ϕμ⋆​\(y\)⊤​Vr​Σr−1\\bm\{\\psi\}\(y\)^\{\\top\}=\\bm\{\\phi\}\_\{\\mu^\{\\star\}\}\(y\)^\{\\top\}V\_\{r\}\\Sigma\_\{r\}^\{\-1\}evaluates the basis at any query pointyy\. Note that[Equation 12](https://arxiv.org/html/2608.06428#S3.E12)is the adaptive realisation of[Equation 3](https://arxiv.org/html/2608.06428#S3.E3), withbk​\(v\)=ℬθ​\(𝐳~​\(v\)\)kb\_\{k\}\(v\)=\\mathcal\{B\}\_\{\\theta\}\(\\tilde\{\\mathbf\{z\}\}\(v\)\)\_\{k\},tk​\(y\)=ψk​\(y\)t\_\{k\}\(y\)=\\psi\_\{k\}\(y\), andb0=0b\_\{0\}=0\(absorbed into the basis\)\.

##### Training objective\.

The Stage II loss combines coefficient regression with three regularisers:

ℒ=ℒc⏟coeff\. loss\+λrec​ℒrec⏟reconstruction\+λorth​ℒorth⏟soft ortho\.\+λdrift​ℒdrift⏟drift\.\\mathcal\{L\}=\\underbrace\{\\mathcal\{L\}\_\{c\}\}\_\{\\text\{coeff\.\\ loss\}\}\+\\lambda\_\{\\mathrm\{rec\}\}\\,\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{rec\}\}\}\_\{\\text\{reconstruction\}\}\+\\lambda\_\{\\mathrm\{orth\}\}\\,\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{orth\}\}\}\_\{\\text\{soft ortho\.\}\}\+\\lambda\_\{\\mathrm\{drift\}\}\\,\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{drift\}\}\}\_\{\\text\{drift\}\}\.\(13\)The coefficient loss is

ℒc=1B​r​∑i∈ℬ‖𝐜^i−𝐜i⋆‖22\+λrel​1B​∑i∈ℬeic\+λworst​τc​log⁡\(1B​∑i∈ℬexp⁡\(eic/τc\)\),\\mathcal\{L\}\_\{c\}=\\frac\{1\}\{Br\}\\sum\_\{i\\in\\mathcal\{B\}\}\\\|\\hat\{\\mathbf\{c\}\}\_\{i\}\-\\mathbf\{c\}\_\{i\}^\{\\star\}\\\|\_\{2\}^\{2\}\+\\lambda\_\{\\mathrm\{rel\}\}\\,\\frac\{1\}\{B\}\\sum\_\{i\\in\\mathcal\{B\}\}e\_\{i\}^\{c\}\+\\lambda\_\{\\mathrm\{worst\}\}\\,\\tau\_\{c\}\\log\\\!\\left\(\\frac\{1\}\{B\}\\sum\_\{i\\in\\mathcal\{B\}\}\\exp\\\!\\left\(e\_\{i\}^\{c\}/\\tau\_\{c\}\\right\)\\right\),\(14\)whereeic=‖𝐜^i−𝐜i⋆‖2/\(‖𝐜i⋆‖2\+εc\)e\_\{i\}^\{c\}=\\\|\\hat\{\\mathbf\{c\}\}\_\{i\}\-\\mathbf\{c\}\_\{i\}^\{\\star\}\\\|\_\{2\}/\(\\\|\\mathbf\{c\}\_\{i\}^\{\\star\}\\\|\_\{2\}\+\\varepsilon\_\{c\}\)is the per\-sample relative error,ℬ\\mathcal\{B\}is a mini\-batch of sizeBB,λrel,λworst≥0\\lambda\_\{\\mathrm\{rel\}\},\\lambda\_\{\\mathrm\{worst\}\}\\geq 0are fixed weights,τc\>0\\tau\_\{c\}\>0is the temperature of the log\-sum\-exp worst\-case term, andεc\>0\\varepsilon\_\{c\}\>0is a small constant preventing division by zero\. A training\-only decoderDω:ℝq→ℝmD\_\{\\omega\}:\\mathbb\{R\}^\{q\}\\rightarrow\\mathbb\{R\}^\{m\}reconstructs the base observations to prevent information collapse:

ℒrec=1B​m​∑i∈ℬ‖Dω​\(𝐳~i\)−𝐬i‖221B​m​∑i∈ℬ‖𝐬i‖22\+εs,\\mathcal\{L\}\_\{\\mathrm\{rec\}\}=\\frac\{\\dfrac\{1\}\{Bm\}\\displaystyle\\sum\_\{i\\in\\mathcal\{B\}\}\\\|D\_\{\\omega\}\(\\tilde\{\\mathbf\{z\}\}\_\{i\}\)\-\\mathbf\{s\}\_\{i\}\\\|\_\{2\}^\{2\}\}\{\\dfrac\{1\}\{Bm\}\\displaystyle\\sum\_\{i\\in\\mathcal\{B\}\}\\\|\\mathbf\{s\}\_\{i\}\\\|\_\{2\}^\{2\}\+\\varepsilon\_\{s\}\},\(15\)where𝐬i=Λm​\(vi\)\\mathbf\{s\}\_\{i\}=\\Lambda\_\{m\}\(v\_\{i\}\)denotes the base observation vector of theii\-th sample andεs\>0\\varepsilon\_\{s\}\>0is a small constant\. The soft weighted\-orthogonality and drift penalties are

ℒorth=1q2​‖M⊤​Wx−1​M−Iq‖F2,ℒdrift=‖M−M0‖F2‖M0‖F2\+εM,\\mathcal\{L\}\_\{\\mathrm\{orth\}\}=\\frac\{1\}\{q^\{2\}\}\\\|M^\{\\top\}W\_\{x\}^\{\-1\}M\-I\_\{q\}\\\|\_\{F\}^\{2\},\\qquad\\mathcal\{L\}\_\{\\mathrm\{drift\}\}=\\frac\{\\\|M\-M\_\{0\}\\\|\_\{F\}^\{2\}\}\{\\\|M\_\{0\}\\\|\_\{F\}^\{2\}\+\\varepsilon\_\{M\}\},\(16\)whereM0M\_\{0\}is the structured initialisation,WxW\_\{x\}is the sensor quadrature matrix, andεM\>0\\varepsilon\_\{M\}\>0is a small constant\. The decoder is discarded at inference; the deployed model consists solely ofMM,ℬθ\\mathcal\{B\}\_\{\\theta\}, and the frozen basisQQ\.

### 3\.3Algorithm and Architecture

The complete training and inference procedures for the fixed and adaptive topological DeepONets are summarized in[Section 3\.2\.2](https://arxiv.org/html/2608.06428#S3.SS2.SSS2)\. Detailed architectural block diagrams of both variants are presented in[Figure 3](https://arxiv.org/html/2608.06428#S3.F3), which summarizes the full data flow of the Adaptive Topological DeepONet, from input observationsΛm​\(v\)\\Lambda\_\{m\}\(v\)through the learned measurement map and Branch Netℬθ\\mathcal\{B\}\_\{\\theta\}to the final coefficient\-weighted output𝒢^​\(v\)​\(y\)\\hat\{\\mathcal\{G\}\}\(v\)\(y\)\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x3.png)Figure 3:Unified architecture of the Fixed and Adaptive Topological DeepONets\. Stage I learns a common reduced output basis from the output snapshots\. In Stage II, the Fixed model constructsqqfeatures using prescribed functionalsFfix𝖳F\_\{\\mathrm\{fix\}\}^\{\\mathsf\{T\}\}, whereas the Adaptive model learns the measurement matrixM𝖳M^\{\\mathsf\{T\}\}\. The adaptive pathway also contains a training\-only decoderDωD\_\{\\omega\}, regularised throughℒrec\\mathcal\{L\}\_\{\\mathrm\{rec\}\}\. Both models predict reduced coefficients and evaluate the output continuously asu^​\(y\)=𝝍​\(y\)𝖳​𝐜^\\widehat\{u\}\(y\)=\\bm\{\\psi\}\(y\)^\{\\mathsf\{T\}\}\\widehat\{\\mathbf\{c\}\}\.

## 4Discrete Error Bounds for Topological DeepONets

The universal approximation result of Ismailov\[[6](https://arxiv.org/html/2608.06428#bib.bib1)\]establishes that continuous operators on compact subsets of Hausdorff locally convex spaces can be approximated uniformly by branch–trunk expansions whose branch networks depend on finitely many continuous linear functionals from the continuous dual\. The result is qualitative and does not distinguish the different sources of approximation error\. We therefore derive a discrete error bound that separates the effects of input measurement compression, output\-basis truncation, and neural\-network approximation\. Let

Kh⊂Vh⊂ℝmK\_\{h\}\\subset V\_\{h\}\\subset\\mathbb\{R\}^\{m\}be a compact set of discrete inputs, and let

𝒢h:Vh⟶ℝn\\mathcal\{G\}\_\{h\}:V\_\{h\}\\longrightarrow\\mathbb\{R\}^\{n\}denote the corresponding discrete solution operator\. Let

Mq:ℝm⟶ℝq,Mq​vh=\[ℓ1,h​\(vh\)⋯ℓq,h​\(vh\)\]⊤,M\_\{q\}:\\mathbb\{R\}^\{m\}\\longrightarrow\\mathbb\{R\}^\{q\},\\qquad M\_\{q\}v\_\{h\}=\\begin\{bmatrix\}\\ell\_\{1,h\}\(v\_\{h\}\)&\\cdots&\\ell\_\{q,h\}\(v\_\{h\}\)\\end\{bmatrix\}^\{\\top\},whereℓj,h∈\(ℝm\)∗\\ell\_\{j,h\}\\in\(\\mathbb\{R\}^\{m\}\)^\{\\ast\}are discrete continuous linear functionals\. Let

Rq:ℝq⟶ℝmR\_\{q\}:\\mathbb\{R\}^\{q\}\\longrightarrow\\mathbb\{R\}^\{m\}be a continuous reconstruction map\. The mapRqR\_\{q\}is an auxiliary reconstruction used only in the error analysis to quantify information lost by the measurement map\. It is distinct from the training\-only decoderDωD\_\{\\omega\}in Section[3\.2\.5](https://arxiv.org/html/2608.06428#S3.SS2.SSS5), which reconstructs the finite base\-observation vector𝒔=Λm​\(v\)\\bm\{s\}=\\Lambda\_\{m\}\(v\)\. The theorem therefore does not assume that the trained decoder reconstructs the full continuum input\. When the base observations form a stable discrete representation ofvhv\_\{h\}, a reconstruction ofvhv\_\{h\}may be composed with the decoded observation vector; otherwise these two reconstruction notions should be kept separate\. To match the weighted output basis used in the algorithm, letWy≻0W\_\{y\}\\succ 0denote the output quadrature matrix and introduce preweighted coordinates𝒈~h=Wy1/2​𝒈h\\widetilde\{\\bm\{g\}\}\_\{h\}=W\_\{y\}^\{1/2\}\\bm\{g\}\_\{h\}andQ~r=Wy1/2​Qr\\widetilde\{Q\}\_\{r\}=W\_\{y\}^\{1/2\}Q\_\{r\}\. Since the algorithm satisfiesQr⊤​Wy​Qr=IrQ\_\{r\}^\{\\top\}W\_\{y\}Q\_\{r\}=I\_\{r\}, the transformed basis obeysQ~r⊤​Q~r=Ir\\widetilde\{Q\}\_\{r\}^\{\\top\}\\widetilde\{Q\}\_\{r\}=I\_\{r\}\. The theorem below is written in these preweighted Euclidean coordinates; equivalently, every output norm may be read as the weighted norm‖v‖Wy=‖Wy1/2​v‖2\\\|v\\\|\_\{W\_\{y\}\}=\\\|W\_\{y\}^\{1/2\}v\\\|\_\{2\}\. For the output representation, let

𝒈¯h∈ℝn,Qr∈ℝn×r,Qr⊤​Qr=Ir,\\overline\{\\bm\{g\}\}\_\{h\}\\in\\mathbb\{R\}^\{n\},\\qquad Q\_\{r\}\\in\\mathbb\{R\}^\{n\\times r\},\\qquad Q\_\{r\}^\{\\top\}Q\_\{r\}=I\_\{r\},whereQrQ\_\{r\}and𝒈¯h\\overline\{\\bm\{g\}\}\_\{h\}now denote the corresponding preweighted quantities\. The corresponding Topological DeepONet approximation is

𝒢^h,θ​\(vh\)=𝒈¯h\+Qr​bθ​\(Mq​vh\),\\widehat\{\\mathcal\{G\}\}\_\{h,\\theta\}\(v\_\{h\}\)=\\overline\{\\bm\{g\}\}\_\{h\}\+Q\_\{r\}b\_\{\\theta\}\(M\_\{q\}v\_\{h\}\),\(17\)where

bθ:ℝq⟶ℝrb\_\{\\theta\}:\\mathbb\{R\}^\{q\}\\longrightarrow\\mathbb\{R\}^\{r\}is the branch network\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x4.png)\(a\)Fixed\-model error budget\.
![Refer to caption](https://arxiv.org/html/2608.06428v1/x5.png)\(b\)Adaptive\-model error budget\.
![Refer to caption](https://arxiv.org/html/2608.06428v1/x6.png)\(c\)Input reconstruction convergence\.
![Refer to caption](https://arxiv.org/html/2608.06428v1/x7.png)\(d\)Samplewise bound verification atq=32q=32\.

Figure 4:Numerical verification of the discrete Topological DeepONet error decomposition for the antiderivative operator\. Panels \(a\) and \(b\) show the operator\-level measurement errorEop,measE\_\{\\mathrm\{op,meas\}\}, its Lipschitz upper boundELipE\_\{\\mathrm\{Lip\}\}, the output truncation errorEoutE\_\{\\mathrm\{out\}\}, neural approximation errorENNE\_\{\\mathrm\{NN\}\}, empirical boundEempE\_\{\\mathrm\{emp\}\}, theorem\-consistent boundEthmE\_\{\\mathrm\{thm\}\}, and total errorEtotE\_\{\\mathrm\{tot\}\}\. Panel \(c\) shows that the input reconstruction error decreases as the measurement dimensionqqincreases\. Panel \(d\) verifies the samplewise inequalityEtot\(i\)≤Ethm\(i\)E\_\{\\mathrm\{tot\}\}^\{\(i\)\}\\leq E\_\{\\mathrm\{thm\}\}^\{\(i\)\}; atq=32q=32, the maximum ratio is0\.12490\.1249, with zero violations\.###### Theorem 4\.1\(Discrete Topological DeepONet error decomposition\)\.

LetKh⊂Vh⊂ℝmK\_\{h\}\\subset V\_\{h\}\\subset\\mathbb\{R\}^\{m\}be compact and assume that

Rq​\(Mq​Kh\)⊂Vh\.R\_\{q\}\(M\_\{q\}K\_\{h\}\)\\subset V\_\{h\}\.Suppose that𝒢h\\mathcal\{G\}\_\{h\}is Lipschitz continuous on

Kh∪Rq​\(Mq​Kh\)K\_\{h\}\\cup R\_\{q\}\(M\_\{q\}K\_\{h\}\)with constantLh\>0L\_\{h\}\>0, namely,

‖𝒢h​\(vh\)−𝒢h​\(wh\)‖2≤Lh​‖vh−wh‖2\\\|\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\-\\mathcal\{G\}\_\{h\}\(w\_\{h\}\)\\\|\_\{2\}\\leq L\_\{h\}\\\|v\_\{h\}\-w\_\{h\}\\\|\_\{2\}\(18\)for allvh,wh∈Kh∪Rq​\(Mq​Kh\)v\_\{h\},w\_\{h\}\\in K\_\{h\}\\cup R\_\{q\}\(M\_\{q\}K\_\{h\}\)\. Define the measurement–reconstruction error by

εrec​\(q\):=supvh∈Kh‖vh−Rq​\(Mq​vh\)‖2,\\varepsilon\_\{\\mathrm\{rec\}\}\(q\):=\\sup\_\{v\_\{h\}\\in K\_\{h\}\}\\\|v\_\{h\}\-R\_\{q\}\(M\_\{q\}v\_\{h\}\)\\\|\_\{2\},\(19\)and define the output\-basis truncation error by

εout​\(r,q\):=supvh∈Kh‖\(In−Qr​Qr⊤\)​\[𝒢h​\(Rq​\(Mq​vh\)\)−𝒈¯h\]‖2\.\\varepsilon\_\{\\mathrm\{out\}\}\(r,q\):=\\sup\_\{v\_\{h\}\\in K\_\{h\}\}\\left\\\|\(I\_\{n\}\-Q\_\{r\}Q\_\{r\}^\{\\top\}\)\\left\[\\mathcal\{G\}\_\{h\}\(R\_\{q\}\(M\_\{q\}v\_\{h\}\)\)\-\\overline\{\\bm\{g\}\}\_\{h\}\\right\]\\right\\\|\_\{2\}\.\(20\)Assume that the activation function inbθb\_\{\\theta\}is continuous and nonpolynomial\. Then, for everyεNN\>0\\varepsilon\_\{\\mathrm\{NN\}\}\>0, there exists a feed\-forward neural networkbθ:ℝq→ℝrb\_\{\\theta\}:\\mathbb\{R\}^\{q\}\\to\\mathbb\{R\}^\{r\}such that

supvh∈Kh∥𝒢h\(vh\)−𝒢^h,θ\(vh\)∥2≤Lhεrec\(q\)\+εout\(r,q\)\+εNN\.\\boxed\{\\sup\_\{v\_\{h\}\\in K\_\{h\}\}\\left\\\|\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\-\\widehat\{\\mathcal\{G\}\}\_\{h,\\theta\}\(v\_\{h\}\)\\right\\\|\_\{2\}\\leq L\_\{h\}\\varepsilon\_\{\\mathrm\{rec\}\}\(q\)\+\\varepsilon\_\{\\mathrm\{out\}\}\(r,q\)\+\\varepsilon\_\{\\mathrm\{NN\}\}\.\}\(21\)

###### Proof\.

BecauseKhK\_\{h\}is compact andMqM\_\{q\}is continuous, the measurement set

Zq:=Mq​Kh⊂ℝqZ\_\{q\}:=M\_\{q\}K\_\{h\}\\subset\\mathbb\{R\}^\{q\}is compact\. Define the reduced coefficient map

cr,q:Zq⟶ℝrc\_\{r,q\}:Z\_\{q\}\\longrightarrow\\mathbb\{R\}^\{r\}by

cr,q​\(z\):=Qr⊤​\[𝒢h​\(Rq​\(z\)\)−𝒈¯h\]\.c\_\{r,q\}\(z\):=Q\_\{r\}^\{\\top\}\\left\[\\mathcal\{G\}\_\{h\}\(R\_\{q\}\(z\)\)\-\\overline\{\\bm\{g\}\}\_\{h\}\\right\]\.\(22\)The continuity ofRqR\_\{q\},𝒢h\\mathcal\{G\}\_\{h\}, andQr⊤Q\_\{r\}^\{\\top\}implies thatcr,qc\_\{r,q\}is continuous onZqZ\_\{q\}\. Therefore, the finite\-dimensional universal approximation theorem implies that, for everyεNN\>0\\varepsilon\_\{\\mathrm\{NN\}\}\>0, there exists a feed\-forward neural networkbθb\_\{\\theta\}satisfying

supz∈Zq‖cr,q​\(z\)−bθ​\(z\)‖2≤εNN\.\\sup\_\{z\\in Z\_\{q\}\}\\\|c\_\{r,q\}\(z\)\-b\_\{\\theta\}\(z\)\\\|\_\{2\}\\leq\\varepsilon\_\{\\mathrm\{NN\}\}\.\(23\)Fixvh∈Khv\_\{h\}\\in K\_\{h\}and set

v~h:=Rq​\(Mq​vh\)\.\\widetilde\{v\}\_\{h\}:=R\_\{q\}\(M\_\{q\}v\_\{h\}\)\.Adding and subtracting𝒢h​\(v~h\)\\mathcal\{G\}\_\{h\}\(\\widetilde\{v\}\_\{h\}\)and the orthogonal projection of the centered reconstructed output gives

𝒢h​\(vh\)−𝒈¯h−Qr​bθ​\(Mq​vh\)\\displaystyle\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\-\\overline\{\\bm\{g\}\}\_\{h\}\-Q\_\{r\}b\_\{\\theta\}\(M\_\{q\}v\_\{h\}\)=𝒢h​\(vh\)−𝒢h​\(v~h\)⏟\(I\)\\displaystyle=\\underbrace\{\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\-\\mathcal\{G\}\_\{h\}\(\\widetilde\{v\}\_\{h\}\)\}\_\{\\mathrm\{\(I\)\}\}\+\(In−Qr​Qr⊤\)​\[𝒢h​\(v~h\)−𝒈¯h\]⏟\(II\)\\displaystyle\\quad\+\\underbrace\{\(I\_\{n\}\-Q\_\{r\}Q\_\{r\}^\{\\top\}\)\\left\[\\mathcal\{G\}\_\{h\}\(\\widetilde\{v\}\_\{h\}\)\-\\overline\{\\bm\{g\}\}\_\{h\}\\right\]\}\_\{\\mathrm\{\(II\)\}\}\+Qr​\[cr,q​\(Mq​vh\)−bθ​\(Mq​vh\)\]⏟\(III\)\.\\displaystyle\\quad\+\\underbrace\{Q\_\{r\}\\left\[c\_\{r,q\}\(M\_\{q\}v\_\{h\}\)\-b\_\{\\theta\}\(M\_\{q\}v\_\{h\}\)\\right\]\}\_\{\\mathrm\{\(III\)\}\}\.\(24\)By the Lipschitz continuity of𝒢h\\mathcal\{G\}\_\{h\},

‖\(I\)‖2\\displaystyle\\\|\\mathrm\{\(I\)\}\\\|\_\{2\}≤Lh​‖vh−v~h‖2\\displaystyle\\leq L\_\{h\}\\\|v\_\{h\}\-\\widetilde\{v\}\_\{h\}\\\|\_\{2\}≤Lh​εrec​\(q\)\.\\displaystyle\\leq L\_\{h\}\\varepsilon\_\{\\mathrm\{rec\}\}\(q\)\.\(25\)By definition,

‖\(II\)‖2≤εout​\(r,q\)\.\\\|\\mathrm\{\(II\)\}\\\|\_\{2\}\\leq\\varepsilon\_\{\\mathrm\{out\}\}\(r,q\)\.SinceQr⊤​Qr=IrQ\_\{r\}^\{\\top\}Q\_\{r\}=I\_\{r\}, the matrixQrQ\_\{r\}is an isometry onℝr\\mathbb\{R\}^\{r\}, and hence

‖\(III\)‖2\\displaystyle\\\|\\mathrm\{\(III\)\}\\\|\_\{2\}=‖cr,q​\(Mq​vh\)−bθ​\(Mq​vh\)‖2\\displaystyle=\\left\\\|c\_\{r,q\}\(M\_\{q\}v\_\{h\}\)\-b\_\{\\theta\}\(M\_\{q\}v\_\{h\}\)\\right\\\|\_\{2\}≤εNN\.\\displaystyle\\leq\\varepsilon\_\{\\mathrm\{NN\}\}\.\(26\)Applying the triangle inequality and taking the supremum overvh∈Khv\_\{h\}\\in K\_\{h\}proves \([21](https://arxiv.org/html/2608.06428#S4.E21)\)\. ∎

###### Corollary 4\.1\(Barron\-rate refinement\)\.

Letν\\nube a probability measure onKhK\_\{h\}, and let

μq:=\(Mq\)\#​ν\\mu\_\{q\}:=\(M\_\{q\}\)\_\{\\\#\}\\nube its pushforward measure on

Zq=Mq​Kh⊂Br0​\(0\)⊂ℝq\.Z\_\{q\}=M\_\{q\}K\_\{h\}\\subset B\_\{r\_\{0\}\}\(0\)\\subset\\mathbb\{R\}^\{q\}\.Assume the hypotheses of Theorem[4\.1](https://arxiv.org/html/2608.06428#S4.Thmtheorem1)\. Suppose that every componentcr,q,jc\_\{r,q,j\}of the reduced mapcr,qc\_\{r,q\}admits an extension

gj:ℝq⟶ℝg\_\{j\}:\\mathbb\{R\}^\{q\}\\longrightarrow\\mathbb\{R\}with finite Barron norm

Cj:=‖gj‖ℬ\.C\_\{j\}:=\\\|g\_\{j\}\\\|\_\{\\mathcal\{B\}\}\.Define

Cr,q:=\(∑j=1rCj2\)1/2\.C\_\{r,q\}:=\\left\(\\sum\_\{j=1\}^\{r\}C\_\{j\}^\{2\}\\right\)^\{1/2\}\.Then, for everyN∈ℕN\\in\\mathbb\{N\}, there exists a two\-layer sigmoidal network

bθ:ℝq⟶ℝr,b\_\{\\theta\}:\\mathbb\{R\}^\{q\}\\longrightarrow\\mathbb\{R\}^\{r\},with at mostNNhidden units for each output component, such that

\(𝔼vh∼ν​‖𝒢h​\(vh\)−𝒈¯h−Qr​bθ​\(Mq​vh\)‖22\)1/2≤Lh​εrec​\(q\)\+εout​\(r,q\)\+2​r0​Cr,qN\.\\boxed\{\\begin\{aligned\} &\\left\(\\mathbb\{E\}\_\{v\_\{h\}\\sim\\nu\}\\left\\\|\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\-\\overline\{\\bm\{g\}\}\_\{h\}\-Q\_\{r\}b\_\{\\theta\}\(M\_\{q\}v\_\{h\}\)\\right\\\|\_\{2\}^\{2\}\\right\)^\{1/2\}\\\\ &\\qquad\\leq L\_\{h\}\\varepsilon\_\{\\mathrm\{rec\}\}\(q\)\+\\varepsilon\_\{\\mathrm\{out\}\}\(r,q\)\+\\frac\{2r\_\{0\}C\_\{r,q\}\}\{\\sqrt\{N\}\}\.\\end\{aligned\}\}\(28\)

###### Proof\.

Forvh∈Khv\_\{h\}\\in K\_\{h\}, use the decomposition \([24](https://arxiv.org/html/2608.06428#S4.E24)\)\. The first two terms satisfy the pointwise estimates

‖\(I\)‖2≤Lh​εrec​\(q\),‖\(II\)‖2≤εout​\(r,q\),\\\|\\mathrm\{\(I\)\}\\\|\_\{2\}\\leq L\_\{h\}\\varepsilon\_\{\\mathrm\{rec\}\}\(q\),\\qquad\\\|\\mathrm\{\(II\)\}\\\|\_\{2\}\\leq\\varepsilon\_\{\\mathrm\{out\}\}\(r,q\),and therefore obey the same bounds inL2​\(ν;ℝn\)L^\{2\}\(\\nu;\\mathbb\{R\}^\{n\}\)\. Applying Barron’s approximation theorem componentwise gives functionsgN,jg\_\{N,j\}, each represented by a sum ofNNsigmoidal ridge functions, such that

𝔼z∼μq​\[\|gj​\(z\)−gN,j​\(z\)\|2\]≤4​r02​Cj2N\.\\mathbb\{E\}\_\{z\\sim\\mu\_\{q\}\}\\left\[\\left\|g\_\{j\}\(z\)\-g\_\{N,j\}\(z\)\\right\|^\{2\}\\right\]\\leq\\frac\{4r\_\{0\}^\{2\}C\_\{j\}^\{2\}\}\{N\}\.\(29\)Define

bθ:=\(gN,1,…,gN,r\)\.b\_\{\\theta\}:=\(g\_\{N,1\},\\ldots,g\_\{N,r\}\)\.BecauseQr⊤​Qr=IrQ\_\{r\}^\{\\top\}Q\_\{r\}=I\_\{r\},

𝔼vh∼ν​‖\(III\)‖22\\displaystyle\\mathbb\{E\}\_\{v\_\{h\}\\sim\\nu\}\\\|\\mathrm\{\(III\)\}\\\|\_\{2\}^\{2\}=∑j=1r𝔼z∼μq​\[\|cr,q,j​\(z\)−gN,j​\(z\)\|2\]\\displaystyle=\\sum\_\{j=1\}^\{r\}\\mathbb\{E\}\_\{z\\sim\\mu\_\{q\}\}\\left\[\\left\|c\_\{r,q,j\}\(z\)\-g\_\{N,j\}\(z\)\\right\|^\{2\}\\right\]≤4​r02N​∑j=1rCj2=4​r02​Cr,q2N\.\\displaystyle\\leq\\frac\{4r\_\{0\}^\{2\}\}\{N\}\\sum\_\{j=1\}^\{r\}C\_\{j\}^\{2\}=\\frac\{4r\_\{0\}^\{2\}C\_\{r,q\}^\{2\}\}\{N\}\.\(30\)Hence,

‖\(III\)‖L2​\(ν\)≤2​r0​Cr,qN\.\\\|\\mathrm\{\(III\)\}\\\|\_\{L^\{2\}\(\\nu\)\}\\leq\\frac\{2r\_\{0\}C\_\{r,q\}\}\{\\sqrt\{N\}\}\.The result follows from Minkowski’s inequality inL2​\(ν;ℝn\)L^\{2\}\(\\nu;\\mathbb\{R\}^\{n\}\)\. ∎

## 5Computational experiments

In this section,[Section 5\.1](https://arxiv.org/html/2608.06428#S5.SS1)presents the convergence analysis of the topological DeepONet and[Section 5\.2](https://arxiv.org/html/2608.06428#S5.SS2)benchmarks the proposed architectures on the antiderivative operator considered in\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]\. The performance of the proposed methods on the Darcy flow problem is examined in[Section 5\.4](https://arxiv.org/html/2608.06428#S5.SS4),[Section 5\.6](https://arxiv.org/html/2608.06428#S5.SS6)presents the benchmark results for the fixed\-time Navier–Stokes equations in the vorticity formulation, and[Section 5\.8](https://arxiv.org/html/2608.06428#S5.SS8)extends the comparison to the time\-evolving trajectory\-prediction operator\.

### 5\.1Convergence study for topological DeepONets

##### Numerical verification of the discrete error decomposition\.

Figure[4](https://arxiv.org/html/2608.06428#S4.F4)evaluates the error terms in Theorem[4\.1](https://arxiv.org/html/2608.06428#S4.Thmtheorem1)for the antiderivative operator using

q∈\{8,16,32,64,128\}\.q\\in\\\{8,16,32,64,128\\\}\.For each test sample, we compute the operator\-level measurement errorEop,measE\_\{\\mathrm\{op,meas\}\}, its Lipschitz upper boundELipE\_\{\\mathrm\{Lip\}\}, the output\-basis truncation errorEoutE\_\{\\mathrm\{out\}\}, the neural approximation errorENNE\_\{\\mathrm\{NN\}\}, the total errorEtotE\_\{\\mathrm\{tot\}\}, and the theorem\-consistent bound

Ethm=ELip\+Eout\+ENN\.E\_\{\\mathrm\{thm\}\}=E\_\{\\mathrm\{Lip\}\}\+E\_\{\\mathrm\{out\}\}\+E\_\{\\mathrm\{NN\}\}\.Figures[4](https://arxiv.org/html/2608.06428#S4.F4)\(a\) and[4](https://arxiv.org/html/2608.06428#S4.F4)\(b\) show that the measurement\-induced error decreases rapidly as the number of functionals increases\. For both the fixed and adaptive models,Eop,measE\_\{\\mathrm\{op,meas\}\}decreases by several orders of magnitude betweenq=8q=8andq=128q=128\. The input reconstruction error in Fig\.[4](https://arxiv.org/html/2608.06428#S4.F4)\(c\) exhibits the same trend, confirming that increasingqqreduces the information lost by the measurement map\. At largeqq, the total error is no longer dominated by measurement compression; instead, the neural approximation error becomes the principal contribution, while the output\-basis truncation error remains comparatively small\. The adaptive model achieves a lower total error at the largest measurement dimensions, reaching approximately10−310^\{\-3\}atq=128q=128, compared with approximately4×10−34\\times 10^\{\-3\}for the fixed model\. The theorem\-consistent bound remains above the observed total error for every test sample\. In particular, atq=32q=32, the maximum samplewise ratio is

maxi⁡Etot\(i\)Ethm\(i\)=0\.1249,\\max\_\{i\}\\frac\{E\_\{\\mathrm\{tot\}\}^\{\(i\)\}\}\{E\_\{\\mathrm\{thm\}\}^\{\(i\)\}\}=0\.1249,with no bound violations, as shown in Fig\.[4](https://arxiv.org/html/2608.06428#S4.F4)\(d\)\. The gap between the total error and the theorem bound reflects the conservativeness of replacing the actual operator perturbation by the global Lipschitz estimate\.

### 5\.2Antiderivative operator benchmark

We next compare the fixed\- and learned\-measurement Topological DeepONets against Vanilla DeepONet and the enhanced Two\-Step DeepONet on the antiderivative operator for the aligned dataset of\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]under matched parameter budgets\. Table[2](https://arxiv.org/html/2608.06428#S5.T2)reports the error statistics over the400400test functions for the cosine\-basis variants, and Figure[5](https://arxiv.org/html/2608.06428#S5.F5)compares the mean and global relativeL2L^\{2\}errors for the cosine\- and Legendre\-basis variants\. The learned\-measurement \(adaptive\) model attains the lowest error in nearly every statistic and outperforms the Two\-Step baseline on72%72\\%of the test functions\.

Table 2:Comparison of Vanilla DeepONet, Two\-Step DeepONet, Fixed\-Measurement DeepONet, and the proposed Learned\-Measurement Topological DeepONet\. Lower values are better for all error metrics\. The win rates for the fixed\- and learned\-measurement models are computed relative to the Two\-Step DeepONet on a per\-function basis\.Notes:Bold values indicate the best result in each error row\. The proposed model uses additional trainable parameters during training because of its learned measurement layer and auxiliary reconstruction decoder, but the decoder is removed at inference\. Consequently, its inference parameter count is essentially matched to that of the Vanilla and Two\-Step baselines\.†The reported value includes the fixed \(non\-trainable\) measurement matrix; excluding it, the inference network size matches the Two\-Step baseline\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x8.png)Figure 5:Bar\-chart comparison of DeepONet variants on the antiderivative operator problem\. \(a\) Cosine\-basis variants withq=40q=40functional measurements\. \(b\) Legendre\-basis variants withq=4q=4functional measurements\. Lower values are better\.
### 5\.3Controlled operator with a known functional representation

To isolate the role of the input representation from the complexity of a PDE solution operator, we consider a controlled operator whose dependence on the input is known exactly\. Leta∈L2​\(0,1\)a\\in L^\{2\}\(0,1\)and define

𝒢​\(a\)​\(y\)=∑r=13ℓr​\(a\)​ψr​\(y\),ℓr​\(a\)=∫01a​\(x\)​ϕr​\(x\)​𝑑x,\\mathcal\{G\}\(a\)\(y\)=\\sum\_\{r=1\}^\{3\}\\ell\_\{r\}\(a\)\\,\\psi\_\{r\}\(y\),\\qquad\\ell\_\{r\}\(a\)=\\int\_\{0\}^\{1\}a\(x\)\\phi\_\{r\}\(x\)\\,dx,\(31\)where

ϕ1​\(x\)=1,ϕ2​\(x\)=2​cos⁡\(2​π​x\),ϕ3​\(x\)=2​sin⁡\(4​π​x\),\\phi\_\{1\}\(x\)=1,\\qquad\\phi\_\{2\}\(x\)=\\sqrt\{2\}\\cos\(2\\pi x\),\\qquad\\phi\_\{3\}\(x\)=\\sqrt\{2\}\\sin\(4\\pi x\),\(32\)and\{ψr\}r=13\\\{\\psi\_\{r\}\\\}\_\{r=1\}^\{3\}is a fixed orthonormal output basis\. The input functions are sampled from a higher\-dimensional random Fourier expansion,

a​\(x\)=∑j=1Jξj​φj​\(x\),ξj∼𝒩​\(0,j−2\),a\(x\)=\\sum\_\{j=1\}^\{J\}\\xi\_\{j\}\\varphi\_\{j\}\(x\),\\qquad\\xi\_\{j\}\\sim\\mathcal\{N\}\(0,j^\{\-2\}\),\(33\)so that only a small subset of the input directions influences the output\. This example is chosen because the exact task\-relevant input subspace,

𝒮true=span⁡\{ϕ1,ϕ2,ϕ3\},\\mathcal\{S\}\_\{\\mathrm\{true\}\}=\\operatorname\{span\}\\\{\\phi\_\{1\},\\phi\_\{2\},\\phi\_\{3\}\\\},is known\. It therefore permits a direct comparison between operator\-relevant functional learning and generic dimensionality\-reduction strategies\. In particular, PCA preserves high\-variance input directions, random projection preserves no task\-specific structure, and point sensors provide only local information\. By contrast, the adaptive topological models learn three global continuous functionals from a prescribed dictionary\. The controlled setting also allows the learned measurement subspace, conditioning, and drift from the structured initialization to be evaluated directly\. We compare the fixed functional representation with adaptive variants that introduce measurement learning, input reconstruction, task\-relevant decoding, orthogonality regularization, drift regularization, and either random or warm initialization\. Additional baselines include random and PCA projections, fixed and optimized sensors, an unconstrained learned bottleneck, and a full\-field MLP\. All compressed models use the same three\-dimensional latent representation, output basis, branch architecture, training data, validation criterion, and test set\. Results are averaged over multiple random seeds\.

Table 3:Controlled operator benchmark averaged over multiple random seeds\. The first column reports the mean samplewise relativeL2L^\{2\}error and its standard deviation across seeds\. The global relativeL2L^\{2\}error and the9595th percentile samplewise error are also reported\. The best value in each column is shown in bold\.As shown in[Table 3](https://arxiv.org/html/2608.06428#S5.T3), all adaptive functional models substantially outperform the fixed functional representation, sensor\-based representations, PCA, and random projection\. The Adaptive Task Decoder attains the lowest mean relative error,8\.83×10−38\.83\\times 10^\{\-3\}, and the lowestP95P\_\{95\}error,1\.63×10−21\.63\\times 10^\{\-2\}\. The unregularized Adaptive Learned Only model gives the lowest global relative error,7\.64×10−37\.64\\times 10^\{\-3\}, while the orthogonally regularized model provides a comparably low mean error of9\.05×10−39\.05\\times 10^\{\-3\}\. These small differences indicate that all three variants reliably identify the task\-relevant functional subspace, although the auxiliary task decoder improves the distribution of samplewise errors\. The task\-relevant decoder is more effective than reconstruction of the complete input\. The latter attains a mean error of1\.10×10−21\.10\\times 10^\{\-2\}, because full\-input reconstruction encourages the three\-dimensional latent representation to preserve variability that does not influence𝒢\\mathcal\{G\}\. The learned bottleneck is also competitive, with a mean error of1\.16×10−21\.16\\times 10^\{\-2\}, confirming that unconstrained task\-dependent compression can discover the relevant low\-dimensional structure\. Nevertheless, the Adaptive Task Decoder reduces the mean error by approximately23\.6%23\.6\\%relative to the learned bottleneck while retaining an explicit representation in terms of continuous input functionals\. The poor performance of PCA, random projections, and point sensors further demonstrates that low dimensionality alone is insufficient\. PCA preserves input variance rather than operator relevance, while three point values or three random projections generally cannot recover the global integral quantities defining the operator\. In contrast, adaptive functional learning reduces the mean error from56\.3%56\.3\\%for the Fixed Topological model to below1%1\\%, showing that the principal gain arises from identifying the operator\-relevant measurement subspace rather than merely increasing network capacity\.[Figure 6](https://arxiv.org/html/2608.06428#S5.F6)provides a complementary mechanistic interpretation\. Because an individual functional basis is identifiable only up to an invertible rotation within its span, the exact and learned bases are aligned before visualization\. The associated projector error and Gram\-matrix condition number quantify recovery of the true functional subspace and the linear independence of the learned measurements\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x9.png)Figure 6:Recovery of the task\-relevant functional subspace for the controlled operator\. \(a\) Exact functional basis and the aligned basis learned by the adaptive functional model for a representative seed\. Since the individual measurement functions are identifiable only up to a rotation within their span, the bases are aligned before visualization\. \(b\) Frobenius error between the learned and exact subspace projectors\. \(c\) Condition number of the learned measurement Gram matrix\. Error bars represent variability across random seeds\.
### 5\.4Heterogeneous Darcy\-flow operator

We consider the Darcy\-flow benchmark commonly used in the neural\-operator literature and in the Fourier Neural Operator study\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\]\. The steady pressure field satisfies

−∇⋅\(a​\(𝒙\)​∇u​\(𝒙\)\)=f​\(𝒙\),𝒙∈Ω,\-\\nabla\\cdot\\left\(a\(\\bm\{x\}\)\\nabla u\(\\bm\{x\}\)\\right\)=f\(\\bm\{x\}\),\\qquad\\bm\{x\}\\in\\Omega,\(34\)wherea​\(𝒙\)a\(\\bm\{x\}\)is the heterogeneous permeability field andu​\(𝒙\)u\(\\bm\{x\}\)is the corresponding pressure solution\. The learning objective is to approximate the nonlinear parameter\-to\-solution operator

𝒢:𝒜⊂𝒱=L∞\(Ω\)⟶𝒰=H01\(Ω\),a⟼u,\\mathcal\{G\}:\\mathcal\{A\}\\subset\\mathcal\{V\}=L^\{\\infty\}\(\\Omega\)\\longrightarrow\\mathcal\{U\}=H\_\{0\}^\{1\}\(\\Omega\),\\qquad a\\longmapsto u,\(35\)where

𝒜=\{a∈L∞​\(Ω\):0<amin≤a​\(𝒙\)≤amax<∞​a\.e\. in​Ω\}\.\\mathcal\{A\}=\\left\\\{a\\in L^\{\\infty\}\(\\Omega\):0<a\_\{\\min\}\\leq a\(\\bm\{x\}\)\\leq a\_\{\\max\}<\\infty\\text\{ a\.e\. in \}\\Omega\\right\\\}\.\(36\)The admissible coefficient set is equipped with the topology inherited fromL∞​\(Ω\)L^\{\\infty\}\(\\Omega\)\. Uniform ellipticity gives well\-posedness and the natural stability of the coefficient\-to\-solution map in this topology\. Although the numerical coefficient fields also belong toL2​\(Ω\)L^\{2\}\(\\Omega\)on the bounded domain, we do not identifyL2​\(Ω\)L^\{2\}\(\\Omega\)as the ambient operator topology and do not claim continuity of the unrestricted coefficient\-to\-solution map from plainL2​\(Ω\)L^\{2\}\(\\Omega\)intoH01​\(Ω\)H\_\{0\}^\{1\}\(\\Omega\)\. The original fields are generated at a resolution of421×421421\\times 421\. For the present study, the data are represented on an85×8585\\times 85grid, and we use800800,100100, and100100realizations for training, validation, and testing, respectively\. For the fixed Topological DeepONet, the permeability field is represented throughqqcontinuous linear functionals

ℓj​\(a\)=∫Ωa​\(𝒙\)​ϕj​\(𝒙\)​d𝒙,j=1,…,q,\\ell\_\{j\}\(a\)=\\int\_\{\\Omega\}a\(\\bm\{x\}\)\\phi\_\{j\}\(\\bm\{x\}\)\\,\\mathrm\{d\}\\bm\{x\},\\qquad j=1,\\ldots,q,\(37\)whereϕj∈L1​\(Ω\)\\phi\_\{j\}\\in L^\{1\}\(\\Omega\)\. For the polynomial and trigonometric dictionaries used here, the atoms are bounded on the bounded domain and therefore belong toL1​\(Ω\)L^\{1\}\(\\Omega\)\. Since the Darcy input space is𝒱=L∞​\(Ω\)\\mathcal\{V\}=L^\{\\infty\}\(\\Omega\),

\|ℓj​\(a\)\|≤‖a‖L∞​\(Ω\)​‖ϕj‖L1​\(Ω\),\|\\ell\_\{j\}\(a\)\|\\leq\\\|a\\\|\_\{L^\{\\infty\}\(\\Omega\)\}\\,\\\|\\phi\_\{j\}\\\|\_\{L^\{1\}\(\\Omega\)\},\(38\)so eachℓj∈𝒱′=\(L∞​\(Ω\)\)′\\ell\_\{j\}\\in\\mathcal\{V\}^\{\\prime\}=\(L^\{\\infty\}\(\\Omega\)\)^\{\\prime\}\. The corresponding functional\-coordinate map is

ℳq​\(a\)=\[ℓ1​\(a\),…,ℓq​\(a\)\]𝖳∈ℝq\.\\mathcal\{M\}\_\{q\}\(a\)=\\left\[\\ell\_\{1\}\(a\),\\ldots,\\ell\_\{q\}\(a\)\\right\]^\{\\mathsf\{T\}\}\\in\\mathbb\{R\}^\{q\}\.\(39\)The implementation supports Legendre, cosine, Chebyshev, and Lagrange functional dictionaries\. In the default setting, the measurement functions are constructed from a total\-degree tensor\-product Legendre dictionary,

ϕi​j​\(x,y\)=P^i​\(x\)​P^j​\(y\),i\+j≤p,\\phi\_\{ij\}\(x,y\)=\\widehat\{P\}\_\{i\}\(x\)\\widehat\{P\}\_\{j\}\(y\),\\qquad i\+j\\leq p,\(40\)whereP^i\\widehat\{P\}\_\{i\}denotes a normalized Legendre polynomial\. The measurements are evaluated numerically using quadrature on the discrete grid\. For the adaptive Topological DeepONet, the learned coordinates are linear combinations of an active set of dictionary measurements:

𝒛ad​\(a\)=Aθ​𝜶​\(a\),\\bm\{z\}\_\{\\mathrm\{ad\}\}\(a\)=A\_\{\\theta\}\\bm\{\\alpha\}\(a\),\(41\)where𝜶​\(a\)\\bm\{\\alpha\}\(a\)contains the active functional coordinates andAθA\_\{\\theta\}is a trainable linear map\. Consequently, each learned coordinate remains a continuous linear functional belonging to the span of the prescribed dictionary\. For all Two\-Step models, the pressure snapshots are projected onto a common rank\-rrbasis obtained from the singular value decomposition of the training outputs\. The branch network predicts the corresponding reduced coefficients, which are mapped back to the full pressure field using the fixed output basis\. The fixed model can therefore be summarized as

a→ℳqℝq→ℬθℝr→fixed output basisu^\.a\\xrightarrow\{\\;\\mathcal\{M\}\_\{q\}\\;\}\\mathbb\{R\}^\{q\}\\xrightarrow\{\\;\\mathcal\{B\}\_\{\\theta\}\\;\}\\mathbb\{R\}^\{r\}\\xrightarrow\{\\;\\text\{fixed output basis\}\\;\}\\widehat\{u\}\.\(42\)Further details on the data preprocessing, dictionary construction, weighted orthonormalization, measurement selection, adaptive\-measurement regularization, reduced output representation, and evaluation metrics are provided in Appendix[B](https://arxiv.org/html/2608.06428#A2)\. The compared architectures and their input dimensions are illustrated in Fig\.[7](https://arxiv.org/html/2608.06428#S5.F7)\. We compare the following models:

1. 1\.full\-field Two\-Step, using all72257225permeability values;
2. 2\.Sensor Two\-Step, usingq=32q=32point evaluations;
3. 3\.Fixed Topological DeepONet, usingq=32q=32prescribed functional coordinates;
4. 4\.Adaptive Topological DeepONet, usingq=32q=32learned functional coordinates;
5. 5\.Vanilla DeepONet, with jointly trained branch and trunk networks\.

The most controlled comparison is between Sensor Two\-Step and Fixed Topological DeepONet\. Both use the same input dimensionq=32q=32, branch architecture, reduced output basis, and144,915144\{,\}915inference parameters\. They differ only in their input representation:

\{a​\(𝒙j\)\}j=132⏟local point evaluationsversus\{ℓj​\(a\)\}j=132⏟global functional coordinates,\\underbrace\{\\left\\\{a\(\\bm\{x\}\_\{j\}\)\\right\\\}\_\{j=1\}^\{32\}\}\_\{\\text\{local point evaluations\}\}\\qquad\\text\{versus\}\\qquad\\underbrace\{\\left\\\{\\ell\_\{j\}\(a\)\\right\\\}\_\{j=1\}^\{32\}\}\_\{\\text\{global functional coordinates\}\},\(43\)with the functional coordinatesℓj​\(a\)\\ell\_\{j\}\(a\)defined in \([37](https://arxiv.org/html/2608.06428#S5.E37)\)\. The complete quantitative results are reported in Table[4](https://arxiv.org/html/2608.06428#S5.T4)\. The corresponding comparisons of global relativeL2L^\{2\}error, training time, and inference parameter count are shown in Fig\.[8](https://arxiv.org/html/2608.06428#S5.F8)\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x10.png)Figure 7:Three operator\-learning architectures compared\. Two\-Step takes the full input field \(72257225values\) as branch input \(solid box\)\. Sensor Two\-Step and Fixed Topological DeepONet both compress to3232dimensions and share an identical downstream architecture \(dashed box\) — differing only in how the3232inputs are formed: point values vs\. global functional measurements\.Fixed Topological DeepONet achieves a global relativeL2L^\{2\}error of5\.88%5\.88\\%, compared with10\.39%10\.39\\%for Sensor Two\-Step\. This corresponds to a relative error reduction of

0\.1038841−0\.05875510\.1038841×100≈43\.4%\.\\frac\{0\.1038841\-0\.0587551\}\{0\.1038841\}\\times 100\\approx 43\.4\\%\.\(44\)The full\-field Two\-Step, Adaptive Topological, and Vanilla DeepONet models achieve global relative errors of8\.31%8\.31\\%,7\.27%7\.27\\%, and9\.21%9\.21\\%, respectively\. The reduced models compress the branch input from72257225values to3232features, corresponding to≈225\.8\\approx 225\.8times input compression\. The results therefore indicate that, for this Darcy dataset, global functional coordinates provide a substantially more informative reduced representation than sparse point evaluations at the same input dimension and parameter count\. The interpretation and limitations of this comparison are discussed further in Appendix[B](https://arxiv.org/html/2608.06428#A2)\.

Table 4:Darcy operator\-learning test performance and computational cost\. The highlighted columns form the controlled comparison: Sensor Two\-Step and Fixed Topological DeepONet use the same input dimension, branch architecture, output decoder, and number of inference parameters\.![Refer to caption](https://arxiv.org/html/2608.06428v1/x11.png)Figure 8:Comparison of DeepONet variants for the Darcy problem\. Abbreviations: S\-2ST = Sensor Two\-Step, Fixed = Fixed Topological DeepONet, Adapt = Adaptive Topological DeepONet, 2ST = full Two\-Step DeepONet, and VAN = Vanilla DeepONet\. Panel \(a\) reports the global relativeL2L^\{2\}error, panel \(b\) the measured training time, and panel \(c\) the inference parameter count\.
### 5\.5Heterogeneous\-discretization Darcy benchmark

We next evaluate whether the learned input representation transfers across spatial discretizations\. The Darcy coefficient fields are observed on a mixture of33×3333\\times 33,49×4949\\times 49,65×6565\\times 65, and85×8585\\times 85grids during training\. At test time, the models are evaluated on the reference85×8585\\times 85grid and on previously unseen57×5757\\times 57,73×7373\\times 73, and97×9797\\times 97grids\. Because the available dataset is stored on a common fine grid, the heterogeneous observations are generated by sampling each realization at the prescribed native resolution\. The Fixed and Adaptive Topological DeepONets evaluate the same continuous linear functionals on every native grid using grid\-dependent quadrature\. Hence, the branch input dimension remains unchanged across discretizations\. We compare these models with an Interpolated Two\-Step baseline, which first maps every native observation to the common85×8585\\times 85grid, and a Sensor Two\-Step baseline using a fixed set of physical sensor locations\.

Table 5:Mean relativeL2L^\{2\}errors for the heterogeneous\-discretization Darcy benchmark\. The models are tested on the reference grid, unseen discretizations, noisy observations, and randomly missing observations\. The best result in each row is shown in bold\.As shown in Table[5](https://arxiv.org/html/2608.06428#S5.T5), both functional models exhibit nearly discretization\-independent accuracy\. The Fixed model achieves mean relative errors of5\.56%5\.56\\%,5\.54%5\.54\\%,5\.54%5\.54\\%, and5\.58%5\.58\\%on the85×8585\\times 85, unseen57×5757\\times 57, unseen73×7373\\times 73, and unseen97×9797\\times 97grids, respectively\. The Adaptive model shows a similarly small variation, with errors between5\.55%5\.55\\%and5\.64%5\.64\\%\. In contrast, the Interpolated Two\-Step and Sensor Two\-Step baselines remain less accurate across all test resolutions\. The functional representations are also insensitive to moderate observation noise\. On the unseen57×5757\\times 57grid, increasing the noise level from0%0\\%to5%5\\%changes the Fixed\-model error only from5\.536%5\.536\\%to5\.539%5\.539\\%\. The distinction becomes more pronounced under missing observations\. With30%30\\%of the native\-grid values removed, the Interpolated Two\-Step error increases from6\.64%6\.64\\%to11\.63%11\.63\\%, whereas the Fixed\-model error increases only from5\.54%5\.54\\%to6\.10%6\.10\\%\. This corresponds to approximately75%75\\%degradation for the interpolated baseline but only10%10\\%for the Fixed Topological DeepONet\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x12.png)Figure 9:Heterogeneous\-discretization Darcy results\. Panel \(a\) compares the mean relativeL2L^\{2\}error on the reference85×8585\\times 85grid and on unseen57×5757\\times 57,73×7373\\times 73, and97×9797\\times 97grids\. Panel \(b\) reports robustness on the unseen57×5757\\times 57grid under additive noise and randomly missing observations\. The functional models retain nearly constant accuracy across discretizations and degrade much less than the interpolation and sensor baselines under incomplete observations\.[Figure 9](https://arxiv.org/html/2608.06428#S5.F9)confirms that the functional coordinates provide a portable representation of the input field\. Since the same continuous functionals are evaluated directly on each native grid, the learned branch network does not depend on a particular discretization\. These results therefore demonstrate both discretization transfer and robustness to imperfect observations, while avoiding the need to reconstruct every input on a common reference mesh\.

### 5\.6Fixed\-time Navier–Stokes benchmark

We consider the two\-dimensional incompressible Navier\-Stokes equations in vorticity form on the periodic unit squareD=\(0,1\)2D=\(0,1\)^\{2\}:

∂ω∂t\+𝒖⋅∇ω=ν​Δ​ω\+f,∇⋅𝒖=0,\\frac\{\\partial\\omega\}\{\\partial t\}\+\\bm\{u\}\\cdot\\nabla\\omega=\\nu\\Delta\\omega\+f,\\qquad\\nabla\\cdot\\bm\{u\}=0,\(45\)whereω\\omegais the scalar vorticity,𝒖=∇⟂\(−Δ\)−1ω\\bm\{u\}=\\nabla^\{\\perp\}\(\-\\Delta\)^\{\-1\}\\omega,ν=10−3\\nu=10^\{\-3\},ffis a prescribed forcing term, and periodic boundary conditions are imposed in both spatial directions\. Following the standard neural\-operator benchmark of\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\], the data consist of 5000 realizations sampled on a64×6464\\times 64grid at 50 stored time levels; a representative vorticity evolution is shown in Fig\.[11](https://arxiv.org/html/2608.06428#S5.F11)\. For the fixed\-time Navier–Stokes benchmark, we learn

𝒢0→10:Lper2​\(D\)⟶Lper2​\(D\),ω​\(⋅,t0\)↦ω​\(⋅,t10\),\\mathcal\{G\}\_\{0\\rightarrow 10\}:L\_\{\\mathrm\{per\}\}^\{2\}\(D\)\\longrightarrow L\_\{\\mathrm\{per\}\}^\{2\}\(D\),\\qquad\\omega\(\\cdot,t\_\{0\}\)\\mapsto\\omega\(\\cdot,t\_\{10\}\),\(46\)whereD=\(0,1\)2D=\(0,1\)^\{2\}\. After discretization on a64×6464\\times 64grid,

𝒢0→10h:ℝ4096⟶ℝ4096\.\\mathcal\{G\}\_\{0\\rightarrow 10\}^\{h\}:\\mathbb\{R\}^\{4096\}\\longrightarrow\\mathbb\{R\}^\{4096\}\.
![Refer to caption](https://arxiv.org/html/2608.06428v1/x13.png)Figure 10:Representative multiscale measurement atoms used in the fixed functional dictionary\. Gaussian measurements capture localized averages, directionally weighted Gaussian atoms emphasize directional spatial variation, and the Laplacian\-of\-Gaussian atom emphasizes localized curvature and vortical structure\.The realizations are divided into 4000 training, 500 validation, and 500 test samples\. After vectorization, both the input and output fields belong toℝ4096\\mathbb\{R\}^\{4096\}\. We compare five operator\-learning formulations: Full\-field Two\-Step, Sensor Two\-Step, Fixed Topological Two\-Step, Adaptive Topological Two\-Step, and Vanilla DeepONet\. The full\-field models receive the complete64×6464\\times 64input\. The compressed models useq=128q=128measurements, corresponding to a32×32\\timesreduction in input dimension\. The Sensor model uses pointwise evaluations

zjsens=ω​\(𝒙sj,t0\),j=1,…,q,z\_\{j\}^\{\\mathrm\{sens\}\}=\\omega\(\\bm\{x\}\_\{s\_\{j\}\},t\_\{0\}\),\\qquad j=1,\\ldots,q,at prescribed sensor locations𝒙s1,…,𝒙sq\\bm\{x\}\_\{s\_\{1\}\},\\ldots,\\bm\{x\}\_\{s\_\{q\}\}, whereas the Fixed Topological model uses distributed linear functionals

zjfix=⟨ω​\(⋅,t0\),ψij⟩,j=1,…,q,z\_\{j\}^\{\\mathrm\{fix\}\}=\\left\\langle\\omega\(\\cdot,t\_\{0\}\),\\psi\_\{i\_\{j\}\}\\right\\rangle,\\qquad j=1,\\ldots,q,\(47\)selected from a localized multiscale dictionary\{ψk\}k=1K\\\{\\psi\_\{k\}\\\}\_\{k=1\}^\{K\}\. Representative multiscale atoms are shown in[Figure 10](https://arxiv.org/html/2608.06428#S5.F10)\. The Adaptive Topological model replaces the fixed selection by

𝐀θ=𝐀0\+Δ​𝐀θ,𝐳ad=𝐡​𝐀θ,\\mathbf\{A\}\_\{\\theta\}=\\mathbf\{A\}\_\{0\}\+\\Delta\\mathbf\{A\}\_\{\\theta\},\\qquad\\mathbf\{z\}^\{\\mathrm\{ad\}\}=\\mathbf\{h\}\\,\\mathbf\{A\}\_\{\\theta\},\(48\)where𝐡=\[⟨ω,ψ1⟩,…,⟨ω,ψK⟩\]\\mathbf\{h\}=\[\\langle\\omega,\\psi\_\{1\}\\rangle,\\ldots,\\langle\\omega,\\psi\_\{K\}\\rangle\]collects allKKdictionary measurements,𝐀0∈ℝK×q\\mathbf\{A\}\_\{0\}\\in\\mathbb\{R\}^\{K\\times q\}is the fixed selection matrix, andΔ​𝐀θ\\Delta\\mathbf\{A\}\_\{\\theta\}is a trainable correction\. The initializationΔ​𝐀θ=0\\Delta\\mathbf\{A\}\_\{\\theta\}=0ensures that the adaptive and fixed models coincide at epoch zero\. For the Two\-Step methods, the target fields are represented in a truncated SVD basis\. If

𝐘c=𝐔​𝚺​𝐕⊤\\mathbf\{Y\}\_\{c\}=\\mathbf\{U\}\\mathbf\{\\Sigma\}\\mathbf\{V\}^\{\\top\}is the SVD of the centered training\-output matrix𝐘c\\mathbf\{Y\}\_\{c\}\(Appendix[C](https://arxiv.org/html/2608.06428#A3)\), then the firstrrright singular vectors form𝐐∈ℝ4096×r\\mathbf\{Q\}\\in\\mathbb\{R\}^\{4096\\times r\}, and

𝐲≈𝐲¯\+𝐐𝐜,\\mathbf\{y\}\\approx\\overline\{\\mathbf\{y\}\}\+\\mathbf\{Q\}\\mathbf\{c\},\(49\)where𝐲¯\\overline\{\\mathbf\{y\}\}is the mean training target field and𝐜\\mathbf\{c\}the reduced coefficient vector\. The branch network predicts the coefficient vector𝐜^\\widehat\{\\mathbf\{c\}\}, after which the target field is reconstructed using \([49](https://arxiv.org/html/2608.06428#S5.E49)\)\. For the common\-backbone comparison, all five models use an 2D Fourier spectral\-layer branch\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\]; Vanilla DeepONet additionally uses a coordinate trunk network as in\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]\. The Two\-Step models are trained using a coefficient\-space mean\-squared error augmented by a relative coefficient loss,

ℒTwoStep=1B​r​∑i=1B‖𝐜^\(i\)−𝐜\(i\)‖22\+λrel​1B​∑i=1B‖𝐜^\(i\)−𝐜\(i\)‖2‖𝐜\(i\)‖2\+ε,\\mathcal\{L\}\_\{\\mathrm\{TwoStep\}\}=\\frac\{1\}\{Br\}\\sum\_\{i=1\}^\{B\}\\left\\\|\\widehat\{\\mathbf\{c\}\}^\{\(i\)\}\-\\mathbf\{c\}^\{\(i\)\}\\right\\\|\_\{2\}^\{2\}\+\\lambda\_\{\\mathrm\{rel\}\}\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\frac\{\\left\\\|\\widehat\{\\mathbf\{c\}\}^\{\(i\)\}\-\\mathbf\{c\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\left\\\|\\mathbf\{c\}^\{\(i\)\}\\right\\\|\_\{2\}\+\\varepsilon\},\(50\)whereBBis the mini\-batch size,λrel≥0\\lambda\_\{\\mathrm\{rel\}\}\\geq 0a fixed weight, andε\>0\\varepsilon\>0a small constant\. The adaptive model includes the additional regularization

λdrift​‖Δ​𝐀θ‖F2\.\\lambda\_\{\\mathrm\{drift\}\}\\left\\\|\\Delta\\mathbf\{A\}\_\{\\theta\}\\right\\\|\_\{F\}^\{2\}\.Performance is reported using the global relativeL2L^\{2\}error

εglobal=‖𝐘^−𝐘‖F‖𝐘‖F,\\varepsilon\_\{\\mathrm\{global\}\}=\\frac\{\\left\\\|\\widehat\{\\mathbf\{Y\}\}\-\\mathbf\{Y\}\\right\\\|\_\{F\}\}\{\\left\\\|\\mathbf\{Y\}\\right\\\|\_\{F\}\},\(51\)where𝐘\\mathbf\{Y\}and𝐘^\\widehat\{\\mathbf\{Y\}\}stack the true and predicted test targets, together with the mean sample\-wise relativeL2L^\{2\}error\. Further details on the multiscale dictionary, functional selection, adaptive parameterization, SVD construction, network architectures, and optimization settings are provided in Appendix[C](https://arxiv.org/html/2608.06428#A3)\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/figures/11_ins_vort.png)Figure 11:Vorticity fields for a representative realization of the two\-dimensional incompressible Navier–Stokes benchmark\. The panels show the initial condition, the reference target field, and the predictions of the Fixed Topological, Adaptive Topological, and Sensor Two\-Step models, respectively\.The empirical cumulative distribution of the sample\-wise relativeL2L^\{2\}error \(Fig\.[12](https://arxiv.org/html/2608.06428#S5.F12)\) reveals a clear advantage of the topological models in the low\-error regime, which is not fully captured by aggregate mean\-error metrics alone\. At the practically relevant2%2\\%threshold, the Adaptive Topological model attains76\.0%76\.0\\%coverage \(380380of the500500test samples\), compared with73\.0%73\.0\\%for Fixed Topological,65\.0%65\.0\\%for Sensor Two\-Step,60\.2%60\.2\\%for Vanilla DeepONet, and44\.8%44\.8\\%for the full\-field Two\-Step model; coverages at the1%1\\%and3%3\\%thresholds are reported in Table[6](https://arxiv.org/html/2608.06428#S5.T6)\. Thus, relative to Vanilla, the adaptive topological representation improves the fraction of highly accurate predictions by15\.815\.8percentage points at the2%2\\%tolerance\. These results indicate that the learned topological measurements improve not only the average predictive accuracy but also the consistency of the model across the test ensemble, while retaining a compressed input representation of only128128measurements instead of the full40964096\-dimensional field\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x14.png)Figure 12:Empirical cumulative distribution function \(ECDF\) of per\-sample relativeL2L^\{2\}errors on the Navier–Stokes test set \(N=500N=500\) for five operator\-learning models\. The vertical dashed line marks the2%2\\%error threshold; filled markers indicate each model’s cumulative fraction at that threshold\. Cumulative fractions at the1%1\\%,2%2\\%, and3%3\\%thresholds are summarised in Table[6](https://arxiv.org/html/2608.06428#S5.T6)\.Table 6:Percentage of test samples below selected relativeL2L^\{2\}\-error thresholds for the fixed\-time Navier–Stokes operator\.
### 5\.7Comparison with a parameter\-matched Fourier neural operator\.

We additionally compare the proposed models with a direct Fourier neural operator \(FNO\)\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\]\. The FNO is a particularly strong baseline for the present fixed\-time Navier–Stokes benchmark because the vorticity equation is solved on a uniform periodic domain\. Fourier convolution layers naturally respect the periodic geometry, provide direct access to global spatial frequencies, and impose a translation\-equivariant spectral inductive bias that is well aligned with periodic vorticity evolution\. This benchmark therefore constitutes a favorable setting for the FNO and provides a stringent comparison for the proposed functional\-coordinate models\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x15.png)Figure 13:Schematic comparison between the tensor\-grid representation used by the Fourier neural operator and the continuous\-functional representation used by the Topological DeepONet for fixed time Navier\-Stokes operator\. The FNO receives the complete64×6464\\times 64field and applies a discrete Fourier transform and spectral convolution on the uniform periodic grid\. The Topological DeepONet instead maps the inputv∈\(𝒱,\{pα\}α∈A\)v\\in\(\\mathcal\{V\},\\\{p\_\{\\alpha\}\\\}\_\{\\alpha\\in A\}\)toq=128q=128coordinates generated by continuous linear functionalsℓj∈𝒱′\\ell\_\{j\}\\in\\mathcal\{V\}^\{\\prime\}\. Numerically, each functional can be evaluated on the available discretization using grid\-dependent quadrature, without requiring the branch input to contain the complete tensor\-grid field\.Table[7](https://arxiv.org/html/2608.06428#S5.T7)reports the mean and sample standard deviation over three independent runs with random seeds12341234,23452345, and34563456\. The FNO contains113,781113\{,\}781trainable parameters, which is within approximately5\.4%5\.4\\%of the120,220120\{,\}220\-parameter Fixed Topological and Sensor Two\-Step models\. The FNO receives the complete64×6464\\times 64vorticity field, corresponding to40964096input values, whereas the Fixed and Adaptive Topological DeepONets use onlyq=128q=128functional coordinates\. The functional models therefore provide a32×32\\timesreduction in branch\-input dimension\.

Table 7:Multi\-seed comparison for the fixed\-time Navier–Stokes benchmark\. Results are reported as mean±\\pmsample standard deviation over three independent runs\. For the Adaptive Topological DeepONet, the reported training time is end to end and includes both the initial Fixed Topological training and the subsequent adaptive optimization\. Lower values are better for prediction error, training time, and peak GPU memory\.As shown in Table[7](https://arxiv.org/html/2608.06428#S5.T7), the FNO achieves the lowest average predictive error, with mean and global relativeL2L^\{2\}errors of0\.832%±0\.172%0\.832\\%\\pm 0\.172\\%and0\.991%±0\.201%0\.991\\%\\pm 0\.201\\%, respectively\. This result is consistent with the strong alignment between the Fourier architecture and the uniform periodic geometry of the benchmark\. Among the DeepONet\-based models, the Adaptive Topological DeepONet performs best, attaining1\.685%±0\.017%1\.685\\%\\pm 0\.017\\%mean relative error and1\.992%±0\.026%1\.992\\%\\pm 0\.026\\%global relative error\. The Fixed Topological model follows with corresponding errors of1\.824%±0\.018%1\.824\\%\\pm 0\.018\\%and2\.110%±0\.030%2\.110\\%\\pm 0\.030\\%\. Although the FNO provides the best mean accuracy, it also exhibits substantially greater sensitivity to random initialization\. Its standard deviation in mean relative error is approximately0\.1720\.172percentage points, nearly ten times the0\.0170\.017\-percentage\-point standard deviation of the Adaptive Topological model\. A similar difference is observed for the global relative error, for which the FNO standard deviation is0\.2010\.201percentage points compared with0\.0260\.026percentage points for the Adaptive model\. The individual FNO mean errors range from approximately0\.682%0\.682\\%to1\.020%1\.020\\%, whereas the Adaptive Topological errors remain within the much narrower interval1\.665%1\.665\\%–1\.698%1\.698\\%\. Thus, the FNO offers superior average accuracy but noticeably greater seed\-to\-seed variability, while the functional\-coordinate models provide more consistent performance across independent training runs\. Because only three seeds are available, this observation should be interpreted as evidence of relative training stability rather than as a definitive statistical conclusion\. The gain in FNO accuracy is also accompanied by greater computational cost\. Its mean training time is2021\.8±98\.92021\.8\\pm 98\.9s, compared with1012\.7±38\.41012\.7\\pm 38\.4s for the complete end\-to\-end Adaptive Topological pipeline\. The FNO therefore requires approximately twice the training time\. Its peak GPU memory is approximately323323MB, compared with about3030MB for the Adaptive Topological model, corresponding to approximately10\.7×10\.7\\timesgreater memory consumption\. Figure[14](https://arxiv.org/html/2608.06428#S5.F14)visualizes these accuracy–cost relationships using the multi\-seed means and one\-standard\-deviation error bars\. The FNO occupies the accuracy\-optimal region but has the largest uncertainty in predictive error and the highest training\-time and memory costs\. The Topological DeepONets occupy a different part of the trade\-off: they use a32×32\\timescompressed functional representation, exhibit much smaller seed\-to\-seed variation, and require substantially less GPU memory\. Consequently, the FNO is preferable when maximum average accuracy on a fixed periodic grid is the principal objective, whereas the Topological DeepONets provide a more compact, memory\-efficient, and empirically stable representation that can be evaluated across discretizations\.

![Refer to caption](https://arxiv.org/html/2608.06428v1/x16.png)Figure 14:Accuracy–cost comparison for the fixed\-time Navier–Stokes benchmark\. The numerical values are reported in Table[7](https://arxiv.org/html/2608.06428#S5.T7)\. Markers in panels \(a\) and \(b\) show averages over three independent random seeds, and error bars denote one sample standard deviation\. \(a\) Mean relativeL2L^\{2\}error versus end\-to\-end training time\. The Adaptive Topological time includes both the initial Fixed Topological training and the subsequent adaptive optimization\. \(b\) Mean relativeL2L^\{2\}error versus inference parameter count\. \(c\) Peak GPU memory\. The FNO achieves the lowest mean error on the uniform periodic benchmark, but it exhibits substantially greater seed\-to\-seed variation and requires the largest training time and memory\. The Fixed and Adaptive Topological models use onlyq=128q=128functional coordinates instead of the full40964096\-value field, corresponding to a32×32\\timesinput compression\.
### 5\.8Time\-evolving Navier–Stokes operator

We consider the same two\-dimensional incompressible Navier–Stokes vorticity problem \([45](https://arxiv.org/html/2608.06428#S5.E45)\) on the periodic unit squareD=\(0,1\)2D=\(0,1\)^\{2\}, withν=10−3\\nu=10^\{\-3\}and prescribed forcingff, but now learn a trajectory\-prediction operator rather than a single\-snapshot forecast\. Let

𝒱=Lper2​\(D\)\\mathcal\{V\}=L\_\{\\mathrm\{per\}\}^\{2\}\(D\)\(52\)denote the space of square\-integrable periodic vorticity fields\. The time\-evolving operator maps an input vorticity history to a future vorticity trajectory:

𝒢:𝒱Tin⟶𝒱Tout,\{ω​\(⋅,tn\)\}n=0Tin−1⟼\{ω​\(⋅,tn\)\}n=TstartTend−1\.\\mathcal\{G\}:\\mathcal\{V\}^\{T\_\{\\mathrm\{in\}\}\}\\longrightarrow\\mathcal\{V\}^\{T\_\{\\mathrm\{out\}\}\},\\qquad\\left\\\{\\omega\(\\cdot,t\_\{n\}\)\\right\\\}\_\{n=0\}^\{T\_\{\\mathrm\{in\}\}\-1\}\\longmapsto\\left\\\{\\omega\(\\cdot,t\_\{n\}\)\\right\\\}\_\{n=T\_\{\\mathrm\{start\}\}\}^\{T\_\{\\mathrm\{end\}\}\-1\}\.\(53\)In the present experiment,

Tin=10,Tstart=10,Tout=40,T\_\{\\mathrm\{in\}\}=10,\\qquad T\_\{\\mathrm\{start\}\}=10,\\qquad T\_\{\\mathrm\{out\}\}=40,\(54\)so that

𝒢10→40:\{ω​\(⋅,t0\),…,ω​\(⋅,t9\)\}⟼\{ω​\(⋅,t10\),…,ω​\(⋅,t49\)\}\.\\mathcal\{G\}\_\{10\\rightarrow 40\}:\\left\\\{\\omega\(\\cdot,t\_\{0\}\),\\ldots,\\omega\(\\cdot,t\_\{9\}\)\\right\\\}\\longmapsto\\left\\\{\\omega\(\\cdot,t\_\{10\}\),\\ldots,\\omega\(\\cdot,t\_\{49\}\)\\right\\\}\.\(55\)The dataset contains50005000realizations sampled on a64×6464\\times 64spatial grid at5050stored time levels\. We use40004000,500500, and500500realizations for training, validation, and testing, respectively\. After discretization, the input and output trajectories belong to

ℝ64×64×10andℝ64×64×40,\\mathbb\{R\}^\{64\\times 64\\times 10\}\\qquad\\text\{and\}\\qquad\\mathbb\{R\}^\{64\\times 64\\times 40\},\(56\)respectively\. For the Fixed Topological Two\-Step model, each input snapshot is represented throughqqcontinuous linear functionals

ℓj​\(ω​\(⋅,tn\)\)=⟨ω​\(⋅,tn\),ψj⟩L2​\(D\)=∫Dω​\(𝒙,tn\)​ψj​\(𝒙\)​d𝒙,j=1,…,q,\\ell\_\{j\}\\\!\\left\(\\omega\(\\cdot,t\_\{n\}\)\\right\)=\\left\\langle\\omega\(\\cdot,t\_\{n\}\),\\psi\_\{j\}\\right\\rangle\_\{L^\{2\}\(D\)\}=\\int\_\{D\}\\omega\(\\bm\{x\},t\_\{n\}\)\\psi\_\{j\}\(\\bm\{x\}\)\\,\\mathrm\{d\}\\bm\{x\},\\qquad j=1,\\ldots,q,\(57\)where

ψj∈Lper2​\(D\)\.\\psi\_\{j\}\\in L\_\{\\mathrm\{per\}\}^\{2\}\(D\)\.\(58\)By the Cauchy–Schwarz inequality,

\|ℓj​\(ω\)\|≤‖ω‖L2​\(D\)​‖ψj‖L2​\(D\),\\left\|\\ell\_\{j\}\(\\omega\)\\right\|\\leq\\\|\\omega\\\|\_\{L^\{2\}\(D\)\}\\\|\\psi\_\{j\}\\\|\_\{L^\{2\}\(D\)\},\(59\)and therefore

ℓj∈𝒱′\.\\ell\_\{j\}\\in\\mathcal\{V\}^\{\\prime\}\.\(60\)The measurement functions are selected from a prescribed localized multiscale dictionary containing centered Gaussian, directional Gaussian\-weighted, and Laplacian\-of\-Gaussian\-type atoms at several spatial scales\. The default implementation uses

measurements at each of the ten input times\. The resulting functional history is

ℳq​\(ω\)=\[ℓj​\(ω​\(⋅,tn\)\)\]n=0,…,Tin−1j=1,…,q∈ℝTin×q\.\\mathcal\{M\}\_\{q\}\(\\omega\)=\\left\[\\ell\_\{j\}\\\!\\left\(\\omega\(\\cdot,t\_\{n\}\)\\right\)\\right\]\_\{\\begin\{subarray\}\{c\}n=0,\\ldots,T\_\{\\mathrm\{in\}\}\-1\\\\ j=1,\\ldots,q\\end\{subarray\}\}\\in\\mathbb\{R\}^\{T\_\{\\mathrm\{in\}\}\\times q\}\.\(62\)A regularized dual synthesis map reconstructs a spatial proxy history from these measurements\. The proxy is then supplied to a three\-dimensional Fourier spectral\-layer branch encoder acting jointly on the two spatial coordinates and the input time coordinate\. The Fixed Topological Two\-Step pipeline is therefore

\{ω​\(⋅,tn\)\}n=09→ℳqℝ10×128→dual synthesisω~q→3D FNO branch𝒄^→output basisω^\.\\left\\\{\\omega\(\\cdot,t\_\{n\}\)\\right\\\}\_\{n=0\}^\{9\}\\xrightarrow\{\\;\\mathcal\{M\}\_\{q\}\\;\}\\mathbb\{R\}^\{10\\times 128\}\\xrightarrow\{\\;\\text\{dual synthesis\}\\;\}\\widetilde\{\\omega\}\_\{q\}\\xrightarrow\{\\;\\text\{3D FNO branch\}\\;\}\\widehat\{\\bm\{c\}\}\\xrightarrow\{\\;\\text\{output basis\}\\;\}\\widehat\{\\omega\}\.\(63\)For the Adaptive Topological Two\-Step model, the fixed measurements are corrected by learned linear combinations of the complete dictionary coordinates\. If

𝒉​\(tn\)=\[ℓ1​\(ω​\(⋅,tn\)\),…,ℓK​\(ω​\(⋅,tn\)\)\]\\bm\{h\}\(t\_\{n\}\)=\\left\[\\ell\_\{1\}\(\\omega\(\\cdot,t\_\{n\}\)\),\\ldots,\\ell\_\{K\}\(\\omega\(\\cdot,t\_\{n\}\)\)\\right\]\(64\)contains all dictionary measurements andℐq\\mathcal\{I\}\_\{q\}denotes the indices of the selected fixed measurements, then

𝒛ad​\(tn\)=𝒉ℐq​\(tn\)\+𝒉​\(tn\)​Δ​Aθ,Δ​Aθ∈ℝK×q\.\\bm\{z\}\_\{\\mathrm\{ad\}\}\(t\_\{n\}\)=\\bm\{h\}\_\{\\mathcal\{I\}\_\{q\}\}\(t\_\{n\}\)\+\\bm\{h\}\(t\_\{n\}\)\\Delta A\_\{\\theta\},\\qquad\\Delta A\_\{\\theta\}\\in\\mathbb\{R\}^\{K\\times q\}\.\(65\)The initializationΔ​Aθ=0\\Delta A\_\{\\theta\}=0makes the adaptive and fixed representations identical before adaptive training\. For all Two\-Step models, the complete output trajectory is flattened and projected onto a rank\-rrbasis obtained from the singular value decomposition of the centered training trajectories\. Writing

Yc=U​Σ​V𝖳,Q=V:,1:r,Y\_\{c\}=U\\Sigma V^\{\\mathsf\{T\}\},\\qquad Q=V\_\{:,1:r\},\(66\)the trajectory is approximated as

𝒚≈𝒚¯\+Q​𝒄\.\\bm\{y\}\\approx\\overline\{\\bm\{y\}\}\+Q\\bm\{c\}\.\(67\)The spectral branch predicts the reduced coefficient vector𝒄^\\widehat\{\\bm\{c\}\}, after which the full64×64×4064\\times 64\\times 40trajectory is reconstructed using the fixed basisQQ\. Further details on the multiscale dictionary, predictive selection, dual synthesis, adaptive parameterization, output SVD construction, and training procedure are provided in Appendix[D](https://arxiv.org/html/2608.06428#A4)\. Test errors for the1010\-to\-4040Navier–Stokes trajectory\-prediction problem are summarized in the preceding discussion, whereas[Table 8](https://arxiv.org/html/2608.06428#S5.T8)reports validation, optimization, timing, and model\-size statistics\. The Adaptive Topological Two\-Step model achieves the best performance across all four reported error measures, attaining a mean relativeL2L^\{2\}error of4\.5729×10−24\.5729\\times 10^\{\-2\}, a global relativeL2L^\{2\}error of4\.8087×10−24\.8087\\times 10^\{\-2\}, an RMSE of4\.5017×10−24\.5017\\times 10^\{\-2\}, and an MAE of3\.2172×10−23\.2172\\times 10^\{\-2\}\. Relative to the Fixed Topological model, adaptive measurements reduce the mean relativeL2L^\{2\}error by approximately5\.3%5\.3\\%\. The improvement is approximately8\.2%8\.2\\%relative to Sensor Two\-Step and4\.2%4\.2\\%relative to the full\-field Two\-Step model\. Compared with Vanilla DeepONet, the adaptive model reduces the mean relativeL2L^\{2\}error by approximately48\.1%48\.1\\%\. The Fixed Topological model also outperforms the Sensor Two\-Step model despite using the same128128measurements at each input time\. Its mean relativeL2L^\{2\}error is4\.8273×10−24\.8273\\times 10^\{\-2\}, compared with4\.9801×10−24\.9801\\times 10^\{\-2\}for the sensor representation\. This result indicates that distributed multiscale functionals provide a more informative compressed description of the vorticity history than an equal number of pointwise samples\. The fixed, adaptive, and sensor models use128128measurements for each of the ten input snapshots, resulting in an effective input dimension of12801280\. In contrast, the full\-field models use all64×64×10=4096064\\times 64\\times 10=40960input values\. The compressed models therefore achieve a32×32\\timesreduction in input dimension\. Despite this substantial compression, the Adaptive Topological model outperforms the full\-field Two\-Step model, while the Fixed Topological model remains competitive with it\. In particular, the adaptive model obtains a mean relativeL2L^\{2\}error of4\.5729×10−24\.5729\\times 10^\{\-2\}, compared with4\.7714×10−24\.7714\\times 10^\{\-2\}for Full Two\-Step\. This result suggests that the learned functional coordinates suppress redundant components of the input history while retaining dynamically relevant multiscale information\.[Table 9](https://arxiv.org/html/2608.06428#S5.T9)and[Figure 15](https://arxiv.org/html/2608.06428#S5.F15)report the percentage of test trajectories whose sample\-wise relativeL2L^\{2\}error falls below selected thresholds\. The Adaptive Topological model achieves the highest coverage at every reported threshold\. In particular,38\.0%38\.0\\%of its test predictions have relative error below4%4\\%, compared with28\.8%28\.8\\%for the Full Two\-Step model,27\.0%27\.0\\%for the Fixed Topological model, and19\.6%19\.6\\%for the Sensor Two\-Step model\. No Vanilla DeepONet prediction falls below the4%4\\%threshold\. At the5%5\\%threshold, the Adaptive Topological model successfully predicts78\.4%78\.4\\%of the test trajectories\. The corresponding coverage is72\.4%72\.4\\%for Fixed Topological,72\.2%72\.2\\%for Full Two\-Step, and64\.4%64\.4\\%for Sensor Two\-Step, whereas Vanilla DeepONet still has no test sample below this threshold\. These results show that adaptive functional measurements improve not only the mean prediction error but also the consistency of the model across individual test realizations\. At the more permissive10%10\\%threshold, all Two\-Step\-based models attain at least99\.0%99\.0\\%coverage\. The Adaptive Topological model remains the best, with99\.4%99\.4\\%of the test samples below10%10\\%, followed by Fixed Topological and Full Two\-Step at99\.2%99\.2\\%, Sensor Two\-Step at99\.0%99\.0\\%, and Vanilla DeepONet at90\.6%90\.6\\%\. Thus, the topological and Two\-Step models produce substantially fewer high\-error trajectories than Vanilla DeepONet\.[Figure 16](https://arxiv.org/html/2608.06428#S5.F16)provides a representative comparison of the predicted time\-evolving vorticity fields\. The complete\-field predictions in the upper row are visually similar because the dominant large\-scale structures are recovered by all models\. To expose the differences that are obscured by the common vorticity scale, the lower row reports localized absolute\-error maps over the shear\-interaction region identified by the dashed boxes\. The Adaptive Topological model produces the smallest localized error and more accurately preserves the position and amplitude of the diagonal shear structure\. The Fixed Topological model remains close to the adaptive result but exhibits a slightly broader and stronger error band\. The Vanilla DeepONet prediction shows the largest deviation, with a more pronounced error concentrated along the interface separating the positive\- and negative\-vorticity regions\. This behavior is consistent with the larger snapshot\-wise and trajectory\-wise relativeL2L^\{2\}errors reported for Vanilla DeepONet\.

Table 8:Validation, optimization, and model\-size statistics for the time\-evolving Navier–Stokes benchmark\. Training times are reported in hours\. The adaptive measurement parameters correspond to the trainable correction matrixΔ​Aθ\\Delta A\_\{\\theta\}\.![Refer to caption](https://arxiv.org/html/2608.06428v1/x17.png)Figure 15:Empirical cumulative distribution functions \(ECDFs\) of the sample\-wise relativeL2L^\{2\}trajectory errors on the Navier–Stokes test set \(Ntest=500N\_\{\\mathrm\{test\}\}=500\) for the five operator\-learning models\. Curves farther toward the upper\-left indicate that a larger fraction of test trajectories is predicted with lower error\. The vertical dashed line marks the5%5\\%relative\-error threshold, and the filled markers indicate the corresponding cumulative fraction for each model\. The percentages of test samples below the4%4\\%,5%5\\%, and10%10\\%thresholds are summarized in Table[9](https://arxiv.org/html/2608.06428#S5.T9)\.Table 9:Percentage of test samples below selected relativeL2L^\{2\}\-error thresholds for the time\-evolving Navier–Stokes benchmark\.![Refer to caption](https://arxiv.org/html/2608.06428v1/figures/time_evolving_ns_clean.png)Figure 16:Qualitative and spectral comparison for the time\-evolving Navier–Stokes benchmark\. Panel\(a\)shows the ground\-truth vorticity field and predictions from the Adaptive Topological, Fixed Topological, and Vanilla DeepONet models at a selected output time\. A common symmetric color scale is used for all full\-field panels\. The dashed rectangles identify a localized shear\-interaction region, while the lower row shows the corresponding ground\-truth crop and absolute prediction errors\|ω^−ω\|\|\\widehat\{\\omega\}\-\\omega\|using a common error scale\. Panel\(b\)compares the isotropic vorticity spectraEω​\(k\)E\_\{\\omega\}\(k\), averaged over the complete predicted trajectory for the same test realization\. The inset magnifies the high\-wavenumber range indicated by the dashed box\.
### 5\.9Distribution\-valued screened Poisson operator

To evaluate the proposed framework beyond normed function spaces, we consider a screened Poisson equation with distribution\-valued forcing\. Let

\(−Δ\+I\)​u=μin​Ω=\(0,1\)2,u=0on​∂Ω,\(\-\\Delta\+I\)u=\\mu\\quad\\text\{in \}\\Omega=\(0,1\)^\{2\},\\qquad u=0\\quad\\text\{on \}\\partial\\Omega,\(68\)with distribution\-valued input

μ=∑i=1Nfai​δ𝒙i\.\\mu=\\sum\_\{i=1\}^\{N\_\{f\}\}a\_\{i\}\\delta\_\{\\bm\{x\}\_\{i\}\}\.\(69\)Thus,

𝒢:ℳ​\(Ω\)→H01​\(Ω\),μ↦u\.\\mathcal\{G\}:\\mathcal\{M\}\(\\Omega\)\\rightarrow H\_\{0\}^\{1\}\(\\Omega\),\\qquad\\mu\\mapsto u\.\(70\)For smooth test functions\{ϕj\}j=1K\\\{\\phi\_\{j\}\\\}\_\{j=1\}^\{K\}, the functional coordinates are evaluated directly from the source list as

ℓj​\(μ\)=⟨μ,ϕj⟩=∑i=1Nfai​ϕj​\(𝒙i\),\\ell\_\{j\}\(\\mu\)=\\langle\\mu,\\phi\_\{j\}\\rangle=\\sum\_\{i=1\}^\{N\_\{f\}\}a\_\{i\}\\phi\_\{j\}\(\\bm\{x\}\_\{i\}\),\(71\)without rasterizing the input onto a common grid\. The source locations are sampled uniformly from\[0\.05,0\.95\]2\[0\.05,0\.95\]^\{2\}, with amplitudesai∼𝒩​\(0,1\)a\_\{i\}\\sim\\mathcal\{N\}\(0,1\)\. The in\-distribution data useNf∈\{1,…,15\}N\_\{f\}\\in\\\{1,\\ldots,15\\\}, while the out\-of\-distribution test set usesNf∈\{16,…,25\}N\_\{f\}\\in\\\{16,\\ldots,25\\\}\. Reference solutions are obtained from a truncated sine\-series expansion and evaluated on a64×6464\\times 64grid\. We compare the proposed functional\-measurement models with two baselines designed for variable\-cardinality or discretized inputs\. First, the*DeepSets source\-list*model represents each realization by the unordered set

𝒮=\{\(xi,1,xi,2,ai\)\}i=1Nf\.\\mathcal\{S\}=\\left\\\{\(x\_\{i,1\},x\_\{i,2\},a\_\{i\}\)\\right\\\}\_\{i=1\}^\{N\_\{f\}\}\.Following the permutation\-invariant Deep Sets architecture\[[30](https://arxiv.org/html/2608.06428#bib.bib30)\], the same encoderϕ\\phiis applied to every source, the resulting features are aggregated by summation, and a second networkρ\\rhopredicts the reduced output coefficients:

𝒄^=ρ​\(∑i=1Nfϕ​\(xi,1,xi,2,ai\)\)\.\\widehat\{\\bm\{c\}\}=\\rho\\left\(\\sum\_\{i=1\}^\{N\_\{f\}\}\\phi\(x\_\{i,1\},x\_\{i,2\},a\_\{i\}\)\\right\)\.\(72\)This construction directly accommodates a variable number of sources and is invariant to their ordering, but its latent coordinates are generic nonlinear learned features rather than prescribed continuous linear functionals of the input measure\. Deep Sets was introduced specifically for permutation\-invariant learning on set\-valued inputs\. Second, the*Rasterized Two\-Step*baseline deposits the signed Dirac sources onto a fixed32×3232\\times 32Cartesian grid using bilinear weights\. The resulting10241024\-dimensional grid vector is passed to a branch network that predicts coefficients in the same POD output basis used by all other models\. This model follows the coefficient\-space Two\-Step DeepONet strategy, in which the output representation and branch regression are trained as separate stages\[[12](https://arxiv.org/html/2608.06428#bib.bib2)\]\. Unlike the functional models, the rasterized baseline depends on a prescribed input grid and introduces an additional approximation through source deposition\. The Two\-Step strategy was proposed to reduce the difficulty of jointly optimizing the DeepONet branch and trunk representations\.

Table 10:In\-distribution and out\-of\-distribution performance for the distribution\-valued screened Poisson benchmark\. Models are trained using1≤Nf≤151\\leq N\_\{f\}\\leq 15\. The OOD set contains16≤Nf≤2516\\leq N\_\{f\}\\leq 25sources\.![Refer to caption](https://arxiv.org/html/2608.06428v1/figures/distributional_screened_poisson_predictions.png)Figure 17:Representative prediction for the distribution\-valued screened Poisson benchmark\. The top row shows the reference solution and the predictions from the Fixed Functional, Adaptive Functional, DeepSets source\-list, and Rasterized Two\-Step models\. The bottom row shows the corresponding absolute\-error fields\. The functional models recover the dominant spatial structures more accurately than the two baseline representations\.The functional\-measurement models substantially outperform the DeepSets and rasterized baselines shown in[Table 10](https://arxiv.org/html/2608.06428#S5.T10)\. The Fixed Functional model gives the lowest mean andP95P\_\{95\}errors, reducing the mean error by approximately55\.9%55\.9\\%relative to Rasterized Two\-Step and65\.2%65\.2\\%relative to DeepSets\. The Adaptive Functional model achieves the lowest global relativeL2L^\{2\}error, although its larger standard deviation andP95P\_\{95\}indicate a small number of high\-error outliers\. These results demonstrate the benefit of evaluating continuous\-dual measurements directly on variable\-cardinality, distribution\-valued inputs\. On the out\-of\-distribution test set \([Table 10](https://arxiv.org/html/2608.06428#S5.T10)\), the Adaptive Functional model achieves the best performance across all reported error statistics, with a mean relativeL2L^\{2\}error of39\.80%39\.80\\%and a global relativeL2L^\{2\}error of42\.37%42\.37\\%\. The Fixed Functional model follows with corresponding errors of42\.28%42\.28\\%and45\.90%45\.90\\%\. Both functional models substantially outperform the Rasterized Two\-Step and DeepSets baselines\. Relative to Rasterized Two\-Step, the Adaptive Functional model reduces the OOD mean error by approximately

0\.6840−0\.39800\.6840×100≈41\.8%\.\\frac\{0\.6840\-0\.3980\}\{0\.6840\}\\times 100\\approx 41\.8\\%\.Relative to DeepSets, the reduction is approximately55\.1%55\.1\\%\. These results indicate that the continuous\-functional representation generalizes more effectively to measures whose cardinality exceeds the entire training range\. Figure[17](https://arxiv.org/html/2608.06428#S5.F17)shows a representative test realization\. The Fixed and Adaptive Functional models reproduce the principal localized solution structures, with sample\-wise relative errors of35\.9%35\.9\\%and30\.5%30\.5\\%, respectively\. In contrast, the DeepSets source\-list prediction is nearly uniform, while the Rasterized Two\-Step model captures only part of the spatial response, leading to substantially larger errors\.

## 6Summary

We developed fixed and adaptive Topological DeepONets as computational realizations of the operator\-learning framework of Ismailov\[[6](https://arxiv.org/html/2608.06428#bib.bib1)\]\. In contrast to conventional DeepONets, whose branch inputs are point evaluations tied to a prescribed sensor grid, the proposed models encode the input through continuous linear functionals drawn from the dual of a Hausdorff locally convex space\. Both variants were combined with an enhanced Two\-Step training strategy\[[12](https://arxiv.org/html/2608.06428#bib.bib2)\]: a weighted singular value decomposition constructs a rank\-stable, discretely orthonormal output basis, and the branch network is trained directly in the corresponding coefficient space\. In the adaptive formulation, the measurement functionals and coefficient map are learned jointly, while a training\-only reconstruction decoder and soft regularization reduce information loss and feature collapse without increasing inference\-time complexity\. On the theoretical side, Theorem[4\.1](https://arxiv.org/html/2608.06428#S4.Thmtheorem1)decomposes the discrete approximation error into measurement\-reconstruction, output\-basis truncation, and neural\-approximation contributions, while Corollary[4\.1](https://arxiv.org/html/2608.06428#S4.Thmcorollary1)provides an explicit Barron\-rate refinement of the network term\. The estimate was verified numerically for the antiderivative operator: all400400test samples satisfied the samplewise bound, with maximum sharpness ratios of0\.940\.94and0\.970\.97for the adaptive and fixed models, respectively\. We further considered a benchmark posed on a genuinely non\-normable locally convex input space\. This experiment demonstrates that the proposed construction remains applicable when the input topology is generated by a family of seminorms and cannot be represented by a single norm, thereby providing a direct computational illustration of the extension beyond the Banach\-space setting\. The computational experiments support several conclusions\. First, distributed functional measurements outperform point sensors under matched representation budgets\. On the Darcy problem, the Fixed Topological DeepONet reduces the global relativeL2L^\{2\}error from10\.39%10\.39\\%to5\.88%5\.88\\%relative to a parameter\-matched point\-sensor model using the same3232\-dimensional input\. In the heterogeneous\-resolution Darcy experiment, the functional models also retain nearly resolution\-independent mean errors of approximately5\.5%5\.5\\%–5\.6%5\.6\\%on previously unseen grids and degrade substantially less than interpolation\-based baselines under missing observations\. Second, learning the measurement functionals produces further gains at essentially unchanged inference cost\. The adaptive model attains the lowest error on the antiderivative benchmark, with mean relativeL2L^\{2\}error2\.12×10−22\.12\\times 10^\{\-2\}, and outperforms the Two\-Step baseline on72%72\\%of the test functions\. It also achieves the highest coverage of accurate predictions among the DeepONet\-based models on both Navier–Stokes problems, with76\.0%76\.0\\%of fixed\-time samples below2%2\\%error and78\.4%78\.4\\%of time\-evolving trajectories below5%5\\%error\. Third, the functional coordinates retain dynamically relevant information despite substantial compression: with a32×32\\timesreduction in input dimension, the adaptive model surpasses the full\-field Two\-Step model on the time\-evolving trajectory\-prediction problem, attaining mean relativeL2L^\{2\}errors of4\.57×10−24\.57\\times 10^\{\-2\}and4\.77×10−24\.77\\times 10^\{\-2\}, respectively\. The comparison with a direct Fourier neural operator \(FNO\)\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\]clarifies the scope of these advantages\. On the uniform periodic fixed\-time Navier–Stokes benchmark, the FNO achieves the lowest three\-seed mean relativeL2L^\{2\}error,0\.832%±0\.172%0\.832\\%\\pm 0\.172\\%, whereas the Adaptive Topological DeepONet attains1\.685%±0\.017%1\.685\\%\\pm 0\.017\\%using only128128functional coordinates\. Thus, the FNO provides higher average accuracy in a setting strongly aligned with its spectral inductive bias, but it exhibits substantially larger seed\-to\-seed variability, uses the complete64×6464\\times 64input field, requires approximately twice the end\-to\-end training time, and consumes about10\.7×10\.7\\timesmore peak GPU memory\. The principal advantage of the proposed framework is therefore not universal superiority over grid\-specific neural operators, but the construction of compact, interpretable, operator\-adapted, and discretization\-portable coordinates in the continuous dual𝒱′\\mathcal\{V\}^\{\\prime\}, including for input spaces that are not normable\. Future work includes physics\-informed\[[29](https://arxiv.org/html/2608.06428#bib.bib9),[14](https://arxiv.org/html/2608.06428#bib.bib35)\]extensions of the topological framework, principled and data\-driven design of measurement dictionaries, applications to variable\-resolution and multimodal experimental data, and an extension of the discrete error analysis of Section[4](https://arxiv.org/html/2608.06428#S4)to fully infinite\-dimensional locally convex settings\.

## Acknowledgments

KS and GEK acknowledge partial support from the U\.S\. Department of Energy \(DOE\), Office of Science, Advanced Scientific Computing Research \(ASCR\) program, through the Scientific Discovery through Advanced Computing \(SciDAC\) Institute ”LEADS: LEarning\-Accelerated Domain Science” \(Subcontract 831126 under DE\-AC05\-76RL01830\)\. GEK further acknowledges support from the ONR Vannevar Bush Faculty Fellowship \(N00014\-22\-1\-2795\)\.

## Code and Reproducibility

The complete implementation, including data\-generation scripts, training configurations, and evaluation code for all four benchmarks, will be made publicly available in a permanent repository upon publication; the archival URL will be inserted in the final version\.

## References

- \[1\]T\. Chen and H\. Chen\(1995\)Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems\.IEEE Transactions on Neural Networks6\(4\),pp\. 911–917\.External Links:[Document](https://dx.doi.org/10.1109/72.392253)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[2\]S\. Cuomo, V\. Schiano Di Cola, F\. Giampaolo, G\. Rozza, M\. Raissi, and F\. Piccialli\(2022\)Scientific machine learning through physics\-informed neural networks: where we are and what’s next\.Journal of Scientific Computing92\(3\),pp\. 88\.External Links:[Document](https://dx.doi.org/10.1007/s10915-022-01939-z)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[3\]T\. O\. de Jong, K\. Shukla, and M\. Lazar\(2025\)Deep operator neural network model predictive control\.IEEE Open Journal of Control Systems\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[4\]G\. Gupta, X\. Xiao, and P\. Bogdan\(2021\)Multiwavelet\-based operator learning for differential equations\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 24048–24062\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[5\]Z\. Hao, Z\. Wang, H\. Su, C\. Ying, Y\. Dong, S\. Liu, Z\. Cheng, J\. Song, and J\. Zhu\(2023\)GNOT: a general neural operator transformer for operator learning\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 12556–12569\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[6\]V\. Ismailov\(2026\)Topological deeponets and a generalization of the chen\-chen operator approximation theorem\.arXiv preprint arXiv:2603\.11972\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12),[Table 1](https://arxiv.org/html/2608.06428#S3.T1),[Table 1](https://arxiv.org/html/2608.06428#S3.T1.13.2),[Table 1](https://arxiv.org/html/2608.06428#S3.T1.2.3.1.3.1),[Remark 3\.1](https://arxiv.org/html/2608.06428#S3.Thmremark1.p1.1),[Remark 3\.1](https://arxiv.org/html/2608.06428#S3.Thmremark1.p1.1.1),[§4](https://arxiv.org/html/2608.06428#S4.p1.15),[§6](https://arxiv.org/html/2608.06428#S6.p1.27)\.
- \[7\]P\. Jin, S\. Meng, and L\. Lu\(2022\)MIONet: learning multiple\-input operators via tensor product\.SIAM Journal on Scientific Computing44\(6\),pp\. A3490–A3514\.External Links:[Document](https://dx.doi.org/10.1137/22M1477751)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12)\.
- \[8\]G\. E\. Karniadakis, I\. G\. Kevrekidis, L\. Lu, P\. Perdikaris, S\. Wang, and L\. Yang\(2021\)Physics\-informed machine learning\.Nature Reviews Physics3,pp\. 422–440\.External Links:[Document](https://dx.doi.org/10.1038/s42254-021-00314-5)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[9\]N\. B\. Kovachki, Z\. Li, B\. Liu, K\. Azizzadenesheli, K\. Bhattacharya, A\. M\. Stuart, and A\. Anandkumar\(2023\)Neural operator: learning maps between function spaces with applications to PDEs\.Journal of Machine Learning Research24\(89\),pp\. 1–97\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12),[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[10\]S\. Lanthaler, S\. Mishra, and G\. E\. Karniadakis\(2021\)Error estimates for DeepONets: a deep learning framework in infinite dimensions\.arXiv preprint arXiv:2102\.09618\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2102.09618),2102\.09618Cited by:[Table 1](https://arxiv.org/html/2608.06428#S3.T1),[Table 1](https://arxiv.org/html/2608.06428#S3.T1.13.2),[Table 1](https://arxiv.org/html/2608.06428#S3.T1.2.3.1.2.1),[Remark 3\.1](https://arxiv.org/html/2608.06428#S3.Thmremark1.p1.1.1)\.
- \[11\]M\. Laudato, L\. Manzari, and K\. Shukla\(2025\)Neural operator modeling of platelet geometry and stress in shear flow\.arXiv preprint arXiv:2503\.12074\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[12\]S\. Lee and Y\. Shin\(2024\)On the training and generalization of deep operator networks\.SIAM Journal on Scientific Computing46\(4\),pp\. C273–C296\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12),[§5\.9](https://arxiv.org/html/2608.06428#S5.SS9.p1.10),[§6](https://arxiv.org/html/2608.06428#S6.p1.27)\.
- \[13\]Z\. Li, N\. Kovachki, K\. Azizzadenesheli, B\. Liu, K\. Bhattacharya, A\. Stuart, and A\. Anandkumar\(2021\)Fourier neural operator for parametric partial differential equations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=c8P9NQVtmnO)Cited by:[§C\.3](https://arxiv.org/html/2608.06428#A3.SS3.p1.1),[§C\.6](https://arxiv.org/html/2608.06428#A3.SS6.p1.5),[§1](https://arxiv.org/html/2608.06428#S1.p1.12),[§1](https://arxiv.org/html/2608.06428#S1.p1.6),[§5\.4](https://arxiv.org/html/2608.06428#S5.SS4.p1.26),[§5\.6](https://arxiv.org/html/2608.06428#S5.SS6.p1.6),[§5\.6](https://arxiv.org/html/2608.06428#S5.SS6.p2.17),[§5\.7](https://arxiv.org/html/2608.06428#S5.SS7.p1.1),[§6](https://arxiv.org/html/2608.06428#S6.p1.27)\.
- \[14\]Z\. Li, H\. Zheng, N\. Kovachki, D\. Jin, H\. Chen, B\. Liu, K\. Azizzadenesheli, and A\. Anandkumar\(2021\)Physics\-informed neural operator for learning partial differential equations\.arXiv preprint arXiv:2111\.03794\.External Links:2111\.03794Cited by:[§6](https://arxiv.org/html/2608.06428#S6.p1.27)\.
- \[15\]L\. Liu and W\. Cai\(2023\)Multiscale DeepONet for nonlinear operators in oscillatory function spaces for building seismic wave responses\.Engineering Structures276,pp\. 115221\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12)\.
- \[16\]L\. Lu, P\. Jin, G\. Pang, Z\. Zhang, and G\. E\. Karniadakis\(2021\)Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators\.Nature Machine Intelligence3\(3\),pp\. 218–229\.External Links:[Document](https://dx.doi.org/10.1038/s42256-021-00302-5)Cited by:[§C\.5](https://arxiv.org/html/2608.06428#A3.SS5.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2608.06428#S1.p1.6),[Figure 2](https://arxiv.org/html/2608.06428#S3.F2),[Figure 2](https://arxiv.org/html/2608.06428#S3.F2.24.12),[§3\.1](https://arxiv.org/html/2608.06428#S3.SS1.p1.10),[§5\.2](https://arxiv.org/html/2608.06428#S5.SS2.p1.3),[§5\.6](https://arxiv.org/html/2608.06428#S5.SS6.p2.17),[§5](https://arxiv.org/html/2608.06428#S5.p1.1)\.
- \[17\]V\. Oommen, K\. Shukla, S\. Desai, R\. Dingreville, and G\. E\. Karniadakis\(2024\)Rethinking materials simulations: blending direct numerical simulations with neural operators\.npj Computational Materials10\(1\),pp\. 145\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[18\]V\. Oommen, K\. Shukla, S\. Goswami, R\. Dingreville, and G\. E\. Karniadakis\(2022\)Learning two\-phase microstructure evolution using neural operators and autoencoder architectures\.npj Computational Materials8\(1\),pp\. 190\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[19\]T\. Pfaff, M\. Fortunato, A\. Sanchez\-Gonzalez, and P\. W\. Battaglia\(2021\)Learning mesh\-based simulation with graph networks\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[20\]M\. Prasthofer, T\. De Ryck, and S\. Mishra\(2022\)Variable\-input deep operator networks\.arXiv preprint arXiv:2205\.11404\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2205.11404),2205\.11404Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12)\.
- \[21\]C\. Rackauckas, Y\. Ma, J\. Martensen, C\. Warner, K\. Zubov, R\. Supekar, D\. Skinner, A\. Ramadhan, and A\. Edelman\(2020\)Universal differential equations for scientific machine learning\.arXiv preprint arXiv:2001\.04385\.External Links:2001\.04385Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[22\]M\. Raissi, P\. Perdikaris, and G\. E\. Karniadakis\(2019\)Physics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational Physics378,pp\. 686–707\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2018.10.045)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[23\]J\. H\. Seidman, G\. Kissas, P\. Perdikaris, and G\. J\. Pappas\(2022\)NOMAD: nonlinear manifold decoders for operator learning\.InAdvances in Neural Information Processing Systems,Vol\.35\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12),[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[24\]L\. Serrano, L\. Le Boudec, A\. K\. Koupaï, T\. X\. Wang, Y\. Yin, J\. Vittaut, and P\. Gallinari\(2023\)Operator learning with neural fields: tackling PDEs on general geometries\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[25\]K\. Shukla, V\. Oommen, A\. Peyvan, M\. Penwarden, N\. Plewacki, L\. Bravo, A\. Ghoshal, R\. M\. Kirby, and G\. E\. Karniadakis\(2024\)Deep neural operators as accurate surrogates for shape optimization\.Engineering Applications of Artificial Intelligence129,pp\. 107615\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[26\]K\. Shukla, J\. Ratchford, L\. Bravo, V\. Oommen, N\. Plewacki, A\. Ghoshal, and G\. Karniadakis\(2024\)Deep operator learning\-based surrogate models for aerothermodynamic analysis of aedc hypersonic waverider\.arXiv preprint arXiv:2405\.13234\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[27\]K\. Shukla, Z\. Zou, T\. Kaeufer, M\. Triantafyllou, and G\. E\. Karniadakis\(2026\)Uncertainty quantification in pinns for turbulent flows: bayesian inference and repulsive ensembles\.arXiv preprint arXiv:2604\.17156\.Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.6)\.
- \[28\]S\. Venturi and T\. Casey\(2023\)SVD perspectives for augmenting DeepONet flexibility and interpretability\.Computer Methods in Applied Mechanics and Engineering403,pp\. 115718\.External Links:[Document](https://dx.doi.org/10.1016/j.cma.2022.115718)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12)\.
- \[29\]S\. Wang, H\. Wang, and P\. Perdikaris\(2021\)Learning the solution operator of parametric partial differential equations with physics\-informed DeepONets\.Science Advances7\(40\),pp\. eabi8605\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.abi8605)Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12),[§1](https://arxiv.org/html/2608.06428#S1.p1.6),[§6](https://arxiv.org/html/2608.06428#S6.p1.27)\.
- \[30\]M\. Zaheer, S\. Kottur, S\. Ravanbakhsh, B\. Poczos, R\. Salakhutdinov, and A\. J\. Smola\(2017\)Deep sets\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 3391–3401\.Cited by:[§5\.9](https://arxiv.org/html/2608.06428#S5.SS9.p1.8)\.
- \[31\]Z\. Zhang, W\. T\. Leung, and H\. Schaeffer\(2022\)BelNet: basis enhanced learning, a mesh\-free neural operator\.arXiv preprint arXiv:2212\.07336\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2212.07336),2212\.07336Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12)\.
- \[32\]Z\. Zhang, W\. T\. Leung, and H\. Schaeffer\(2023\)A discretization\-invariant extension and analysis of some deep operator networks\.arXiv preprint arXiv:2307\.09738\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2307.09738),2307\.09738Cited by:[§1](https://arxiv.org/html/2608.06428#S1.p1.12)\.

## Appendix ADetails of the antiderivative convergence study

### A\.1Discretization and data generation

The interval\[0,1\]\[0,1\]is discretized using

uniformly spaced points

xj=jm−1,j=0,…,m−1\.x\_\{j\}=\\frac\{j\}\{m\-1\},\\qquad j=0,\\ldots,m\-1\.Each input function is represented by

vh=\(v​\(x0\),…,v​\(xm−1\)\)⊤∈ℝm\.v\_\{h\}=\\bigl\(v\(x\_\{0\}\),\\ldots,v\(x\_\{m\-1\}\)\\bigr\)^\{\\top\}\\in\\mathbb\{R\}^\{m\}\.The discrete antiderivative is computed using the cumulative composite trapezoidal rule:

\[𝒢h​\(vh\)\]0=0,\\bigl\[\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\\bigr\]\_\{0\}=0,and

\[𝒢h​\(vh\)\]j=∑k=0j−1xk\+1−xk2​\(v​\(xk\)\+v​\(xk\+1\)\),j=1,…,m−1\.\\bigl\[\\mathcal\{G\}\_\{h\}\(v\_\{h\}\)\\bigr\]\_\{j\}=\\sum\_\{k=0\}^\{j\-1\}\\frac\{x\_\{k\+1\}\-x\_\{k\}\}\{2\}\\left\(v\(x\_\{k\}\)\+v\(x\_\{k\+1\}\)\\right\),\\qquad j=1,\\ldots,m\-1\.The input functions are generated as random truncated Fourier series,

v\(i\)​\(x\)=a0\(i\)\+∑k=1K\[ak\(i\)​cos⁡\(2​π​k​x\)\+bk\(i\)​sin⁡\(2​π​k​x\)\],v^\{\(i\)\}\(x\)=a\_\{0\}^\{\(i\)\}\+\\sum\_\{k=1\}^\{K\}\\left\[a\_\{k\}^\{\(i\)\}\\cos\(2\\pi kx\)\+b\_\{k\}^\{\(i\)\}\\sin\(2\\pi kx\)\\right\],with

The coefficients are sampled independently according to

ak\(i\),bk\(i\)∼𝒩​\(0,k−2\),k=1,…,K,a\_\{k\}^\{\(i\)\},b\_\{k\}^\{\(i\)\}\\sim\\mathcal\{N\}\(0,k^\{\-2\}\),\\qquad k=1,\\ldots,K,Herek−2k^\{\-2\}denotes the variance, so the corresponding standard deviation isk−1k^\{\-1\}\. while

a0\(i\)∼𝒩​\(0,0\.252\)\.a\_\{0\}^\{\(i\)\}\\sim\\mathcal\{N\}\(0,0\.25^\{2\}\)\.Each realization is scaled by

αi∼𝒰​\(0\.7,1\.3\),\\alpha\_\{i\}\\sim\\mathcal\{U\}\(0\.7,1\.3\),so that the final input is

v~\(i\)​\(x\)=αi​v\(i\)​\(x\)\.\\widetilde\{v\}^\{\(i\)\}\(x\)=\\alpha\_\{i\}v^\{\(i\)\}\(x\)\.For each input, the corresponding output is

yh\(i\)=𝒢h​\(vh\(i\)\)\.y\_\{h\}^\{\(i\)\}=\\mathcal\{G\}\_\{h\}\(v\_\{h\}^\{\(i\)\}\)\.A total of28002800aligned input–output pairs are generated and divided into

Ntrain=2000,Nval=400,Ntest=400\.N\_\{\\mathrm\{train\}\}=2000,\\qquad N\_\{\\mathrm\{val\}\}=400,\\qquad N\_\{\\mathrm\{test\}\}=400\.The same random permutation is applied to both the input and output arrays, and the resulting dataset is reused for every value ofqq\.

### A\.2Fixed and adaptive measurements

For the fixed Topological DeepONet, the measurement map is

Mq:ℝm⟶ℝq,Mq​vh=Φq⊤​vh,M\_\{q\}:\\mathbb\{R\}^\{m\}\\longrightarrow\\mathbb\{R\}^\{q\},\\qquad M\_\{q\}v\_\{h\}=\\Phi\_\{q\}^\{\\top\}v\_\{h\},where

Φq=\[ϕ1,h⋯ϕq,h\]∈ℝm×q\\Phi\_\{q\}=\\begin\{bmatrix\}\\phi\_\{1,h\}&\\cdots&\\phi\_\{q,h\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{m\\times q\}contains the firstqqdiscrete Legendre modes after QR orthonormalization:

Φq⊤​Φq=Iq\.\\Phi\_\{q\}^\{\\top\}\\Phi\_\{q\}=I\_\{q\}\.The corresponding reconstruction is

Rq​\(Mq​vh\)=Φq​Φq⊤​vh\.R\_\{q\}\(M\_\{q\}v\_\{h\}\)=\\Phi\_\{q\}\\Phi\_\{q\}^\{\\top\}v\_\{h\}\.The adaptive measurement basis is initialized withΦq\\Phi\_\{q\}and updated according to

Φ~q=qf⁡\(Φq\+Δ​Φq\),\\widetilde\{\\Phi\}\_\{q\}=\\operatorname\{qf\}\\left\(\\Phi\_\{q\}\+\\Delta\\Phi\_\{q\}\\right\),whereqf\\operatorname\{qf\}denotes the orthonormal factor from a reduced QR decomposition\. The correction is regularized using

λdrift​‖Δ​Φq‖F2,λdrift=10−4\.\\lambda\_\{\\mathrm\{drift\}\}\\\|\\Delta\\Phi\_\{q\}\\\|\_\{F\}^\{2\},\\qquad\\lambda\_\{\\mathrm\{drift\}\}=10^\{\-4\}\.

### A\.3Reduced output representation and training

Let

y¯h=1Ntrain​∑i=1Ntrainyh\(i\)\\overline\{y\}\_\{h\}=\\frac\{1\}\{N\_\{\\mathrm\{train\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{train\}\}\}y\_\{h\}^\{\(i\)\}be the training\-output mean\. Let

Qr∈ℝm×r,Qr⊤​Qr=Ir,Q\_\{r\}\\in\\mathbb\{R\}^\{m\\times r\},\\qquad Q\_\{r\}^\{\\top\}Q\_\{r\}=I\_\{r\},contain the firstr=64r=64POD modes of the centered training\-output matrix\. The exact reduced coefficients are

cr\(i\)=Qr⊤​\(yh\(i\)−y¯h\)\.c\_\{r\}^\{\(i\)\}=Q\_\{r\}^\{\\top\}\\left\(y\_\{h\}^\{\(i\)\}\-\\overline\{y\}\_\{h\}\\right\)\.The learned operator is

𝒢^h,θ​\(vh\)=y¯h\+Qr​bθ​\(Mq​vh\),\\widehat\{\\mathcal\{G\}\}\_\{h,\\theta\}\(v\_\{h\}\)=\\overline\{y\}\_\{h\}\+Q\_\{r\}b\_\{\\theta\}\(M\_\{q\}v\_\{h\}\),where

bθ:ℝq⟶ℝrb\_\{\\theta\}:\\mathbb\{R\}^\{q\}\\longrightarrow\\mathbb\{R\}^\{r\}is a feed\-forward network with four hidden layers, width128128, and GELU activations\. The fixed model is trained with AdamW using learning rate

The adaptive model is initialized from the trained fixed branch network and optimized with learning rate

The batch size is6464, and the checkpoint with the smallest global relativeL2L^\{2\}validation error is retained\.

### A\.4Measurement\-dimension convergence protocol

The measurement dimension is varied according to

q∈\{8,16,32,64,128\}\.q\\in\\\{8,16,32,64,128\\\}\.The spatial discretization, POD rank, network architecture, dataset, and optimization parameters are held fixed\. Therefore, the experiment is a convergence study with respect to the measurement dimensionqq, not a spatial mesh\-convergence study\.

### A\.5Error decomposition

For a test inputvh\(i\)v\_\{h\}^\{\(i\)\}, define the reconstructed input

v~h,q\(i\)=Rq​\(Mq​vh\(i\)\)\\widetilde\{v\}\_\{h,q\}^\{\(i\)\}=R\_\{q\}\(M\_\{q\}v\_\{h\}^\{\(i\)\}\)and its exact discrete output

y~h,q\(i\)=𝒢h​\(v~h,q\(i\)\)\.\\widetilde\{y\}\_\{h,q\}^\{\(i\)\}=\\mathcal\{G\}\_\{h\}\\left\(\\widetilde\{v\}\_\{h,q\}^\{\(i\)\}\\right\)\.The propagated measurement error is

Emeas\(i\)=‖yh\(i\)−y~h,q\(i\)‖2‖yh\(i\)‖2\.E\_\{\\mathrm\{meas\}\}^\{\(i\)\}=\\frac\{\\left\\\|y\_\{h\}^\{\(i\)\}\-\\widetilde\{y\}\_\{h,q\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\\|y\_\{h\}^\{\(i\)\}\\\|\_\{2\}\}\.The output projection error is

Eout\(i\)=‖y~h,q\(i\)−y¯h−Qr​Qr⊤​\(y~h,q\(i\)−y¯h\)‖2‖yh\(i\)‖2\.E\_\{\\mathrm\{out\}\}^\{\(i\)\}=\\frac\{\\left\\\|\\widetilde\{y\}\_\{h,q\}^\{\(i\)\}\-\\overline\{y\}\_\{h\}\-Q\_\{r\}Q\_\{r\}^\{\\top\}\\left\(\\widetilde\{y\}\_\{h,q\}^\{\(i\)\}\-\\overline\{y\}\_\{h\}\\right\)\\right\\\|\_\{2\}\}\{\\\|y\_\{h\}^\{\(i\)\}\\\|\_\{2\}\}\.The neural approximation error is

ENN\(i\)=‖y¯h\+Qr​Qr⊤​\(y~h,q\(i\)−y¯h\)−𝒢^h,θ​\(vh\(i\)\)‖2‖yh\(i\)‖2\.E\_\{\\mathrm\{NN\}\}^\{\(i\)\}=\\frac\{\\left\\\|\\overline\{y\}\_\{h\}\+Q\_\{r\}Q\_\{r\}^\{\\top\}\\left\(\\widetilde\{y\}\_\{h,q\}^\{\(i\)\}\-\\overline\{y\}\_\{h\}\\right\)\-\\widehat\{\\mathcal\{G\}\}\_\{h,\\theta\}\\left\(v\_\{h\}^\{\(i\)\}\\right\)\\right\\\|\_\{2\}\}\{\\\|y\_\{h\}^\{\(i\)\}\\\|\_\{2\}\}\.The total prediction error is

Etot\(i\)=‖yh\(i\)−𝒢^h,θ​\(vh\(i\)\)‖2‖yh\(i\)‖2\.E\_\{\\mathrm\{tot\}\}^\{\(i\)\}=\\frac\{\\left\\\|y\_\{h\}^\{\(i\)\}\-\\widehat\{\\mathcal\{G\}\}\_\{h,\\theta\}\\left\(v\_\{h\}^\{\(i\)\}\\right\)\\right\\\|\_\{2\}\}\{\\\|y\_\{h\}^\{\(i\)\}\\\|\_\{2\}\}\.The computable upper bound is

Ebound\(i\)=Emeas\(i\)\+Eout\(i\)\+ENN\(i\)\.E\_\{\\mathrm\{bound\}\}^\{\(i\)\}=E\_\{\\mathrm\{meas\}\}^\{\(i\)\}\+E\_\{\\mathrm\{out\}\}^\{\(i\)\}\+E\_\{\\mathrm\{NN\}\}^\{\(i\)\}\.The triangle inequality gives

Etot\(i\)≤Ebound\(i\)\.E\_\{\\mathrm\{tot\}\}^\{\(i\)\}\\leq E\_\{\\mathrm\{bound\}\}^\{\(i\)\}\.The reported convergence curves are test\-set means,

E¯α​\(q\)=1Ntest​∑i=1NtestEα\(i\)​\(q\),\\overline\{E\}\_\{\\alpha\}\(q\)=\\frac\{1\}\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}E\_\{\\alpha\}^\{\(i\)\}\(q\),where

α∈\{meas,out,NN,tot,bound\}\.\\alpha\\in\\\{\\mathrm\{meas\},\\mathrm\{out\},\\mathrm\{NN\},\\mathrm\{tot\},\\mathrm\{bound\}\\\}\.The sharpness of the sample\-wise estimate is measured by

ρmax=maxi⁡Etot\(i\)Ebound\(i\)\.\\rho\_\{\\max\}=\\max\_\{i\}\\frac\{E\_\{\\mathrm\{tot\}\}^\{\(i\)\}\}\{E\_\{\\mathrm\{bound\}\}^\{\(i\)\}\}\.

## Appendix BAdditional details for the Darcy experiment

### B\.1Discrete data representation

Each permeability and pressure field is stored as a vector inℝ7225\\mathbb\{R\}^\{7225\}, obtained by flattening an85×8585\\times 85grid\. The data matrices have dimensions

Atrain,Utrain\\displaystyle A\_\{\\mathrm\{train\}\},U\_\{\\mathrm\{train\}\}∈ℝ800×7225,\\displaystyle\\in\\mathbb\{R\}^\{800\\times 7225\},\(73\)Aval,Uval\\displaystyle A\_\{\\mathrm\{val\}\},U\_\{\\mathrm\{val\}\}∈ℝ100×7225,\\displaystyle\\in\\mathbb\{R\}^\{100\\times 7225\},Atest,Utest\\displaystyle A\_\{\\mathrm\{test\}\},U\_\{\\mathrm\{test\}\}∈ℝ100×7225\.\\displaystyle\\in\\mathbb\{R\}^\{100\\times 7225\}\.All normalization statistics are computed using only the training set\.

### B\.2Reduced output basis

Let

Utrain=W​Σ​V𝖳U\_\{\\mathrm\{train\}\}=W\\Sigma V^\{\\mathsf\{T\}\}\(74\)be the singular\-value decomposition of the training pressure snapshots\. The firstrrright singular vectors define

Q=V:,1:r\.Q=V\_\{:,1:r\}\.\(75\)Each solution is represented by

𝒄\(i\)=Q𝖳​𝒖\(i\),\\bm\{c\}^\{\(i\)\}=Q^\{\\mathsf\{T\}\}\\bm\{u\}^\{\(i\)\},\(76\)and reconstructed as

𝒖^\(i\)=Q​𝒄^\(i\)\.\\widehat\{\\bm\{u\}\}^\{\(i\)\}=Q\\widehat\{\\bm\{c\}\}^\{\(i\)\}\.\(77\)Thus, all Two\-Step models predict the same reduced pressure coefficients and use the same fixed output basis for reconstruction\.

### B\.3Input representations

The full\-field Two\-Step model uses the complete discretized permeability field,

𝒂∈ℝ7225\.\\bm\{a\}\\in\\mathbb\{R\}^\{7225\}\.\(78\)The Sensor Two\-Step model usesq=32q=32pointwise observations,

𝒔​\(a\)=\[a​\(𝒙1\),…,a​\(𝒙32\)\]𝖳\.\\bm\{s\}\(a\)=\\left\[a\(\\bm\{x\}\_\{1\}\),\\ldots,a\(\\bm\{x\}\_\{32\}\)\\right\]^\{\\mathsf\{T\}\}\.\(79\)The Fixed Topological DeepONet uses the same number of functional coordinates,

ℓ​\(a\)=\[ℓ1​\(a\),…,ℓ32​\(a\)\]𝖳,\\bm\{\\ell\}\(a\)=\\left\[\\ell\_\{1\}\(a\),\\ldots,\\ell\_\{32\}\(a\)\\right\]^\{\\mathsf\{T\}\},\(80\)with

ℓj​\(a\)=∫Ωa​\(𝒙\)​ϕj​\(𝒙\)​d𝒙\.\\ell\_\{j\}\(a\)=\\int\_\{\\Omega\}a\(\\bm\{x\}\)\\phi\_\{j\}\(\\bm\{x\}\)\\,\\mathrm\{d\}\\bm\{x\}\.\(81\)The continuum input space for the Darcy operator is𝒱=L∞​\(Ω\)\\mathcal\{V\}=L^\{\\infty\}\(\\Omega\), restricted to the uniformly elliptic admissible set𝒜\\mathcal\{A\}defined in \([36](https://arxiv.org/html/2608.06428#S5.E36)\)\. Takingϕj∈L1​\(Ω\)\\phi\_\{j\}\\in L^\{1\}\(\\Omega\), each measurement is a continuous linear functional on this ambient space because

\|ℓj​\(a\)\|≤‖a‖L∞​\(Ω\)​‖ϕj‖L1​\(Ω\)\.\\left\|\\ell\_\{j\}\(a\)\\right\|\\leq\\\|a\\\|\_\{L^\{\\infty\}\(\\Omega\)\}\\\|\\phi\_\{j\}\\\|\_\{L^\{1\}\(\\Omega\)\}\.\(82\)The fact that the sampled coefficient fields also lie inL2​\(Ω\)L^\{2\}\(\\Omega\)is used only at the level of their finite\-dimensional numerical representation; it does not change the continuum topology assigned to the Darcy operator\. On the discrete grid, the functional is approximated by

ℓj​\(a\)≈∑m=17225wm​a​\(𝒙m\)​ϕj​\(𝒙m\),\\ell\_\{j\}\(a\)\\approx\\sum\_\{m=1\}^\{7225\}w\_\{m\}a\(\\bm\{x\}\_\{m\}\)\\phi\_\{j\}\(\\bm\{x\}\_\{m\}\),\(83\)wherewmw\_\{m\}are quadrature weights associated with the grid points𝒙m\\bm\{x\}\_\{m\}\.

### B\.4Functional dictionaries

The implementation supports Legendre, cosine, Chebyshev, and Lagrange measurement dictionaries\. In the default Darcy configuration, the fixed measurements are constructed from a total\-degree Legendre dictionary\. After mapping the physical coordinates to the reference squareΩ^=\[−1,1\]2\\widehat\{\\Omega\}=\[\-1,1\]^\{2\}, the Legendre atoms are defined by

ϕi​jleg​\(x,y\)=P^i​\(x\)​P^j​\(y\),i\+j≤p,\\phi\_\{ij\}^\{\\mathrm\{leg\}\}\(x,y\)=\\widehat\{P\}\_\{i\}\(x\)\\widehat\{P\}\_\{j\}\(y\),\\qquad i\+j\\leq p,\(84\)where

P^n​\(x\)=2​n\+12​Pn​\(x\)\\widehat\{P\}\_\{n\}\(x\)=\\sqrt\{\\frac\{2n\+1\}\{2\}\}\\,P\_\{n\}\(x\)\(85\)is the normalized Legendre polynomial of degreenn\. The alternative tensor\-product dictionaries are

ϕi​jcos​\(x,y\)\\displaystyle\\phi\_\{ij\}^\{\\mathrm\{cos\}\}\(x,y\)=ci​\(x\)​cj​\(y\),\\displaystyle=c\_\{i\}\(x\)c\_\{j\}\(y\),\(86\)ϕi​jcheb​\(x,y\)\\displaystyle\\phi\_\{ij\}^\{\\mathrm\{cheb\}\}\(x,y\)=Ti​\(x\)​Tj​\(y\),\\displaystyle=T\_\{i\}\(x\)T\_\{j\}\(y\),ϕi​jlag​\(x,y\)\\displaystyle\\phi\_\{ij\}^\{\\mathrm\{lag\}\}\(x,y\)=Li​\(x\)​Lj​\(y\),\\displaystyle=L\_\{i\}\(x\)L\_\{j\}\(y\),where

c0​\(x\)=1,ck​\(x\)=2​cos⁡\(k​π​x\),c\_\{0\}\(x\)=1,\\qquad c\_\{k\}\(x\)=\\sqrt\{2\}\\cos\(k\\pi x\),\(87\)TiT\_\{i\}denotes a first\-kind Chebyshev polynomial, andLiL\_\{i\}denotes a Lagrange cardinal polynomial constructed on either Chebyshev or uniformly spaced interpolation nodes\. For the polynomial dictionaries, the index set may be chosen either as a total\-degree set,

ℐp=\{\(i,j\)∈ℕ02:i\+j≤p\},\\mathcal\{I\}\_\{p\}=\\left\\\{\(i,j\)\\in\\mathbb\{N\}\_\{0\}^\{2\}:i\+j\\leq p\\right\\\},\(88\)or as a full tensor\-product set\. The total\-degree construction is used in the default experiment\.

### B\.5Discrete weighted orthonormalization

Let

Φm​j=ϕj​\(𝒙m\)\\Phi\_\{mj\}=\\phi\_\{j\}\(\\bm\{x\}\_\{m\}\)\(89\)denote the discrete dictionary matrix, and let

Wq=diag⁡\(w1,…,w7225\)W\_\{q\}=\\operatorname\{diag\}\(w\_\{1\},\\ldots,w\_\{7225\}\)\(90\)contain the quadrature weights\. To improve conditioning and permit a fair comparison among different basis families, the raw dictionary is orthonormalized with respect to the weighted discrete inner product\. The weighted Gram matrix is

G=Φ𝖳​Wq​Φ,G=\\Phi^\{\\mathsf\{T\}\}W\_\{q\}\\Phi,\(91\)and the orthonormalized dictionary is

Φorth=Φ​G−1/2\.\\Phi\_\{\\mathrm\{orth\}\}=\\Phi G^\{\-1/2\}\.\(92\)The corresponding measurement matrix is

Mbase=Wq​Φorth\.M\_\{\\mathrm\{base\}\}=W\_\{q\}\\Phi\_\{\\mathrm\{orth\}\}\.\(93\)For a discretized permeability vector𝒂\\bm\{a\}, the complete dictionary coordinate vector is therefore

𝜶​\(a\)=Mbase𝖳​𝒂\.\\bm\{\\alpha\}\(a\)=M\_\{\\mathrm\{base\}\}^\{\\mathsf\{T\}\}\\bm\{a\}\.\(94\)

### B\.6Active dictionary and fixed measurements

The full dictionary may contain more functions than are needed by the branch network\. The dictionary coordinates are first ranked according to their variance over the training set\. An active subset is then retained so that it captures a prescribed fraction of the total measurement variance, subject to a minimum active dimension\. Let

𝜶act​\(a\)∈ℝKact\\bm\{\\alpha\}\_\{\\mathrm\{act\}\}\(a\)\\in\\mathbb\{R\}^\{K\_\{\\mathrm\{act\}\}\}\(95\)denote the standardized coordinates associated with the active dictionary\. The Fixed Topological DeepONet uses the firstq=32q=32selected coordinates,

ℓfix​\(a\)=\[α1​\(a\),…,α32​\(a\)\]𝖳\.\\bm\{\\ell\}\_\{\\mathrm\{fix\}\}\(a\)=\\left\[\\alpha\_\{1\}\(a\),\\ldots,\\alpha\_\{32\}\(a\)\\right\]^\{\\mathsf\{T\}\}\.\(96\)These coordinates are fixed before neural\-network training\.

### B\.7Adaptive functional measurements

The Adaptive Topological DeepONet begins from the active dictionary coordinates and learnsq=32q=32linear combinations,

𝒛ad​\(a\)=Aθ​𝜶act​\(a\),Aθ∈ℝ32×Kact\.\\bm\{z\}\_\{\\mathrm\{ad\}\}\(a\)=A\_\{\\theta\}\\bm\{\\alpha\}\_\{\\mathrm\{act\}\}\(a\),\\qquad A\_\{\\theta\}\\in\\mathbb\{R\}^\{32\\times K\_\{\\mathrm\{act\}\}\}\.\(97\)Thekk\-th learned coordinate can be written as

zkad​\(a\)\\displaystyle z\_\{k\}^\{\\mathrm\{ad\}\}\(a\)=∑j=1Kact\(Aθ\)k​j​αj​\(a\)\\displaystyle=\\sum\_\{j=1\}^\{K\_\{\\mathrm\{act\}\}\}\(A\_\{\\theta\}\)\_\{kj\}\\alpha\_\{j\}\(a\)\(98\)=∫Ωa​\(𝒙\)​ψkθ​\(𝒙\)​d𝒙,\\displaystyle=\\int\_\{\\Omega\}a\(\\bm\{x\}\)\\psi\_\{k\}^\{\\theta\}\(\\bm\{x\}\)\\,\\mathrm\{d\}\\bm\{x\},where

ψkθ​\(𝒙\)=∑j=1Kact\(Aθ\)k​j​ϕj​\(𝒙\)∈span⁡\{ϕ1,…,ϕKact\}\.\\psi\_\{k\}^\{\\theta\}\(\\bm\{x\}\)=\\sum\_\{j=1\}^\{K\_\{\\mathrm\{act\}\}\}\(A\_\{\\theta\}\)\_\{kj\}\\phi\_\{j\}\(\\bm\{x\}\)\\in\\operatorname\{span\}\\left\\\{\\phi\_\{1\},\\ldots,\\phi\_\{K\_\{\\mathrm\{act\}\}\}\\right\\\}\.\(99\)Thus, every adaptive coordinate remains a continuous linear functional in𝒱′\\mathcal\{V\}^\{\\prime\}, while its measurement function is learned within the span of the active dictionary\. A training\-only decoder reconstructs the active dictionary coordinates from𝒛ad\\bm\{z\}\_\{\\mathrm\{ad\}\}\. The associated reconstruction loss discourages the learned map from discarding excessive input information\. In addition, a Gram regularization term promotes nondegenerate and approximately decorrelated learned measurements\.

### B\.8Two\-Step branch prediction

For the full\-field, sensor, fixed\-functional, and adaptive\-functional Two\-Step models, the branch network predicts the standardized reduced coefficients

𝒄^​\(a\)∈ℝr\.\\widehat\{\\bm\{c\}\}\(a\)\\in\\mathbb\{R\}^\{r\}\.\(100\)The predicted pressure field is reconstructed using

𝒖^​\(a\)=Q​𝒄^​\(a\)\.\\widehat\{\\bm\{u\}\}\(a\)=Q\\widehat\{\\bm\{c\}\}\(a\)\.\(101\)The corresponding input\-to\-output pipelines are

𝒂⟶𝒄^⟶𝒖^\\bm\{a\}\\longrightarrow\\widehat\{\\bm\{c\}\}\\longrightarrow\\widehat\{\\bm\{u\}\}\(102\)for the full\-field Two\-Step model,

𝒔​\(a\)⟶𝒄^⟶𝒖^\\bm\{s\}\(a\)\\longrightarrow\\widehat\{\\bm\{c\}\}\\longrightarrow\\widehat\{\\bm\{u\}\}\(103\)for the sensor model,

ℓfix​\(a\)⟶𝒄^⟶𝒖^\\bm\{\\ell\}\_\{\\mathrm\{fix\}\}\(a\)\\longrightarrow\\widehat\{\\bm\{c\}\}\\longrightarrow\\widehat\{\\bm\{u\}\}\(104\)for the fixed Topological DeepONet, and

𝜶act​\(a\)⟶Aθ​𝜶act​\(a\)⟶𝒄^⟶𝒖^\\bm\{\\alpha\}\_\{\\mathrm\{act\}\}\(a\)\\longrightarrow A\_\{\\theta\}\\bm\{\\alpha\}\_\{\\mathrm\{act\}\}\(a\)\\longrightarrow\\widehat\{\\bm\{c\}\}\\longrightarrow\\widehat\{\\bm\{u\}\}\(105\)for the adaptive model\.

### B\.9Evaluation metrics

For theii\-th test sample, the relativeL2L^\{2\}error is

ei=‖𝒖^\(i\)−𝒖\(i\)‖2‖𝒖\(i\)‖2\.e\_\{i\}=\\frac\{\\left\\\|\\widehat\{\\bm\{u\}\}^\{\(i\)\}\-\\bm\{u\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\left\\\|\\bm\{u\}^\{\(i\)\}\\right\\\|\_\{2\}\}\.\(106\)The mean relativeL2L^\{2\}error is

Emean=1Ntest​∑i=1Ntestei\.E\_\{\\mathrm\{mean\}\}=\\frac\{1\}\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}e\_\{i\}\.\(107\)The global relativeL2L^\{2\}error is

Eglobal=‖Utest−U^test‖F‖Utest‖F\.E\_\{\\mathrm\{global\}\}=\\frac\{\\left\\\|U\_\{\\mathrm\{test\}\}\-\\widehat\{U\}\_\{\\mathrm\{test\}\}\\right\\\|\_\{F\}\}\{\\left\\\|U\_\{\\mathrm\{test\}\}\\right\\\|\_\{F\}\}\.\(108\)The root\-mean\-square error is

RMSE=1Ntest​Nh​∑i=1Ntest∑m=1Nh\(u^m\(i\)−um\(i\)\)2,\\mathrm\{RMSE\}=\\sqrt\{\\frac\{1\}\{N\_\{\\mathrm\{test\}\}N\_\{h\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{m=1\}^\{N\_\{h\}\}\\left\(\\widehat\{u\}\_\{m\}^\{\(i\)\}\-u\_\{m\}^\{\(i\)\}\\right\)^\{2\}\},\(109\)whereNh=7225N\_\{h\}=7225is the number of grid points\. The mean absolute error is

MAE=1Ntest​Nh​∑i=1Ntest∑m=1Nh\|u^m\(i\)−um\(i\)\|\.\\mathrm\{MAE\}=\\frac\{1\}\{N\_\{\\mathrm\{test\}\}N\_\{h\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{m=1\}^\{N\_\{h\}\}\\left\|\\widehat\{u\}\_\{m\}^\{\(i\)\}\-u\_\{m\}^\{\(i\)\}\\right\|\.\(110\)

### B\.10Interpretation of the sensor comparison

Sensor Two\-Step and Fixed Topological DeepONet are matched in input dimension, branch architecture, output basis, and inference parameter count\. Their comparison therefore isolates the effect of replacing local point evaluations with global functional coordinates\. The comparison should be interpreted as an equal*representation\-budget*comparison, rather than necessarily an equal physical sensing\-cost comparison, because a global functional may require access to the full discretized permeability field\. The adaptive model uses the same final coordinate dimension as the fixed and sensor models, but its coordinates are learned as linear combinations of the active functional dictionary\. Consequently, its comparison with the fixed model evaluates whether adapting the measurement functions within a common dictionary span improves the reduced input representation\.

## Appendix CAdditional Details for the Fixed\-Time Navier–Stokes Benchmark

### C\.1Governing equations

We consider the two\-dimensional incompressible Navier–Stokes equations on the periodic unit square

written in vorticity form as

∂ω∂t\+𝒖⋅∇ω=ν​Δ​ω\+f,\(𝒙,t\)∈D×\(0,T\],\\frac\{\\partial\\omega\}\{\\partial t\}\+\\bm\{u\}\\cdot\\nabla\\omega=\\nu\\Delta\\omega\+f,\\qquad\(\\bm\{x\},t\)\\in D\\times\(0,T\],\(111\)subject to the incompressibility constraint

∇⋅𝒖=0\.\\nabla\\cdot\\bm\{u\}=0\.\(112\)Here,ω=ω​\(𝒙,t\)\\omega=\\omega\(\\bm\{x\},t\)denotes the scalar vorticity,𝒖=\(u,v\)\\bm\{u\}=\(u,v\)is the velocity field,ν\>0\\nu\>0is the kinematic viscosity, andf=f​\(𝒙\)f=f\(\\bm\{x\}\)is a prescribed forcing term\. The velocity and vorticity are related by

ω=∂v∂x−∂u∂y\.\\omega=\\frac\{\\partial v\}\{\\partial x\}\-\\frac\{\\partial u\}\{\\partial y\}\.\(113\)Introducing the streamfunctionψ\\psi, one may write

𝒖=∇⟂ψ=\(∂yψ−∂xψ\),−Δ​ψ=ω\.\\bm\{u\}=\\nabla^\{\\perp\}\\psi=\\begin\{pmatrix\}\\partial\_\{y\}\\psi\\\\ \-\\partial\_\{x\}\\psi\\end\{pmatrix\},\\qquad\-\\Delta\\psi=\\omega\.\(114\)Equivalently,

𝒖=∇⟂\(−Δ\)−1ω\.\\bm\{u\}=\\nabla^\{\\perp\}\(\-\\Delta\)^\{\-1\}\\omega\.\(115\)Periodic boundary conditions are imposed in both spatial directions:

ω​\(x\+1,y,t\)=ω​\(x,y,t\),ω​\(x,y\+1,t\)=ω​\(x,y,t\),\\omega\(x\+1,y,t\)=\\omega\(x,y,t\),\\qquad\\omega\(x,y\+1,t\)=\\omega\(x,y,t\),\(116\)together with the initial condition

ω​\(𝒙,0\)=ω0​\(𝒙\)\.\\omega\(\\bm\{x\},0\)=\\omega\_\{0\}\(\\bm\{x\}\)\.\(117\)

### C\.2Fixed\-time operator\-learning problem

The objective is to learn a fixed\-time solution operator rather than an entire temporal trajectory\. Lettint\_\{\\mathrm\{in\}\}denote the input time and lettout\>tint\_\{\\mathrm\{out\}\}\>t\_\{\\mathrm\{in\}\}denote the target time\. The exact solution operator is

𝒢tin→tout:ω​\(⋅,tin\)⟼ω​\(⋅,tout\)\.\\mathcal\{G\}\_\{t\_\{\\mathrm\{in\}\}\\rightarrow t\_\{\\mathrm\{out\}\}\}:\\omega\(\\cdot,t\_\{\\mathrm\{in\}\}\)\\longmapsto\\omega\(\\cdot,t\_\{\\mathrm\{out\}\}\)\.\(118\)In the present experiment,

tin=t0,tout=t10,t\_\{\\mathrm\{in\}\}=t\_\{0\},\\qquad t\_\{\\mathrm\{out\}\}=t\_\{10\},\(119\)and therefore

𝒢0→10:ω​\(⋅,t0\)⟼ω​\(⋅,t10\)\.\\mathcal\{G\}\_\{0\\rightarrow 10\}:\\omega\(\\cdot,t\_\{0\}\)\\longmapsto\\omega\(\\cdot,t\_\{10\}\)\.\(120\)Because the implementation uses zero\-based snapshot indices,input\-index=0andtarget\-index=10define a ten\-snapshot\-interval forecast\. For theii\-th realization, we define

a\(i\)​\(𝒙\)=ω\(i\)​\(𝒙,t0\),u\(i\)​\(𝒙\)=ω\(i\)​\(𝒙,t10\)\.a^\{\(i\)\}\(\\bm\{x\}\)=\\omega^\{\(i\)\}\(\\bm\{x\},t\_\{0\}\),\\qquad u^\{\(i\)\}\(\\bm\{x\}\)=\\omega^\{\(i\)\}\(\\bm\{x\},t\_\{10\}\)\.\(121\)The supervised dataset is

𝒟=\{\(a\(i\),u\(i\)\)\}i=1N,u\(i\)=𝒢0→10​\(a\(i\)\)\.\\mathcal\{D\}=\\left\\\{\\left\(a^\{\(i\)\},u^\{\(i\)\}\\right\)\\right\\\}\_\{i=1\}^\{N\},\\qquad u^\{\(i\)\}=\\mathcal\{G\}\_\{0\\rightarrow 10\}\\left\(a^\{\(i\)\}\\right\)\.\(122\)A neural operator𝒢θ\\mathcal\{G\}\_\{\\theta\}is trained such that

u^\(i\)=𝒢θ​\(a\(i\)\)≈u\(i\)\.\\widehat\{u\}^\{\(i\)\}=\\mathcal\{G\}\_\{\\theta\}\\left\(a^\{\(i\)\}\\right\)\\approx u^\{\(i\)\}\.\(123\)

### C\.3Dataset and preprocessing

We use the standard two\-dimensional Navier–Stokes vorticity dataset employed in neural\-operator studies\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\]\. Each realization is stored on a uniform periodic grid with

and contains 50 temporal snapshots\. The complete data tensor may be written as

𝐖∈ℝ5000×64×64×50,\\mathbf\{W\}\\in\\mathbb\{R\}^\{5000\\times 64\\times 64\\times 50\},\(124\)up to the axis ordering used by the MATLAB or HDF5 data file\. For the fixed\-time experiment, only the input and target snapshots are extracted:

𝐗\(i\)\\displaystyle\\mathbf\{X\}^\{\(i\)\}=𝐖\(i\)​\(:,:,0\),\\displaystyle=\\mathbf\{W\}^\{\(i\)\}\(:,:,0\),\(125\)𝐘\(i\)\\displaystyle\\mathbf\{Y\}^\{\(i\)\}=𝐖\(i\)​\(:,:,10\)\.\\displaystyle=\\mathbf\{W\}^\{\(i\)\}\(:,:,10\)\.\(126\)Thus,

𝐗,𝐘∈ℝN×64×64\.\\mathbf\{X\},\\mathbf\{Y\}\\in\\mathbb\{R\}^\{N\\times 64\\times 64\}\.\(127\)After vectorization,

𝐱\(i\)=vec⁡\(𝐗\(i\)\)∈ℝ4096,𝐲\(i\)=vec⁡\(𝐘\(i\)\)∈ℝ4096\.\\mathbf\{x\}^\{\(i\)\}=\\operatorname\{vec\}\\left\(\\mathbf\{X\}^\{\(i\)\}\\right\)\\in\\mathbb\{R\}^\{4096\},\\qquad\\mathbf\{y\}^\{\(i\)\}=\\operatorname\{vec\}\\left\(\\mathbf\{Y\}^\{\(i\)\}\\right\)\\in\\mathbb\{R\}^\{4096\}\.\(128\)The 5000 realizations are partitioned into

Ntrain=4000,Nval=500,Ntest=500\.N\_\{\\mathrm\{train\}\}=4000,\\qquad N\_\{\\mathrm\{val\}\}=500,\\qquad N\_\{\\mathrm\{test\}\}=500\.\(129\)The viscosity is

Table 11:Configuration of the fixed\-time Navier–Stokes benchmark\.
### C\.4Reduced output representation

For the Two\-Step models, the target fields are represented in a low\-dimensional, data\-adaptive basis\. Let

𝐲¯=1Ntrain​∑i=1Ntrain𝐲\(i\)\\overline\{\\mathbf\{y\}\}=\\frac\{1\}\{N\_\{\\mathrm\{train\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{train\}\}\}\\mathbf\{y\}^\{\(i\)\}\(130\)denote the mean target field, and define the centered output matrix

𝐘c=\[\(𝐲\(1\)−𝐲¯\)⊤⋮\(𝐲\(Ntrain\)−𝐲¯\)⊤\]\.\\mathbf\{Y\}\_\{c\}=\\begin\{bmatrix\}\(\\mathbf\{y\}^\{\(1\)\}\-\\overline\{\\mathbf\{y\}\}\)^\{\\top\}\\\\ \\vdots\\\\ \(\\mathbf\{y\}^\{\(N\_\{\\mathrm\{train\}\}\)\}\-\\overline\{\\mathbf\{y\}\}\)^\{\\top\}\\end\{bmatrix\}\.\(131\)The singular value decomposition is

𝐘c=𝐔​𝚺​𝐕⊤\.\\mathbf\{Y\}\_\{c\}=\\mathbf\{U\}\\mathbf\{\\Sigma\}\\mathbf\{V\}^\{\\top\}\.\(132\)The firstrrright singular vectors define

𝐐=\[𝐯1⋯𝐯r\]∈ℝ4096×r\.\\mathbf\{Q\}=\\begin\{bmatrix\}\\mathbf\{v\}\_\{1\}&\\cdots&\\mathbf\{v\}\_\{r\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{4096\\times r\}\.\(133\)The rank is chosen according to

∑j=1rσj2∑jσj2≥ηSVD,\\frac\{\\sum\_\{j=1\}^\{r\}\\sigma\_\{j\}^\{2\}\}\{\\sum\_\{j\}\\sigma\_\{j\}^\{2\}\}\\geq\\eta\_\{\\mathrm\{SVD\}\},\(134\)whereηSVD∈\(0,1\)\\eta\_\{\\mathrm\{SVD\}\}\\in\(0,1\)is a prescribed energy threshold, subject to a prescribed maximum rank\. Each target field is approximated as

𝐲\(i\)≈𝐲¯\+𝐐𝐜\(i\),\\mathbf\{y\}^\{\(i\)\}\\approx\\overline\{\\mathbf\{y\}\}\+\\mathbf\{Q\}\\mathbf\{c\}^\{\(i\)\},\(135\)where

𝐜\(i\)=𝐐⊤​\(𝐲\(i\)−𝐲¯\)\.\\mathbf\{c\}^\{\(i\)\}=\\mathbf\{Q\}^\{\\top\}\\left\(\\mathbf\{y\}^\{\(i\)\}\-\\overline\{\\mathbf\{y\}\}\\right\)\.\(136\)Hence, the Two\-Step models learn the reduced map

𝐜^\(i\)=ℬθ​\(ℳ​\(a\(i\)\)\),\\widehat\{\\mathbf\{c\}\}^\{\(i\)\}=\\mathcal\{B\}\_\{\\theta\}\\left\(\\mathcal\{M\}\\left\(a^\{\(i\)\}\\right\)\\right\),\(137\)whereℳ\\mathcal\{M\}denotes the input representation associated with the particular model\.

### C\.5Benchmark models

##### Full\-field Two\-Step\.

The Full\-field Two\-Step model receives the complete input field:

ℳfull​\(a\)=a∈ℝ64×64\.\\mathcal\{M\}\_\{\\mathrm\{full\}\}\(a\)=a\\in\\mathbb\{R\}^\{64\\times 64\}\.\(138\)The branch predicts

𝐜^=ℬθfull​\(a\),\\widehat\{\\mathbf\{c\}\}=\\mathcal\{B\}^\{\\mathrm\{full\}\}\_\{\\theta\}\(a\),\(139\)and the output is reconstructed through

𝐲^=𝐲¯\+𝐐​𝐜^\.\\widehat\{\\mathbf\{y\}\}=\\overline\{\\mathbf\{y\}\}\+\\mathbf\{Q\}\\widehat\{\\mathbf\{c\}\}\.\(140\)

##### Sensor Two\-Step\.

Let

𝒮=\{𝒙s1,…,𝒙sq\}\\mathcal\{S\}=\\left\\\{\\bm\{x\}\_\{s\_\{1\}\},\\ldots,\\bm\{x\}\_\{s\_\{q\}\}\\right\\\}\(141\)be a prescribed set ofq=128q=128sensor locations\. The measurements are

zjsens=a​\(𝒙sj\),j=1,…,q\.z\_\{j\}^\{\\mathrm\{sens\}\}=a\(\\bm\{x\}\_\{s\_\{j\}\}\),\\qquad j=1,\\ldots,q\.\(142\)Thus,

ℳsens​\(a\)=\[a​\(𝒙s1\)⋯a​\(𝒙sq\)\]⊤∈ℝq\.\\mathcal\{M\}\_\{\\mathrm\{sens\}\}\(a\)=\\begin\{bmatrix\}a\(\\bm\{x\}\_\{s\_\{1\}\}\)&\\cdots&a\(\\bm\{x\}\_\{s\_\{q\}\}\)\\end\{bmatrix\}^\{\\top\}\\in\\mathbb\{R\}^\{q\}\.\(143\)The corresponding compression ratio is

CR=4096128=32\.\\mathrm\{CR\}=\\frac\{4096\}\{128\}=32\.\(144\)When a Fourier spectral\-layer branch is employed, the point values are placed on a sparse grid,

S​\(𝒙\)=\{a​\(𝒙\),𝒙∈𝒮,0,otherwise,S\(\\bm\{x\}\)=\\begin\{cases\}a\(\\bm\{x\}\),&\\bm\{x\}\\in\\mathcal\{S\},\\\\ 0,&\\text\{otherwise\},\\end\{cases\}\(145\)and are accompanied by the observation mask

M​\(𝒙\)=\{1,𝒙∈𝒮,0,otherwise\.M\(\\bm\{x\}\)=\\begin\{cases\}1,&\\bm\{x\}\\in\\mathcal\{S\},\\\\ 0,&\\text\{otherwise\}\.\\end\{cases\}\(146\)

##### Fixed Topological Two\-Step\.

Let

Ψ=\{ψk\}k=1K\\Psi=\\left\\\{\\psi\_\{k\}\\right\\\}\_\{k=1\}^\{K\}\(147\)denote a dictionary of localized multiscale test functions\. For each input fieldaa, define

hk​\(a\)=⟨a,ψk⟩=∫Da​\(𝒙\)​ψk​\(𝒙\)​𝑑𝒙\.h\_\{k\}\(a\)=\\left\\langle a,\\psi\_\{k\}\\right\\rangle=\\int\_\{D\}a\(\\bm\{x\}\)\\psi\_\{k\}\(\\bm\{x\}\)\\,d\\bm\{x\}\.\(148\)A subset ofq=128q=128predictive atoms is selected,

ℐ=\{i1,…,iq\},\\mathcal\{I\}=\\\{i\_\{1\},\\ldots,i\_\{q\}\\\},\(149\)and the fixed measurements are

zjfix=⟨a,ψij⟩,j=1,…,q\.z\_\{j\}^\{\\mathrm\{fix\}\}=\\left\\langle a,\\psi\_\{i\_\{j\}\}\\right\\rangle,\\qquad j=1,\\ldots,q\.\(150\)The fixed measurement operator is therefore

ℳfix​\(a\)=\[⟨a,ψi1⟩⋮⟨a,ψiq⟩\]∈ℝq\.\\mathcal\{M\}\_\{\\mathrm\{fix\}\}\(a\)=\\begin\{bmatrix\}\\langle a,\\psi\_\{i\_\{1\}\}\\rangle\\\\ \\vdots\\\\ \\langle a,\\psi\_\{i\_\{q\}\}\\rangle\\end\{bmatrix\}\\in\\mathbb\{R\}^\{q\}\.\(151\)For a fair comparison, these measurements are synthesized into a spatial proxy,

a~fix​\(𝒙\)=∑j=1qzjfix​ψ~ij​\(𝒙\),\\widetilde\{a\}\_\{\\mathrm\{fix\}\}\(\\bm\{x\}\)=\\sum\_\{j=1\}^\{q\}z\_\{j\}^\{\\mathrm\{fix\}\}\\widetilde\{\\psi\}\_\{i\_\{j\}\}\(\\bm\{x\}\),\(152\)whereψ~ij\\widetilde\{\\psi\}\_\{i\_\{j\}\}denotes the associated synthesis atom\.

##### Adaptive Topological Two\-Step\.

Let

𝒉​\(a\)=\[h1​\(a\)⋯hK​\(a\)\]\\bm\{h\}\(a\)=\\begin\{bmatrix\}h\_\{1\}\(a\)&\\cdots&h\_\{K\}\(a\)\\end\{bmatrix\}\(153\)contain all dictionary measurements, and let

𝐀0∈ℝK×q\\mathbf\{A\}\_\{0\}\\in\\mathbb\{R\}^\{K\\times q\}\(154\)denote the fixed selection matrix\. The adaptive measurement matrix is

𝐀θ=𝐀0\+Δ​𝐀θ\.\\mathbf\{A\}\_\{\\theta\}=\\mathbf\{A\}\_\{0\}\+\\Delta\\mathbf\{A\}\_\{\\theta\}\.\(155\)The adaptive measurements are

𝒛ad=𝒉​\(a\)​𝐀θ∈ℝq\.\\bm\{z\}^\{\\mathrm\{ad\}\}=\\bm\{h\}\(a\)\\mathbf\{A\}\_\{\\theta\}\\in\\mathbb\{R\}^\{q\}\.\(156\)At initialization,

Δ​𝐀θ=0,\\Delta\\mathbf\{A\}\_\{\\theta\}=0,\(157\)so the adaptive model coincides exactly with the fixed model at epoch zero\. Equivalently, the learned functionals are

ϕjθ=∑k=1K\(𝐀θ\)k​j​ψk,\\phi\_\{j\}^\{\\theta\}=\\sum\_\{k=1\}^\{K\}\(\\mathbf\{A\}\_\{\\theta\}\)\_\{kj\}\\psi\_\{k\},\(158\)with

zjad=⟨a,ϕjθ⟩\.z\_\{j\}^\{\\mathrm\{ad\}\}=\\langle a,\\phi\_\{j\}^\{\\theta\}\\rangle\.\(159\)A drift penalty controls departure from the fixed measurement system:

ℒdrift=λdrift​‖Δ​𝐀θ‖F2\.\\mathcal\{L\}\_\{\\mathrm\{drift\}\}=\\lambda\_\{\\mathrm\{drift\}\}\\left\\\|\\Delta\\mathbf\{A\}\_\{\\theta\}\\right\\\|\_\{F\}^\{2\}\.\(160\)

##### Vanilla DeepONet\.

The Vanilla DeepONet\[[16](https://arxiv.org/html/2608.06428#bib.bib3)\]uses a branch network to encode the input field and a trunk network to encode the output coordinate:

𝒃θ​\(a\)=\[b1​\(a\),…,bp​\(a\)\],\\bm\{b\}\_\{\\theta\}\(a\)=\\begin\{bmatrix\}b\_\{1\}\(a\),\\ldots,b\_\{p\}\(a\)\\end\{bmatrix\},\(161\)and

𝒕η​\(𝒙\)=\[t1​\(𝒙\),…,tp​\(𝒙\)\]\.\\bm\{t\}\_\{\\eta\}\(\\bm\{x\}\)=\\begin\{bmatrix\}t\_\{1\}\(\\bm\{x\}\),\\ldots,t\_\{p\}\(\\bm\{x\}\)\\end\{bmatrix\}\.\(162\)The predicted target field is

u^​\(𝒙;a\)=∑ℓ=1pbℓ​\(a\)​tℓ​\(𝒙\)\+b0\.\\widehat\{u\}\(\\bm\{x\};a\)=\\sum\_\{\\ell=1\}^\{p\}b\_\{\\ell\}\(a\)t\_\{\\ell\}\(\\bm\{x\}\)\+b\_\{0\}\.\(163\)In the common\-backbone comparison, the branch is implemented using a 2D Fourier spectral\-layer encoder, while the trunk is a coordinate multilayer perceptron\.

### C\.62D Fourier spectral\-layer branch architecture

For models using a Fourier spectral\-layer branch\[[13](https://arxiv.org/html/2608.06428#bib.bib4)\], one spectral layer is written schematically as

vℓ\+1​\(𝒙\)=σ​\[Wℓ​vℓ​\(𝒙\)\+ℱ−1​\(Rℓ​\(𝒌\)​ℱ​\[vℓ\]​\(𝒌\)\)​\(𝒙\)\],v\_\{\\ell\+1\}\(\\bm\{x\}\)=\\sigma\\left\[W\_\{\\ell\}v\_\{\\ell\}\(\\bm\{x\}\)\+\\mathcal\{F\}^\{\-1\}\\left\(R\_\{\\ell\}\(\\bm\{k\}\)\\mathcal\{F\}\[v\_\{\\ell\}\]\(\\bm\{k\}\)\\right\)\(\\bm\{x\}\)\\right\],\(164\)whereℱ\\mathcal\{F\}denotes the discrete Fourier transform,RℓR\_\{\\ell\}is a learned complex spectral multiplier,WℓW\_\{\\ell\}is a local linear transformation, andσ\\sigmais a nonlinear activation\. For the present fixed\-time problem, the branch acts on the two spatial dimensions only\.

### C\.7Training objective

Let𝐜^\(i\)\\widehat\{\\mathbf\{c\}\}^\{\(i\)\}and𝐜\(i\)\\mathbf\{c\}^\{\(i\)\}denote the predicted and exact output coefficients\. The coefficient\-space mean squared error is

ℒcoef=1B​r​∑i=1B‖𝐜^\(i\)−𝐜\(i\)‖22,\\mathcal\{L\}\_\{\\mathrm\{coef\}\}=\\frac\{1\}\{Br\}\\sum\_\{i=1\}^\{B\}\\left\\\|\\widehat\{\\mathbf\{c\}\}^\{\(i\)\}\-\\mathbf\{c\}^\{\(i\)\}\\right\\\|\_\{2\}^\{2\},\(165\)and the relative coefficient loss is

ℒrel=1B​∑i=1B‖𝐜^\(i\)−𝐜\(i\)‖2‖𝐜\(i\)‖2\+ε\.\\mathcal\{L\}\_\{\\mathrm\{rel\}\}=\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\frac\{\\left\\\|\\widehat\{\\mathbf\{c\}\}^\{\(i\)\}\-\\mathbf\{c\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\left\\\|\\mathbf\{c\}^\{\(i\)\}\\right\\\|\_\{2\}\+\\varepsilon\}\.\(166\)The Two\-Step objective is

ℒTwoStep=ℒcoef\+λrel​ℒrel\.\\mathcal\{L\}\_\{\\mathrm\{TwoStep\}\}=\\mathcal\{L\}\_\{\\mathrm\{coef\}\}\+\\lambda\_\{\\mathrm\{rel\}\}\\mathcal\{L\}\_\{\\mathrm\{rel\}\}\.\(167\)For the adaptive model,

ℒAdaptive=ℒTwoStep\+λdrift​‖Δ​𝐀θ‖F2\.\\mathcal\{L\}\_\{\\mathrm\{Adaptive\}\}=\\mathcal\{L\}\_\{\\mathrm\{TwoStep\}\}\+\\lambda\_\{\\mathrm\{drift\}\}\\left\\\|\\Delta\\mathbf\{A\}\_\{\\theta\}\\right\\\|\_\{F\}^\{2\}\.\(168\)

### C\.8Evaluation metrics

The principal metric is the global relativeL2L^\{2\}error:

εglobal=‖𝐘^−𝐘‖F‖𝐘‖F\.\\varepsilon\_\{\\mathrm\{global\}\}=\\frac\{\\left\\\|\\widehat\{\\mathbf\{Y\}\}\-\\mathbf\{Y\}\\right\\\|\_\{F\}\}\{\\left\\\|\\mathbf\{Y\}\\right\\\|\_\{F\}\}\.\(169\)For theii\-th test sample, the relativeL2L^\{2\}error is

εi=‖𝐲^\(i\)−𝐲\(i\)‖2‖𝐲\(i\)‖2\.\\varepsilon\_\{i\}=\\frac\{\\left\\\|\\widehat\{\\mathbf\{y\}\}^\{\(i\)\}\-\\mathbf\{y\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\left\\\|\\mathbf\{y\}^\{\(i\)\}\\right\\\|\_\{2\}\}\.\(170\)The mean sample\-wise relative error is

ε¯=1Ntest​∑i=1Ntestεi\.\\overline\{\\varepsilon\}=\\frac\{1\}\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}\\varepsilon\_\{i\}\.\(171\)The comparison is designed to determine whether distributed linear functionals preserve more predictive information than point sensors at the same feature dimensionq=128q=128, whether adaptive functionals improve on their fixed counterparts, and how closely compressed\-input models approach models that observe the complete64×6464\\times 64field\.

## Appendix DAdditional details for the time\-evolving Navier–Stokes experiment

### D\.1Discrete trajectory representation

Each vorticity snapshot is stored on a64×6464\\times 64periodic grid and is therefore represented by a vector in

Each realization contains5050stored snapshots\. The firstTin=10T\_\{\\mathrm\{in\}\}=10snapshots define the input history,

𝒙\(i\)=\[ω\(i\)​\(⋅,t0\),…,ω\(i\)​\(⋅,t9\)\],\\bm\{x\}^\{\(i\)\}=\\left\[\\omega^\{\(i\)\}\(\\cdot,t\_\{0\}\),\\ldots,\\omega^\{\(i\)\}\(\\cdot,t\_\{9\}\)\\right\],\(172\)while the remainingTout=40T\_\{\\mathrm\{out\}\}=40snapshots define the target trajectory,

𝒚\(i\)=\[ω\(i\)​\(⋅,t10\),…,ω\(i\)​\(⋅,t49\)\]\.\\bm\{y\}^\{\(i\)\}=\\left\[\\omega^\{\(i\)\}\(\\cdot,t\_\{10\}\),\\ldots,\\omega^\{\(i\)\}\(\\cdot,t\_\{49\}\)\\right\]\.\(173\)Thus,

𝒙\(i\)∈ℝ64×64×10,𝒚\(i\)∈ℝ64×64×40\.\\bm\{x\}^\{\(i\)\}\\in\\mathbb\{R\}^\{64\\times 64\\times 10\},\\qquad\\bm\{y\}^\{\(i\)\}\\in\\mathbb\{R\}^\{64\\times 64\\times 40\}\.\(174\)The realizations are divided into

Ntrain=4000,Nval=500,Ntest=500\.N\_\{\\mathrm\{train\}\}=4000,\\qquad N\_\{\\mathrm\{val\}\}=500,\\qquad N\_\{\\mathrm\{test\}\}=500\.\(175\)The implementation constructs the input and output windows directly from the first ten and subsequent forty snapshots, respectively\.

### D\.2Input and output function spaces

Let

𝒱=Lper2​\(D\),D=\(0,1\)2,\\mathcal\{V\}=L\_\{\\mathrm\{per\}\}^\{2\}\(D\),\\qquad D=\(0,1\)^\{2\},\(176\)denote the space of square\-integrable periodic vorticity fields\. The time\-evolving operator is

𝒢:𝒱10⟶𝒱40,\\mathcal\{G\}:\\mathcal\{V\}^\{10\}\\longrightarrow\\mathcal\{V\}^\{40\},\(177\)with

𝒢​\(ω​\(⋅,t0\),…,ω​\(⋅,t9\)\)=\(ω​\(⋅,t10\),…,ω​\(⋅,t49\)\)\.\\mathcal\{G\}\\left\(\\omega\(\\cdot,t\_\{0\}\),\\ldots,\\omega\(\\cdot,t\_\{9\}\)\\right\)=\\left\(\\omega\(\\cdot,t\_\{10\}\),\\ldots,\\omega\(\\cdot,t\_\{49\}\)\\right\)\.\(178\)Equivalently, using discrete\-time sequence spaces,

𝒢:ℓ2​\(\{t0,…,t9\};Lper2​\(D\)\)⟶ℓ2​\(\{t10,…,t49\};Lper2​\(D\)\)\.\\mathcal\{G\}:\\ell^\{2\}\\left\(\\\{t\_\{0\},\\ldots,t\_\{9\}\\\};L\_\{\\mathrm\{per\}\}^\{2\}\(D\)\\right\)\\longrightarrow\\ell^\{2\}\\left\(\\\{t\_\{10\},\\ldots,t\_\{49\}\\\};L\_\{\\mathrm\{per\}\}^\{2\}\(D\)\\right\)\.\(179\)

### D\.3Training\-data standardization

The input histories are standardized using the scalar mean and standard deviation computed from the training inputs:

ωstd=ω−μinσin\.\\omega\_\{\\mathrm\{std\}\}=\\frac\{\\omega\-\\mu\_\{\\mathrm\{in\}\}\}\{\\sigma\_\{\\mathrm\{in\}\}\}\.\(180\)The sameμin\\mu\_\{\\mathrm\{in\}\}andσin\\sigma\_\{\\mathrm\{in\}\}are used for training, validation, and testing\. No validation or test statistics are used during normalization\. The same scalar standardization procedure is applied to the synthesized topological proxy fields\.

### D\.4Multiscale functional dictionary

For each input snapshot, the topological representations are constructed from a dictionary

𝒟K=\{ψ1,…,ψK\}⊂Lper2​\(D\)\.\\mathcal\{D\}\_\{K\}=\\left\\\{\\psi\_\{1\},\\ldots,\\psi\_\{K\}\\right\\\}\\subset L\_\{\\mathrm\{per\}\}^\{2\}\(D\)\.\(181\)The default dictionary size is

K=γ​q,γ=4,q=128,K=\\gamma q,\\qquad\\gamma=4,\\qquad q=128,\(182\)so thatK=512K=512\. Let

𝒄=\(cx,cy\)\\bm\{c\}=\(c\_\{x\},c\_\{y\}\)denote a dictionary center and let

dx​\(x,cx\)=min⁡\(\|x−cx\|,1−\|x−cx\|\),d\_\{x\}\(x,c\_\{x\}\)=\\min\\left\(\|x\-c\_\{x\}\|,1\-\|x\-c\_\{x\}\|\\right\),dy​\(y,cy\)=min⁡\(\|y−cy\|,1−\|y−cy\|\)d\_\{y\}\(y,c\_\{y\}\)=\\min\\left\(\|y\-c\_\{y\}\|,1\-\|y\-c\_\{y\}\|\\right\)denote periodic coordinate distances\. Define

rper2=dx2\+dy2\.r\_\{\\mathrm\{per\}\}^\{2\}=d\_\{x\}^\{2\}\+d\_\{y\}^\{2\}\.\(183\)For each grid\-relative scale

s∈𝒮=\{1\.5,3,6,12\},s\\in\\mathcal\{S\}=\\\{1\.5,3,6,12\\\},\(184\)the corresponding normalized spatial width is

σs=max⁡\(smax⁡\(Nx,Ny\),1max⁡\(Nx,Ny\)\)\.\\sigma\_\{s\}=\\max\\left\(\\frac\{s\}\{\\max\(N\_\{x\},N\_\{y\}\)\},\\frac\{1\}\{\\max\(N\_\{x\},N\_\{y\}\)\}\\right\)\.\(185\)For the64×6464\\times 64grid,

σs∈\{1\.564,364,664,1264\}\.\\sigma\_\{s\}\\in\\left\\\{\\frac\{1\.5\}\{64\},\\frac\{3\}\{64\},\\frac\{6\}\{64\},\\frac\{12\}\{64\}\\right\\\}\.\(186\)The Gaussian envelope is

g𝒄,σ​\(𝒙\)=exp⁡\(−rper22​σ2\)\.g\_\{\\bm\{c\},\\sigma\}\(\\bm\{x\}\)=\\exp\\left\(\-\\frac\{r\_\{\\mathrm\{per\}\}^\{2\}\}\{2\\sigma^\{2\}\}\\right\)\.\(187\)The dictionary contains the following atom families:

ψ𝒄,σ\(0\)​\(𝒙\)\\displaystyle\\psi\_\{\\bm\{c\},\\sigma\}^\{\(0\)\}\(\\bm\{x\}\)=g𝒄,σ​\(𝒙\),\\displaystyle=g\_\{\\bm\{c\},\\sigma\}\(\\bm\{x\}\),\(188\)ψ𝒄,σ\(x\)​\(𝒙\)\\displaystyle\\psi\_\{\\bm\{c\},\\sigma\}^\{\(x\)\}\(\\bm\{x\}\)=dx​\(x,cx\)σ​g𝒄,σ​\(𝒙\),\\displaystyle=\\frac\{d\_\{x\}\(x,c\_\{x\}\)\}\{\\sigma\}g\_\{\\bm\{c\},\\sigma\}\(\\bm\{x\}\),ψ𝒄,σ\(y\)​\(𝒙\)\\displaystyle\\psi\_\{\\bm\{c\},\\sigma\}^\{\(y\)\}\(\\bm\{x\}\)=dy​\(y,cy\)σ​g𝒄,σ​\(𝒙\),\\displaystyle=\\frac\{d\_\{y\}\(y,c\_\{y\}\)\}\{\\sigma\}g\_\{\\bm\{c\},\\sigma\}\(\\bm\{x\}\),ψ𝒄,σ\(Δ\)​\(𝒙\)\\displaystyle\\psi\_\{\\bm\{c\},\\sigma\}^\{\(\\Delta\)\}\(\\bm\{x\}\)=\(rper2σ2−2\)​g𝒄,σ​\(𝒙\)\.\\displaystyle=\\left\(\\frac\{r\_\{\\mathrm\{per\}\}^\{2\}\}\{\\sigma^\{2\}\}\-2\\right\)g\_\{\\bm\{c\},\\sigma\}\(\\bm\{x\}\)\.A constant atom is also included\. The implementation labels the second and third families asdxanddy\. Since unsigned periodic distances are used, these should be interpreted as directionally weighted Gaussian atoms rather than exact signed Gaussian derivatives\. The dictionary construction and scale conversion are implemented directly in the code\.

### D\.5Atom normalization

Each nonconstant atom is centered by subtracting its discrete mean:

ψj←ψj−1Nh​∑m=1Nhψj​\(𝒙m\),\\psi\_\{j\}\\leftarrow\\psi\_\{j\}\-\\frac\{1\}\{N\_\{h\}\}\\sum\_\{m=1\}^\{N\_\{h\}\}\\psi\_\{j\}\(\\bm\{x\}\_\{m\}\),\(189\)whereNh=4096N\_\{h\}=4096\. It is then normalized using the discrete Euclidean norm:

ψj←ψj\(∑m=1Nhψj​\(𝒙m\)2\)1/2\.\\psi\_\{j\}\\leftarrow\\frac\{\\psi\_\{j\}\}\{\\left\(\\sum\_\{m=1\}^\{N\_\{h\}\}\\psi\_\{j\}\(\\bm\{x\}\_\{m\}\)^\{2\}\\right\)^\{1/2\}\}\.\(190\)The assembled dictionary matrix is normalized columnwise once more for numerical robustness\.

### D\.6Functional measurements

Let

Ψ=\[𝝍1⋯𝝍K\]∈ℝ4096×K\\Psi=\\begin\{bmatrix\}\\bm\{\\psi\}\_\{1\}&\\cdots&\\bm\{\\psi\}\_\{K\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{4096\\times K\}\(191\)denote the discrete dictionary matrix \(denotedΨ\\Psihere to avoid a clash with the mini\-batch sizeBBused in the training losses\)\. For thenn\-th input snapshot, the complete dictionary\-coordinate vector is

𝒉​\(tn\)=Ψ𝖳​𝝎​\(tn\)∈ℝK\.\\bm\{h\}\(t\_\{n\}\)=\\Psi^\{\\mathsf\{T\}\}\\bm\{\\omega\}\(t\_\{n\}\)\\in\\mathbb\{R\}^\{K\}\.\(192\)Equivalently,

hj​\(tn\)=⟨ω​\(⋅,tn\),ψj⟩h,h\_\{j\}\(t\_\{n\}\)=\\left\\langle\\omega\(\\cdot,t\_\{n\}\),\\psi\_\{j\}\\right\\rangle\_\{h\},\(193\)where⟨⋅,⋅⟩h\\langle\\cdot,\\cdot\\rangle\_\{h\}denotes the discrete grid inner product used by the implementation\. For a batch ofNNtrajectories, the full measurement tensor has shape

H∈ℝN×Tin×K\.H\\in\\mathbb\{R\}^\{N\\times T\_\{\\mathrm\{in\}\}\\times K\}\.\(194\)The implementation computes these measurements by multiplying each flattened input snapshot by the dictionary matrix\.

### D\.7Predictive selection of fixed measurements

The fixed topological model selectsq=128q=128atoms from the complete dictionary\. Each dictionary coordinate is standardized over the training set and scored according to its cross\-covariance with the reduced output coefficients\. Let

H~∈ℝNtrain×Tin×K\\widetilde\{H\}\\in\\mathbb\{R\}^\{N\_\{\\mathrm\{train\}\}\\times T\_\{\\mathrm\{in\}\}\\times K\}\(195\)denote the standardized measurement tensor and let

C∈ℝNtrain×rC\\in\\mathbb\{R\}^\{N\_\{\\mathrm\{train\}\}\\times r\}\(196\)contain the centered output coefficients\. For atomjjand input timetnt\_\{n\}, define

𝜸n,j=1Ntrain​H~:,n,j𝖳​C\.\\bm\{\\gamma\}\_\{n,j\}=\\frac\{1\}\{N\_\{\\mathrm\{train\}\}\}\\widetilde\{H\}\_\{:,n,j\}^\{\\mathsf\{T\}\}C\.\(197\)The time\-aggregated predictive score is

sj=\[1Tin​∑n=0Tin−1‖𝜸n,j‖22\]1/2\.s\_\{j\}=\\left\[\\frac\{1\}\{T\_\{\\mathrm\{in\}\}\}\\sum\_\{n=0\}^\{T\_\{\\mathrm\{in\}\}\-1\}\\left\\\|\\bm\{\\gamma\}\_\{n,j\}\\right\\\|\_\{2\}^\{2\}\\right\]^\{1/2\}\.\(198\)Theqqatoms with the largest scores define the fixed measurement index set

ℐq\.\\mathcal\{I\}\_\{q\}\.\(199\)This selection is performed using training data only\.

### D\.8Dual synthesis and proxy reconstruction

Let

Ψq=Ψ:,ℐq∈ℝ4096×q\\Psi\_\{q\}=\\Psi\_\{:,\\mathcal\{I\}\_\{q\}\}\\in\\mathbb\{R\}^\{4096\\times q\}\(200\)contain the selected atoms\. The regularized dual synthesis matrix is

Dq=Ψq​\(Ψq𝖳​Ψq\+εdual​I\)−1,εdual=10−5\.D\_\{q\}=\\Psi\_\{q\}\\left\(\\Psi\_\{q\}^\{\\mathsf\{T\}\}\\Psi\_\{q\}\+\\varepsilon\_\{\\mathrm\{dual\}\}I\\right\)^\{\-1\},\\qquad\\varepsilon\_\{\\mathrm\{dual\}\}=10^\{\-5\}\.\(201\)For the selected measurements

𝒛​\(tn\)=𝒉ℐq​\(tn\)∈ℝq,\\bm\{z\}\(t\_\{n\}\)=\\bm\{h\}\_\{\\mathcal\{I\}\_\{q\}\}\(t\_\{n\}\)\\in\\mathbb\{R\}^\{q\},\(202\)the synthesized proxy snapshot is

𝝎~q​\(tn\)=Dq​𝒛​\(tn\)\.\\widetilde\{\\bm\{\\omega\}\}\_\{q\}\(t\_\{n\}\)=D\_\{q\}\\bm\{z\}\(t\_\{n\}\)\.\(203\)The complete proxy history belongs to

ℝ64×64×10\.\\mathbb\{R\}^\{64\\times 64\\times 10\}\.\(204\)This proxy history is standardized using training\-set proxy statistics and then passed to the common three\-dimensional spectral encoder\.

### D\.9Sensor representation

The sensor model usesq=128q=128approximately uniformly distributed spatial locations,

\{𝒙s1,…,𝒙sq\}\.\\left\\\{\\bm\{x\}\_\{s\_\{1\}\},\\ldots,\\bm\{x\}\_\{s\_\{q\}\}\\right\\\}\.\(205\)For every input time,

zjsens​\(tn\)=ω​\(𝒙sj,tn\)\.z\_\{j\}^\{\\mathrm\{sens\}\}\(t\_\{n\}\)=\\omega\(\\bm\{x\}\_\{s\_\{j\}\},t\_\{n\}\)\.\(206\)The measured values are embedded into a sparse field

ωsparse∈ℝ64×64×10,\\omega\_\{\\mathrm\{sparse\}\}\\in\\mathbb\{R\}^\{64\\times 64\\times 10\},\(207\)and a binary mask

msens∈ℝ64×64×10m\_\{\\mathrm\{sens\}\}\\in\\mathbb\{R\}^\{64\\times 64\\times 10\}\(208\)indicates the sensor locations\. The FNO branch therefore receives the two\-channel tensor

\[ωsparse,msens\]∈ℝ64×64×10×2\.\\left\[\\omega\_\{\\mathrm\{sparse\}\},m\_\{\\mathrm\{sens\}\}\\right\]\\in\\mathbb\{R\}^\{64\\times 64\\times 10\\times 2\}\.\(209\)This representation is constructed explicitly in the sensor\-proxy routine\.

### D\.10Adaptive topological measurements

The adaptive model starts from the fixed selected coordinates and adds a trainable correction formed from all dictionary coordinates:

𝒛ad​\(tn\)=𝒉ℐq​\(tn\)\+𝒉​\(tn\)​Δ​Aθ,\\bm\{z\}\_\{\\mathrm\{ad\}\}\(t\_\{n\}\)=\\bm\{h\}\_\{\\mathcal\{I\}\_\{q\}\}\(t\_\{n\}\)\+\\bm\{h\}\(t\_\{n\}\)\\Delta A\_\{\\theta\},\(210\)where

Δ​Aθ∈ℝK×q\.\\Delta A\_\{\\theta\}\\in\\mathbb\{R\}^\{K\\times q\}\.\(211\)The correction matrix is initialized as

Δ​Aθ=0,\\Delta A\_\{\\theta\}=0,\(212\)so that the adaptive and fixed coordinates coincide at initialization\. The adaptive proxy is reconstructed using the same fixed dual matrix:

𝝎~ad​\(tn\)=Dq​𝒛ad​\(tn\)\.\\widetilde\{\\bm\{\\omega\}\}\_\{\\mathrm\{ad\}\}\(t\_\{n\}\)=D\_\{q\}\\bm\{z\}\_\{\\mathrm\{ad\}\}\(t\_\{n\}\)\.\(213\)Thus, the adaptive model changes the measurement coordinates while retaining the same synthesis space and spectral\-encoder branch architecture\.

### D\.11Common three\-dimensional Fourier spectral encoder

All four Two\-Step models use the same three\-dimensional FNO encoder acting over the two spatial dimensions and the input\-time dimension\. The encoder receives a tensor of the form

X∈ℝNx×Ny×Tin×Cin,X\\in\\mathbb\{R\}^\{N\_\{x\}\\times N\_\{y\}\\times T\_\{\\mathrm\{in\}\}\\times C\_\{\\mathrm\{in\}\}\},\(214\)whereCin=1C\_\{\\mathrm\{in\}\}=1for the full\-field and topological models andCin=2C\_\{\\mathrm\{in\}\}=2for the sensor model\. Normalized coordinate channels

are concatenated to the input\. Each spectral layer applies a truncated three\-dimensional Fourier convolution and a local pointwise convolution:

vℓ\+1=σ​\[ℱ−1​\(Rℓ⊙ℱ​\(vℓ\)\)\+Wℓ​vℓ\]\.v\_\{\\ell\+1\}=\\sigma\\left\[\\mathcal\{F\}^\{\-1\}\\left\(R\_\{\\ell\}\\odot\\mathcal\{F\}\(v\_\{\\ell\}\)\\right\)\+W\_\{\\ell\}v\_\{\\ell\}\\right\]\.\(216\)The default spectral truncation is

\(kx,ky,kt\)=\(12,12,6\),\(k\_\{x\},k\_\{y\},k\_\{t\}\)=\(12,12,6\),\(217\)with width3232and four spectral layers\. After the last layer, global averaging over the spatial and input\-time coordinates produces the branch latent vector\.

### D\.12Reduced output trajectory basis

Let

Ytrain∈ℝ4000×\(4096⋅40\)Y\_\{\\mathrm\{train\}\}\\in\\mathbb\{R\}^\{4000\\times\(4096\\cdot 40\)\}\(218\)contain the flattened training output trajectories\. The mean trajectory is

𝒚¯=1Ntrain​∑i=1Ntrain𝒚\(i\),\\overline\{\\bm\{y\}\}=\\frac\{1\}\{N\_\{\\mathrm\{train\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{train\}\}\}\\bm\{y\}^\{\(i\)\},\(219\)and the centered data matrix is

Yc=Ytrain−𝟏​𝒚¯𝖳\.Y\_\{c\}=Y\_\{\\mathrm\{train\}\}\-\\bm\{1\}\\overline\{\\bm\{y\}\}^\{\\mathsf\{T\}\}\.\(220\)The economy singular\-value decomposition is

Yc=U​Σ​V𝖳\.Y\_\{c\}=U\\Sigma V^\{\\mathsf\{T\}\}\.\(221\)The retained rank is selected so that

∑j=1rσj2∑jσj2≥ηSVD,\\frac\{\\sum\_\{j=1\}^\{r\}\\sigma\_\{j\}^\{2\}\}\{\\sum\_\{j\}\\sigma\_\{j\}^\{2\}\}\\geq\\eta\_\{\\mathrm\{SVD\}\},\(222\)subject to

r≤rmax\.r\\leq r\_\{\\max\}\.\(223\)In the default configuration,

ηSVD=0\.999,rmax=128\.\\eta\_\{\\mathrm\{SVD\}\}=0\.999,\\qquad r\_\{\\max\}=128\.\(224\)The output basis is

Q=V:,1:r∈ℝ\(4096⋅40\)×r\.Q=V\_\{:,1:r\}\\in\\mathbb\{R\}^\{\(4096\\cdot 40\)\\times r\}\.\(225\)The coefficient vector is

𝒄\(i\)=Q𝖳​\(𝒚\(i\)−𝒚¯\),\\bm\{c\}^\{\(i\)\}=Q^\{\\mathsf\{T\}\}\\left\(\\bm\{y\}^\{\(i\)\}\-\\overline\{\\bm\{y\}\}\\right\),\(226\)and the reconstructed trajectory is

𝒚^\(i\)=𝒚¯\+Q​𝒄^\(i\)\.\\widehat\{\\bm\{y\}\}^\{\(i\)\}=\\overline\{\\bm\{y\}\}\+Q\\widehat\{\\bm\{c\}\}^\{\(i\)\}\.\(227\)The SVD basis is constructed from the training outputs only\.

### D\.13Coefficient standardization and Two\-Step loss

Each retained coefficient is standardized independently using its training mean and standard deviation:

cj,std=cj−μcjσcj\.c\_\{j,\\mathrm\{std\}\}=\\frac\{c\_\{j\}\-\\mu\_\{c\_\{j\}\}\}\{\\sigma\_\{c\_\{j\}\}\}\.\(228\)The branch network predicts standardized coefficients\. Before evaluating the training loss, the implementation maps both predictions and targets back to the physical coefficient scale:

𝒄^=𝒄^std⊙𝝈c\+𝝁c\.\\widehat\{\\bm\{c\}\}=\\widehat\{\\bm\{c\}\}\_\{\\mathrm\{std\}\}\\odot\\bm\{\\sigma\}\_\{c\}\+\\bm\{\\mu\}\_\{c\}\.\(229\)The Two\-Step loss is

ℒTwoStep\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{TwoStep\}\}=1B​r​∑i=1B‖𝒄^\(i\)−𝒄\(i\)‖22\\displaystyle=\\frac\{1\}\{Br\}\\sum\_\{i=1\}^\{B\}\\left\\\|\\widehat\{\\bm\{c\}\}^\{\(i\)\}\-\\bm\{c\}^\{\(i\)\}\\right\\\|\_\{2\}^\{2\}\(230\)\+λrel​1B​∑i=1B‖𝒄^\(i\)−𝒄\(i\)‖2‖𝒄\(i\)‖2\+ε\.\\displaystyle\\quad\+\\lambda\_\{\\mathrm\{rel\}\}\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\frac\{\\left\\\|\\widehat\{\\bm\{c\}\}^\{\(i\)\}\-\\bm\{c\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\left\\\|\\bm\{c\}^\{\(i\)\}\\right\\\|\_\{2\}\+\\varepsilon\}\.Computing the loss after undoing coefficient standardization restores the physical POD/SVD energy weighting and avoids overemphasizing weak high\-index modes\. For the adaptive model, the drift penalty

ℒdrift=λdrift​‖Δ​Aθ‖F2\\mathcal\{L\}\_\{\\mathrm\{drift\}\}=\\lambda\_\{\\mathrm\{drift\}\}\\left\\\|\\Delta A\_\{\\theta\}\\right\\\|\_\{F\}^\{2\}\(231\)is added to the coefficient\-space loss\.

### D\.14Adaptive training schedule

The adaptive branch is initialized from the trained fixed topological branch\. Training then proceeds in two stages\. In the first stage, the FNO branch is frozen and onlyΔ​Aθ\\Delta A\_\{\\theta\}is optimized:

minΔ​Aθ⁡ℒTwoStep\+ℒdrift\.\\min\_\{\\Delta A\_\{\\theta\}\}\\mathcal\{L\}\_\{\\mathrm\{TwoStep\}\}\+\\mathcal\{L\}\_\{\\mathrm\{drift\}\}\.\(232\)In the second stage, the measurement correction and the FNO branch are optimized jointly:

minΔ​Aθ,θFNO⁡ℒTwoStep\+ℒdrift\.\\min\_\{\\Delta A\_\{\\theta\},\\theta\_\{\\mathrm\{FNO\}\}\}\\mathcal\{L\}\_\{\\mathrm\{TwoStep\}\}\+\\mathcal\{L\}\_\{\\mathrm\{drift\}\}\.\(233\)The default schedule uses10001000measurement\-only epochs followed by20002000joint epochs\.

### D\.15Vanilla DeepONet representation

Vanilla DeepONet uses the complete standardized input history\. Its branch network is the same three\-dimensional FNO encoder, producing a latent vector

𝒃∈ℝp\.\\bm\{b\}\\in\\mathbb\{R\}^\{p\}\.\(234\)The trunk network receives a space\-time coordinate

𝝃=\(x,y,t\)∈\[0,1\]3\\bm\{\\xi\}=\(x,y,t\)\\in\[0,1\]^\{3\}\(235\)and returns

𝝉​\(𝝃\)∈ℝp\.\\bm\{\\tau\}\(\\bm\{\\xi\}\)\\in\\mathbb\{R\}^\{p\}\.\(236\)The predicted vorticity is

ω^​\(𝝃\)=𝒃𝖳​𝝉​\(𝝃\)\+b0\.\\widehat\{\\omega\}\(\\bm\{\\xi\}\)=\\bm\{b\}^\{\\mathsf\{T\}\}\\bm\{\\tau\}\(\\bm\{\\xi\}\)\+b\_\{0\}\.\(237\)Unlike the four Two\-Step models, Vanilla DeepONet predicts space\-time values directly and does not use the common output SVD basis\.

### D\.16Evaluation metrics

For test realizationii, the sample\-wise relativeL2L^\{2\}error over the entire predicted trajectory is

ei=‖𝒚^\(i\)−𝒚\(i\)‖2‖𝒚\(i\)‖2\.e\_\{i\}=\\frac\{\\left\\\|\\widehat\{\\bm\{y\}\}^\{\(i\)\}\-\\bm\{y\}^\{\(i\)\}\\right\\\|\_\{2\}\}\{\\left\\\|\\bm\{y\}^\{\(i\)\}\\right\\\|\_\{2\}\}\.\(238\)The mean sample\-wise relative error is

Emean=1Ntest​∑i=1Ntestei\.E\_\{\\mathrm\{mean\}\}=\\frac\{1\}\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}e\_\{i\}\.\(239\)The global relativeL2L^\{2\}error is

Eglobal=‖Y^−Y‖F‖Y‖F\.E\_\{\\mathrm\{global\}\}=\\frac\{\\left\\\|\\widehat\{Y\}\-Y\\right\\\|\_\{F\}\}\{\\left\\\|Y\\right\\\|\_\{F\}\}\.\(240\)The root\-mean\-square error is

RMSE=1Ntest​Nh​Tout​∑i=1Ntest∑m=1Nh∑n=1Tout\(ω^m,n\(i\)−ωm,n\(i\)\)2\.\\mathrm\{RMSE\}=\\sqrt\{\\frac\{1\}\{N\_\{\\mathrm\{test\}\}N\_\{h\}T\_\{\\mathrm\{out\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{m=1\}^\{N\_\{h\}\}\\sum\_\{n=1\}^\{T\_\{\\mathrm\{out\}\}\}\\left\(\\widehat\{\\omega\}\_\{m,n\}^\{\(i\)\}\-\\omega\_\{m,n\}^\{\(i\)\}\\right\)^\{2\}\}\.\(241\)The mean absolute error is

MAE=1Ntest​Nh​Tout​∑i=1Ntest∑m=1Nh∑n=1Tout\|ω^m,n\(i\)−ωm,n\(i\)\|\.\\mathrm\{MAE\}=\\frac\{1\}\{N\_\{\\mathrm\{test\}\}N\_\{h\}T\_\{\\mathrm\{out\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathrm\{test\}\}\}\\sum\_\{m=1\}^\{N\_\{h\}\}\\sum\_\{n=1\}^\{T\_\{\\mathrm\{out\}\}\}\\left\|\\widehat\{\\omega\}\_\{m,n\}^\{\(i\)\}\-\\omega\_\{m,n\}^\{\(i\)\}\\right\|\.\(242\)These are the metrics reported by the implementation\.

### D\.17Interpretation of the model comparison

The Full Two\-Step, Sensor Two\-Step, Fixed Topological Two\-Step, and Adaptive Topological Two\-Step models use a common output SVD basis and the same three\-dimensional FNO encoder family\. Their principal difference is the representation of the input history\. The full\-field model uses

4096×104096\\times 10\(243\)scalar input values\. The sensor and topological models use

measurements, corresponding to a compression ratio

4096×10128×10=32\.\\frac\{4096\\times 10\}\{128\\times 10\}=32\.\(245\)The fixed topological and sensor comparisons should be interpreted as equal representation\-budget comparisons\. They are not necessarily equal physical sensing\-cost comparisons, because a global functional measurement may require access to the full spatial vorticity field\. The adaptive model retains the same final measurement dimensionq=128q=128per input time but learns corrections within the span of the complete dictionary\. Its comparison with the fixed model therefore evaluates whether adapting the measurement functionals improves prediction while preserving the same compressed coordinate dimension\.

相似文章

非线性算子及其导数的通用逼近

arXiv cs.LG

本文证明了在无限维空间中非线性算子及其导数的首个通用逼近定理,将经典结果扩展到DeepONet和PCA-Net等算子学习架构。