The Frame Kernel Method for Multiscale Operator Learning
Summary
The paper presents the Frame Kernel Method, a novel multiscale operator learning approach for surrogate modeling of PDEs that uses kernel frame approximations and achieves higher accuracy than popular neural operators while enabling multiscale decomposition.
View Cached Full Text
Cached at: 08/27/26, 09:32 AM
# The Frame Kernel Method for Multiscale Operator Learning
Source: [https://arxiv.org/html/2608.25084](https://arxiv.org/html/2608.25084)
## The Frame Kernel Method for Multiscale Operator LearningThanks:Submitted to the editors on August 17, 2026\.
Branden Frieden††thanks:Kahlert School of Computing, University of Utah, Salt Lake City, UT 84112, USA\.M\. Keith Ballard33footnotemark:3Robert M\. Kirby22footnotemark:2Varun Shankar††thanks:Kahlert School of Computing, University of Utah, Salt Lake City, UT 84112, USA\. Corresponding author\.
###### Abstract
We present a natively multiscale operator learning method for the surrogate modeling of \(numerical solvers for\) multiscale partial differential equations \(PDEs\)\. The primary novelty of our method lies in a novel multiscale kernel frame function approximation technique\. Leveraging this new kernel frame technique, we cast the operator learning problem as one of learning frame coefficients of output functions as a function of frame coefficients of input functions\. The generalization step then automatically allows for a multiscale decomposition of the output functions\. Our method is applicable to both tensor\-product grids and point clouds\. We present interpolation proofs, error estimates, and numerical convergence rates for our frame approximation\. We the demonstrate the applicability of our method for the surrogate modeling of inherently multiscale PDEs\. The new multiscale frame kernel method is significantly more accurate than popular neural operators on challenging problems from the literature, while simultaneously admitting an a posteriori multiscale decomposition upon generalization\.
###### keywords
Frame Kernel Method, multiscale kernel\-frame approximation, scientific machine learning, multiscale operator learning
###### Funding\.
BF, RMK, and VS were supported by Air Force Office of Scientific Research \(AFOSR\) LRIR grant FA9550\-25\-1\-0042\. RW was supported by AFOSR LRIR grant 24RXCOR005\. VS was also supported by the U\.S\. National Science Foundation, Division of Mathematical Sciences, under award 2505986\.
## 1Introduction
Many applications require thousands or millions of solves for design, optimization, uncertainty quantification, inverse problems, control, and digital twins\. Surrogate models replace these repeated solves with rapidly evaluable approximations trained from simulation or experimental data\[[47](https://arxiv.org/html/2608.25084#bib.bib14),[28](https://arxiv.org/html/2608.25084#bib.bib15),[5](https://arxiv.org/html/2608.25084#bib.bib16),[45](https://arxiv.org/html/2608.25084#bib.bib17),[9](https://arxiv.org/html/2608.25084#bib.bib18),[29](https://arxiv.org/html/2608.25084#bib.bib19)\]\. Operator learning goes further by approximating maps between function spaces\. Neural operators parametrize the map using neural networks and come in multiple varieties\. Examples include deep operator networks \(DeepONets\)\[[36](https://arxiv.org/html/2608.25084#bib.bib20)\]; Fourier neural operators \(FNOs\), which parameterize nonlocal layers in Fourier space\[[33](https://arxiv.org/html/2608.25084#bib.bib21),[30](https://arxiv.org/html/2608.25084#bib.bib22)\]; kernel neural operators \(KNOs\), which use trainable closed\-form kernels and quadrature\[[35](https://arxiv.org/html/2608.25084#bib.bib23)\]; and Transolver and Transolver\+\+, which use physics\-aware attention on general, including million\-point, geometries\[[60](https://arxiv.org/html/2608.25084#bib.bib24),[38](https://arxiv.org/html/2608.25084#bib.bib25)\]\. Basis\-to\-Basis \(B2B\) operator learning makes the representation problem explicit by learning input and output bases together with a coefficient map\[[26](https://arxiv.org/html/2608.25084#bib.bib26)\]\.
Kernel methods offer a compelling alternative to neural operators\. Batlle et al\. combine observation and recovery maps with finite\-dimensional kernel regression and show that standard kernels can compete with widely used neural architectures\[[4](https://arxiv.org/html/2608.25084#bib.bib27)\]\. Turnage et al\. develop a complementary weighted least\-squares theory with optimal sampling measures, stability guarantees, and explicit finite\-dimensional operator spaces\[[55](https://arxiv.org/html/2608.25084#bib.bib28)\]\. These approaches build on scalar\-, vector\-, and operator\-valued kernel theory\[[40](https://arxiv.org/html/2608.25084#bib.bib29),[27](https://arxiv.org/html/2608.25084#bib.bib30)\], while retaining a crucial modeling choice: the user selects the input and output approximation spaces \(much like in B2B or DeepONets\)\. In each case, the representation directly controls the difficulty of the learned map\.
This choice becomes central for multiscale PDEs\. Heterogeneous media, singular perturbations, turbulence, porous flow, composite materials, and high\-frequency waves generate interacting spatial scales that may differ by orders of magnitude\[[25](https://arxiv.org/html/2608.25084#bib.bib31),[16](https://arxiv.org/html/2608.25084#bib.bib32),[17](https://arxiv.org/html/2608.25084#bib.bib13),[18](https://arxiv.org/html/2608.25084#bib.bib33)\]\. Operator\-learning methods have only recently begun to encode comparable structure\. Multiwavelet and wavelet neural operators use repeated multiresolution transforms\[[24](https://arxiv.org/html/2608.25084#bib.bib34),[54](https://arxiv.org/html/2608.25084#bib.bib35)\]; hierarchical attention neural operators address spectral bias against fine scales\[[34](https://arxiv.org/html/2608.25084#bib.bib36)\]; M2NO combines multiwavelets with algebraic multigrid\[[32](https://arxiv.org/html/2608.25084#bib.bib37)\]; locally subspace\-informed neural operators learn GMsFEM spaces\[[46](https://arxiv.org/html/2608.25084#bib.bib38)\]; and MscaleFNO uses parallel rescaled Fourier branches for highly oscillatory operator maps\[[61](https://arxiv.org/html/2608.25084#bib.bib39)\]\. These methods all recognize that multiscale surrogates require representations that expose the relevant scales\.
Kernel operator learning lets us impose such a representation directly without architecture tuning, but the appropriate multiscale space is not obvious\. Fourier bases provide global spectral efficiency and fast transforms on regular grids, yet encode scale through frequency and fit periodic or rectangular domains most naturally\[[53](https://arxiv.org/html/2608.25084#bib.bib40),[11](https://arxiv.org/html/2608.25084#bib.bib41),[33](https://arxiv.org/html/2608.25084#bib.bib21)\]\. Wavelets and multiwavelets localize in space and frequency and support sparse multiresolution representations, but require refinement structures, filter banks, or mesh\-dependent machinery\[[14](https://arxiv.org/html/2608.25084#bib.bib42),[39](https://arxiv.org/html/2608.25084#bib.bib43),[13](https://arxiv.org/html/2608.25084#bib.bib44)\]\. Hierarchical finite elements, splines, and multigrid bases handle complex PDE discretizations but usually inherit mesh\-connectivity and conformity requirements\[[7](https://arxiv.org/html/2608.25084#bib.bib12),[8](https://arxiv.org/html/2608.25084#bib.bib45),[18](https://arxiv.org/html/2608.25084#bib.bib33)\]\. Radial basis function \(RBF\) and kernel spaces work directly on scattered sites and general geometries, but a single length scale does not provide an explicit multiscale representation\[[58](https://arxiv.org/html/2608.25084#bib.bib46),[19](https://arxiv.org/html/2608.25084#bib.bib3)\]\.
We seek coordinates tied directly to multiple*spatial length scales*, rather than frequency separation alone\. For implementation simplicity, we also require one local construction that applies without modification to tensor\-product grids, vertices of general simulation meshes, and unstructured point clouds; interpolates the sampled data exactly; and exposes each scale after prediction\.
We introduce a multiscale kernel frame, built from compactly supported radial basis functions centered on nested subsets of the observation sites, that satisfies these requirements\. Each level uses a different center density and support radius, and all levels form one finite redundant dictionary\. While earlier work developed multiscale reproducing kernels, tight\-frame expansions, and multilevel radial basis function algorithms with changing scales\[[43](https://arxiv.org/html/2608.25084#bib.bib48),[44](https://arxiv.org/html/2608.25084#bib.bib49),[59](https://arxiv.org/html/2608.25084#bib.bib50),[23](https://arxiv.org/html/2608.25084#bib.bib51),[31](https://arxiv.org/html/2608.25084#bib.bib52),[3](https://arxiv.org/html/2608.25084#bib.bib47)\], our construction differs algebraically and in purpose: it requires neither a refinement equation, an orthogonal or tight basis, a dyadic lattice, nor sequential residual correction\. Instead, we determine all scale coefficients simultaneously\. A kernel translate at every observation site in the finest block guarantees exact interpolation, while coarser blocks add redundancy and produce a scale\-indexed representation\. The same construction works on regular grids and scattered points without mesh connectivity\.
We establish the frame property and exact interpolation of its canonical minimum\-norm reconstruction, then derive Sobolev error estimates from scattered\-zeros inequalities\. The resulting baseline native\-space rates do not explain the substantially faster convergence observed for smooth targets\. An exact induced\-kernel interpretation identifies the frame kernel as a finite multiscale autocorrelation and motivates a doubled\-native\-space conjecture consistent with the observed orders\.
We next combine the frame with the kernel operator\-learning framework of Batlle et al\.\[[4](https://arxiv.org/html/2608.25084#bib.bib27)\], which we call the*vanilla kernel method*\(VKM\)\. Our*frame kernel method*\(FKM\) represents every input and output in a multiscale frame, normalizes input coefficients level by level, and organizes the predicted output coefficients by frame level\. FKM thus maps between explicit multiscale coordinates rather than raw nodal values\. It uses the same training data, yet our benchmarks show reductions of11–44orders of magnitude in prediction error relative to VKM and competing operator\-learning methods, including neural operators\. Because the predicted coefficients retain their frame\-level organization, FKM also provides an a posteriori scale\-indexed decomposition of every prediction\.
The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2608.25084#S2)defines the multiscale kernel frame, its nested center hierarchy, the required numerical linear algebra, and the FKM\. Section[3](https://arxiv.org/html/2608.25084#S3)proves exact interpolation, derives the baseline Sobolev estimate, and states the doubled\-order conjecture\. Section[4](https://arxiv.org/html/2608.25084#S4)presents the frame\-approximation \(Section[4\.1](https://arxiv.org/html/2608.25084#S4.SS1)\) and operator\-learning results \(Section[4\.2](https://arxiv.org/html/2608.25084#S4.SS2)\), including scale\-resolved reconstructions \(Section[4\.2\.3](https://arxiv.org/html/2608.25084#S4.SS2.SSS3)\)\. Section[5](https://arxiv.org/html/2608.25084#S5)summarizes the results and discusses future directions\.
## 2Methods
### 2\.1Multiscale kernel\-frame approximation
LetΩ⊂ℝd\\Omega\\subset\\mathbb\{R\}^\{d\}, and letX=\{xi\}i=1M⊂ΩX=\\\{x\_\{i\}\\\}\_\{i=1\}^\{M\}\\subset\\Omegadenote a set of sampling sites with observationsui≈u\(xi\)u\_\{i\}\\approx u\(x\_\{i\}\), collected in𝐮=\[u1,…,uM\]⊤∈ℝM\\mathbf\{u\}=\[u\_\{1\},\\ldots,u\_\{M\}\]^\{\\top\}\\in\\mathbb\{R\}^\{M\}\. We construct a nested hierarchy of primary center sets,
C0=X⊃C1⊃⋯⊃CJ,Cj⊂X\.C\_\{0\}=X\\supset C\_\{1\}\\supset\\cdots\\supset C\_\{J\},\\qquad C\_\{j\}\\subset X\.\(1\)Level00is the finest level, while increasingjjcorresponds to progressively coarser center sets\. When additional boundary coverage is needed, we augmentCjC\_\{j\}with a setGjG\_\{j\}of auxiliary centers and write
Ξj=Cj∪Gj\.\\Xi\_\{j\}=C\_\{j\}\\cup G\_\{j\}\.The primary setsCjC\_\{j\}remain nested, whereas the auxiliary centers need not belong toXXor satisfy a nesting relation\. If no boundary augmentation is used, thenGj=∅G\_\{j\}=\\varnothingandΞj=Cj\\Xi\_\{j\}=C\_\{j\}\.
At leveljj, letρj\>0\\rho\_\{j\}\>0be a support radius and define
ψj,k\(x\)=ϕ\(‖x−ξj,k‖2ρj\),ξj,k∈Ξj,k=1,…,Nj,\\psi\_\{j,k\}\(x\)=\\phi\\\!\\left\(\\frac\{\\\|x\-\\xi\_\{j,k\}\\\|\_\{2\}\}\{\\rho\_\{j\}\}\\right\),\\qquad\\xi\_\{j,k\}\\in\\Xi\_\{j\},\\qquad k=1,\\ldots,N\_\{j\},\(2\)whereNj=\|Ξj\|N\_\{j\}=\|\\Xi\_\{j\}\|\. We chooseϕ\\phifrom the Wendland family of compactly supported positive\-definite radial basis functions\[[56](https://arxiv.org/html/2608.25084#bib.bib1),[58](https://arxiv.org/html/2608.25084#bib.bib46),[19](https://arxiv.org/html/2608.25084#bib.bib3)\]\. The multiscale approximant is
u^\(x\)=∑j=0J∑k=1Njcj,kψj,k\(x\)\.\\widehat\{u\}\(x\)=\\sum\_\{j=0\}^\{J\}\\sum\_\{k=1\}^\{N\_\{j\}\}c\_\{j,k\}\\psi\_\{j,k\}\(x\)\.\(3\)Outer level weights could be introduced by replacingψj,k\\psi\_\{j,k\}withwjψj,kw\_\{j\}\\psi\_\{j,k\}\. We omit them here because the principal scale dependence is already encoded throughρj\\rho\_\{j\}\.
BecauseC0=XC\_\{0\}=Xand the coarser levels contribute additional basis functions, the total number of functions,
N=∑j=0JNj,N=\\sum\_\{j=0\}^\{J\}N\_\{j\},typically exceeds the numberMMof observations\. The representation is therefore redundant: a sampled function may be reproduced by more than one coefficient vector\. This redundancy is intentional\. The finest level provides sufficient resolution at the sampling sites, while the additional levels allow the representation to distribute information across multiple spatial scales\.
The natural mathematical language for such a redundant spanning family is that of a frame\. A collection\{an\}n=1N⊂ℝM\\\{a\_\{n\}\\\}\_\{n=1\}^\{N\}\\subset\\mathbb\{R\}^\{M\}is a frame forℝM\\mathbb\{R\}^\{M\}if there exist constants0<a≤b<∞0<a\\leq b<\\inftysuch that
a‖v‖22≤∑n=1N\|⟨v,an⟩\|2≤b‖v‖22,v∈ℝM\.a\\\|v\\\|\_\{2\}^\{2\}\\leq\\sum\_\{n=1\}^\{N\}\|\\langle v,a\_\{n\}\\rangle\|^\{2\}\\leq b\\\|v\\\|\_\{2\}^\{2\},\\qquad v\\in\\mathbb\{R\}^\{M\}\.\(4\)In finite dimensions, the lower frame bound is equivalent to the vectors spanningℝM\\mathbb\{R\}^\{M\}, whileaaandbbquantify the stability of that spanning property\[[15](https://arxiv.org/html/2608.25084#bib.bib4),[12](https://arxiv.org/html/2608.25084#bib.bib5)\]\. In the present construction, the frame vectors are the sampled basis functions
aj,k=\[ψj,k\(x1\),…,ψj,k\(xM\)\]⊤,a\_\{j,k\}=\[\\psi\_\{j,k\}\(x\_\{1\}\),\\ldots,\\psi\_\{j,k\}\(x\_\{M\}\)\]^\{\\top\},namely, the columns of the full evaluation matrix defined below\. Theorem[1](https://arxiv.org/html/2608.25084#Thmtheorem1)shows that their union forms a frame for the discrete data space and that the resulting minimum\-norm approximation interpolates the observations\.
##### Scale selection
We choose
ρj=η2js0,\\rho\_\{j\}=\\eta\\,2^\{j\}s\_\{0\},\(5\)whereη\>0\\eta\>0controls support overlap ands0s\_\{0\}is a representative local spacing on the finest level\. This quantity is used only to set the support radii and is distinct from the fill distance used in the convergence analysis\. On a tensor\-product grid, we takes0s\_\{0\}to be the minimum grid spacing\. For a general point cloud, we use the median nearest\-neighbor distance,
s0=medianx∈Xminx′∈X∖\{x\}‖x−x′‖2,s\_\{0\}=\\operatorname\{median\}\_\{x\\in X\}\\min\_\{x^\{\\prime\}\\in X\\setminus\\\{x\\\}\}\\\|x\-x^\{\\prime\}\\\|\_\{2\},\(6\)which provides a robust estimate of the local sampling scale\[[58](https://arxiv.org/html/2608.25084#bib.bib46),[10](https://arxiv.org/html/2608.25084#bib.bib2)\]\.
##### Discrete evaluation operators
For each level, define the sparse evaluation matrix
Bj∈ℝM×Nj,\(Bj\)i,k=ϕ\(‖xi−ξj,k‖2ρj\)\.B\_\{j\}\\in\\mathbb\{R\}^\{M\\times N\_\{j\}\},\\qquad\(B\_\{j\}\)\_\{i,k\}=\\phi\\\!\\left\(\\frac\{\\\|x\_\{i\}\-\\xi\_\{j,k\}\\\|\_\{2\}\}\{\\rho\_\{j\}\}\\right\)\.\(7\)Compact support implies\(Bj\)i,k=0\(B\_\{j\}\)\_\{i,k\}=0whenever‖xi−ξj,k‖2\>ρj\\\|x\_\{i\}\-\\xi\_\{j,k\}\\\|\_\{2\}\>\\rho\_\{j\}\. The full evaluation matrix is the block concatenation
A=\[B0B1⋯BJ\]∈ℝM×N\.A=\[\\,B\_\{0\}\\;B\_\{1\}\\;\\cdots\\;B\_\{J\}\\,\]\\in\\mathbb\{R\}^\{M\\times N\}\.\(8\)We assemble each block using radius queries implemented with a k\-d\-tree range\-search structure\[[6](https://arxiv.org/html/2608.25084#bib.bib7)\]\.
For query sitesXq=\{xq,i\}i=1QX\_\{q\}=\\\{x\_\{q,i\}\\\}\_\{i=1\}^\{Q\}, the same construction gives
Bqj∈ℝQ×Nj,\(Bqj\)i,k=ψj,k\(xq,i\),B\_\{qj\}\\in\\mathbb\{R\}^\{Q\\times N\_\{j\}\},\\qquad\(B\_\{qj\}\)\_\{i,k\}=\\psi\_\{j,k\}\(x\_\{q,i\}\),and
Aq=\[Bq0Bq1⋯BqJ\]\.A\_\{q\}=\[\\,B\_\{q0\}\\;B\_\{q1\}\\;\\cdots\\;B\_\{qJ\}\\,\]\.\(9\)For a coefficient vector𝐜\\mathbf\{c\}, evaluation at the query sites is thenu^\(Xq\)=Aq𝐜\\widehat\{u\}\(X\_\{q\}\)=A\_\{q\}\\mathbf\{c\}\.
#### 2\.1\.1Nested center selection
The primary center sets satisfyCj\+1⊂Cj⊂XC\_\{j\+1\}\\subset C\_\{j\}\\subset X\. Auxiliary centers may subsequently be added near the boundary to improve geometric coverage\.
##### Dyadic thinning on tensor\-product grids
Assume thatXXis add\-dimensional tensor product grid with sizesn1,…,ndn\_\{1\},\\ldots,n\_\{d\}and a fixed flattening order consistent with Cartesian product indexing\. At leveljj, letsj=2js\_\{j\}=2^\{j\}and define
ℐj=\{\(i1,…,id\):iℓ∈\{1,1\+sj,1\+2sj,…\},ℓ=1,…,d\}\.\\mathcal\{I\}\_\{j\}=\\left\\\{\(i\_\{1\},\\ldots,i\_\{d\}\):i\_\{\\ell\}\\in\\\{1,1\+s\_\{j\},1\+2s\_\{j\},\\ldots\\\},\\quad\\ell=1,\\ldots,d\\right\\\}\.\(10\)The corresponding primary center set is
Cj=\{x\(i\):i∈ℐj\}\.C\_\{j\}=\\\{x\(i\):i\\in\\mathcal\{I\}\_\{j\}\\\}\.Because every index selected with stride2j\+12^\{j\+1\}is also selected with stride2j2^\{j\}, we haveCj\+1⊂CjC\_\{j\+1\}\\subset C\_\{j\}\. The center spacing therefore grows geometrically withjj, consistently with the support\-radius schedule \([5](https://arxiv.org/html/2608.25084#S2.E5)\)\. This construction is the grid analogue of multiresolution subsampling used in multilevel kernel approximation\[[58](https://arxiv.org/html/2608.25084#bib.bib46),[19](https://arxiv.org/html/2608.25084#bib.bib3)\]\.
##### Farthest\-point thinning on point clouds
For a general point cloud, we construct the nested hierarchy using farthest\-first traversal, also known as Gonzalez traversal\[[22](https://arxiv.org/html/2608.25084#bib.bib6)\]\. Beginning from an arbitrary pointp1∈Xp\_\{1\}\\in X, define
pm\+1∈argmaxx∈Xmin1≤r≤m‖x−pr‖2\.p\_\{m\+1\}\\in\\operatorname\*\{arg\\,max\}\_\{x\\in X\}\\min\_\{1\\leq r\\leq m\}\\\|x\-p\_\{r\}\\\|\_\{2\}\.\(11\)This produces a single orderingpm=xπ\(m\)p\_\{m\}=x\_\{\\pi\(m\)\}of the points inXX\. We compute the ordering once and define each primary center set as a prefix,
Cj=\{p1,…,pmj\},M=m0\>m1\>⋯\>mJ,C\_\{j\}=\\\{p\_\{1\},\\ldots,p\_\{m\_\{j\}\}\\\},\\qquad M=m\_\{0\}\>m\_\{1\}\>\\cdots\>m\_\{J\},\(12\)so thatCj\+1⊂CjC\_\{j\+1\}\\subset C\_\{j\}automatically\. The target cardinalities may be chosen to mimic geometric coarsening, for examplemj≈M/2jdm\_\{j\}\\approx M/2^\{jd\}, or to produce a prescribed geometric increase in separation distance\.
Farthest\-first traversal provides good coverage of the observed point cloud but does not guarantee comparable coverage outside its convex hull\. When evaluation is performed on a bounding box, poorly covered boundary layers can dominate the error\. We therefore augment each primary setCjC\_\{j\}, when necessary, with auxiliary samplesGjG\_\{j\}placed on the boundary of the evaluation region\[[20](https://arxiv.org/html/2608.25084#bib.bib61)\]\. These auxiliary centers are included in the level dictionaryΞj=Cj∪Gj\\Xi\_\{j\}=C\_\{j\}\\cup G\_\{j\}but are not part of the nested primary hierarchy\.
#### 2\.1\.2Linear algebra
Collect the coefficients levelwise as
𝐜=\[𝐜0⊤,…,𝐜J⊤\]⊤∈ℝN,𝐜j∈ℝNj\.\\mathbf\{c\}=\[\\mathbf\{c\}\_\{0\}^\{\\top\},\\ldots,\\mathbf\{c\}\_\{J\}^\{\\top\}\]^\{\\top\}\\in\\mathbb\{R\}^\{N\},\\qquad\\mathbf\{c\}\_\{j\}\\in\\mathbb\{R\}^\{N\_\{j\}\}\.Thenu^\(X\)=A𝐜\\widehat\{u\}\(X\)=A\\mathbf\{c\}\. BecauseAAis wide, the interpolation equations generally admit multiple solutions\. We select the canonical minimum\-norm coefficient vector
𝐜\(𝐮\)=argmin𝐳∈ℝN‖𝐳‖2subject toA𝐳=𝐮\.\\mathbf\{c\}\(\\mathbf\{u\}\)=\\operatorname\*\{arg\\,min\}\_\{\\mathbf\{z\}\\in\\mathbb\{R\}^\{N\}\}\\\|\\mathbf\{z\}\\\|\_\{2\}\\quad\\text\{subject to\}\\quad A\\mathbf\{z\}=\\mathbf\{u\}\.\(13\)Theorem[1](https://arxiv.org/html/2608.25084#Thmtheorem1)guarantees thatAAhas full row rank, so \([13](https://arxiv.org/html/2608.25084#S2.E13)\) is feasible for every𝐮∈ℝM\\mathbf\{u\}\\in\\mathbb\{R\}^\{M\}and has a unique solution\. We compute this solution using a thin QR factorization ofA⊤A^\{\\top\},
A⊤=QR,Q∈ℝN×M,R∈ℝM×M,A^\{\\top\}=QR,\\qquad Q\\in\\mathbb\{R\}^\{N\\times M\},\\qquad R\\in\\mathbb\{R\}^\{M\\times M\},\(14\)whereQ⊤Q=IMQ^\{\\top\}Q=I\_\{M\}andRRis nonsingular and upper triangular\. Solving
R⊤𝐲=𝐮R^\{\\top\}\\mathbf\{y\}=\\mathbf\{u\}and setting
𝐜\(𝐮\)=Q𝐲\\mathbf\{c\}\(\\mathbf\{u\}\)=Q\\mathbf\{y\}\(15\)yields the minimum\-norm interpolating coefficients\. Equivalently,
𝐜\(𝐮\)=A⊤\(AA⊤\)−1𝐮\.\\mathbf\{c\}\(\\mathbf\{u\}\)=A^\{\\top\}\(AA^\{\\top\}\)^\{\-1\}\\mathbf\{u\}\.Thus, a single factorization ofA⊤A^\{\\top\}can be reused for every function sampled on the same sites\. In our experiments, this construction required no additional regularization\. Reconstruction at the observation and query sites is additive across levels:
u^\(X\)=∑j=0JBj𝐜j,u^\(Xq\)=∑j=0JBqj𝐜j\.\\widehat\{u\}\(X\)=\\sum\_\{j=0\}^\{J\}B\_\{j\}\\mathbf\{c\}\_\{j\},\\qquad\\widehat\{u\}\(X\_\{q\}\)=\\sum\_\{j=0\}^\{J\}B\_\{qj\}\\mathbf\{c\}\_\{j\}\.\(16\)
### 2\.2The frame kernel method for multiscale operator learning
We now use the multiscale frame representation within a kernel method for operator learning\[[4](https://arxiv.org/html/2608.25084#bib.bib27)\]\. Let
𝒢:𝒰→𝒱\\mathcal\{G\}:\\mathcal\{U\}\\rightarrow\\mathcal\{V\}be an operator, and suppose that the training data consist of pairs
\{\(uℓ,vℓ\)\}ℓ=1NT,vℓ=𝒢\(uℓ\)\.\\\{\(u\_\{\\ell\},v\_\{\\ell\}\)\\\}\_\{\\ell=1\}^\{N\_\{T\}\},\\qquad v\_\{\\ell\}=\\mathcal\{G\}\(u\_\{\\ell\}\)\.The input and output functions may be defined on different domains and sampled at different sets of sites\. We therefore construct separate input and output frames with evaluation matrices
Au∈ℝMu×Nu,Av∈ℝMv×Nv\.A\_\{u\}\\in\\mathbb\{R\}^\{M\_\{u\}\\times N\_\{u\}\},\\qquad A\_\{v\}\\in\\mathbb\{R\}^\{M\_\{v\}\\times N\_\{v\}\}\.The two frames may differ in their numbers of levels, centers, and coefficients\.
The vanilla kernel method \(VKM\) learns a map directly from sampled input functions to sampled output functions\. The proposed*frame kernel method*\(FKM\) instead uses input\-frame coefficients as kernel features and organizes the output coefficients according to the levels of the multiscale frame\. For each training pair, we compute the canonical coefficient vectors
𝐜uℓ=Au†uℓ\(Xu\),𝐜vℓ=Av†vℓ\(Xv\),\\mathbf\{c\}\_\{u\_\{\\ell\}\}=A\_\{u\}^\{\\dagger\}u\_\{\\ell\}\(X\_\{u\}\),\\qquad\\mathbf\{c\}\_\{v\_\{\\ell\}\}=A\_\{v\}^\{\\dagger\}v\_\{\\ell\}\(X\_\{v\}\),\(17\)using the QR procedure in \([14](https://arxiv.org/html/2608.25084#S2.E14)\)–\([15](https://arxiv.org/html/2608.25084#S2.E15)\)\. Partition these vectors according to their frame levels:
𝐜uℓ=\[𝐜uℓ,0⊤,…,𝐜uℓ,Ju⊤\]⊤,𝐜uℓ,j∈ℝNu,j,\\mathbf\{c\}\_\{u\_\{\\ell\}\}=\[\\mathbf\{c\}\_\{u\_\{\\ell\},0\}^\{\\top\},\\ldots,\\mathbf\{c\}\_\{u\_\{\\ell\},J\_\{u\}\}^\{\\top\}\]^\{\\top\},\\qquad\\mathbf\{c\}\_\{u\_\{\\ell\},j\}\\in\\mathbb\{R\}^\{N\_\{u,j\}\},and
𝐜vℓ=\[𝐜vℓ,0⊤,…,𝐜vℓ,Jv⊤\]⊤,𝐜vℓ,j∈ℝNv,j\.\\mathbf\{c\}\_\{v\_\{\\ell\}\}=\[\\mathbf\{c\}\_\{v\_\{\\ell\},0\}^\{\\top\},\\ldots,\\mathbf\{c\}\_\{v\_\{\\ell\},J\_\{v\}\}^\{\\top\}\]^\{\\top\},\\qquad\\mathbf\{c\}\_\{v\_\{\\ell\},j\}\\in\\mathbb\{R\}^\{N\_\{v,j\}\}\.
##### Scale\-dependent normalization of the input coefficients
The coefficient magnitudes can differ substantially across input\-frame levels because the levels use different center densities and support radii\. If kernel distances are computed directly from the raw coefficient vectors, levels with larger coefficients can dominate the induced geometry\. We therefore normalize each input block by a single scalar estimated from the training set\. Define
Dj=\[𝐜u1,j,…,𝐜uNT,j\]∈ℝNu,j×NT\.D\_\{j\}=\[\\mathbf\{c\}\_\{u\_\{1\},j\},\\ldots,\\mathbf\{c\}\_\{u\_\{N\_\{T\}\},j\}\]\\in\\mathbb\{R\}^\{N\_\{u,j\}\\times N\_\{T\}\}\.\(18\)For leveljj, set
βj=\(1Nu,jNT∑ℓ=1NT∥𝐜uℓ,j∥22\+ε\)−1/2,\\beta\_\{j\}=\\left\(\\frac\{1\}\{N\_\{u,j\}N\_\{T\}\}\\sum\_\{\\ell=1\}^\{N\_\{T\}\}\\\|\\mathbf\{c\}\_\{u\_\{\\ell\},j\}\\\|\_\{2\}^\{2\}\+\\varepsilon\\right\)^\{\-1/2\},\(19\)whereε\>0\\varepsilon\>0prevents division by zero\. The normalized input coefficients are
𝐜~uℓ,j=βj𝐜uℓ,j,𝐜~uℓ=\[𝐜~uℓ,0⊤,…,𝐜~uℓ,Ju⊤\]⊤\.\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\\ell\},j\}=\\beta\_\{j\}\\mathbf\{c\}\_\{u\_\{\\ell\},j\},\\qquad\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\\ell\}\}=\[\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\\ell\},0\}^\{\\top\},\\ldots,\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\\ell\},J\_\{u\}\}^\{\\top\}\]^\{\\top\}\.\(20\)If
sj=1Nu,jNT∑ℓ=1NT‖𝐜uℓ,j‖22,s\_\{j\}=\\frac\{1\}\{N\_\{u,j\}N\_\{T\}\}\\sum\_\{\\ell=1\}^\{N\_\{T\}\}\\\|\\mathbf\{c\}\_\{u\_\{\\ell\},j\}\\\|\_\{2\}^\{2\},then
1Nu,jNT∑ℓ=1NT‖𝐜~uℓ,j‖22=sjsj\+ε\.\\frac\{1\}\{N\_\{u,j\}N\_\{T\}\}\\sum\_\{\\ell=1\}^\{N\_\{T\}\}\\\|\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\\ell\},j\}\\\|\_\{2\}^\{2\}=\\frac\{s\_\{j\}\}\{s\_\{j\}\+\\varepsilon\}\.\(21\)Hence, wheneversj≫εs\_\{j\}\\gg\\varepsilon, the transformation gives each level unit empirical mean\-squared coefficient magnitude per degree of freedom\. The factorsβj\\beta\_\{j\}are estimated from the training inputs only and then applied unchanged to validation, test, and prediction inputs\. No centering or covariance transformation is performed\.
##### Levelwise multi\-output kernel regression
Let
κ:ℝNu×ℝNu→ℝ\\kappa:\\mathbb\{R\}^\{N\_\{u\}\}\\times\\mathbb\{R\}^\{N\_\{u\}\}\\rightarrow\\mathbb\{R\}be a positive\-definite kernel, whereNu=∑j=0JuNu,jN\_\{u\}=\\sum\_\{j=0\}^\{J\_\{u\}\}N\_\{u,j\}, and define the training Gram matrix
Kℓm=κ\(𝐜~uℓ,𝐜~um\),K∈ℝNT×NT\.K\_\{\\ell m\}=\\kappa\(\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\\ell\}\},\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{m\}\}\),\\qquad K\\in\\mathbb\{R\}^\{N\_\{T\}\\times N\_\{T\}\}\.\(22\)For output leveljj, collect the corresponding training coefficients rowwise as
Yj=\[𝐜v1,j⊤𝐜vNT,j⊤\]∈ℝNT×Nv,j\.Y\_\{j\}=\\begin\{bmatrix\}\\mathbf\{c\}\_\{v\_\{1\},j\}^\{\\top\}\\\\ \\vdots\\\\ \\mathbf\{c\}\_\{v\_\{N\_\{T\}\},j\}^\{\\top\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{N\_\{T\}\\times N\_\{v,j\}\}\.\(23\)For each output level, we solve
Wj=\(K\+λINT\)−1Yj,j=0,…,Jv,W\_\{j\}=\(K\+\\lambda I\_\{N\_\{T\}\}\)^\{\-1\}Y\_\{j\},\\qquad j=0,\\ldots,J\_\{v\},\(24\)whereλ\>0\\lambda\>0is the regularization parameter\. Because the same Gram matrix and regularization parameter are used at every level, these solves are algebraically equivalent to a single multi\-output kernel ridge regression with the concatenated response matrix
Y=\[Y0⋯YJv\]\.Y=\[\\,Y\_\{0\}\\;\\cdots\\;Y\_\{J\_\{v\}\}\\,\]\.We retain the levelwise form because it exposes the scale\-indexed coefficient blocks used in reconstruction, while allowing the factorization ofK\+λINTK\+\\lambda I\_\{N\_\{T\}\}to be reused for every right\-hand side\.
For a new input functionu∗u\_\{\*\}, we first compute its frame coefficients𝐜u∗\\mathbf\{c\}\_\{u\_\{\*\}\}and apply the training\-set normalization factors,
𝐜~u∗,j=βj𝐜u∗,j,𝐜~u∗=\[𝐜~u∗,0⊤,…,𝐜~u∗,Ju⊤\]⊤\.\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\*\},j\}=\\beta\_\{j\}\\mathbf\{c\}\_\{u\_\{\*\},j\},\\qquad\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\*\}\}=\[\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\*\},0\}^\{\\top\},\\ldots,\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\*\},J\_\{u\}\}^\{\\top\}\]^\{\\top\}\.\(25\)Define
𝐤∗=\[κ\(𝐜~u∗,𝐜~u1\)κ\(𝐜~u∗,𝐜~uNT\)\]∈ℝNT\.\\mathbf\{k\}\_\{\*\}=\\begin\{bmatrix\}\\kappa\(\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\*\}\},\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{1\}\}\)\\\\ \\vdots\\\\ \\kappa\(\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{\*\}\},\\widetilde\{\\mathbf\{c\}\}\_\{u\_\{N\_\{T\}\}\}\)\\end\{bmatrix\}\\in\\mathbb\{R\}^\{N\_\{T\}\}\.\(26\)The predicted coefficient block at output leveljjis
𝐜^v∗,j=Wj⊤𝐤∗=Yj⊤\(K\+λINT\)−1𝐤∗\.\\widehat\{\\mathbf\{c\}\}\_\{v\_\{\*\},j\}=W\_\{j\}^\{\\top\}\\mathbf\{k\}\_\{\*\}=Y\_\{j\}^\{\\top\}\(K\+\\lambda I\_\{N\_\{T\}\}\)^\{\-1\}\\mathbf\{k\}\_\{\*\}\.\(27\)The predicted coefficients remain organized by frame level and are combined during reconstruction\. The distinction from VKM therefore lies in the multiscale representation of the inputs and outputs, rather than in a different algebraic form of the shared multi\-output kernel ridge\-regression solve\.
Finally, the predicted output function is synthesized from the levelwise coefficient blocks:
𝒢^\(u∗\)\(x\)=∑j=0Jv∑k=1Nv,j\(𝐜^v∗,j\)kψj,kv\(x\)\.\\widehat\{\\mathcal\{G\}\}\(u\_\{\*\}\)\(x\)=\\sum\_\{j=0\}^\{J\_\{v\}\}\\sum\_\{k=1\}^\{N\_\{v,j\}\}\(\\widehat\{\\mathbf\{c\}\}\_\{v\_\{\*\},j\}\)\_\{k\}\\,\\psi^\{v\}\_\{j,k\}\(x\)\.\(28\)At output query sitesXv,qX\_\{v,q\}, this becomes
𝒢^\(u∗\)\(Xv,q\)=∑j=0JvBqjv𝐜^v∗,j\.\\widehat\{\\mathcal\{G\}\}\(u\_\{\*\}\)\(X\_\{v,q\}\)=\\sum\_\{j=0\}^\{J\_\{v\}\}B^\{v\}\_\{qj\}\\widehat\{\\mathbf\{c\}\}\_\{v\_\{\*\},j\}\.\(29\)Relative to VKM, FKM changes the representation used by the regression stage: it forms kernel similarities from normalized multiscale input\-frame coefficients, organizes the predicted output coefficients by frame level, and synthesizes the resulting function from those levelwise blocks\.
## 3Exact interpolation and convergence rates
### 3\.1Frame property and exact interpolation
The sampled basis functions
aj,k=\[ψj,k\(x1\),…,ψj,k\(xM\)\]⊤a\_\{j,k\}=\[\\psi\_\{j,k\}\(x\_\{1\}\),\\ldots,\\psi\_\{j,k\}\(x\_\{M\}\)\]^\{\\top\}are the columns of the multiscale evaluation matrixAA\. The following result shows that they form a finite\-dimensional frame forℝM\\mathbb\{R\}^\{M\}and that the redundant approximation interpolates arbitrary data on the observation set\[[15](https://arxiv.org/html/2608.25084#bib.bib4),[12](https://arxiv.org/html/2608.25084#bib.bib5)\]\. The proof uses only strict positive definiteness at the finest level\.
We distinguish the numberM=\|X\|M=\|X\|of sampling sites from the numberNj=\|Ξj\|N\_\{j\}=\|\\Xi\_\{j\}\|of frame functions at leveljj\. If auxiliary boundary centers are present, thenN0N\_\{0\}may exceedMM\. LetB0,X∈ℝM×MB\_\{0,X\}\\in\\mathbb\{R\}^\{M\\times M\}denote the submatrix ofB0B\_\{0\}associated with the primary centersC0=XC\_\{0\}=X; if level zero contains no auxiliary centers, thenB0,X=B0B\_\{0,X\}=B\_\{0\}\.
###### Theorem 1\(Frame property and exact interpolation\)\.
LetX=\{xi\}i=1M⊂ΩX=\\\{x\_\{i\}\\\}\_\{i=1\}^\{M\}\\subset\\Omegaconsist of distinct sites, letC0=X⊃C1⊃⋯⊃CJC\_\{0\}=X\\supset C\_\{1\}\\supset\\cdots\\supset C\_\{J\}, and suppose thatϕ\(∥⋅∥2/ρ0\)\\phi\(\\\|\\cdot\\\|\_\{2\}/\\rho\_\{0\}\)is strictly positive definite onℝd\\mathbb\{R\}^\{d\}\. Then the columns ofA=\[B0⋯BJ\]A=\[\\,B\_\{0\}\\;\\cdots\\;B\_\{J\}\\,\]form a finite\-dimensional frame forℝM\\mathbb\{R\}^\{M\}, with
σmin\(B0,X\)2‖v‖22≤‖A⊤v‖22≤‖A‖22‖v‖22,v∈ℝM\.\\sigma\_\{\\min\}\(B\_\{0,X\}\)^\{2\}\\\|v\\\|\_\{2\}^\{2\}\\leq\\\|A^\{\\top\}v\\\|\_\{2\}^\{2\}\\leq\\\|A\\\|\_\{2\}^\{2\}\\\|v\\\|\_\{2\}^\{2\},\\qquad v\\in\\mathbb\{R\}^\{M\}\.\(30\)Consequently, every𝐮∈ℝM\\mathbf\{u\}\\in\\mathbb\{R\}^\{M\}admits coefficientscj,kc\_\{j,k\}satisfyingu^\(xi\)=ui\\widehat\{u\}\(x\_\{i\}\)=u\_\{i\},i=1,…,Mi=1,\\ldots,M\. IfA⊤=QRA^\{\\top\}=QRis a thin QR factorization, then
𝐜⋆=QR−⊤𝐮\\mathbf\{c\}^\{\\star\}=QR^\{\-\\top\}\\mathbf\{u\}is the unique minimum\-ℓ2\\ell^\{2\}\-norm coefficient vector producing this interpolant\.
###### Proof\.
SinceC0=XC\_\{0\}=X, the matrixB0,XB\_\{0,X\}is, up to a column permutation, the kernel matrix
\(B0,X\)i,k=ϕ\(‖xi−xk‖2ρ0\)\.\(B\_\{0,X\}\)\_\{i,k\}=\\phi\\\!\\left\(\\frac\{\\\|x\_\{i\}\-x\_\{k\}\\\|\_\{2\}\}\{\\rho\_\{0\}\}\\right\)\.Strict positive definiteness and distinctness of the sites makeB0,XB\_\{0,X\}nonsingular\[[56](https://arxiv.org/html/2608.25084#bib.bib1),[58](https://arxiv.org/html/2608.25084#bib.bib46)\]\. Hence
‖A⊤v‖22=∑j=0J‖Bj⊤v‖22≥‖B0,X⊤v‖22≥σmin\(B0,X\)2‖v‖22,\\\|A^\{\\top\}v\\\|\_\{2\}^\{2\}=\\sum\_\{j=0\}^\{J\}\\\|B\_\{j\}^\{\\top\}v\\\|\_\{2\}^\{2\}\\geq\\\|B\_\{0,X\}^\{\\top\}v\\\|\_\{2\}^\{2\}\\geq\\sigma\_\{\\min\}\(B\_\{0,X\}\)^\{2\}\\\|v\\\|\_\{2\}^\{2\},and the upper bound follows from‖A⊤v‖2≤‖A‖2‖v‖2\\\|A^\{\\top\}v\\\|\_\{2\}\\leq\\\|A\\\|\_\{2\}\\\|v\\\|\_\{2\}\. ThusAAhas full row rank\. SinceA=R⊤Q⊤A=R^\{\\top\}Q^\{\\top\},
AQR−⊤=IM,AQR^\{\-\\top\}=I\_\{M\},so𝐜⋆=QR−⊤𝐮\\mathbf\{c\}^\{\\star\}=QR^\{\-\\top\}\\mathbf\{u\}interpolates the data\. Moreover,
QR−⊤=A⊤\(AA⊤\)−1=A†,QR^\{\-\\top\}=A^\{\\top\}\(AA^\{\\top\}\)^\{\-1\}=A^\{\\dagger\},which is the Moore–Penrose pseudoinverse and hence the minimum\-norm right inverse\[[21](https://arxiv.org/html/2608.25084#bib.bib54)\]\.
The finest primary block therefore guarantees exact interpolation, while the remaining columns add redundancy and allow the canonical QR solution to distribute the representation across scales\.
### 3\.2Convergence rates
To express the error directly in terms of the numberMMof sampling sites, define the fill distance
hM:=hXM,Ω=supx∈Ωminxi∈XM‖x−xi‖2\.h\_\{M\}:=h\_\{X\_\{M\},\\Omega\}=\\sup\_\{x\\in\\Omega\}\\min\_\{x\_\{i\}\\in X\_\{M\}\}\\\|x\-x\_\{i\}\\\|\_\{2\}\.\(31\)We assume that the sampling sets under consideration satisfy
hM≤ChM−1/d,h\_\{M\}\\leq C\_\{h\}M^\{\-1/d\},\(32\)withChC\_\{h\}independent ofMM\. No assumption is made on the separation radius or mesh ratio\. Thus, the analysis below depends only on coverage ofΩ\\OmegathroughhMh\_\{M\}, and not on a lower bound for pairwise distances between sampling sites\.
For theC2ℓC^\{2\\ell\}Wendland kernelϕd,ℓ\\phi\_\{d,\\ell\}, the native space is norm\-equivalent toHτℓ\(ℝd\)H^\{\\tau\_\{\\ell\}\}\(\\mathbb\{R\}^\{d\}\), where
τℓ:=d2\+ℓ\+12,\\tau\_\{\\ell\}:=\\frac\{d\}\{2\}\+\\ell\+\\frac\{1\}\{2\},and the casesℓ=1,2,3\\ell=1,2,3correspond to theC2C^\{2\},C4C^\{4\}, andC6C^\{6\}kernels\[[57](https://arxiv.org/html/2608.25084#bib.bib53),[58](https://arxiv.org/html/2608.25084#bib.bib46),[48](https://arxiv.org/html/2608.25084#bib.bib57)\]\. Translation and positive scaling preserve this Sobolev regularity, so every finite multiscale interpolant belongs toHτℓ\(Ω\)H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\.
LetEM:Hτℓ\(Ω\)→ℝME\_\{M\}:H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\\to\\mathbb\{R\}^\{M\}denote sampling onXMX\_\{M\}, letTM:ℝN\(M\)→Hτℓ\(Ω\)T\_\{M\}:\\mathbb\{R\}^\{N\(M\)\}\\to H^\{\\tau\_\{\\ell\}\}\(\\Omega\)denote synthesis by the multiscale frame functions, and write
ℐM=TMAM†EM\.\\mathcal\{I\}\_\{M\}=T\_\{M\}A\_\{M\}^\{\\dagger\}E\_\{M\}\.Becauseτℓ\>d/2\\tau\_\{\\ell\}\>d/2, the sampling mapEME\_\{M\}is bounded by Sobolev embedding\. For each fixed discretization,AM†A\_\{M\}^\{\\dagger\}is a bounded finite\-dimensional map andTMT\_\{M\}is bounded because it synthesizes a finite collection of functions inHτℓ\(Ω\)H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\. HenceℐM:Hτℓ\(Ω\)→Hτℓ\(Ω\)\\mathcal\{I\}\_\{M\}:H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\\to H^\{\\tau\_\{\\ell\}\}\(\\Omega\)is bounded for each fixedMM\[[1](https://arxiv.org/html/2608.25084#bib.bib55)\]\. Define
Λℓ,J\(M\):=sup0≠u∈Hτℓ\(Ω\)‖ℐMu‖Hτℓ\(Ω\)‖u‖Hτℓ\(Ω\)\.\\Lambda\_\{\\ell,J\}\(M\):=\\sup\_\{0\\neq u\\in H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\\frac\{\\\|\\mathcal\{I\}\_\{M\}u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\}\{\\\|u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\}\.\(33\)ThusΛℓ,J\(M\)<∞\\Lambda\_\{\\ell,J\}\(M\)<\\inftyfor every fixed discretization; uniform boundedness inMMis a separate stability question\.
###### Theorem 2\(Sobolev error estimate\)\.
LetΩ⊂ℝd\\Omega\\subset\\mathbb\{R\}^\{d\}be a bounded Lipschitz domain satisfying an interior cone condition, and letXM⊂ΩX\_\{M\}\\subset\\Omegabe sets ofMMdistinct sampling sites satisfying \([32](https://arxiv.org/html/2608.25084#S3.E32)\)\. Letϕ=ϕd,ℓ\\phi=\\phi\_\{d,\\ell\}be theC2ℓC^\{2\\ell\}Wendland kernel\. Then there existh∗\>0h\_\{\*\}\>0andC\>0C\>0, independent ofMManduu, such that, wheneverhM≤h∗h\_\{M\}\\leq h\_\{\*\}, everyu∈Hτℓ\(Ω\)u\\in H^\{\\tau\_\{\\ell\}\}\(\\Omega\)satisfies
∥u−ℐMu∥Hμ\(Ω\)≤CM−\(τℓ−μ\)/d\(1\+Λℓ,J\(M\)\)∥u∥Hτℓ\(Ω\),0≤μ≤τℓ\.\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{H^\{\\mu\}\(\\Omega\)\}\\leq CM^\{\-\(\\tau\_\{\\ell\}\-\\mu\)/d\}\\bigl\(1\+\\Lambda\_\{\\ell,J\}\(M\)\\bigr\)\\\|u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\},\\qquad 0\\leq\\mu\\leq\\tau\_\{\\ell\}\.\(34\)IfΛℓ,J\(M\)\\Lambda\_\{\\ell,J\}\(M\)remains bounded asM→∞M\\to\\infty, then
∥u−ℐMu∥Hμ\(Ω\)=𝒪\(M−\(τℓ−μ\)/d\)\.\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{H^\{\\mu\}\(\\Omega\)\}=\\mathcal\{O\}\\\!\\left\(M^\{\-\(\\tau\_\{\\ell\}\-\\mu\)/d\}\\right\)\.
###### Proof\.
Sete:=u−ℐMue:=u\-\\mathcal\{I\}\_\{M\}u\. Theorem[1](https://arxiv.org/html/2608.25084#Thmtheorem1)givese\|XM=0e\|\_\{X\_\{M\}\}=0\. ForhM≤h∗h\_\{M\}\\leq h\_\{\*\}, the scattered\-zeros inequality gives
‖e‖Hμ\(Ω\)≤ChMτℓ−μ‖e‖Hτℓ\(Ω\)\\\|e\\\|\_\{H^\{\\mu\}\(\\Omega\)\}\\leq Ch\_\{M\}^\{\\tau\_\{\\ell\}\-\\mu\}\\\|e\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}withCCindependent ofMMandee\[[42](https://arxiv.org/html/2608.25084#bib.bib58),[58](https://arxiv.org/html/2608.25084#bib.bib46)\]\. Using \([32](https://arxiv.org/html/2608.25084#S3.E32)\),
∥e∥Hμ\(Ω\)≤CM−\(τℓ−μ\)/d∥e∥Hτℓ\(Ω\)\.\\\|e\\\|\_\{H^\{\\mu\}\(\\Omega\)\}\\leq CM^\{\-\(\\tau\_\{\\ell\}\-\\mu\)/d\}\\\|e\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\.Finally, the definition ofΛℓ,J\(M\)\\Lambda\_\{\\ell,J\}\(M\)and the triangle inequality give
‖e‖Hτℓ\(Ω\)≤\(1\+Λℓ,J\(M\)\)‖u‖Hτℓ\(Ω\),\\\|e\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\\leq\\bigl\(1\+\\Lambda\_\{\\ell,J\}\(M\)\\bigr\)\\\|u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\},which proves \([34](https://arxiv.org/html/2608.25084#S3.E34)\)\.
##### Example
Ford=2d=2,τℓ=ℓ\+3/2\\tau\_\{\\ell\}=\\ell\+3/2\. Under bounded stability, the baselineL2L^\{2\}rates are
C2:𝒪\(\(M1/2\)−5/2\),C4:𝒪\(\(M1/2\)−7/2\),C6:𝒪\(\(M1/2\)−9/2\)\.C^\{2\}:\\ \\mathcal\{O\}\\\!\\left\(\(M^\{1/2\}\)^\{\-5/2\}\\right\),\\qquad C^\{4\}:\\ \\mathcal\{O\}\\\!\\left\(\(M^\{1/2\}\)^\{\-7/2\}\\right\),\\qquad C^\{6\}:\\ \\mathcal\{O\}\\\!\\left\(\(M^\{1/2\}\)^\{\-9/2\}\\right\)\.\(35\)The correspondingH1H^\{1\}rates are
C2:𝒪\(\(M1/2\)−3/2\),C4:𝒪\(\(M1/2\)−5/2\),C6:𝒪\(\(M1/2\)−7/2\)\.C^\{2\}:\\ \\mathcal\{O\}\\\!\\left\(\(M^\{1/2\}\)^\{\-3/2\}\\right\),\\qquad C^\{4\}:\\ \\mathcal\{O\}\\\!\\left\(\(M^\{1/2\}\)^\{\-5/2\}\\right\),\\qquad C^\{6\}:\\ \\mathcal\{O\}\\\!\\left\(\(M^\{1/2\}\)^\{\-7/2\}\\right\)\.The observed discrete relativeℓ2\\ell\_\{2\}orders of approximately5\.35\.3,7\.37\.3, and8\.758\.75, measured againstM1/2M^\{1/2\}, for theC2C^\{2\},C4C^\{4\}, andC6C^\{6\}kernels are numerically close to twice the baseline exponents\. These results suggest superconvergence beyond Theorem[2](https://arxiv.org/html/2608.25084#Thmtheorem2), while the theorem itself concerns continuous Sobolev norms\.
### 3\.3A route to superconvergence
The baseline estimate uses only exact interpolation and a scattered\-zeros inequality\. For smooth targets, the observed rates instead suggest the classical doubling phenomenon associated with source conditions in kernel interpolation\[[49](https://arxiv.org/html/2608.25084#bib.bib59),[51](https://arxiv.org/html/2608.25084#bib.bib60)\]\. The minimum\-coefficient\-norm property alone, however, does not place the QR interpolant in the standard native\-space orthogonal\-projection framework used by classical doubling arguments\.
The QR approximation nevertheless has an exact kernel interpretation\. LetΨM\(x\)∈ℝN\(M\)\\Psi\_\{M\}\(x\)\\in\\mathbb\{R\}^\{N\(M\)\}collect all multiscale basis functions and define
KM\(x,y\):=ΨM\(x\)⊤ΨM\(y\)=∑j=0J\(M\)∑k=1Njψj,k\(x\)ψj,k\(y\)\.K\_\{M\}\(x,y\):=\\Psi\_\{M\}\(x\)^\{\\top\}\\Psi\_\{M\}\(y\)=\\sum\_\{j=0\}^\{J\(M\)\}\\sum\_\{k=1\}^\{N\_\{j\}\}\\psi\_\{j,k\}\(x\)\\psi\_\{j,k\}\(y\)\.\(36\)This feature\-map construction defines a positive\-semidefinite kernel\[[2](https://arxiv.org/html/2608.25084#bib.bib56)\]\. Since
𝐜⋆=A⊤\(AA⊤\)−1𝐮M,KM\(XM,XM\)=AA⊤,\\mathbf\{c\}^\{\\star\}=A^\{\\top\}\(AA^\{\\top\}\)^\{\-1\}\\mathbf\{u\}\_\{M\},\\qquad K\_\{M\}\(X\_\{M\},X\_\{M\}\)=AA^\{\\top\},the QR interpolant satisfies
ℐMu\(x\)=KM\(x,XM\)KM\(XM,XM\)−1𝐮M\.\\mathcal\{I\}\_\{M\}u\(x\)=K\_\{M\}\(x,X\_\{M\}\)K\_\{M\}\(X\_\{M\},X\_\{M\}\)^\{\-1\}\\mathbf\{u\}\_\{M\}\.\(37\)Thus, the multiscale QR approximation is itself kernel interpolation with a finite, discretization\-dependent multiscale autocorrelation kernel\. Unfortunately, this identity does not by itself imply a doubled rate\. To see the obstruction, write
A=\[B0,XC\],D:=B0,X−1C,𝐚0:=B0,X−1𝐮M,A=\[\\,B\_\{0,X\}\\;C\\,\],\\qquad D:=B\_\{0,X\}^\{\-1\}C,\\qquad\\mathbf\{a\}\_\{0\}:=B\_\{0,X\}^\{\-1\}\\mathbf\{u\}\_\{M\},whereCCcontains every frame column not belonging to the primary finest block\. LetT0T\_\{0\}andTcT\_\{c\}denote synthesis by the corresponding finest and remaining basis functions\. Writing the coefficient vector as\(𝐚,𝐛\)\(\\mathbf\{a\},\\mathbf\{b\}\), the interpolation constraint is
𝐚\+D𝐛=𝐚0\.\\mathbf\{a\}\+D\\mathbf\{b\}=\\mathbf\{a\}\_\{0\}\.Hence the minimum\-norm problem reduces to
min𝐛‖𝐚0−D𝐛‖22\+‖𝐛‖22,\\min\_\{\\mathbf\{b\}\}\\\|\\mathbf\{a\}\_\{0\}\-D\\mathbf\{b\}\\\|\_\{2\}^\{2\}\+\\\|\\mathbf\{b\}\\\|\_\{2\}^\{2\},whose normal equations are
\(I\+D⊤D\)𝐛=D⊤𝐚0\.\(I\+D^\{\\top\}D\)\\mathbf\{b\}=D^\{\\top\}\\mathbf\{a\}\_\{0\}\.Substituting𝐚=𝐚0−D𝐛\\mathbf\{a\}=\\mathbf\{a\}\_\{0\}\-D\\mathbf\{b\}intoT0𝐚\+Tc𝐛T\_\{0\}\\mathbf\{a\}\+T\_\{c\}\\mathbf\{b\}gives
ℐMu−𝒮Mu=\(Tc−T0D\)\(I\+D⊤D\)−1D⊤𝐚0,\\mathcal\{I\}\_\{M\}u\-\\mathcal\{S\}\_\{M\}u=\(T\_\{c\}\-T\_\{0\}D\)\(I\+D^\{\\top\}D\)^\{\-1\}D^\{\\top\}\\mathbf\{a\}\_\{0\},\(38\)where𝒮Mu:=T0𝐚0\\mathcal\{S\}\_\{M\}u:=T\_\{0\}\\mathbf\{a\}\_\{0\}is the ordinary finest\-level interpolant\. Although
‖\(I\+D⊤D\)−1D⊤‖2≤12,\\\|\(I\+D^\{\\top\}D\)^\{\-1\}D^\{\\top\}\\\|\_\{2\}\\leq\\frac\{1\}\{2\},this estimate contains no factor ofM−τℓ/dM^\{\-\\tau\_\{\\ell\}/d\}and controls coefficient space rather than theHτℓH^\{\\tau\_\{\\ell\}\}norm\. A superconvergence proof therefore requires additional function\-space stability and approximation estimates\.
Let
WM:=range\(ℐM\)⊂Hτℓ\(Ω\)\.W\_\{M\}:=\\operatorname\{range\}\(\\mathcal\{I\}\_\{M\}\)\\subset H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\.WithℐM=TMAM†EM\\mathcal\{I\}\_\{M\}=T\_\{M\}A\_\{M\}^\{\\dagger\}E\_\{M\}andEMTM=AME\_\{M\}T\_\{M\}=A\_\{M\}, the Moore–Penrose identityAM†AMAM†=AM†A\_\{M\}^\{\\dagger\}A\_\{M\}A\_\{M\}^\{\\dagger\}=A\_\{M\}^\{\\dagger\}gives
ℐM2=TMAM†AMAM†EM=ℐM\.\\mathcal\{I\}\_\{M\}^\{2\}=T\_\{M\}A\_\{M\}^\{\\dagger\}A\_\{M\}A\_\{M\}^\{\\dagger\}E\_\{M\}=\\mathcal\{I\}\_\{M\}\.HenceℐM\\mathcal\{I\}\_\{M\}is a projector ontoWMW\_\{M\}\. The following result isolates sufficient conditions for doubling\.
###### Theorem 3\(Conditional doubled\-rate estimate\)\.
LetΩ⊂ℝd\\Omega\\subset\\mathbb\{R\}^\{d\}be a bounded Lipschitz domain satisfying an interior cone condition, and letXM⊂ΩX\_\{M\}\\subset\\Omegabe sets ofMMdistinct sampling sites satisfying \([32](https://arxiv.org/html/2608.25084#S3.E32)\)\. LetℐM\\mathcal\{I\}\_\{M\}be the associated multiscale QR interpolants\. Suppose\(ℬℓ,0\(Ω\),∥⋅∥ℬℓ,0\(Ω\)\)\(\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\),\\\|\\cdot\\\|\_\{\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)\}\)is a normed space continuously embedded inH2τℓ\(Ω\)H^\{2\\tau\_\{\\ell\}\}\(\\Omega\)and that there are constantsCstabC\_\{\\mathrm\{stab\}\}andCappC\_\{\\mathrm\{app\}\}, independent ofMM, such that
‖ℐM‖Hτℓ\(Ω\)→Hτℓ\(Ω\)\\displaystyle\\\|\\mathcal\{I\}\_\{M\}\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\\to H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}≤Cstab,\\displaystyle\\leq C\_\{\\mathrm\{stab\}\},\(39\)infw∈WM‖u−w‖Hτℓ\(Ω\)\\displaystyle\\inf\_\{w\\in W\_\{M\}\}\\\|u\-w\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}≤CappM−τℓ/d∥u∥ℬℓ,0\(Ω\)\\displaystyle\\leq C\_\{\\mathrm\{app\}\}M^\{\-\\tau\_\{\\ell\}/d\}\\\|u\\\|\_\{\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)\}\(40\)for everyu∈ℬℓ,0\(Ω\)u\\in\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)\. Then, for all sufficiently largeMM,
∥u−ℐMu∥L2\(Ω\)≤CM−2τℓ/d∥u∥ℬℓ,0\(Ω\),\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{L^\{2\}\(\\Omega\)\}\\leq CM^\{\-2\\tau\_\{\\ell\}/d\}\\\|u\\\|\_\{\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)\},\(41\)whereCCis independent ofMManduu\.
###### Proof\.
For anyw∈WMw\\in W\_\{M\}, the projector property gives
u−ℐMu=\(I−ℐM\)\(u−w\)\.u\-\\mathcal\{I\}\_\{M\}u=\(I\-\\mathcal\{I\}\_\{M\}\)\(u\-w\)\.Therefore, by \([39](https://arxiv.org/html/2608.25084#S3.E39)\) and \([40](https://arxiv.org/html/2608.25084#S3.E40)\),
∥u−ℐMu∥Hτℓ\(Ω\)≤\(1\+Cstab\)CappM−τℓ/d∥u∥ℬℓ,0\(Ω\)\.\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\\leq\(1\+C\_\{\\mathrm\{stab\}\}\)C\_\{\\mathrm\{app\}\}M^\{\-\\tau\_\{\\ell\}/d\}\\\|u\\\|\_\{\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)\}\.The error vanishes onXMX\_\{M\}\. For all sufficiently largeMM, the scattered\-zeros inequality withμ=0\\mu=0and \([32](https://arxiv.org/html/2608.25084#S3.E32)\) give
∥u−ℐMu∥L2\(Ω\)≤ChMτℓ∥u−ℐMu∥Hτℓ\(Ω\)≤CM−τℓ/d∥u−ℐMu∥Hτℓ\(Ω\)\.\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{L^\{2\}\(\\Omega\)\}\\leq Ch\_\{M\}^\{\\tau\_\{\\ell\}\}\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\\leq CM^\{\-\\tau\_\{\\ell\}/d\}\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{H^\{\\tau\_\{\\ell\}\}\(\\Omega\)\}\.Combining the two estimates proves \([41](https://arxiv.org/html/2608.25084#S3.E41)\)\.
Hereℬℓ,0\(Ω\)\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)denotes a localized doubled source class for which a global doubling estimate is meaningful; on bounded domains, additional smoothness alone need not supply the localization or boundary compatibility required by classical superconvergence theory\[[49](https://arxiv.org/html/2608.25084#bib.bib59),[51](https://arxiv.org/html/2608.25084#bib.bib60)\]\. For the convolution kernel
κd,ℓ:=ϕd,ℓ∗ϕd,ℓ,\\kappa\_\{d,\\ell\}:=\\phi\_\{d,\\ell\}\*\\phi\_\{d,\\ell\},the corresponding whole\-space native norm is equivalent to theH2τℓH^\{2\\tau\_\{\\ell\}\}norm because
κ^d,ℓ=\|ϕ^d,ℓ\|2,ϕ^d,ℓ\(ω\)≍\(1\+‖ω‖22\)−τℓ\\widehat\{\\kappa\}\_\{d,\\ell\}=\|\\widehat\{\\phi\}\_\{d,\\ell\}\|^\{2\},\\qquad\\widehat\{\\phi\}\_\{d,\\ell\}\(\\omega\)\\asymp\(1\+\\\|\\omega\\\|\_\{2\}^\{2\}\)^\{\-\\tau\_\{\\ell\}\}\[[58](https://arxiv.org/html/2608.25084#bib.bib46),[48](https://arxiv.org/html/2608.25084#bib.bib57)\]\.
Theorem[3](https://arxiv.org/html/2608.25084#Thmtheorem3)reduces the observed doubling phenomenon to two concrete properties of the multiscale rangeWMW\_\{M\}: uniformHτℓH^\{\\tau\_\{\\ell\}\}stability of the QR projector and anHτℓH^\{\\tau\_\{\\ell\}\}Jackson estimate on the localized doubled source class\. We state these properties as the central conjecture\.
###### Conjecture 3\.1\(Uniform stability and doubled\-space approximation\)\.
LetXM⊂ΩX\_\{M\}\\subset\\Omegabe sets ofMMdistinct sampling sites satisfying \([32](https://arxiv.org/html/2608.25084#S3.E32)\), with no separation\-radius assumption, and let
C0,M=XM⊃C1,M⊃⋯⊃CJ\(M\),MC\_\{0,M\}=X\_\{M\}\\supset C\_\{1,M\}\\supset\\cdots\\supset C\_\{J\(M\),M\}be the nested hierarchy constructed as in Section[2\.1\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS1)\. Lets0,Ms\_\{0,M\}denote the representative finest\-level spacing used in \([6](https://arxiv.org/html/2608.25084#S2.E6)\), assumes0,M≍M−1/ds\_\{0,M\}\\asymp M^\{\-1/d\}, and set
sj,M=2js0,M,ρj,M=ηsj,Ms\_\{j,M\}=2^\{j\}s\_\{0,M\},\\qquad\\rho\_\{j,M\}=\\eta\\,s\_\{j,M\}for fixedη\>0\\eta\>0\. For the multiscale QR interpolant built from theC2ℓC^\{2\\ell\}Wendland kernel, the uniform stability estimate \([39](https://arxiv.org/html/2608.25084#S3.E39)\) and approximation estimate \([40](https://arxiv.org/html/2608.25084#S3.E40)\) hold onℬℓ,0\(Ω\)\\mathcal\{B\}\_\{\\ell,0\}\(\\Omega\)\.
##### Example
If Conjecture[3\.1](https://arxiv.org/html/2608.25084#Thmtheorem3.Thmconjecture1)holds, Theorem[3](https://arxiv.org/html/2608.25084#Thmtheorem3)gives
∥u−ℐMu∥L2\(Ω\)=𝒪\(M−2τℓ/d\)\.\\\|u\-\\mathcal\{I\}\_\{M\}u\\\|\_\{L^\{2\}\(\\Omega\)\}=\\mathcal\{O\}\(M^\{\-2\\tau\_\{\\ell\}/d\}\)\.Ford=2d=2, this predicts
C2:𝒪\(\(M1/2\)−5\),C4:𝒪\(\(M1/2\)−7\),C6:𝒪\(\(M1/2\)−9\)\.C^\{2\}:\\ \\mathcal\{O\}\(\(M^\{1/2\}\)^\{\-5\}\),\\qquad C^\{4\}:\\ \\mathcal\{O\}\(\(M^\{1/2\}\)^\{\-7\}\),\\qquad C^\{6\}:\\ \\mathcal\{O\}\(\(M^\{1/2\}\)^\{\-9\}\)\.\(42\)The observed discrete relativeℓ2\\ell\_\{2\}orders are numerically close to these predictions\.
## 4Numerical Results
We now present numerical results for both the multiscale frame approximant \(Section[4\.1](https://arxiv.org/html/2608.25084#S4.SS1)\) and the FKM \(Section[4\.2](https://arxiv.org/html/2608.25084#S4.SS2)\)\. First, in Section[4\.1](https://arxiv.org/html/2608.25084#S4.SS1), we present traditional numerical convergence studies of the multiscale frame approximation for target functions of different smoothness, with comparisons to our theorems and the doubling conjecture\. We also study the influence of design parameters such as the kernel support size and the density of the design matrix\. Then, in Section[4\.2](https://arxiv.org/html/2608.25084#S4.SS2), we present results on several standard operator learning benchmarks, with comparisons against neural operators and the VKM\. Finally, in Section[4\.2\.3](https://arxiv.org/html/2608.25084#S4.SS2.SSS3), we present a qualitative analysis of the*a posteriori*multiscale decompositions obtained from a subset of the operator learning benchmark problems\.
Unless otherwise stated, all errors are reported using the relativeℓ2\\ell\_\{2\}erroreℓ2=‖y−y^‖2‖y‖2e\_\{\\ell\_\{2\}\}=\\frac\{\\\|y\-\\hat\{y\}\\\|\_\{2\}\}\{\\\|y\\\|\_\{2\}\}, whereyydenotes the ground truth solution vector andy^\\hat\{y\}the prediction\.
### 4\.1Frame Approximation Results
Figure 1:Example of a randomly sampled point cloud of 1000 data sites for decomposition and 700 evaluation sites\.We first evaluate the ability of the proposed frame representation to accurately reconstruct functions from their multiscale coefficients\. Here, we restrict ourselves to synthetic test functions and study the influence of key design parameters within the frame\. To assess reconstruction fidelity across different regularity classes, we consider two families of test functions with varying smoothness: analytic \(Cω\(ℝd\)C^\{\\omega\}\(\\mathbb\{R\}^\{d\}\)\) and finitely smooth \(C2\(ℝd\)C^\{2\}\(\\mathbb\{R\}^\{d\}\)\)\. The corresponding functions in both two and three dimensions are shown in Figure[9](https://arxiv.org/html/2608.25084#A1.F9)of the appendix\. For the discussion that follows, letΩ=\[0,1\]d\\Omega=\[0,1\]^\{d\}withd∈\{2,3\}d\\in\\\{2,3\\\}, and letx∈Ωx\\in\\Omega\. As shown in Figure[1](https://arxiv.org/html/2608.25084#S4.F1), the decomposition and evaluation sites are sampled independently and uniformly fromΩ\\Omega; no minimum\-separation constraint is imposed\. The target functions are:
##### Analytic \(CωC^\{\\omega\}\)
f\(x\)=exp\(−\(∑i=1d\(xi−0\.5\)2\)cd\),cd=\{0\.2,d=2,0\.8,d=3\.\\displaystyle f\(x\)=\\exp\\\!\\left\(\\frac\{\-\\left\(\\sum\_\{i=1\}^\{d\}\(x\_\{i\}\-0\.5\)^\{2\}\\right\)\}\{c\_\{d\}\}\\right\),\\qquad c\_\{d\}=\\begin\{cases\}0\.2,&d=2,\\\\ 0\.8,&d=3\.\\end\{cases\}\(43\)
##### Finitely smooth \(C2C^\{2\}\)
f\(x\)=\(∑i=1d\(xi−0\.5\)2\)3/2\.\\displaystyle\\ f\(x\)=\\left\(\\sum\_\{i=1\}^\{d\}\(x\_\{i\}\-0\.5\)^\{2\}\\right\)^\{3/2\}\.\(44\)Together, these functions provide a controlled setting for comparing kernel choices and evaluating the effect of the decomposition parameters on reconstruction accuracy\.
#### 4\.1\.1Convergence as a Function of the Number of Data SitesMM
We use a design\-matrix density ofδ=0\.2\\delta=0\.2as the reference density in the remaining approximation experiments; in the appendix, section[B](https://arxiv.org/html/2608.25084#A2)reports a detailed density sweep\. We then investigate how reconstruction accuracy changes as the number of decomposition sitesMMincreases under two protocols:
1. 1\.For eachMM, choose the finest\-level support radiusρ0\\rho\_\{0\}so that the resulting design\-matrix density is approximatelyδ=0\.2\\delta=0\.2\.
2. 2\.Chooseρ0\\rho\_\{0\}at the largest value ofMMso that the design\-matrix density is approximately0\.20\.2, and then hold this support radius fixed for all smaller values ofMM\.
All experiments use three decomposition scales \(J=2J=2\) for simplicity, although an arbitrary number of scales can be used in practice\. Farthest\-point thinning is performed with target cardinalitiesmj≈M/2jdm\_\{j\}\\approx M/2^\{jd\}, producing scale\-indexed fine, intermediate, and coarse frame levels\.
##### Fixed density
Figure 2:ℓ2\\ell\_\{2\}error at evaluation sites as a function of the number of data sitesMMwhile holding the design matrix density at approximately0\.20\.2\(2D\)\.Figure 3:ℓ2\\ell\_\{2\}error at evaluation sites as a function of the number of decomposition points while holding the design matrix density at approximately0\.20\.2\(3D\)\.We first consider the fixed\-density case\. Letδ\\deltadenote the fraction of nonzero entries in the design matrix\. Holdingδ\\deltaapproximately constant fixes the fraction of active kernel evaluations and therefore provides a convenient way to compare sparsity across resolutions\. In this protocol, the support radiusρ0\\rho\_\{0\}is adjusted separately for eachMMto attain the target density; we make no assumption here about an asymptotic scaling law forρ0\\rho\_\{0\}withMM\.
Withδ\\deltafixed at approximately0\.20\.2, Figures[2](https://arxiv.org/html/2608.25084#S4.F2)and[3](https://arxiv.org/html/2608.25084#S4.F3)show that the evaluation\-site error decreases rapidly asMMincreases\. For the analytic targets, the observed slopes are numerically close to the doubled rates conjectured in Section[3\.3](https://arxiv.org/html/2608.25084#S3.SS3)\. TheC2C^\{2\}targets exhibit faster than expected rates over the tested range, which we interpret as preasymptotic behavior\. Note the differentyy\-axis scales for the analytic and finitely smooth targets in both 2D and 3D\. TheC6C^\{6\}kernel in 3D shows a leveling off under the fixed\-density protocol\. Becauseρ0\\rho\_\{0\}is retuned asMMchanges in this experiment, we examine fixed support separately below\.
##### Fixed support
Figure 4:ℓ2\\ell\_\{2\}error as a function of the number of points in the domain\. The kernel support radius \(ρ\\rho\) is chosen on the finest scale to produce a design matrix density of approximately0\.20\.2, and this value is then held fixed as the number of points changes\.Figure 5:ℓ2\\ell\_\{2\}error as a function of the number of points in the domain\. The kernel support radius \(ρ\\rho\) is chosen on the finest scale to produce a design matrix density of approximately0\.20\.2, and this value is then held fixed as the number of points changes\.Next, we examine convergence with a fixed finest\-level support radiusρ0\\rho\_\{0\}\. We first chooseρ0\\rho\_\{0\}at the largest value ofMMso that the design matrix has density approximately0\.20\.2, and then use this same value ofρ0\\rho\_\{0\}for all remaining resolutions; consequently,δ\\deltais allowed to vary withMM\. The evaluation\-site errors in Figures[4](https://arxiv.org/html/2608.25084#S4.F4)and[5](https://arxiv.org/html/2608.25084#S4.F5)again decrease withMM, rapidly for the analytic target and more slowly for the finitely smooth target\. Under this fixed\-support protocol, the observed slopes for the analytic targets are numerically close to the rates predicted by the doubling conjecture in Section[3\.3](https://arxiv.org/html/2608.25084#S3.SS3)\.
### 4\.2Operator Learning Results
We also evaluated the FKM on a set of operator learning benchmark problems spanning both standard PDEs and problems with inherent multiscale structure\. The standard PDE benchmarks allow comparison against existing operator learning methods, while the multiscale benchmarks test the effect of an explicitly scale\-indexed representation\. We used aC4C^\{4\}Matérn kernel for the operator interpolant and aC4C^\{4\}Wendland kernel for the frame\.
In selecting a design\-matrix density for the operator\-learning problems, we found that an initial value ofη≈2\\eta\\approx 2in \([5](https://arxiv.org/html/2608.25084#S2.E5)\) provided a useful balance between computational efficiency and reconstruction accuracy on unit domains\. We then adjusted the density for each benchmark; the values used are reported in Tables[1](https://arxiv.org/html/2608.25084#S4.T1)and[2](https://arxiv.org/html/2608.25084#S4.T2)\. The resulting levelwise frame components are illustrated in Figures[6](https://arxiv.org/html/2608.25084#S4.F6)–[8](https://arxiv.org/html/2608.25084#S4.F8)\. Unlike the approximation convergence studies, these experiments report generalization errors at fixed spatial resolutions\. We compare against our own implementations of Geo\-FNO, Transolver, and VKM\.
#### 4\.2\.1Base PDEs
To compare our approach to the state\-of\-the\-art methods, we first present results for the following problems that use a mix of gridded data, triangular meshes, and point clouds\.
2D Lid\-Driven Cavity Flow:Maps an initial velocity field that is zero in the interior, with a rightward\-moving lid, to the velocity field atT=5T=5s, i\.e\.,𝒢:u\(x,0\)→u\(x,T\)\\mathcal\{G\}:u\(x,0\)\\to u\(x,T\), for a square cavity with a randomly initialized lid velocity\[[50](https://arxiv.org/html/2608.25084#bib.bib9)\]\(triangular mesh\)\.
2D Darcy Flow \(PWC\):Maps a piecewise constant \(PWC\) permeability fieldc:\[0,1\]2→ℝc:\[0,1\]^\{2\}\\to\\mathbb\{R\}to the pressure fieldu:\[0,1\]2→ℝu:\[0,1\]^\{2\}\\to\\mathbb\{R\}via the Darcy flow equation\[[37](https://arxiv.org/html/2608.25084#bib.bib8)\]\(gridded data\)\.
2D Darcy Flow on a Triangle:Maps a boundary\-condition field to the pressure fieldh\(x,y\)h\(x,y\)on a triangular domain through Darcy’s equation\[[37](https://arxiv.org/html/2608.25084#bib.bib8)\]\. The permeability is set to0\.10\.1, the forcing is set to−1\.0\-1\.0, and the boundary\-condition field is sampled from a Gaussian random field \(triangular mesh\)\.
3D Reaction\-Variable\-Coefficient\-Diffusion on a Point Cloud:Maps a spatially constant initial chemical concentrationc\(y,0\):ℝ3→ℝc\(y,0\):\\mathbb\{R\}^\{3\}\\to\\mathbb\{R\}, whose scalar value is sampled uniformly at random for each realization, to the concentrationc\(y,0\.5\)c\(y,0\.5\)inside the unit ball, sampled on a point cloud, as described in\[[50](https://arxiv.org/html/2608.25084#bib.bib9)\]\. Due to a discontinuity in the reaction term, the final solution exhibits a sharp gradient across the planey1=0y\_\{1\}=0\.
Table 1:Comparison of operator learning methods on standard PDE benchmarks\. Errors are reported as relativeℓ2\\ell\_\{2\}errors\. Here, “Density” refers to the density of the design matrix,κ\\kappadenotes the target condition number of the operator\-learning Gram matrixKK,λ\\lambdais the regularization parameter,NTN\_\{T\}is the number of training functions, andMMis the number of spatial locations\.Table[1](https://arxiv.org/html/2608.25084#S4.T1)demonstrates that the FKM achieves the lowest generalization error on three of the four benchmark problems, with absolutely no changes to the input or output data\. The only exception is the piecewise\-constant Darcy problem, where discontinuities violate the smoothness assumptions of the kernel\-frame representation\. Nevertheless, even on this problem, the FKM remains competitive with the single\-scale VKM while offering an*a posteriori*multiscale decomposition\.
The largest improvement is observed for the cavity\-flow dataset, where FKM reduces the prediction error by more than two orders of magnitude relative to the closest competing method\. On the triangular Darcy problem, the reduction is nearly two orders of magnitude\. For reaction–diffusion, FKM improves on VKM by a factor of approximately two and on Geo\-FNO and Transolver by more than two orders of magnitude\. These results indicate that the scale\-indexed kernel\-frame representation can substantially improve prediction accuracy across several operator\-learning problems\.
#### 4\.2\.2Multiscale PDEs
To understand the impacts of multiscale structure within the PDEs themselves, we further evaluated our method on the following benchmarks:
Flow Past a Cylinder \(Laminar\):Predict the incompressible velocity fieldu\(x,y,t=10s\)u\(x,y,t=10\\,\\mathrm\{s\}\)from the initial uniform fieldu\(x,y,0\)u\(x,y,0\)\. We impose a no\-slip condition on the cylinder, zero pressure at the outlet, and prescribed velocity on the top, bottom, and inlet boundaries\. We choose Reynolds numbersRe∈\[25,64\]Re\\in\[25,64\]such that the flow regime is laminar\[[50](https://arxiv.org/html/2608.25084#bib.bib9)\]\. The geometry introduces distinct length scales associated with the cylinder radius and the domain length\.
Flow Past a Cylinder \(Vortex Shedding\):Predict the incompressible velocity fieldu\(x,y,t=10s\)u\(x,y,t=10\\,\\mathrm\{s\}\)from the initial uniform velocity fieldu\(x,y,0\)u\(x,y,0\)\. We impose a no\-slip condition on the cylinder, zero pressure at the outlet, and prescribed velocity on the top, bottom, and inlet boundaries\. We choose Reynolds numbersRe∈\[112,199\]Re\\in\[112,199\]such that the flow regime exhibits vortex shedding\[[50](https://arxiv.org/html/2608.25084#bib.bib9)\]\. The resulting vortex train introduces additional spatial structure between the cylinder and domain length scales\.
3D Compressible Navier–Stokes \(NS\), Mach 1\.0:Predict the velocity fieldv:\[0,1\]3→ℝv:\[0,1\]^\{3\}\\to\\mathbb\{R\}from an initial velocity fieldv0:\[0,1\]3→ℝv\_\{0\}:\[0,1\]^\{3\}\\to\\mathbb\{R\}using the compressible Navier–Stokes equations\. A Mach number of1\.01\.0introduces strong compressibility effects and additional spatial structure\[[52](https://arxiv.org/html/2608.25084#bib.bib10)\]\. The associated loss of smoothness makes this a challenging benchmark\.
3D Species Transport:Learn the operator mapping inlet velocity fieldsuinletu\_\{\\text\{inlet\}\}to the velocity fieldu\(y,z=0\.5\)u\(y,z=0\.5\)at final timet=0\.5t=0\.5for turbulent mixing of air and methane in a static Kenics mixer with three twisted blades\. Air enters fory<0y<0and methane fory\>0y\>0\[[41](https://arxiv.org/html/2608.25084#bib.bib11)\]\. We model this as an incompressible flow problem with thek−ωk\-\\omega\-SST turbulence model; turbulence here creates multiple length scales\.
Table 2:Comparison of operator learning methods on multiscale PDE benchmarks\. Errors are reported as relativeℓ2\\ell\_\{2\}errors\. Here, “Density” refers to the density of the design matrix,κ\\kappadenotes the target condition number of the operator\-learning Gram matrixKK,λ\\lambdais the regularization parameter,NTN\_\{T\}is the number of training functions, andMMis the number of output points\.Table[2](https://arxiv.org/html/2608.25084#S4.T2)demonstrates that FKM achieves the lowest prediction error on three of the four multiscale benchmark problems\. The only exception is the compressible Navier–Stokes benchmark, where FKM remains competitive with the best\-performing method despite the increased difficulty of the problem and the limited training set of only 50 functions\.
The largest improvement is observed for the laminar cylinder\-flow problem, where FKM reduces the prediction error by more than two orders of magnitude relative to the closest competing method\. For vortex shedding, FKM reduces the error by a factor of about two relative to Transolver, the closest competitor\. On species transport, FKM achieves approximately half the error of VKM and larger improvements over Geo\-FNO and Transolver\. These results suggest that organizing the learned representation by kernel\-frame level can be advantageous for PDEs with spatial structure across multiple length scales\.
#### 4\.2\.3Qualitative Multiscale Analysis
While the improved prediction accuracies demonstrate the effectiveness of FKM, they do not fully capture the representational information provided by the frame\. The predicted coefficients retain their frame\-level organization, so the contribution from each level can be synthesized separately before the contributions are summed to recover the full prediction\. Because the frame is redundant and its levels are generally nonorthogonal, these levelwise components should be interpreted as scale\-indexed frame contributions rather than disjoint spectral bands\. Figures[6](https://arxiv.org/html/2608.25084#S4.F6),[7](https://arxiv.org/html/2608.25084#S4.F7), and[8](https://arxiv.org/html/2608.25084#S4.F8)illustrate these components for selected operator\-learning benchmarks\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 6:Scale\-indexed frame decomposition of a predicted solution to the cylinder\-flow problem with vortex shedding at three levels\. The bottom right image shows the combined reconstruction of the problem\.##### Flow past a cylinder
Figure[6](https://arxiv.org/html/2608.25084#S4.F6)presents the levelwise frame contributions for a single component of the velocity field in the vortex\-shedding problem\. In this example, the coarsest\-level contribution contains the primary wake and large downstream vortical structures, while the finer\-level contributions contain increasingly localized structure near the cylinder and in the wake\. Their sum gives the combined reconstruction shown in the bottom\-right panel\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 7:Scale\-indexed frame decomposition of a predicted solution to the reaction–diffusion problem at three levels\. The bottom right image shows the combined reconstruction\.
##### Reaction–Diffusion
Figure[7](https://arxiv.org/html/2608.25084#S4.F7)illustrates the levelwise frame contributions to the concentration field in the reaction–diffusion problem, shown by coloring the collocation points directly\. The coarsest\-level contribution contains the dominant global concentration variation and transition interface, while the intermediate and finest levels contain progressively more localized variations\. The components are nonorthogonal and should therefore be read as frame\-level contributions rather than a strict frequency decomposition\.
\(a\)Fine
\(b\)Medium
\(c\)Coarse
\(d\)Combined
Figure 8:Scale\-indexed frame decomposition of a predicted solution to the species\-transport problem for thezz\-component of velocity at three levels\. The bottom right image shows the combined reconstruction\.
##### Species Transport
Finally, Figure[8](https://arxiv.org/html/2608.25084#S4.F8)illustrates the levelwise frame contributions to thezz\-component of the velocity field in the species\-transport problem\. The coarsest\-level contribution contains the dominant transport pattern, while the intermediate and finest levels show increasingly localized vortical and mixing\-induced structure downstream of the blades\. As above, these panels visualize the contributions associated with the frame levels and do not imply an orthogonal separation of physical scales\.
Similar scale\-indexed behavior is observed across the remaining benchmark problems in appendix section[D](https://arxiv.org/html/2608.25084#A4)\. These levelwise outputs provide additional interpretability and a representation that may be useful for coupling operator learning with multiscale numerical methods beyond end\-to\-end solution prediction\.
## 5Conclusions and Future Work
We presented a novel multiscale kernel frame approximation technique that allows for the decomposition of target functions into multiple spatial length scales\. We proved that the frame is interpolatory, established formal error estimates, and presented a doubling conjecture that matches observed numerical convergence rates\. We then leveraged the frame as input and output approximation spaces within a kernel method for operator learning, leading to the Frame Kernel Method \(FKM\)\. We found that the FKM produced more accurate predictions than several standard operator learning architectures without any changes to the data\.
For future work, we aim to prove the doubling conjecture and thereby establish the doubled convergence rates\. The sparse QR factorization currently limits our 3D implementation to approximatelyO\(105\)O\(10^\{5\}\)points\. To allow for scaling to millions of spatial locations, we will explore sketching approaches, preconditioning techniques within iterative least squares solvers, and block QR decompositions that take advantage of the block structure of our approximant\. Finally, we aim to incorporate the FKM within multiscale\-multigrid solvers designed for challenging problems arising from cohesive zone formulations and nonlinear elasticity formulations for solids, textiles, and composites\.
## References
- \[1\]R\. A\. Adams and J\. J\. F\. Fournier\(2003\)Sobolev spaces\.2 edition,Pure and Applied Mathematics, Vol\.140,Academic Press,Amsterdam\.External Links:ISBN 978\-0\-12\-044143\-3Cited by:[§3\.2](https://arxiv.org/html/2608.25084#S3.SS2.p3.2)\.
- \[2\]N\. Aronszajn\(1950\)Theory of reproducing kernels\.Transactions of the American Mathematical Society68\(3\),pp\. 337–404\.External Links:[Document](https://dx.doi.org/10.1090/S0002-9947-1950-0051437-7)Cited by:[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p2.2)\.
- \[3\]S\. Avesani, R\. Kempf, M\. Multerer, and H\. Wendland\(2024\)Multiscale scattered data analysis in samplet coordinates\.arXiv preprint arXiv:2409\.14791\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2409.14791)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p6.1)\.
- \[4\]P\. Batlle, M\. Darcy, B\. Hosseini, and H\. Owhadi\(2024\)Kernel methods are competitive for operator learning\.Journal of Computational Physics496,pp\. 112549\.External Links:ISSN 0021\-9991,[Document](https://dx.doi.org/10.1016/j.jcp.2023.112549),[Link](https://www.sciencedirect.com/science/article/pii/S0021999123006447)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p2.1),[§1](https://arxiv.org/html/2608.25084#S1.p8.1),[§2\.2](https://arxiv.org/html/2608.25084#S2.SS2.p1.1)\.
- \[5\]P\. Benner, S\. Gugercin, and K\. Willcox\(2015\)A survey of projection\-based model reduction methods for parametric dynamical systems\.SIAM Review57\(4\),pp\. 483–531\.External Links:[Document](https://dx.doi.org/10.1137/130932715)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[6\]J\. L\. Bentley\(1975\)Multidimensional binary search trees used for associative searching\.Communications of the ACM18\(9\),pp\. 509–517\.External Links:[Document](https://dx.doi.org/10.1145/361002.361007)Cited by:[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS0.Px2.p1.3)\.
- \[7\]S\. C\. Brenner and L\. R\. Scott\(2008\)The mathematical theory of finite element methods\.3 edition,Springer\.External Links:[Document](https://dx.doi.org/10.1007/978-0-387-75934-0)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[8\]W\. L\. Briggs, V\. E\. Henson, and S\. F\. McCormick\(2000\)A multigrid tutorial\.2 edition,Society for Industrial and Applied Mathematics\.External Links:[Document](https://dx.doi.org/10.1137/1.9780898719505)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[9\]S\. L\. Brunton, B\. R\. Noack, and P\. Koumoutsakos\(2020\)Machine learning for fluid mechanics\.Annual Review of Fluid Mechanics52,pp\. 477–508\.External Links:[Document](https://dx.doi.org/10.1146/annurev-fluid-010719-060214)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[10\]M\. D\. Buhmann\(2003\)Radial basis functions: theory and implementations\.Cambridge University Press,Cambridge\.External Links:ISBN 9780521633383Cited by:[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS0.Px1.p1.3)\.
- \[11\]C\. Canuto, M\. Y\. Hussaini, A\. Quarteroni, and T\. A\. Zang\(2006\)Spectral methods: fundamentals in single domains\.Springer\.External Links:[Document](https://dx.doi.org/10.1007/978-3-540-30726-6)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[12\]O\. Christensen\(2016\)An introduction to frames and riesz bases\.2 edition,Birkhäuser,New York\.External Links:ISBN 9783319256560Cited by:[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.p4.2),[§3\.1](https://arxiv.org/html/2608.25084#S3.SS1.p1.2)\.
- \[13\]A\. Cohen\(2003\)Numerical analysis of wavelet methods\.Elsevier\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[14\]I\. Daubechies\(1992\)Ten lectures on wavelets\.Society for Industrial and Applied Mathematics\.External Links:[Document](https://dx.doi.org/10.1137/1.9781611970104)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[15\]R\. J\. Duffin and A\. C\. Schaeffer\(1952\)A class of nonharmonic fourier series\.Transactions of the American Mathematical Society72\(2\),pp\. 341–366\.External Links:[Document](https://dx.doi.org/10.1090/S0002-9947-1952-0047179-6)Cited by:[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.p4.2),[§3\.1](https://arxiv.org/html/2608.25084#S3.SS1.p1.2)\.
- \[16\]W\. E and B\. Engquist\(2003\)The heterogeneous multiscale methods\.Communications in Mathematical Sciences1\(1\),pp\. 87–132\.External Links:[Document](https://dx.doi.org/10.4310/CMS.2003.v1.n1.a8)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[17\]W\. E\(2011\)Principles of multiscale modeling\.Cambridge University Press\.External Links:[Document](https://dx.doi.org/10.1017/CBO9780511975616)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[18\]Y\. Efendiev, J\. Galvis, and T\. Y\. Hou\(2013\)Generalized multiscale finite element methods \(GMsFEM\)\.Journal of Computational Physics251,pp\. 116–135\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2013.04.045)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1),[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[19\]G\. E\. Fasshauer and M\. J\. McCourt\(2015\)Kernel\-based approximation methods using MATLAB\.World Scientific,Singapore\.External Links:ISBN 9789814630139Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1),[§2\.1\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS1.Px1.p1.3),[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.p2.2)\.
- \[20\]B\. Fornberg, \{T\. A\. Driscoll, G\. Wright, and R\. Charles\(2002\)Observations on the behavior of radial basis function approximations near boundaries\.Computers and Mathematics with Applications43\(3\-5\),pp\. 473–490\(English\)\.External Links:[Document](https://dx.doi.org/10.1016/S0898-1221%2801%2900299-1)Cited by:[§2\.1\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS1.Px2.p2.1)\.
- \[21\]G\. H\. Golub and C\. F\. Van Loan\(2013\)Matrix computations\.4 edition,Johns Hopkins University Press,Baltimore, MD\.External Links:ISBN 978\-1\-4214\-0794\-4Cited by:[§3\.1](https://arxiv.org/html/2608.25084#S3.SS1.p3.5.1)\.
- \[22\]T\. F\. Gonzalez\(1985\)Clustering to minimize the maximum intercluster distance\.Theoretical Computer Science38,pp\. 293–306\.External Links:[Document](https://dx.doi.org/10.1016/0304-3975%2885%2990224-5)Cited by:[§2\.1\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS1.Px2.p1.1)\.
- \[23\]M\. Griebel, C\. Rieger, and B\. Zwicknagl\(2015\)Multiscale approximation and reproducing kernel hilbert space methods\.SIAM Journal on Numerical Analysis53\(2\),pp\. 852–873\.External Links:[Document](https://dx.doi.org/10.1137/130932144)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p6.1)\.
- \[24\]G\. Gupta, X\. Xiao, and P\. Bogdan\(2021\)Multiwavelet\-based operator learning for differential equations\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 24048–24062\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[25\]T\. Y\. Hou and X\. Wu\(1997\)A multiscale finite element method for elliptic problems in composite materials and porous media\.Journal of Computational Physics134\(1\),pp\. 169–189\.External Links:[Document](https://dx.doi.org/10.1006/jcph.1997.5682)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[26\]T\. Ingebrand, A\. J\. Thorpe, S\. Goswami, K\. Kumar, and U\. Topcu\(2025\)Basis\-to\-basis operator learning using function encoders\.Computer Methods in Applied Mechanics and Engineering435,pp\. 117646\.External Links:[Document](https://dx.doi.org/10.1016/j.cma.2024.117646)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[27\]H\. Kadri, E\. Duflos, P\. Preux, S\. Canu, A\. Rakotomamonjy, and J\. Audiffren\(2016\)Operator\-valued kernels for learning from functional response data\.Journal of Machine Learning Research17\(20\),pp\. 1–54\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p2.1)\.
- \[28\]M\. C\. Kennedy and A\. O’Hagan\(2001\)Bayesian calibration of computer models\.Journal of the Royal Statistical Society: Series B63\(3\),pp\. 425–464\.External Links:[Document](https://dx.doi.org/10.1111/1467-9868.00294)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[29\]D\. Kochkov, J\. A\. Smith, A\. Alieva, Q\. Wang, M\. P\. Brenner, and S\. Hoyer\(2021\)Machine learning–accelerated computational fluid dynamics\.Proceedings of the National Academy of Sciences118\(21\),pp\. e2101784118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2101784118)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[30\]N\. Kovachki, Z\. Li, B\. Liu, K\. Azizzadenesheli, K\. Bhattacharya, A\. Stuart, and A\. Anandkumar\(2023\)Neural operator: learning maps between function spaces with applications to PDEs\.Journal of Machine Learning Research24\(89\),pp\. 1–97\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[31\]Q\. T\. Le Gia, I\. H\. Sloan, and H\. Wendland\(2017\)Zooming from global to local: a multiscale RBF approach\.Advances in Computational Mathematics43,pp\. 581–606\.External Links:[Document](https://dx.doi.org/10.1007/s10444-016-9498-4)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p6.1)\.
- \[32\]Z\. Li, Z\. Lai, X\. Zhang, and W\. Wang\(2024\)M2NO: multiresolution operator learning with multiwavelet\-based algebraic multigrid method\.arXiv preprint arXiv:2406\.04822\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.04822)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[33\]Z\. Li, N\. Kovachki, K\. Azizzadenesheli, B\. Liu, K\. Bhattacharya, A\. Stuart, and A\. Anandkumar\(2021\)Fourier neural operator for parametric partial differential equations\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1),[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[34\]X\. Liu, B\. Xu, S\. Cao, and L\. Zhang\(2024\)Mitigating spectral bias for the multiscale operator learning\.Journal of Computational Physics506,pp\. 112944\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2024.112944)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[35\]M\. Lowery, J\. Turnage, Z\. Morrow, J\. D\. Jakeman, A\. Narayan, S\. Zhe, and V\. Shankar\(2026\)Kernel neural operators \(KNOs\) for scalable, memory\-efficient, geometrically\-flexible operator learning\.Transactions on Machine Learning Research\.Note:arXiv:2407\.00809Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[36\]L\. Lu, P\. Jin, G\. Pang, Z\. Zhang, and G\. E\. Karniadakis\(2021\)Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators\.Nature Machine Intelligence3,pp\. 218–229\.External Links:[Document](https://dx.doi.org/10.1038/s42256-021-00302-5)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[37\]L\. Lu, X\. Meng, S\. Cai, Z\. Mao, S\. Goswami, Z\. Zhang, and G\. E\. Karniadakis\(2022\)A comprehensive and fair comparison of two neural operators \(with practical extensions\) based on fair data\.Computer Methods in Applied Mechanics and Engineering393,pp\. 114778\.External Links:[Document](https://dx.doi.org/10.1016/j.cma.2022.114778),ISSN 0045\-7825Cited by:[§4\.2\.1](https://arxiv.org/html/2608.25084#S4.SS2.SSS1.p3.1),[§4\.2\.1](https://arxiv.org/html/2608.25084#S4.SS2.SSS1.p4.1)\.
- \[38\]H\. Luo, H\. Wu, H\. Zhou, L\. Xing, Y\. Di, J\. Wang, and M\. Long\(2025\)Transolver\+\+: an accurate neural solver for PDEs on million\-scale geometries\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 41432–41449\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[39\]S\. Mallat\(2009\)A wavelet tour of signal processing: the sparse way\.3 edition,Academic Press\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[40\]C\. A\. Micchelli and M\. Pontil\(2005\)On learning vector\-valued functions\.Neural Computation17\(1\),pp\. 177–204\.External Links:[Document](https://dx.doi.org/10.1162/0899766052530802)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p2.1)\.
- \[41\]C\. Morales Ubal, N\. Beishuizen, L\. Kusch, and J\. van Oijen\(2024\)Adjoint\-based design optimization of a kenics static mixer\.Results in Engineering21,pp\. 101856\.Cited by:[§4\.2\.2](https://arxiv.org/html/2608.25084#S4.SS2.SSS2.p5.1)\.
- \[42\]F\. J\. Narcowich, J\. D\. Ward, and H\. Wendland\(2005\)Sobolev bounds on functions with scattered zeros, with applications to radial basis function surface fitting\.Mathematics of Computation74\(250\),pp\. 743–763\.External Links:[Document](https://dx.doi.org/10.1090/S0025-5718-04-01708-9)Cited by:[§3\.2](https://arxiv.org/html/2608.25084#S3.SS2.p4.2.1)\.
- \[43\]R\. Opfer\(2006\)Multiscale kernels\.Advances in Computational Mathematics25,pp\. 357–380\.External Links:[Document](https://dx.doi.org/10.1007/s10444-004-7622-3)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p6.1)\.
- \[44\]R\. Opfer\(2006\)Tight frame expansions of multiscale reproducing kernels in sobolev spaces\.Applied and Computational Harmonic Analysis20,pp\. 357–374\.External Links:[Document](https://dx.doi.org/10.1016/j.acha.2005.05.003)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p6.1)\.
- \[45\]B\. Peherstorfer, K\. Willcox, and M\. Gunzburger\(2018\)Survey of multifidelity methods in uncertainty propagation, inference, and optimization\.SIAM Review60\(3\),pp\. 550–591\.External Links:[Document](https://dx.doi.org/10.1137/16M1082469)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[46\]A\. Rudikov, V\. Fanaskov, S\. Stepanov, B\. Shan, E\. Muravleva, Y\. Efendiev, and I\. Oseledets\(2025\)Locally subspace\-informed neural operators for efficient multiscale PDE solving\.arXiv preprint arXiv:2505\.16030\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.16030)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[47\]J\. Sacks, W\. J\. Welch, T\. J\. Mitchell, and H\. P\. Wynn\(1989\)Design and analysis of computer experiments\.Statistical Science4\(4\),pp\. 409–423\.External Links:[Document](https://dx.doi.org/10.1214/ss/1177012413)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[48\]R\. Schaback and H\. Wendland\(2006\)Kernel techniques: from machine learning to meshless methods\.Acta Numerica15,pp\. 543–639\.External Links:[Document](https://dx.doi.org/10.1017/S0962492906270016)Cited by:[§3\.2](https://arxiv.org/html/2608.25084#S3.SS2.p2.2),[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p5.3)\.
- \[49\]R\. Schaback\(1999\)Improved error bounds for scattered data interpolation by radial basis functions\.Mathematics of Computation68\(225\),pp\. 201–216\.External Links:[Document](https://dx.doi.org/10.1090/S0025-5718-99-01009-1)Cited by:[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p5.1)\.
- \[50\]R\. Sharma, M\. Lowery, H\. Owhadi, and V\. Shankar\(2026\)Fluids you can trust: property\-preserving operator learning for incompressible flows\.Journal of Machine Learning Research27,pp\. 1–59\.Cited by:[§4\.2\.1](https://arxiv.org/html/2608.25084#S4.SS2.SSS1.p2.1),[§4\.2\.1](https://arxiv.org/html/2608.25084#S4.SS2.SSS1.p5.1),[§4\.2\.2](https://arxiv.org/html/2608.25084#S4.SS2.SSS2.p2.1),[§4\.2\.2](https://arxiv.org/html/2608.25084#S4.SS2.SSS2.p3.1)\.
- \[51\]I\. H\. Sloan and V\. Kaarnioja\(2025\)Doubling the rate: improved error bounds for orthogonal projection with application to interpolation\.BIT Numerical Mathematics65\.Note:Article 10External Links:[Document](https://dx.doi.org/10.1007/s10543-024-01049-2)Cited by:[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p5.1)\.
- \[52\]M\. Takamoto, T\. Praditia, R\. Leiteritz, D\. MacKinlay, F\. Alesiani, D\. Pflüger, and M\. Niepert\(2022\)PDEBench: an extensive benchmark for scientific machine learning\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 1596–1611\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/0a9747136d411fb83f0cf81820d44afb-Abstract-Datasets_and_Benchmarks.html)Cited by:[§4\.2\.2](https://arxiv.org/html/2608.25084#S4.SS2.SSS2.p4.1)\.
- \[53\]L\. N\. Trefethen\(2000\)Spectral methods in MATLAB\.Society for Industrial and Applied Mathematics\.External Links:[Document](https://dx.doi.org/10.1137/1.9780898719598)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1)\.
- \[54\]T\. Tripura and S\. Chakraborty\(2023\)Wavelet neural operator for solving parametric partial differential equations in computational mechanics problems\.Computer Methods in Applied Mechanics and Engineering404,pp\. 115783\.External Links:[Document](https://dx.doi.org/10.1016/j.cma.2022.115783)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
- \[55\]J\. Turnage, M\. Lowery, J\. Jakeman, Z\. Morrow, A\. Narayan, and V\. Shankar\(2025\)An optimal weighted least\-squares method for operator learning\.arXiv preprint arXiv:2512\.11168\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.11168)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p2.1)\.
- \[56\]H\. Wendland\(1995\)Piecewise polynomial, positive definite and compactly supported radial functions of minimal degree\.Advances in Computational Mathematics4\(1\),pp\. 389–396\.External Links:[Document](https://dx.doi.org/10.1007/BF02123482)Cited by:[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.p2.2),[§3\.1](https://arxiv.org/html/2608.25084#S3.SS1.p3.2.1)\.
- \[57\]H\. Wendland\(1998\)Error estimates for interpolation by compactly supported radial basis functions of minimal degree\.Journal of Approximation Theory93\(2\),pp\. 258–272\.External Links:[Document](https://dx.doi.org/10.1006/jath.1997.3137)Cited by:[§3\.2](https://arxiv.org/html/2608.25084#S3.SS2.p2.2)\.
- \[58\]H\. Wendland\(2005\)Scattered data approximation\.Cambridge University Press\.External Links:[Document](https://dx.doi.org/10.1017/CBO9780511617539)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS0.Px1.p1.3),[§2\.1\.1](https://arxiv.org/html/2608.25084#S2.SS1.SSS1.Px1.p1.3),[§2\.1](https://arxiv.org/html/2608.25084#S2.SS1.p2.2),[§3\.1](https://arxiv.org/html/2608.25084#S3.SS1.p3.2.1),[§3\.2](https://arxiv.org/html/2608.25084#S3.SS2.p2.2),[§3\.2](https://arxiv.org/html/2608.25084#S3.SS2.p4.2.1),[§3\.3](https://arxiv.org/html/2608.25084#S3.SS3.p5.3)\.
- \[59\]H\. Wendland\(2010\)Multiscale analysis in sobolev spaces on bounded domains\.Numerische Mathematik116,pp\. 493–517\.External Links:[Document](https://dx.doi.org/10.1007/s00211-010-0313-8)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p6.1)\.
- \[60\]H\. Wu, H\. Luo, H\. Wang, J\. Wang, and M\. Long\(2024\)Transolver: a fast transformer solver for PDEs on general geometries\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 53681–53705\.Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p1.1)\.
- \[61\]Z\. You, Z\. Xu, and W\. Cai\(2026\)MscaleFNO: multi\-scale fourier neural operator learning for oscillatory functions and wave scattering problems\.Journal of Computational Physics547,pp\. 114530\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2025.114530)Cited by:[§1](https://arxiv.org/html/2608.25084#S1.p3.1)\.
## Appendix ATarget functions




Figure 9:Synthetic test functions used throughout the numerical experiments\. The top row shows theC2C^\{2\}function in two and three dimensions, while the bottom row shows theCωC^\{\\omega\}function in two and three dimensions\.Figure[9](https://arxiv.org/html/2608.25084#A1.F9)shows the target functions used for testing numerical convergence rates\.
## Appendix BSelection of Kernel Support Size
The kernel support size determines the locality of the multiscale frame and directly influences the sparsity of the design matrix\. Rather than varying the support radius directly, we vary the design matrix density, which provides a domain\-independent measure of the average kernel support\. Figures[10](https://arxiv.org/html/2608.25084#A2.F10)and[11](https://arxiv.org/html/2608.25084#A2.F11)illustrate the tradeoff between reconstruction accuracy at the data sites and at unseen evaluation sites\.
At low densities, the kernel supports are highly localized, with few points to fit in a given kernel, producing near\-exact reconstruction at the data sites but relatively poor generalization\. Increasing the density enlarges the kernel supports, allowing neighboring kernels to overlap more substantially and reducing the evaluation\-site error\.
For the analytic function, the evaluation\-site error decreases by more than five orders of magnitude as the density increases\. TheC2C^\{2\}function exhibits the same overall trend, although the tradeoff is more pronounced\. As the density approaches one, the kernels become nearly global, effectively eliminating the locality and multiscale structure of the frame\. Although this dense representation achieves excellent reconstruction accuracy, it sacrifices the sparsity and computational advantages of the multiscale formulation\. These results motivate the use of a design matrix density of20%20\\%throughout the experiments presented in the main text, as it provides a favorable balance between reconstruction accuracy, sparsity, and preservation of the multiscale structure\.
Figure 10:Reconstruction accuracy as a function of design matrix density for the test families in two dimensions\. Results are shown for the analytic \(CωC^\{\\omega\}\) and finitely smooth \(C2C^\{2\}\) functions\.Figure 11:Reconstruction accuracy as a function of design matrix density for the test families in three dimensions\. Results are shown for the analytic \(CωC^\{\\omega\}\) and finitely smooth \(C2C^\{2\}\) functions\.
## Appendix CInterpolation Plots
Figure 12:ℓ2\\ell\_\{2\}error at the decomposition \(data\) sites as a function of the number of decomposition points while holding the design matrix density at approximately0\.20\.2\(2D\)\.Figure 13:ℓ2\\ell\_\{2\}error at the decomposition \(data\) sites as a function of the number of decomposition points while holding the design matrix density at approximately0\.20\.2\(3D\)\.Figure 14:ℓ2\\ell\_\{2\}error as a function of the number of points in the domain\. The kernel support size \(ρ\\rho\) is chosen on the finest scale to produce a design matrix density of approximately0\.20\.2, and this value is then held fixed as the number of points changes\.Figure 15:ℓ2\\ell\_\{2\}error as a function of the number of points in the domain\. The kernel support size \(ρ\\rho\) is chosen on the finest scale to produce a design matrix density of approximately0\.20\.2, and this value is then held fixed as the number of points changes\.These results show interpolation errors withδ\\deltafixed at 20% for the constant density case, andρ\\rhochosen to achieve that density on the finest scale for the case of constant support radius\. As shown in Figures[12](https://arxiv.org/html/2608.25084#A3.F12)and[14](https://arxiv.org/html/2608.25084#A3.F14), the reconstruction error at the 2D data sites increases slightly asMMgrows, but stays close to machine precision, validating our theoretical result\. Sinceδ\\deltaandρ\\rhoare held constant, increasingMMalso increases the average number of neighboring points contained within each kernel\. Figures[13](https://arxiv.org/html/2608.25084#A3.F13)and[15](https://arxiv.org/html/2608.25084#A3.F15)show the trend holding in 3D as well\.
## Appendix DOperator Learning Multiscale Plots
We show multiscale decompositions \(upon generalization\) for some of the other test problems in our operator learning benchmark suite\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 16:Multiscale decomposition of the laminar cylinder flow problem at three different scales\. The bottom right image shows the combined reconstruction\.##### Cylinder Flow \(Laminar\)
In Figure[16](https://arxiv.org/html/2608.25084#A4.F16), the coarse scale captures the dominant wake structure downstream of the cylinder; the medium scale highlights localized wake disturbances and flow deflection caused by the cylinder; and the fine scale isolates the smallest flow features concentrated near the cylinder surface\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 17:Multiscale decomposition of the lid\-driven cavity flow problem at three different scales\. The bottom right image shows the combined reconstruction\.
##### Cavity Flow
In Figure[17](https://arxiv.org/html/2608.25084#A4.F17), the coarse scale captures the dominant behavior throughout most of the domain, while the medium and fine scales resolve increasingly localized features concentrated near the top boundary and the upper corners\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 18:Multiscale decomposition of the piecewise\-constant Darcy flow problem at three different scales\. The bottom right image shows the combined reconstruction\.
##### Darcy PWC
In Figure[18](https://arxiv.org/html/2608.25084#A4.F18), the coarse scale captures the overall smooth global trend trend of the solution, while the intermediate and fine scales progressively refine its shape by resolving the localized features needed to accurately reproduce the full solution\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 19:Multiscale decomposition of the triangular\-domain Darcy flow problem at three different scales\. The bottom right image shows the combined reconstruction\.
##### Darcy Triangle
In Figure[19](https://arxiv.org/html/2608.25084#A4.F19), the coarse scale captures the dominant behavior throughout the interior of the domain, while the medium scale resolves localized boundary effects and emerging interior features\. The fine scale further refines these localized structures, capturing the smallest variations near the boundary and throughout the interior\.
\(a\)Fine
\(b\)Medium
\(c\)Coarse
\(d\)Combined
Figure 20:Multiscale decomposition of the species transport velocities in thexxdirection at three different scales\. The bottom right image shows the combined reconstruction\.
##### Species Transport Alongxx
In Figure[20](https://arxiv.org/html/2608.25084#A4.F20), the coarse scale captures only weak large\-scale variations, while most of the solution structure is resolved at the finest scale\. This behavior reflects the relatively small extent of the flow in thexx\-direction, for which the coarse scale support is comparatively large\. Despite this, the finer scales clearly capture the localized mixing induced by the blades as the gases are transported through the pipe\.
\(a\)Fine
\(b\)Medium
\(c\)Coarse
##### Species transport alongyy

\(d\)Combined
Figure 21:Multiscale decomposition of the species transport velocities in theyydirection at three different scales\. The bottom right image shows the combined reconstruction
##### Species Transport Alongyy
In Figure[21](https://arxiv.org/html/2608.25084#A4.F21), the coarse scale captures only a few broad disturbances, primarily along the center of the pipe where the largest spatial scales are present\. The intermediate and fine scales resolve the flow induced by the blades, capturing the localized mixing as the gases are driven upward from the bottom of the pipe and downward from the top, producing the characteristic rotational motion within the cross\-section\.Similar Articles
Neural means and kernel corrections for operator learning
This paper presents a method combining neural network means with exact Matérn kernel corrections for operator learning in PDEs, achieving competitive or improved performance on public benchmarks like structural mechanics and OCO-2 radiative transfer emulation.
PIKFNO: An Interpretable Neural Operator Based on Physics Informed Kernel Function
PIKFNO is a new interpretable neural operator framework that integrates physics-informed kernel functions from governing equations to enhance predictive accuracy and interpretability with limited training data.
Physics-Integrated Operator Learning via Gaussian Splatting Representations
This paper introduces a physics-integrated operator learning framework using a feed-forward Gaussian splatting representation to directly incorporate PDE operators, reducing errors in long-horizon autoregressive predictions for spatiotemporal systems.
FRAME: Learning the Adaptation Domain with a Mixture of Fractional-Fourier Experts
Proposes FRAME, a mixture-of-experts adapter that uses learnable fractional-Fourier orders to interpolate between spatial and spectral domains, improving parameter-efficient fine-tuning performance on LLMs across multiple benchmarks.
A Trainable-by-Parts Operator Learning Framework: Bridging DeepONet and Karhunen-Loeve Expansions for Large-Scale Applications
Proposes KL-DNN, a scalable operator learning framework that uses Karhunen-Loève expansions to handle large-scale PDE problems, achieving lower errors and two-order-of-magnitude speedup over DeepONet on a 3D carbon storage problem.