Shared Physics Responses Recover Hidden Rankings in Neural Operator Libraries

arXiv cs.LG Papers

Summary

This paper presents a method to rank neural operator models during deployment using shared physics responses, achieving high accuracy without ground-truth reference solutions for scientific computing applications.

arXiv:2608.20441v1 Announce Type: new Abstract: Selecting the optimal neural-operator prediction during deployment is challenging when high-fidelity reference solutions are unavailable. We demonstrate that under a squared Hilbert-space loss, ranking a finite model library depends strictly on the low-dimensional span of candidate differences, allowing us to score all models simultaneously using a single anchor-based linearized response of the governing equation. This shared physical diagnostic accurately recovered over 99.6\% of pairwise preferences and 99.0\% of optimal checkpoints across diverse Fourier and convolutional operator libraries for fluid, reaction-diffusion, and wave dynamics. Furthermore, the corrected physical proxy frequently outperformed the best individual candidates, and we establish computable sufficient conditions that rigorously certify exact decisions for strongly monotone discretizations. By exploiting the local dynamical response rather than raw defect magnitude, this framework enables the reliable and highly efficient deployment of scientific surrogates without requiring ground-truth data.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:29 AM

# Shared Physics Responses Recover Hidden Rankings in Neural Operator Libraries
Source: [https://arxiv.org/html/2608.20441](https://arxiv.org/html/2608.20441)
Hanbing LiangAffiliation:Nanophotonics and Biophotonics Key Laboratory of Jilin Province, School of Physics, Changchun University of Science and Technology, Changchun 130022, P\.R\. ChinaFujun LiuAffiliation:Nanophotonics and Biophotonics Key Laboratory of Jilin Province, School of Physics, Changchun University of Science and Technology, Changchun 130022, P\.R\. ChinaAffiliation:Correspondence:[fjliu@cust\.edu\.cn](mailto:[email protected])

###### Abstract

Selecting the optimal neural\-operator prediction during deployment is challenging when high\-fidelity reference solutions are unavailable\. We demonstrate that under a squared Hilbert\-space loss, ranking a finite model library depends strictly on the low\-dimensional span of candidate differences, allowing us to score all models simultaneously using a single anchor\-based linearized response of the governing equation\. This shared physical diagnostic accurately recovered over 99\.6% of pairwise preferences and 99\.0% of optimal checkpoints across diverse Fourier and convolutional operator libraries for fluid, reaction\-diffusion, and wave dynamics\. Furthermore, the corrected physical proxy frequently outperformed the best individual candidates, and we establish computable sufficient conditions that rigorously certify exact decisions for strongly monotone discretizations\. By exploiting the local dynamical response rather than raw defect magnitude, this framework enables the reliable and highly efficient deployment of scientific surrogates without requiring ground\-truth data\.

## 1Introduction

Partial differential equations \(PDEs\) govern fundamental physical phenomena, yet relying exclusively on high\-fidelity numerical solvers imposes severe computational bottlenecks\. Machine\-learning surrogates mitigate this burden by learning mappings from problem inputs to solution fields across diverse families of physical systems\[[1](https://arxiv.org/html/2608.20441#bib.bib1),[2](https://arxiv.org/html/2608.20441#bib.bib2),[3](https://arxiv.org/html/2608.20441#bib.bib3),[4](https://arxiv.org/html/2608.20441#bib.bib4)\]\. While different architectures are continuously evaluated across public benchmarks\[[5](https://arxiv.org/html/2608.20441#bib.bib5),[6](https://arxiv.org/html/2608.20441#bib.bib6),[7](https://arxiv.org/html/2608.20441#bib.bib7)\], a practical scientific workflow often generates multiple plausible predictions from various models, training runs, or checkpoints for the exact same input\. Because the optimal candidate frequently changes depending on the specific physical input and the scientific quantity of interest, practitioners face a fundamental deployment dilemma\.

During benchmarking, an available high\-fidelity solution effortlessly ranks these competing candidates\. During active deployment, however, this reference solution is inherently unavailable, and computing it merely to select a surrogate would entirely forfeit the computational advantage of using neural operators\. The central question is therefore whether the governing physics can reliably recover the hidden ranking of a finite candidate library without first computing the missing reference\.

Addressing this question requires recognizing how a finite candidate library alters the necessary information landscape\. Under a fixed squared Hilbert\-space task loss, subtracting the losses of two candidates eliminates every component of the unknown solution that cannot distinguish between them\. All pairwise preferences subsequently depend exclusively on the span of candidate differences, whose dimension is at most the library size minus one\. This mathematical reduction establishes that choosing among a finite set requires significantly less information than reconstructing an arbitrary physical solution, meaning a physical diagnostic only needs to recover this low\-dimensional decision coordinate\.

Existing diagnostic tools target related but fundamentally distinct objectives\. Ground\-truth\-free model selection and validation errors typically rely on historical examples or ensemble statistics rather than reconstructing the specific ordering induced by the governing physics for the current input\[[8](https://arxiv.org/html/2608.20441#bib.bib8),[9](https://arxiv.org/html/2608.20441#bib.bib9),[10](https://arxiv.org/html/2608.20441#bib.bib10)\]\. Furthermore, while the raw equation residual measures a physical defect, its magnitude alone fails to capture how that defect propagates into the requested output\. Physics\-informed correction methods use the governing equations to improve individual predictions\[[11](https://arxiv.org/html/2608.20441#bib.bib11),[12](https://arxiv.org/html/2608.20441#bib.bib12),[13](https://arxiv.org/html/2608.20441#bib.bib13),[14](https://arxiv.org/html/2608.20441#bib.bib14),[15](https://arxiv.org/html/2608.20441#bib.bib15),[16](https://arxiv.org/html/2608.20441#bib.bib16)\], while a posteriori estimators translate residuals into global error bounds\[[17](https://arxiv.org/html/2608.20441#bib.bib17),[18](https://arxiv.org/html/2608.20441#bib.bib18),[19](https://arxiv.org/html/2608.20441#bib.bib19),[20](https://arxiv.org/html/2608.20441#bib.bib20),[21](https://arxiv.org/html/2608.20441#bib.bib21),[22](https://arxiv.org/html/2608.20441#bib.bib22)\], and complementary verification methods place rigorous bounds over the continuous domain\[[23](https://arxiv.org/html/2608.20441#bib.bib23)\]\. In contrast, our objective is the relative ordering of a fixed finite library\. By leveraging the candidate\-difference quotient, our framework allows a single anchor\-derived response to be reused across the entire library, conditioning the scores directly on the current physical input and evaluating all contrasts simultaneously\.

We formulate this paradigm as reference\-free finite\-library ranking, where “reference\-free” specifically denotes that the scoring process utilizes no deployment reference solution or reference\-derived labels\. If the hidden high\-fidelity solution were accessible, the chosen scientific loss would cleanly induce pairwise preferences and identify the best candidate\. Our goal is to recover those precise decisions relying solely on the input, the candidate predictions, and the governing PDE system\. This operational challenge spans two distinct deployment scales: query\-time routing to select an optimal candidate for a single input, and panel\-level aggregation to freeze a specific checkpoint for repeated future use\.

To solve this ranking problem, our approach distills these decisions into a single shared physics calculation\. The candidate library first defines a shared library anchor, at which we evaluate the complete physical defect before propagating it once through a linearized physics solve\. Mapping this corrected state to the scientific task yields a shared task proxy, denoted byqq, which subsequently estimates the hidden reference ordering by measuring its distance to every candidate\. Although the numerical implementation computes a full\-state correction, our theory explicitly identifies the smaller projection of this response that dictates library decisions\. Unlike candidate\-wise correction strategies, this shared framework avoids solving a separate physics problem for every prediction while remaining deeply task\-aware\.

We theoretically prove that this shared comparison coordinate is both sufficient and minimal among bounded linear observations for globally exact ordering, linking it to the physics response through an exact conditional error identity\. Because ranking reliability intrinsically depends on the candidate margin, we demonstrate that a decision is trustworthy when the remaining physical uncertainty is strictly bounded by the separation between candidates\. For a strongly monotone reaction\-diffusion operator, this geometric principle translates into computable sufficient conditions that rigorously certify same\-grid pairwise or top\-1 decisions\. We validate this framework comprehensively across separately trained Fourier and convolutional neural\-operator libraries for Burgers and reaction\-diffusion equations under controlled input shifts\. Notably, in an operator\-mismatched two\-dimensional compressible\-flow probe, the shared proxy recovered all 10/10 candidate comparisons although it outperformed the best neural\-operator prediction in only 4/10 cases, showing that accurate ranking can persist without superior reconstruction or exact operator–data alignment\. By extending our analysis to Sine\-Gordon dynamics and public PDEBench models, we systematically distinguish ranking from direct proxy use, confirming that governing\-physics responses can successfully replace missing ground truth in scientific surrogate deployment\.

## 2Results

When several surrogate predictions disagree, deployment provides the candidate fields and governing equation but not the high\-fidelity solution needed to rank them\. A finite library changes what must be recovered from that missing solution\. Under a fixed squared Hilbert\-space task loss, only target directions spanned by candidate differences can change their relative errors\. This comparison space has dimension at most the library size minus one, even when the physical state is very high\-dimensional\. Ranking can therefore require much less information than reconstructing the hidden solution\.

We exploit this reduction by constructing a library anchorwwfrom the candidates and applying one linearized physics correction\. Mapping the corrected state to the scientific task gives a proxyqq, which is compared with every candidate output \(Fig\.[1](https://arxiv.org/html/2608.20441#S2.F1)a–c\)\. These shared scores select a prediction for the current input and can also be averaged over an unlabeled panel to rank reusable models\.

![Refer to caption](https://arxiv.org/html/2608.20441v1/Fig1.png)Figure 1:One shared physics response turns a finite surrogate library into a reference\-free decision\.a, Candidate statesv1,…,vNv\_\{1\},\\ldots,v\_\{N\}are aggregated by a reference\-free rule𝒞\\mathcal\{C\}into the library anchorww\.b, One linearized physics response gives the corrected states=w\+δws=w\+\\delta\_\{w\}, which the task mapQaQ\_\{a\}converts to the proxyqq\.c, DistancesSi​\(q\)=∥q−Qa​\(vi\)∥M2S\_\{i\}\(q\)=\\lVert q\-Q\_\{a\}\(v\_\{i\}\)\\rVert\_\{M\}^\{2\}rank the candidates and selectvk⋆v\_\{k^\{\\star\}\}; the dashed branch denotes optional direct use ofqq\.Theorem[1](https://arxiv.org/html/2608.20441#Thmtheorem1)below makes the information reduction exact\. The remaining questions are whether one physical response can recover the required coordinate in learned surrogate libraries, which features of the library make recovery reliable, and when the resulting choice can be certified without revealing the reference\.

### 2\.1Ranking a finite library requires less information than solution recovery

The framework treats ranking as its primary output rather than as a by\-product of reconstructing the hidden solution\. This distinction is exact for a fixed library: only the parts of the target that change the candidates’ relative errors can affect their ordering\. We now formalize that smaller information requirement\.

Let\(ℋ,⟨⋅,⋅⟩M\)\(\\mathcal\{H\},\\left\\langle\\cdot,\\cdot\\right\\rangle\_\{M\}\)be a real Hilbert task space, lety1,…,yN∈ℋy\_\{1\},\\ldots,y\_\{N\}\\in\\mathcal\{H\}be fixed candidate outputs and lett∈ℋt\\in\\mathcal\{H\}denote the unavailable task truth\. Candidateiihas squared loss

Ri​\(t\)=‖t−yi‖M2\.R\_\{i\}\(t\)=\\left\\lVert t\-y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\.\(1\)For an ordered pair, defineGi​j​\(t\)=Rj​\(t\)−Ri​\(t\)G\_\{ij\}\(t\)=R\_\{j\}\(t\)\-R\_\{i\}\(t\), so thatGi​j​\(t\)\>0G\_\{ij\}\(t\)\>0favours candidateii\. Expanding the two losses gives

Gi​j​\(t\)=‖yj‖M2−‖yi‖M2\+2​⟨t,yi−yj⟩M\.G\_\{ij\}\(t\)=\\left\\lVert y\_\{j\}\\right\\rVert\_\{M\}^\{2\}\-\\left\\lVert y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\+2\\left\\langle t,y\_\{i\}\-y\_\{j\}\\right\\rangle\_\{M\}\.\(2\)Thus the hidden target enters every pairwise decision only through inner products with candidate differences\. Define

𝒟=span⁡\{yi−y1:2≤i≤N\}\.\\mathcal\{D\}=\\operatorname\{span\}\\\{y\_\{i\}\-y\_\{1\}:2\\leq i\\leq N\\\}\.\(3\)Targets that differ by an element of𝒟⟂M\\mathcal\{D\}^\{\\perp\_\{M\}\}belong to the same decision class; the quotientℋ/𝒟⟂M\\mathcal\{H\}/\\mathcal\{D\}^\{\\perp\_\{M\}\}is canonically represented byP𝒟M​t∈𝒟P\_\{\\mathcal\{D\}\}^\{M\}t\\in\\mathcal\{D\}\. The same collapse extends to general bounded linear observations\.

###### Theorem 1\(Finite\-library decision quotient and minimal bounded\-linear observation\)\.

The coordinateP𝒟M​tP\_\{\\mathcal\{D\}\}^\{M\}tdetermines every exact pair gap, every three\-way pair label \(candidateii, tie or candidatejj\) and the complete weak ordering, anddim𝒟≤N−1\\dim\\mathcal\{D\}\\leq N\-1\. It also determines top\-1 after fixing a candidate\-index tie rule that depends only on the minimizer set\.

LetWWbe a real normed space and letT:ℋ→WT:\\mathcal\{H\}\\to Wbe bounded and linear\. The exact gap vector, the three\-way pair\-label vector and the complete weak ordering are determined byT​tTtfor everyt∈ℋt\\in\\mathcal\{H\}if and only if

kerT⊆𝒟⟂M\.\\ker T\\subseteq\\mathcal\{D\}^\{\\perp\_\{M\}\}\.\(4\)Letτ\\tauchoose an elementτ⁡\(S\)∈S\\tau\(S\)\\in Sfrom every nonempty candidate subsetSSand defineιτ​\(t\)=τ⁡\(arg⁡mini​Ri​\(t\)\)\\iota\_\{\\tau\}\(t\)=\\tau\(\\arg\\min\_\{i\}R\_\{i\}\(t\)\)\. For this fixed minimizer\-set tie rule, the same equivalence holds for globally exact recovery ofιτ​\(t\)\\iota\_\{\\tau\}\(t\)when targets range over all ofℋ\\mathcal\{H\}\. IfWWis finite\-dimensional, any such observation satisfies

dimW≥dim𝒟\.\\dim W\\geq\\dim\\mathcal\{D\}\.\(5\)

###### Proof\.

Equation \([2](https://arxiv.org/html/2608.20441#S2.E2)\) proves sufficiency\. For necessity, suppose thath∈ker⁡Th\\in\\ker Tbuth∉𝒟⟂Mh\\notin\\mathcal\{D\}^\{\\perp\_\{M\}\}\. At least one candidate difference has nonzero inner product withhh\. Perturbing the midpoint of that pair by±ε​h\\pm\\varepsilon hproduces opposite strict signs with the same observation, which rules out exact sign recovery and hence exact gap or complete\-order recovery\. For top\-1, setai=⟨h,yi⟩Ma\_\{i\}=\\left\\langle h,y\_\{i\}\\right\\rangle\_\{M\}\. Theaia\_\{i\}are not all equal\. For sufficiently large positiveα\\alpha, every minimizer att=α​ht=\\alpha hlies in the maximum\-projection candidate set; for sufficiently large negativeα\\alpha, every minimizer lies in the disjoint minimum\-projection set\. Both targets have the same observation\. ThereforeT\|𝒟T\|\_\{\\mathcal\{D\}\}must be injective, which also gives Eq\. \([5](https://arxiv.org/html/2608.20441#S2.E5)\)\. ∎

Theorem[1](https://arxiv.org/html/2608.20441#Thmtheorem1)establishes minimality among bounded linear observations for globally exact decisions over all targets\. On a restricted physical solution set, some candidate boundaries may never be exposed, and exact gaps, pair signs and top\-1 can require different amounts of information\. The corresponding exposure conditions, duplicate\-candidate cases and a continuous\-statistic dimension result are given in Supplementary Section S1\.

The theorem isolates the target information that a library decision can use\. A physical solver may still return a full\-state response, but the library reads only its comparison projection\. The difficulty of a choice is then governed by the margin between competing candidates\. The empirical libraries contained both ambiguous and well\-separated decisions: among 1,920 input–library combinations, 336 had a best–runner\-up gap no larger than10−610^\{\-6\}, whereas the remaining gaps spanned a broad range \(Supplementary Section S9\)\.

Equation \([2](https://arxiv.org/html/2608.20441#S2.E2)\) also identifies the relevant notion of error\. A large state error orthogonal to𝒟\\mathcal\{D\}is invisible to every squared\-loss decision, whereas a small error aligned with a near\-tie candidate difference can reverse a ranking\. The next question is therefore whether one reference\-free physical response can recover this decision coordinate relative to the margins found in actual libraries\.

### 2\.2A shared physics response recovers hidden neural\-operator choices

Theorem[1](https://arxiv.org/html/2608.20441#Thmtheorem1)identifies what must be known to rank a finite library\. We next construct a shared physical response and show how its error enters precisely that comparison coordinate\. For a physical inputaa, letv1,…,vNv\_\{1\},\\ldots,v\_\{N\}be the candidate states\. A reference\-free library rule𝒞\\mathcal\{C\}defines an anchor

w=𝒞⁡\(v1,…,vN\)\.w=\\mathcal\{C\}\(v\_\{1\},\\ldots,v\_\{N\}\)\.\(6\)Let𝒜a​\(v\)=0\\mathcal\{A\}\_\{a\}\(v\)=0denote the assembled discrete physical problem, including its initial and boundary conditions\. WithLw=D​𝒜a​\(w\)L\_\{w\}=D\\mathcal\{A\}\_\{a\}\(w\), we compute one inexact shared response

Lw​δw=−𝒜a​\(w\)\+ℓ,L\_\{w\}\\delta\_\{w\}=\-\\mathcal\{A\}\_\{a\}\(w\)\+\\ell,\(7\)whereℓ\\ellrecords the algebraic solve defect\. The corrected state iss=w\+δws=w\+\\delta\_\{w\}\. For the principal linear tasks,q=Qa​\(s\)q=Q\_\{a\}\(s\)andyi=Qa​\(vi\)y\_\{i\}=Q\_\{a\}\(v\_\{i\}\); every candidate is scored against the same task proxy,

Si​\(q\)=‖q−yi‖M2\.S\_\{i\}\(q\)=\\left\\lVert q\-y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\.\(8\)No numerical reference enters Eqs\. \([6](https://arxiv.org/html/2608.20441#S2.E6)\)–\([8](https://arxiv.org/html/2608.20441#S2.E8)\)\.

The role of the shared solve is given by an exact conditional representation\. Letuusolve the governing equation, sete=u−we=u\-w, and define

ℛA​\(e\)=𝒜a​\(w\+e\)−𝒜a​\(w\)−Lw​e\.\\mathcal\{R\}\_\{A\}\(e\)=\\mathcal\{A\}\_\{a\}\(w\+e\)\-\\mathcal\{A\}\_\{a\}\(w\)\-L\_\{w\}e\.\(9\)
###### Proposition 1\(Conditional physics\-to\-decision error representation\)\.

LetXX,YYandℋ\\mathcal\{H\}be real Hilbert spaces, let𝒜a:Ω⊂X→Y\\mathcal\{A\}\_\{a\}:\\Omega\\subset X\\to YandQa:Ω→ℋQ\_\{a\}:\\Omega\\to\\mathcal\{H\}, and suppose𝒜a​\(u\)=0\\mathcal\{A\}\_\{a\}\(u\)=0\. Assume that𝒜a\\mathcal\{A\}\_\{a\}andQaQ\_\{a\}are Fréchet differentiable where used, thatLw:X→YL\_\{w\}:X\\to Yis a bounded isomorphism, and that the required line segments lie inΩ\\Omega\. Then

δw−e=Lw−1​\[ℓ\+ℛA​\(e\)\]\.\\delta\_\{w\}\-e=L\_\{w\}^\{\-1\}\\bigl\[\\ell\+\\mathcal\{R\}\_\{A\}\(e\)\\bigr\]\.\(10\)IfQaQ\_\{a\}is nonlinear, define

qlin=Qa​\(w\)\+D​Qa​\(w\)​δw,qeval=Qa​\(w\+δw\),q\_\{\\rm lin\}=Q\_\{a\}\(w\)\+DQ\_\{a\}\(w\)\\delta\_\{w\},\\qquad q\_\{\\rm eval\}=Q\_\{a\}\(w\+\\delta\_\{w\}\),\(11\)andℛQ​\(h\)=Qa​\(w\+h\)−Qa​\(w\)−D​Qa​\(w\)​h\\mathcal\{R\}\_\{Q\}\(h\)=Q\_\{a\}\(w\+h\)\-Q\_\{a\}\(w\)\-DQ\_\{a\}\(w\)h\. Then

P𝒟M​\(qlin−t\)\\displaystyle P\_\{\\mathcal\{D\}\}^\{M\}\(q\_\{\\rm lin\}\-t\)=P𝒟M​D​Qa​\(w\)​Lw−1​\[ℓ\+ℛA​\(e\)\]−P𝒟M​ℛQ​\(e\),\\displaystyle=P\_\{\\mathcal\{D\}\}^\{M\}DQ\_\{a\}\(w\)L\_\{w\}^\{\-1\}\\bigl\[\\ell\+\\mathcal\{R\}\_\{A\}\(e\)\\bigr\]\-P\_\{\\mathcal\{D\}\}^\{M\}\\mathcal\{R\}\_\{Q\}\(e\),\(12\)P𝒟M​\(qeval−t\)\\displaystyle P\_\{\\mathcal\{D\}\}^\{M\}\(q\_\{\\rm eval\}\-t\)=P𝒟M​D​Qa​\(w\)​Lw−1​\[ℓ\+ℛA​\(e\)\]\+P𝒟M​\[ℛQ​\(δw\)−ℛQ​\(e\)\],\\displaystyle=P\_\{\\mathcal\{D\}\}^\{M\}DQ\_\{a\}\(w\)L\_\{w\}^\{\-1\}\\bigl\[\\ell\+\\mathcal\{R\}\_\{A\}\(e\)\\bigr\]\+P\_\{\\mathcal\{D\}\}^\{M\}\\bigl\[\\mathcal\{R\}\_\{Q\}\(\\delta\_\{w\}\)\-\\mathcal\{R\}\_\{Q\}\(e\)\\bigr\],\(13\)wheret=Qa​\(u\)t=Q\_\{a\}\(u\)\. IfQaQ\_\{a\}is bounded and linear, the two proxies coincide and

P𝒟M​\(q−t\)=P𝒟M​Qa​Lw−1​\[ℓ\+ℛA​\(e\)\]\.P\_\{\\mathcal\{D\}\}^\{M\}\(q\-t\)=P\_\{\\mathcal\{D\}\}^\{M\}Q\_\{a\}L\_\{w\}^\{\-1\}\\bigl\[\\ell\+\\mathcal\{R\}\_\{A\}\(e\)\\bigr\]\.\(14\)If both derivatives are locally Hölder continuous with the same exponent0<α≤10<\\alpha\\leq 1along the required segments, these identities yield local bounds in‖ℓ‖\\\|\\ell\\\|,‖u−w‖1\+α\\\|u\-w\\\|^\{1\+\\alpha\}and, for an evaluated nonlinear task,‖δw‖1\+α\\\|\\delta\_\{w\}\\\|^\{1\+\\alpha\}\. For any resulting task proxyqq,

Gi​j​\(t\)−Gi​j​\(q\)=2​⟨t−q,yi−yj⟩MG\_\{ij\}\(t\)\-G\_\{ij\}\(q\)=2\\left\\langle t\-q,y\_\{i\}\-y\_\{j\}\\right\\rangle\_\{M\}\(15\)remains exact\.

###### Proof\.

Substitute𝒜a​\(u\)=0\\mathcal\{A\}\_\{a\}\(u\)=0into the first\-order expansion atwwand compare it with Eq\. \([7](https://arxiv.org/html/2608.20441#S2.E7)\); this gives Eq\. \([10](https://arxiv.org/html/2608.20441#S2.E10)\)\. ExpandingQaQ\_\{a\}atwwgives Eqs\. \([12](https://arxiv.org/html/2608.20441#S2.E12)\)–\([13](https://arxiv.org/html/2608.20441#S2.E13)\)\. The last identity is Eq\. \([2](https://arxiv.org/html/2608.20441#S2.E2)\) evaluated atttandqq\. ∎

The unprojected identities and their complete algebraic derivation are given in Supplementary Proposition 7\.

These identities explain how the physical correction reaches the decision coordinate\. A computable reference\-free certificate additionally requires a way to bound the anchor\-to\-solution distance\. NonlinearQaQ\_\{a\}does not alter the final squared\-Hilbert comparison geometry; it adds a physics\-to\-task remainder\. A non\-Hilbert or truth\-normalized pair loss changes the comparison geometry itself\.

For linear tasks, the same structure has an adjoint interpretation\. Define

Ji​j​\(x\)=‖Qa​x−yj‖M2−‖Qa​x−yi‖M2\.J\_\{ij\}\(x\)=\\left\\lVert Q\_\{a\}x\-y\_\{j\}\\right\\rVert\_\{M\}^\{2\}\-\\left\\lVert Q\_\{a\}x\-y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\.\(16\)The quadratic terms inQa​xQ\_\{a\}xcancel\. IfLw∗​zi​j=2​Qa∗​\(yi−yj\)L\_\{w\}^\{\*\}z\_\{ij\}=2Q\_\{a\}^\{\*\}\(y\_\{i\}\-y\_\{j\}\)andℓ=0\\ell=0, then the shared solve satisfies

Ji​j​\(w\)−⟨zi​j,𝒜a​\(w\)⟩Y=Ji​j​\(w\+δw\)\.J\_\{ij\}\(w\)\-\\langle z\_\{ij\},\\mathcal\{A\}\_\{a\}\(w\)\\rangle\_\{Y\}=J\_\{ij\}\(w\+\\delta\_\{w\}\)\.\(17\)One corrected primal state therefore realizes the full family of pairwise first\-order adjoint point estimates\. The adjoint machinery is classical; the finite\-library structure is that the candidate set induces this affine family of goals and one shared state evaluates all of them\. Rigorous uncertainty still requires localization or directional support bounds\. For the quadratic Burgers operator, the remainder is exactly bilinear:

𝒜a​\(w\+e\)=𝒜a​\(w\)\+Lw​e\+𝐁⁡\(e,e\),δw−e=Lw−1​𝐁​\(e,e\)\\mathcal\{A\}\_\{a\}\(w\+e\)=\\mathcal\{A\}\_\{a\}\(w\)\+L\_\{w\}e\+\\mathbf\{B\}\(e,e\),\\qquad\\delta\_\{w\}\-e=L\_\{w\}^\{\-1\}\\mathbf\{B\}\(e,e\)\(18\)for an exact solve\. This gives a PDE\-specific local quadratic realization of Proposition[1](https://arxiv.org/html/2608.20441#Thmproposition1)\.

We tested this construction on eight libraries formed by crossing two equations, two architectures and two seeded ensembles, evaluated under controlled input shifts\. One shared response selected a candidate within10−610^\{\-6\}of the best library loss in 99\.01% of cases and recovered 99\.59% of non\-tied pairwise preferences \(Fig\.[2](https://arxiv.org/html/2608.20441#S2.F2)a,b\)\. Recovery remained high in all eight libraries and among the closest candidate comparisons \(Supplementary Section S9\)\. The raw residual recovered 67\.83% of non\-tied pairs and 42\.66% of tolerant top\-1 choices\.

The ranking behaviour extended to Sine–Gordon\. Across eight libraries and 240 inputs, the proxy recovered 95\.57% of exact top\-1 decisions \(Supplementary Table 6 and Supplementary Fig\. 2\)\.

Figure 2:One shared physics response recovers hidden neural\-operator choices\.a, The same eight candidate states are shown in three horizontal lanes, ordered from best to worst by hidden\-reference task error, distance to the shared proxy and assembled\-defect energy\. Lines connect candidate identities across lanes\. Within each frame, the black curve showsviv\_\{i\}on a common state scale, whereas the magenta curve showsvi−wv\_\{i\}\-wabout the grey zero line on a separate magnified scale shared by all candidates\.b, Tolerant top\-1 and non\-tie pairwise recovery in eight libraries\. Thin segments join the raw\-residual and shared\-response results within each library; large markers show pooled recovery with input\-clustered 95% confidence intervals\.The gap geometry in Fig\.[3](https://arxiv.org/html/2608.20441#S2.F3)a shows what supplied the signal\. With the correct response, proxy gaps track hidden\-reference gaps in both sign and magnitude\. Assigning the response to the wrong input disperses this relation, while reversing its direction produces a predominantly reversed gap pattern\. The recovered preference therefore depended on the current physical instance and on propagating its defect in the correct direction, not simply on static library geometry\.

![Refer to caption](https://arxiv.org/html/2608.20441v1/Fig3.png)Figure 3:Shared physics recovers candidate gaps and yields low\-regret choices\.a, Hidden\-reference pair gaps plotted against gaps estimated from the shared, input\-shuffled and reversed responses \(49,323 non\-tie pairs; common normalized axes, bins and density scale\)\. Matching signs indicate correct ordering; the diagonal indicates agreement in gap magnitude\.b, Empirical cumulative distributions of span\-normalized selection regret for 1,920 primary input–library cases\. Zero denotes selection of the hidden\-reference best candidate; curves nearer the upper left are better\.c, The same regret measure for a separate reaction–diffusion population \(1,920 cases\), comparing the shared proxy, raw residual and two candidate\-wise variational majorants; the two majorant curves nearly coincide\.Top\-1 recovery alone masks the severity of incorrect choices\. Across the 1,920 Burgers and reaction–diffusion input–library cases, the shared proxy selected the exact winner in 1,899 cases; when it missed, its normalized regret remained close to zero \(Fig\.[3](https://arxiv.org/html/2608.20441#S2.F3)b\)\. Here regret is the selected candidate’s excess hidden loss divided by the loss span from the best to the worst library member\. The comparison baselines select by assembled\-defect energy \(the raw residual baseline\), distance to the candidate mean or medoid, or input\-independent validation loss; all produced broader regret distributions\.

We also compared the shared response with candidate\-wise error estimators \(Fig\.[3](https://arxiv.org/html/2608.20441#S2.F3)c\)\. A common\-energy dual\-norm estimator and a candidate\-adapted, operator\-aware majorant both gave valid a posteriori bounds and improved pairwise ranking to 82\.7% in the reaction–diffusion cases\. Both nevertheless remained below the shared proxy on the same candidates\. These estimators bound individual state errors, whereas the shared proxy estimates the direction of a candidate comparison\. In this reaction–diffusion comparison, these absolute\-error estimators preserved the pairwise ordering less reliably than the shared task proxy; detailed pairwise, top\-1 and regret results are given in Supplementary Table 18\.

Ranking candidates by the task\-space magnitude of their own linearized corrections produced nearly the same decisions as the shared construction\. In a separate 192\-case control with 16 candidates per library, the shared selector and the candidate\-specific correction\-magnitude selector chose the same checkpoint in 191 cases and had nearly identical pairwise accuracy\. Thus one library\-level response retained almost all of the decision information fromNNcandidate\-specific linearized error proxies while requiring only one correction solve\. The difference in solve count was visible in the component timings: fromN=2N=2toN=16N=16, shared\-response time remained nearly constant, whereas candidate\-wise time grew approximately in proportion to library size\. AtN=16N=16, the candidate\-wise component was 16\.0 times slower for Burgers and 15\.3 times slower for reaction–diffusion \(Supplementary Table 12\)\.

### 2\.3Reliable ranking depends on margins, library composition and admissibility

High average accuracy can hide sensitivity to library composition\. In a reaction–diffusion CNO library, two extreme candidates pulled the arithmetic anchor away from the main cluster and reduced exact top\-1 recovery to 67 of 240 inputs \(Fig\.[4](https://arxiv.org/html/2608.20441#S2.F4)a\)\. For those inputs, a candidate\-valued medoid remained inside the main cluster and recovered all 240 choices, linking the failure to the position of the anchor rather than to an uninformative candidate library\. Across 16 further libraries constructed from separate training, validation and evaluation data, both anchors selected the exact winner in 3,800 of 3,840 cases\. Candidate composition therefore matters, but neither anchor was uniformly better; the anchor rule must suit the intended candidate family\.

![Refer to caption](https://arxiv.org/html/2608.20441v1/Fig4.png)Figure 4:Library composition, candidate separation and admissibility shape reliable recovery\.a, Eight candidate states, their arithmetic mean and medoid on the common signed\-log scalesgn⁡\(u\)​ln⁡\(1\+\|u\|/0\.05\)\\operatorname\{sgn\}\(u\)\\ln\(1\+\|u\|/0\.05\)\. Extreme states displace the mean, whereas the medoid remains candidate\-valued\.b, Non\-tie pair recovery across hidden\-gap deciles for the shared response, raw residual, validation choice, mean and medoid \(49,323 pairs\)\.c, Paired minimum pressures before and after backtracking for 35 shock cases; the horizontal zero line marks physical admissibility\.Reliable ranking therefore requires the shared anchor to produce an informative physical response\. A second condition is geometric: proxy error must be smaller than the separation between competing candidates\. If only

‖P𝒟M​\(t−q\)‖M≤ε\\left\\lVert P\_\{\\mathcal\{D\}\}^\{M\}\(t\-q\)\\right\\rVert\_\{M\}\\leq\\varepsilon\(19\)is known, then Eq\. \([15](https://arxiv.org/html/2608.20441#S2.E15)\) gives

\|Gi​j​\(t\)−Gi​j​\(q\)\|≤2​ε​‖yi−yj‖M\.\|G\_\{ij\}\(t\)\-G\_\{ij\}\(q\)\|\\leq 2\\varepsilon\\left\\lVert y\_\{i\}\-y\_\{j\}\\right\\rVert\_\{M\}\.\(20\)This bound is sharp under ball uncertainty: for a nonzero pair direction, equality is attained when the projected error is parallel or antiparallel to that direction\. Near\-tie decisions are therefore intrinsically fragile\. More generally, for any fixedk\>0k\>0, anO⁡\(ρk\)O\(\\rho^\{k\}\)proxy error asρ↓0\\rho\\downarrow 0, with a constant independent ofρ\\rho, does not guarantee correct ranking when a nonzero true best–runner\-up gap is alsoΘ⁡\(ρk\)\\Theta\(\\rho^\{k\}\)on the same family; the directional reversal construction is given in Supplementary Section S1\. The 49,323 pair decisions follow this margin dependence\. Shared recovery was already 97\.8% in the smallest hidden\-gap decile and approached 100% as gaps widened\. Raw residual improved more slowly, while validation and static\-centre baselines did not acquire the same high\-margin recovery \(Fig\.[4](https://arxiv.org/html/2608.20441#S2.F4)b\)\. The same margin logic applies to algebraic solve error, which must be controlled on the scale of the decision being resolved \(Supplementary Section S1\)\.

The shared response has a second boundary in the one\-dimensional compressible\-flow shock probe\. An undamped update could leave the physically admissible set, whereas positivity and residual backtracking increased the number of feasible cases from 22 of 35 to 35 of 35 \(Fig\.[4](https://arxiv.org/html/2608.20441#S2.F4)c\)\. The global certificate developed below provides a different reliability route when strong monotonicity is available\.

### 2\.4Accurate ranking does not require accurate solution reconstruction

The shared proxy need not reproduce the hidden solution everywhere to rank the candidates correctly\. It only needs to be accurate in directions that distinguish one candidate from another\. This separation explains why strong ranking can coexist with an imperfect corrected prediction, and why a small full\-task error can still flip a near\-tied choice\.

The finite\-library quotient gives an exact orthogonal decomposition of the task error,

t−q=P𝒟M\(t−q\)\+P𝒟⟂MM\(t−q\)\.t\-q=P\_\{\\mathcal\{D\}\}^\{M\}\(t\-q\)\+P\_\{\\mathcal\{D\}^\{\\perp\_\{M\}\}\}^\{M\}\(t\-q\)\.\(21\)Only the first term changes squared\-loss comparisons\. Consequently, when𝒟⟂M\\mathcal\{D\}^\{\\perp\_\{M\}\}is non\-trivial, correct comparison does not imply a uniformly accurate task proxy: the orthogonal component can be arbitrarily large without changing any candidate decision\. Conversely, whenever a candidate pair is distinct, an arbitrarily small full\-task error can cross an arbitrarily small candidate\-bisector margin and reverse a decision\. Full\-task accuracy and decision accuracy are therefore neither equivalent nor ordered without a margin condition\.

This distinction concerns the final loss geometry, not whether the state\-to\-task map is linear\. IfQaQ\_\{a\}is nonlinear butt=Qa​\(u\)t=Q\_\{a\}\(u\), the proxyqqand candidate outputsyi=Qa​\(vi\)y\_\{i\}=Q\_\{a\}\(v\_\{i\}\)all lie in one fixed Hilbert task space, Eq\. \([15](https://arxiv.org/html/2608.20441#S2.E15)\) remains exact\. NonlinearQaQ\_\{a\}changes the remainders that connect a physical state correction toqq, as in Proposition[1](https://arxiv.org/html/2608.20441#Thmproposition1)\. By contrast, a truth\-normalized nRMSE, a generalLpL^\{p\}loss or a nonlinear quantity\-of\-interest loss need not have affine pair contrasts\. Such losses require task\-specific structure or a local tangent and remainder analysis\.

The tested systems exhibited several regimes allowed by this geometry\. Across the public PDEBench population, the correction reduced full\-trajectory error by a factor of 9\.86 and comparison\-subspace error by a factor of 16\.1\. The larger contraction in the comparison subspace shows preferential improvement in decision\-relevant directions \(Supplementary Fig\. 5\)\. In Sine–Gordon, the median weighted full\-trajectory error norm fell by a factor of 1,835\.7\. Here the correction entered a full\-proxy regime:qqwas accurate not only in the comparison subspace but also across the complete task\. The correction paths and task definitions are given in Supplementary Fig\. 4 and Supplementary Section S12\. The task loss ofqqwas also no greater than that of the best candidate in all 1,920 primary Burgers/reaction–diffusion cases and all 99 inputs under the fixed PDEBench terminal squared\-L2L^\{2\}task\. In these populations, ranking and direct improvement coincided\.

Notably, the ranking signal remained accurate under operator–data mismatch\. In the two\-dimensional compressible\-flow probe, the proxy recovered the correct candidate comparison in all ten cases and reduced comparison\-subspace error by a median factor of 22\.2, although direct proxy use improved on the best candidate in only four cases \(Supplementary Fig\. 5\)\. This behaviour directly realizes the separation permitted by the exact geometry:

correct comparison⇏best corrected state\.\\text\{correct comparison\}\\;\\not\\Rightarrow\\;\\text\{best corrected state\}\.
Direct use of a corrected state therefore requires accuracy in the full task space, whereas selection depends only on the candidate\-difference coordinate\.

### 2\.5Shared scores support routing and reusable model selection

Once a hidden\-reference ordering can be estimated, it can be used at different points in a deployment workflow\. At query time, the candidates for the current input can be rescored against the task\-matched proxy and the selected candidate returned\. At a slower timescale, scores can be averaged over an unlabeled selection panel to rank the candidate models and choose one checkpoint for repeated use\. These two uses share the candidate scores but answer different questions: routing preserves input\- and task\-specific choice, whereas panel\-level selection trades that flexibility for a single reusable model \(Fig\.[5](https://arxiv.org/html/2608.20441#S2.F5)\)\.

Figure 5:Routing and panel\-level model selection use comparison at different timescales\.a, Hidden\-reference winner and task\-matched proxy route for each of 99 inputs; colours distinguish checkpoints and black ticks mark mismatches\.b, Empirical cumulative distributions of span\-normalized selection regret for the same inputs\. Regret is excess hidden loss divided by the three\-model loss span; curves nearer the upper left are better\.c, Empirical cumulative distributions of the same regret for panel\-level selection\. The left facet contains eight libraries with eight panel evaluations per library atB=24B=24; the right contains 16 additional libraries with eight evaluations per library atB=144B=144\. Methods are identified in the panel legends\.Three public U\-Net checkpoints for the two\-field PDEBench diffusion–reaction system provide a test beyond the trained libraries above\[[24](https://arxiv.org/html/2608.20441#bib.bib24),[5](https://arxiv.org/html/2608.20441#bib.bib5)\]\. These released checkpoints were selected within the benchmark’s validation workflow\. On the terminal two\-field task, one checkpoint was best for every input and the shared score recovered every choice, whereas raw residual recovered none of the top choices \(Supplementary Table 7\)\. Because that ordering was constant across the 99 inputs, the terminal task tests recovery on one fixed public library rather than input\-adaptive routing\. We therefore rescored the same trajectories under a restricted\-quadrant high\-band spectral task, for which the preferred checkpoint became input dependent\. This spectral distance lies outside the exact squared\-Hilbert identity and probes task transfer beyond the exact comparison theorem\.

For this high\-band spectral objective, the preferred checkpoint changed across inputs \(Fig\.[5](https://arxiv.org/html/2608.20441#S2.F5)a\)\. Rescoring the same proxy in this task recovered 88 of 99 choices, compared with 69 for retaining the best fixed terminal model, 23 for medoid centrality and 7 for the raw residual; it closed 86\.7% of the fixed\-to\-oracle risk gap \(Fig\.[5](https://arxiv.org/html/2608.20441#S2.F5)b\)\. Because the terminal and spectral tasks use the same models, inputs and proxy trajectories, the winner crossover isolates task dependence rather than a change in the candidate library\. Additional task objectives are compared in Supplementary Section S12\.

Panel\-level selection asks whether the instance\-wise signal survives aggregation\. For an unlabeled panel of sizeBB, candidate scores are averaged to select one reusable checkpoint\. Small panels were not always sufficient\. AtB=24B=24, performance varied sharply among libraries, including one Burgers–CNO library in which only one of eight panels selected a top\-two model\. Pooled across the eight libraries, shared scoring produced lower normalized regret than the raw residual baseline, validation\-loss selection and mean or medoid centrality \(Fig\.[5](https://arxiv.org/html/2608.20441#S2.F5)c\)\. This pooled advantage came from the reaction–diffusion libraries and was not observed in either Burgers architecture group\. Candidate\-specific correction made the same selection and incurred the same regret as the shared response in all 64 panel evaluations\. Strong query\-wise ranking therefore need not translate into reliable selection from a small selection panel\.

Across 16 additional libraries evaluated with a selection panel ofB=144B=144, the shared score chose a model within the study tolerance of the best in 122 of 128 panels and a top\-two model in all 128\. It also outperformed the raw residual baseline, validation and candidate\-only centrality baselines \(Fig\.[5](https://arxiv.org/html/2608.20441#S2.F5)c\), with normalized regret concentrated near zero\.

### 2\.6Strong monotonicity certifies decisions against a same\-grid operator reference

The preceding results estimate decisions without access to a numerical reference\. For one zero\-Dirichlet, strongly monotone cubic reaction–diffusion discretization, the residual at any returned state supplies a global error enclosure and hence a computable decision guarantee\. This result does not require the local linearization assumptions of Proposition[1](https://arxiv.org/html/2608.20441#Thmproposition1)\.

LetVhV\_\{h\}contain the interior grid degrees of freedom with weighted inner product

⟨v,z⟩h=h2​∑kvk​zk\.\\langle v,z\\rangle\_\{h\}=h^\{2\}\\sum\_\{k\}v\_\{k\}z\_\{k\}\.\(22\)LetKhK\_\{h\}be self\-adjoint positive definite and define

Ah​\(v\)=Kh​v\+λ​v\+γ​v⊙3−fh,λ,γ\>0\.A\_\{h\}\(v\)=K\_\{h\}v\+\\lambda v\+\\gamma v^\{\\odot 3\}\-f\_\{h\},\\qquad\\lambda,\\gamma\>0\.\(23\)LetQh:Vh→YQ\_\{h\}:V\_\{h\}\\to Ybe linear into a finite\-dimensional Hilbert task space\. We identifyYYwith the task spaceℋ\\mathcal\{H\}above and⟨⋅,⋅⟩Y=⟨⋅,⋅⟩M\\langle\\cdot,\\cdot\\rangle\_\{Y\}=\\left\\langle\\cdot,\\cdot\\right\\rangle\_\{M\}; the state\-space metricMsM\_\{s\}below is distinct from the fixed task metricMM\. Define the weighted adjoint by

⟨Qh​v,z⟩Y=⟨v,Qh∗​z⟩h\.\\langle Q\_\{h\}v,z\\rangle\_\{Y\}=\\langle v,Q\_\{h\}^\{\*\}z\\rangle\_\{h\}\.\(24\)For any returned states∈Vhs\\in V\_\{h\}, define

Ms=Kh\+λ​I\+3​γ4​diag⁡\(s⊙2\),M\_\{s\}=K\_\{h\}\+\\lambda I\+\\frac\{3\\gamma\}\{4\}\\operatorname\{diag\}\(s^\{\\odot 2\}\),\(25\)and write‖z‖Ms2=⟨Ms​z,z⟩h\\\|z\\\|\_\{M\_\{s\}\}^\{2\}=\\langle M\_\{s\}z,z\\rangle\_\{h\}and‖ρ‖Ms−12=⟨ρ,Ms−1​ρ⟩h\\\|\\rho\\\|\_\{M\_\{s\}^\{\-1\}\}^\{2\}=\\langle\\rho,M\_\{s\}^\{\-1\}\\rho\\rangle\_\{h\}\.

###### Theorem 2\(Global strongly monotone reaction–diffusion decision certificate\)\.

EquationAh​\(uh\)=0A\_\{h\}\(u\_\{h\}\)=0has a unique solution,MsM\_\{s\}is self\-adjoint positive definite, and

‖s−uh‖Ms≤‖Ah​\(s\)‖Ms−1\.\\\|s\-u\_\{h\}\\\|\_\{M\_\{s\}\}\\leq\\\|A\_\{h\}\(s\)\\\|\_\{M\_\{s\}^\{\-1\}\}\.\(26\)Fort=Qh​uht=Q\_\{h\}u\_\{h\},q=Qh​sq=Q\_\{h\}sanddi​j=yi−yjd\_\{ij\}=y\_\{i\}\-y\_\{j\},

\|Gi​j​\(t\)−Gi​j​\(q\)\|≤2​‖Ah​\(s\)‖Ms−1​‖Qh∗​di​j‖Ms−1=:Bi​j\.\|G\_\{ij\}\(t\)\-G\_\{ij\}\(q\)\|\\leq 2\\\|A\_\{h\}\(s\)\\\|\_\{M\_\{s\}^\{\-1\}\}\\\|Q\_\{h\}^\{\*\}d\_\{ij\}\\\|\_\{M\_\{s\}^\{\-1\}\}=:B\_\{ij\}\.\(27\)Hence\|Gi​j​\(q\)\|\>Bi​j\|G\_\{ij\}\(q\)\|\>B\_\{ij\}certifies a strict pair sign\. If proxy winnerkksatisfies

Rj​\(q\)−Rk​\(q\)\>Bk​jfor every​j≠k,R\_\{j\}\(q\)\-R\_\{k\}\(q\)\>B\_\{kj\}\\quad\\text\{for every \}j\\neq k,\(28\)then it is the unique same\-grid operator winner\. More generally, for a scalar toleranceτ≥0\\tau\\geq 0, if

Rk​\(q\)−Rj​\(q\)\+Bk​j≤τfor every​j≠k,R\_\{k\}\(q\)\-R\_\{j\}\(q\)\+B\_\{kj\}\\leq\\tau\\quad\\text\{for every \}j\\neq k,\(29\)then its same\-grid regret is at mostτ\\tau\. Failure of a sufficient condition is abstention\.

Letℬh=Qh∗​𝒟\\mathcal\{B\}\_\{h\}=Q\_\{h\}^\{\*\}\\mathcal\{D\}andr=dimℬh≤N−1r=\\dim\\mathcal\{B\}\_\{h\}\\leq N\-1\. One residual inverse action andrrcomparison\-basis inverse actions determine every pair bound through a Gram matrix\.

If a dataset target satisfiestdata=t\+ξt\_\{\\rm data\}=t\+\\xiandP𝒟M​ξ∈𝒰discP\_\{\\mathcal\{D\}\}^\{M\}\\xi\\in\\mathcal\{U\}\_\{\\rm disc\}for a compact set𝒰disc⊂𝒟\\mathcal\{U\}\_\{\\rm disc\}\\subset\\mathcal\{D\}, defineσ𝒰​\(d\)=supz∈𝒰⟨z,d⟩M\\sigma\_\{\\mathcal\{U\}\}\(d\)=\\sup\_\{z\\in\\mathcal\{U\}\}\\left\\langle z,d\\right\\rangle\_\{M\}\. Then

Gi​j\(tdata\)∈\[\\displaystyle G\_\{ij\}\(t\_\{\\rm data\}\)\\in\\bigl\[Gi​j​\(q\)−Bi​j−2​σ𝒰disc​\(−di​j\),\\displaystyle G\_\{ij\}\(q\)\-B\_\{ij\}\-2\\sigma\_\{\\mathcal\{U\}\_\{\\rm disc\}\}\(\-d\_\{ij\}\),Gi​j\(q\)\+Bi​j\+2σ𝒰disc\(di​j\)\]\.\\displaystyle G\_\{ij\}\(q\)\+B\_\{ij\}\+2\\sigma\_\{\\mathcal\{U\}\_\{\\rm disc\}\}\(d\_\{ij\}\)\\bigr\]\.\(30\)In particular,‖P𝒟M​ξ‖M≤Δ𝒟\\\|P\_\{\\mathcal\{D\}\}^\{M\}\\xi\\\|\_\{M\}\\leq\\Delta\_\{\\mathcal\{D\}\}adds2​Δ𝒟​‖di​j‖M2\\Delta\_\{\\mathcal\{D\}\}\\\|d\_\{ij\}\\\|\_\{M\}toBi​jB\_\{ij\}\.

###### Proof\.

The scalar inequality

\(a3−b3\)​\(a−b\)≥34​a2​\(a−b\)2\(a^\{3\}\-b^\{3\}\)\(a\-b\)\\geq\\frac\{3\}\{4\}a^\{2\}\(a\-b\)^\{2\}is sharp and follows froma2\+a​b\+b2−3​a2/4=\(b\+a/2\)2a^\{2\}\+ab\+b^\{2\}\-3a^\{2\}/4=\(b\+a/2\)^\{2\}\. Applied componentwise, it gives

⟨Ah​\(s\)−Ah​\(uh\),s−uh⟩h≥‖s−uh‖Ms2\.\\langle A\_\{h\}\(s\)\-A\_\{h\}\(u\_\{h\}\),s\-u\_\{h\}\\rangle\_\{h\}\\geq\\\|s\-u\_\{h\}\\\|\_\{M\_\{s\}\}^\{2\}\.Strict convexity and coercivity of the discrete energy give existence and uniqueness\. Cauchy–Schwarz in theMsM\_\{s\}metric yields Eq\. \([26](https://arxiv.org/html/2608.20441#S2.E26)\); the squared\-Hilbert gap identity and weighted adjoint relation yield Eq\. \([27](https://arxiv.org/html/2608.20441#S2.E27)\)\. The pair, strict and tolerant conditions follow by one\-sided interval arithmetic\. Expanding everyQh∗​di​jQ\_\{h\}^\{\*\}d\_\{ij\}in a basis ofQh∗​𝒟Q\_\{h\}^\{\*\}\\mathcal\{D\}gives the Gram reduction\. Finally, support\-function addition transfers the same\-grid interval totdata=t\+ξt\_\{\\rm data\}=t\+\\xi\. ∎

The certificate uses directional structure rather than a separate inverse action for every pair\. A basis ofQh∗​𝒟Q\_\{h\}^\{\*\}\\mathcal\{D\}builds one Gram matrix from at mostN−1N\-1comparison\-basis actions; every pair bound is then a quadratic form in that basis\. The complete discrete proof, solve\-defect inflation and weighted\-adjoint implementation are given in Supplementary Section S11\.

Across 24 library–input combinations and 672 candidate pairs, every pairwise, strict top\-1 andτ=10−6\\tau=10^\{\-6\}tolerant top\-1 margin exceeded its bound; the largest bound\-to\-margin ratio was 0\.09494 \(Fig\.[6](https://arxiv.org/html/2608.20441#S2.F6)a,b\)\.

Figure 6:Bound\-to\-margin ratios remain below the certification threshold\.a, Empirical cumulative distribution of 672 ratiosBi​j/\|Gi​j​\(q\)\|B\_\{ij\}/\|G\_\{ij\}\(q\)\|; the vertical line marks the sufficient threshold of one\.b, Largest ratio in each of 24 cases formed by six physical inputs and four libraries\. All ratios are evaluated against the unique zero of the stated same\-grid discrete operator\.The theorem concerns the unique zero of Eq\. \([23](https://arxiv.org/html/2608.20441#S2.E23)\) on the stated grid\. Agreement with a dataset or continuum target additionally requires a discrepancy bound such as Eq\. \([30](https://arxiv.org/html/2608.20441#S2.E30)\)\. Cross\-resolution results are given in Supplementary Section S10\.

## 3Discussion

Neural operators are conventionally evaluated as individual predictors, yet a deployed scientific workflow frequently involves multiple plausible models\. Once a finite surrogate library is established, the immediate operational challenge shifts from full\-state reconstruction to candidate selection\. Under a fixed squared Hilbert loss, the unknown solution influences candidate preferences exclusively through the span of candidate differences\. This comparison subspace has a dimension of at most one less than the library size, regardless of the ambient state dimension\. Recognizing this distinction is critical because it isolates the precise fraction of an unavailable reference required to make a reliable deployment decision\.

The shared physics response translates this geometric reduction into a practical computational diagnostic\. By propagating the complete constraint defect of a single library anchor through a full\-state linearized solve, the method yields a corrected task output to score every candidate simultaneously\. While defect correction, adjoint analysis, and a posteriori estimation supply the underlying numerical ingredients, the finite\-library quotient identifies the specific family of relative decisions that a single response can resolve\. In the reaction\-diffusion census, candidate\-wise variational estimators raised non\-tie pair recovery from 59\.8% to 82\.7%, whereas the shared task proxy achieved 100% recovery\. Furthermore, candidate\-specific correction\-magnitude proxies produced almost identical choices to the shared calculation, confirming that library\-level physics propagation preserves the necessary decision signals\.

Despite this shared computational foundation, ranking and direct correction solve fundamentally different problems\. Returning the corrected proxy demands accuracy across the entire task space, whereas candidate selection only requires accurate comparison coordinates and sufficient separation between predictions\. In several smooth systems, the proxy proved highly effective for both purposes; however, the compressible\-flow probe demonstrated that correct relative comparisons can persist even when the direct\-use advantage diminishes\. These scores subsequently support two distinct deployment modes\. Query\-time routing preserves the current input and task context, while panel\-level aggregation identifies a single optimal model for repeated use\. Both deployment strategies require only one post\-inference physics response for the entire library, shifting the computational burden from a one\-per\-candidate scaling to a single shared solve\. The overall operational cost will therefore depend primarily on the governing operator, the chosen discretization, and the broader serving workflow\.

The reliable application of this shared response is naturally bounded by three operational constraints\. First, small candidate margins inherently amplify both proxy and algebraic errors\. Second, library composition directly influences the anchor state from which the physical response is computed\. Because neither the arithmetic mean nor the candidate medoid demonstrated universal superiority across the tested datasets, anchor construction remains intrinsically library dependent\. Third, shock\-dominated dynamics may necessitate numerical damping to maintain physical admissibility\. Consequently, panel\-level aggregation heavily relies on whether a small, unlabeled evaluation panel accurately represents the eventual deployment population, a vulnerability highlighted by the retained Burgers\-CNO failure case\.

Moving beyond empirical proxy performance, the framework can yield rigorous theoretical guarantees for well\-structured systems\. For the strongly monotone reaction\-diffusion discretization, the residual provides an exact\-arithmetic bound for the unique zero of the corresponding same\-grid operator, thereby certifying pairwise or top\-1 decisions whenever the required margins are present\. Transferring these guarantees to a dataset or continuum target will require an additional operator\-data discrepancy bound\. Future extensions of this work include integrating mixed and evolving candidate libraries, accommodating non\-Hilbert objectives, utilizing multifidelity data, and conducting more rigorous stress tests under severe shocks or simulator mismatches\. Across all such settings, the central pursuit remains unchanged: identifying precisely which features of a hidden physical solution are strictly necessary to distinguish among available machine learning predictions\.

## 4Methods

### Finite libraries and decision targets

Letaadenote a deployment input and let\{ℳi\}i=1N\\\{\\mathcal\{M\}\_\{i\}\\\}\_\{i=1\}^\{N\}be the finite candidate library\. The high\-fidelity reference is unavailable when the library is queried\. Candidateiireturns a state trajectoryvi=ℳi​\(a\)v\_\{i\}=\\mathcal\{M\}\_\{i\}\(a\)and a task output

yi=Qa​\(vi\)∈ℋ,y\_\{i\}=Q\_\{a\}\(v\_\{i\}\)\\in\\mathcal\{H\},\(31\)where\(ℋ,⟨⋅,⋅⟩M\)\(\\mathcal\{H\},\\left\\langle\\cdot,\\cdot\\right\\rangle\_\{M\}\)is a fixed Hilbert space after all outputs have been lifted to a common grid\. Ifuau\_\{a\}is the numerical reference andt=Qa​\(ua\)t=Q\_\{a\}\(u\_\{a\}\), the hidden squared task loss isRi​\(t\)=‖t−yi‖M2R\_\{i\}\(t\)=\\left\\lVert t\-y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\. The selector is specified without access to the referencett\. Throughout, “reference\-free” means that scoring uses no deployment reference solution or reference\-derived label; the governing operator, problem data and task remain available\. In the Burgers and reaction–diffusion comparisons, each library contains candidates from one architecture; the two\-dimensional compressible\-flow analysis instead compares one public FNO with one public U\-Net\. The score depends on the lifted outputs rather than the model parameters, so architecture matching is an experimental choice rather than a mathematical requirement\. Every experimental state\-to\-task map is linear after the fixed common\-grid lift: it is either the identity map or a terminal\-state trace\. The derivative notation below retains the framework’s extension to smooth nonlinear task maps\. The theorem\-backed analyses use these maps with squared\-Hilbert losses\. Some task\-transfer distances applied to the same outputs are nonlinear and are reported only as empirical transfer tests; they do not inherit the exact squared\-Hilbert comparison identity\.

For an unordered pairi<ji<j, define the diagnostic and truth gaps

Gi​jq=Sj−Si,Gi​jt=Rj​\(t\)−Ri​\(t\)\.G\_\{ij\}^\{q\}=S\_\{j\}\-S\_\{i\},\\qquad G\_\{ij\}^\{t\}=R\_\{j\}\(t\)\-R\_\{i\}\(t\)\.\(32\)ThusGi​j\>0G\_\{ij\}\>0means that candidateiiis preferred\. A per\-case toleranceτ\\taumaps each gap to one of three labels:iipreferred, tie, orjjpreferred\. A strict pair comparison setsτ=0\\tau=0and tests the resulting ordering for every unordered pair\. “Non\-tie pairwise accuracy” instead conditions on non\-ties of the hidden\-reference label, whereas “all\-pair three\-way accuracy” retains the tie class\. Exact top\-1 requires the same minimizing checkpoint; tolerant top\-1 accepts a selected checkpoint whose hidden loss is within the chosen distance of the minimum\. For the eight Burgers and reaction–diffusion libraries, this distance is10−610^\{\-6\}\. These comparisons use non\-tie pairwise and tolerant top\-1 recovery\. Their all\-pair sensitivity analysis uses a10−610^\{\-6\}hidden\-reference tolerance, a1\.01×10−81\.01\\times 10^\{\-8\}propagated\-score numerical tolerance \(the fixed10−1010^\{\-10\}absolute floor plus a10−810^\{\-8\}unit\-scale relative floor\), and a10−1210^\{\-12\}raw\-score tolerance; a matched\-tolerance sensitivity analysis applies10−610^\{\-6\}to both scores\. The 16\-library comparison uses strict pairwise and exact top\-1, and the Sine–Gordon study uses strict pairwise and exact top\-1 with numerically unresolved outcomes counted as errors\. The public PDEBench study also uses strict pairwise and exact top\-1 for one fixed library\.

### Study design and statistical units

We crossed two PDEs, two architectures and two independently seeded ensembles to form eight candidate libraries\. Each library contained the epoch\-20 and epoch\-40 checkpoints from four training runs, giving 64 checkpoints from 32 runs\. Eligibility required finite metrics, in\-distribution relativeL2L^\{2\}error at most one, exact hard initial or boundary conditions where applicable, and completion of the full eight\-candidate library\. Each PDE contributed 240 controlled\-shift inputs, giving 480 distinct physical inputs and 1,920 input–library cases; none overlapped with training, validation or development data\. Candidate eligibility, the anchor rule, task metric, tie tolerances, selectors and evaluation inputs were specified before reference evaluation\. Within each input–library case, every candidate was scored against the same library anchor and metric\. Confidence intervals use 10,000 input\-clustered bootstrap replicates\.

Each eight\-candidate library yields 28 dependent pair rows per input\. Confidence intervals therefore resample inputs while preserving the coupling of architectures and ensembles evaluated on the same input\. For the 16\-library comparison, resampling likewise preserves matched libraries and shared inputs\. Its hierarchical estimator first averages each metric over inputs within each library and then gives each of the 16 libraries equal weight; it therefore differs from the pooled 3,800/3,840 case count\. For Sine–Gordon, the cluster is the shared input identifier across architectures, libraries, and pairs\. A Sine–Gordon pair was unresolved when the absolute fine\-grid loss gap did not exceed the sum of the candidates’ coarse\-to\-fine loss changes; the top\-1 rule applied the analogous check to the best–runner\-up gap\. These outcomes were counted as errors\. The public PDEBench comparison contains one fixed three\-candidate library and is therefore summarized by exact counts\.

### Shared correction and candidate scores

Let𝒜a​\(v\)\\mathcal\{A\}\_\{a\}\(v\)be the assembled discrete or continuous constraint defect, including initial and boundary components rather than only the interior PDE residual; components enforced exactly appear as zero rows\. A deterministic, reference\-free library\-anchor rule𝒞\\mathcal\{C\}produces

w=𝒞⁡\(v1,…,vN\)\.w=\\mathcal\{C\}\(v\_\{1\},\\ldots,v\_\{N\}\)\.\(33\)The eight Burgers and reaction–diffusion libraries used the arithmetic candidate mean\. The 16\-library, Sine–Gordon and public PDEBench comparisons used the geometric medoid: the candidate state minimizing the sum of distances to all candidates in the fixed anchor metric\. The two choices define distinct variants of the framework\. Burgers and reaction–diffusion medoid distances use the Euclidean norm over all stored state entries on the diagnostic grid\. Sine–Gordon uses trapezoidal time–spaceL2L^\{2\}distance over displacement and velocity, with velocity weight one\. PDEBench uses the corresponding physical time–space distance over both released fields\. Exact medoid ties follow the fixed candidate order\. The anchor rule is part of the method specification and is chosen without the deployment reference\. The composition study below shows that no single anchor was uniformly preferable across the tested libraries; the response must be used in a library regime for which its anchor and linearization remain informative\.

Atww, formrw=𝒜a​\(w\)r\_\{w\}=\\mathcal\{A\}\_\{a\}\(w\)and the state\-dependent JacobianLw=D​𝒜a​\(w\)L\_\{w\}=D\\mathcal\{A\}\_\{a\}\(w\)\. The shared correction satisfies

Lw​δw=−rw\+ℓ,L\_\{w\}\\delta\_\{w\}=\-r\_\{w\}\+\\ell,\(34\)whereℓ\\ellis the algebraic solve defect and vanishes for an exact solve\. The corresponding corrected state, linearized task estimate and evaluated task estimate are

s=w\+δw,qlin=Qa​\(w\)\+D​Qa​\(w\)​δw,qeval=Qa​\(s\),s=w\+\\delta\_\{w\},\\qquad q\_\{\\rm lin\}=Q\_\{a\}\(w\)\+DQ\_\{a\}\(w\)\\delta\_\{w\},\\qquad q\_\{\\rm eval\}=Q\_\{a\}\(s\),\(35\)respectively\. Every task map used here is linear, soq≡qlin=qevalq\\equiv q\_\{\\rm lin\}=q\_\{\\rm eval\}exactly; we retain the symbolqqfor the task proxy andssfor the full corrected state or trajectory\. For a nonlinearQaQ\_\{a\}, the linearized and evaluated proxies remain distinct\. The propagated score is Eq\. \([8](https://arxiv.org/html/2608.20441#S2.E8)\)\. Equation \([34](https://arxiv.org/html/2608.20441#S4.E34)\) is solved once per input and library; the resultingqqis shared by allNNcandidate comparisons\. This implementation is a full\-state linearized solve, not an explicitly reduced\-order solver; the finite\-library quotient specifies which component of its output is consumed by ranking\. Figure[1](https://arxiv.org/html/2608.20441#S2.F1)summarizes this shared physics\-to\-decision path, and Fig\.[2](https://arxiv.org/html/2608.20441#S2.F2)reports its recovery and comparisons\.

The raw residual baseline evaluates each candidate’s assembled defect energy,

Siraw=‖𝒜a​\(vi\)‖Y,a2,S\_\{i\}^\{\\mathrm\{raw\}\}=\\left\\lVert\\mathcal\{A\}\_\{a\}\(v\_\{i\}\)\\right\\rVert\_\{Y,a\}^\{2\},\(36\)using the fixed residual\-product norm\. This baseline uses the same candidates, inputs and hidden references as the propagated method, while scoring the defect directly instead of forming a task\-space estimate\. Its exact block, time, and quadrature weights are defined in Supplementary Section S6\. A higher\-cost control solvesD​𝒜a​\(vi\)​δi=−𝒜a​\(vi\)D\\mathcal\{A\}\_\{a\}\(v\_\{i\}\)\\delta\_\{i\}=\-\\mathcal\{A\}\_\{a\}\(v\_\{i\}\)for every candidate and assigns

qicand=Qa​\(vi\)\+D​Qa​\(vi\)​δi,Sicand=‖qicand−yi‖M2\.q\_\{i\}^\{\\mathrm\{cand\}\}=Q\_\{a\}\(v\_\{i\}\)\+DQ\_\{a\}\(v\_\{i\}\)\\delta\_\{i\},\\qquad S\_\{i\}^\{\\mathrm\{cand\}\}=\\left\\lVert q\_\{i\}^\{\\mathrm\{cand\}\}\-y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\.\(37\)Thus the control uses one candidate\-derived correction per candidate rather than one common correction for the library\. It ranks candidates by increasingSicandS\_\{i\}^\{\\mathrm\{cand\}\}and selectsi^cand=arg⁡mini⁡Sicand\\widehat\{i\}\_\{\\mathrm\{cand\}\}=\\arg\\min\_\{i\}S\_\{i\}^\{\\mathrm\{cand\}\}; evaluation uses the same pair and top\-1 conventions defined above\. This score is the task\-space norm of the candidate\-specific correction—becauseyi=Qa​\(vi\)y\_\{i\}=Q\_\{a\}\(v\_\{i\}\), it reduces to∥D​Qa​\(vi\)​δi∥M2\\lVert DQ\_\{a\}\(v\_\{i\}\)\\delta\_\{i\}\\rVert\_\{M\}^\{2\}—and serves as a candidate\-specific linearized error proxy\. Comparing it with the shared score tests whether one library\-level correction retains the decisions ofNNseparate corrections\.

Candidate\-only baselines rank candidates by their in\-distribution validation loss or by squared task\-space distance to the candidate mean or geometric medoid\. The validation ordering is independent of the deployment input, whereas the two centrality scores depend on the current candidate outputs but do not use the governing equation\. For any selector evaluated against a finite set of hidden losses, normalized regret is the selected excess loss divided by the best\-to\-worst loss span, with a denominator floor of10−810^\{\-8\}\.

The component\-cost comparison used nestedN=2,4,8,16N=2,4,8,16libraries, two architectures, three input strata and ten repeated CPU measurements\. Timers began after candidate outputs were available and covered only the shared or candidate\-specific physics calculation\. The reported ratio therefore compares one shared response withNNseparate responses after candidate inference; the complete timing design is reported in Supplementary Section S9\.

Algorithm 1Shared physics ranking for one deployment input1:Candidate states

v1,…,vNv\_\{1\},\\ldots,v\_\{N\}; complete constraint

𝒜a\\mathcal\{A\}\_\{a\}; linear task map

QaQ\_\{a\}; anchor rule

𝒞\\mathcal\{C\}, lift, and metric

MM\.

2:Lift all candidates to their shared state space\.

3:

w←𝒞⁡\(v1,…,vN\)w\\leftarrow\\mathcal\{C\}\(v\_\{1\},\\ldots,v\_\{N\}\)\.

4:Evaluate

rw←𝒜a​\(w\)r\_\{w\}\\leftarrow\\mathcal\{A\}\_\{a\}\(w\), retaining initial/boundary rows\.

5:Compute

δw\\delta\_\{w\}from

D​𝒜a​\(w\)​δw≃−rwD\\mathcal\{A\}\_\{a\}\(w\)\\delta\_\{w\}\\simeq\-r\_\{w\}once and record

ℓ←D​𝒜a​\(w\)​δw\+rw\\ell\\leftarrow D\\mathcal\{A\}\_\{a\}\(w\)\\delta\_\{w\}\+r\_\{w\}\.

6:Form

s←w\+δws\\leftarrow w\+\\delta\_\{w\}and

q←Qa​\(w\)\+D​Qa​\(w\)​δw=Qa​\(s\)q\\leftarrow Q\_\{a\}\(w\)\+DQ\_\{a\}\(w\)\\delta\_\{w\}=Q\_\{a\}\(s\)\.

7:for

i=1,…,Ni=1,\\ldots,Ndo

8:

yi←Qa​\(vi\)y\_\{i\}\\leftarrow Q\_\{a\}\(v\_\{i\}\);

Si←‖q−yi‖M2S\_\{i\}\\leftarrow\\left\\lVert q\-y\_\{i\}\\right\\rVert\_\{M\}^\{2\}\.

9:endfor

10:returnthe ranked candidates, pair labels, or

arg⁡mini⁡Si\\arg\\min\_\{i\}S\_\{i\}\.

### Viscous Burgers

On the periodic line, the equation isut−ν​ux​x\+u​ux=fu\_\{t\}\-\\nu u\_\{xx\}\+uu\_\{x\}=fwith prescribed initial state\. The candidate output is the trajectory and the task output is its terminal state\. The complete defect is

𝒜a​\(w\)=\(w⁡\(0\)−u0,wt−ν​wx​x\+w​wx−f\),\\mathcal\{A\}\_\{a\}\(w\)=\\bigl\(w\(0\)\-u\_\{0\},\\;w\_\{t\}\-\\nu w\_\{xx\}\+ww\_\{x\}\-f\\bigr\),\(38\)andLw​h=\(h⁡\(0\),ht−ν​hx​x\+∂x\(w​h\)\)L\_\{w\}h=\(h\(0\),h\_\{t\}\-\\nu h\_\{xx\}\+\\partial\_\{x\}\(wh\)\)\. The implementation uses periodic Fourier differentiation\[[25](https://arxiv.org/html/2608.20441#bib.bib25)\], a time\-marching tangent solve, and a physicalL2L^\{2\}terminal metric after Fourier lifting to the common grid\.

### Cubic reaction–diffusion

OnΩ=\(0,L\)2\\Omega=\(0,L\)^\{2\}with zero Dirichlet boundary,

−∇⋅\(a∇u\)\+λu\+γu3=f,u\|∂Ω=0\.\-\\nabla\\\!\\cdot\(a\\nabla u\)\+\\lambda u\+\\gamma u^\{3\}=f,\\qquad u\|\_\{\\partial\\Omega\}=0\.\(39\)States include boundary nodes\. Boundary entries of𝒜a\\mathcal\{A\}\_\{a\}are the Dirichlet defect, and interior entries are a conservative variable\-coefficient finite\-difference defect\. The correction on the boundary is fixed exactly by the Dirichlet rows; its known contribution is transferred to the interior right\-hand side before solving the boundary\-eliminated system\. The interior Jacobian contains the reaction factorλ\+3​γ​w2\\lambda\+3\\gamma w^\{2\}\. Candidate and corrected fields are compared after common interpolation in a fixed tensor\-product physical quadrature metric\. The task output is the complete stationary solution field on this common grid, including boundary nodes\. Thus reaction–diffusion tests the identity\-task case, while the Burgers and Sine–Gordon terminal tasks exercise task compression\.

### Reaction–diffusion decision certificate

The certificate calculations use four reaction–diffusion groups, each with eight candidates, and two inputs from each of three shift strata\. This gives 24 library–input combinations and 672 unordered pairs\. The bounds depend on the candidates, forcing, coefficients, operator, anchor rule and tolerance\.

All certificate operators act on the 128\-by\-128 interior space after verifying exact zero boundary values\. With the identity task and itsh2h^\{2\}\-weighted inner product, the corrected proxy iss=w\+δws=w\+\\delta\_\{w\}andq=sq=s\. For eachss, we formedMsM\_\{s\}from Eq\. \([25](https://arxiv.org/html/2608.20441#S2.E25)\), computed the residual dual norm, and solved a basis ofℬh=Qh∗​𝒟\\mathcal\{B\}\_\{h\}=Q\_\{h\}^\{\*\}\\mathcal\{D\}, requiringr=dimℬh≤N−1r=\\dim\\mathcal\{B\}\_\{h\}\\leq N\-1spanning right\-hand sides\. Approximate inverse actions were inflated by their algebraic solve defects and an analytic coercivity lower bound\. Certificate quantities were evaluated in float64; the theorem itself is stated in exact arithmetic\.

### Candidate\-wise residual\-estimator controls

The residual\-estimator comparison used all 240 reaction–diffusion inputs, four eight\-candidate libraries and two state realizations: the stored candidate fields transferred to the comparison grid and the same checkpoints rerun on that grid\. This produced 1,920 input–library–scenario cases\. The controls compared the raw product residual, a common\-energy dual norm based onHh=Kh\+λ​IH\_\{h\}=K\_\{h\}\+\\lambda I, a candidate\-adaptedL2L^\{2\}majorant and the shared corrected proxy\. Each variational selector applies an inverse operator to all eight candidate residuals; the shared score uses the single correction already defined above\. Component timings begin after candidate outputs are available and exclude inference, data movement, reference computation and end\-to\-end serving\. The complete estimator definitions and results appear in Supplementary Section S13\.

### Sine–Gordon

The third PDE is the periodic phase\-space system

ut=v,vt=c2​ux​x−sin⁡u\+f\.u\_\{t\}=v,\\qquad v\_\{t\}=c^\{2\}u\_\{xx\}\-\\sin u\+f\.\(40\)Its complete defect is defined by the initial displacement/velocity errors and all steps of the fixed velocity\-Verlet map\[[26](https://arxiv.org/html/2608.20441#bib.bib26)\]\. The exact Jacobian of that discrete map is block lower triangular, so Eq\. \([34](https://arxiv.org/html/2608.20441#S4.E34)\) is solved by a forward sweep\. The task is terminal displacement and velocity, compared in the physical phase\-spaceL2L^\{2\}metric with velocity weight one after periodic Fourier lifting\.

### PDEBench two\-field diffusion–reaction

The released PDEBench models predict trajectories for the homogeneous\-Neumann system

ut=u−u3−k−v\+Du​Δ​u,vt=u−v\+Dv​Δ​v,u\_\{t\}=u\-u^\{3\}\-k\-v\+D\_\{u\}\\Delta u,\\qquad v\_\{t\}=u\-v\+D\_\{v\}\\Delta v,\(41\)with\(Du,Dv,k\)=\(10−3,5×10−3,5×10−3\)\(D\_\{u\},D\_\{v\},k\)=\(10^\{\-3\},5\\times 10^\{\-3\},5\\times 10^\{\-3\}\)\. The candidate trajectories begin with ten observed frames on the released128×128128\\times 128cell\-centred grid\. Every checkpoint receives the same 20\-channel history and predicts one two\-field next frame\. Irrespective of its training curriculum, each checkpoint is evaluated by the same 91\-step autoregressive rollout: the oldest frame is dropped and the prediction appended until the 101\-frame trajectory is complete\. From the last observed frame onward, the complete defect contains the known initial row and every trapezoidal\-map defect\. Its exact linearization is solved forward in time with a preconditioned sparse solve\. The raw and propagated diagnostics therefore use identical candidate trajectories and time horizons\. The task is the terminal two\-field state under cell\-area\-weighted squaredL2L^\{2\}loss\.

The task\-transfer analysis uses the same 99 inputs, three public checkpoints and 91 predicted future frames\. The already computed trajectoryssis scored under ten objective\-specific distances\. These comprised terminal and trajectory squaredL2L^\{2\}, RMSE, a normalized\-RMSE plug\-in, conservation, maximum and boundary errors, and three fixed restricted\-quadrant spectral bands\. Only the two squared\-L2L^\{2\}tasks lie directly within the exact Hilbert comparison identity\. The normalized score uses the root\-mean\-square magnitude of its own reference argument, so candidate\-to\-truth and candidate\-to\-proxy evaluations do not share one fixed denominator\. The ten metrics therefore provide correlated views of the same 99 inputs\.

### Compressible\-flow boundary studies

The one\-dimensional shock probe used 35 samples from an official PDEBench held\-out file and three released checkpoints: two FNOs with different history contracts and one U\-Net\. Candidate states were transformed to conservative variables and combined by an arithmetic mean\. The comparison operator was a separately implemented fixed\-substep Rusanov–SSPRK2 map rather than the adaptive HLLC data generator, so the dataset truth has a nonzero operator defect\. A reference\-free backtracking rule tested multipliers00,0\.250\.25,0\.50\.5,0\.750\.75and11and chose the largest value that preserved positive density and pressure without increasing that defect\. Inadmissible unit steps remained errors in strict decision counts\. Supplementary Section S15 gives the complete results\.

The two\-dimensional probe used one public FNO and one public push\-forward\-20 U\-Net for periodic PDEBench compressible\-flow data at64×6464\\times 64resolution\. We used ten trajectories\. Both candidates received the same ten observed frames and were rolled autoregressively for eleven further frames\. Primitive variables were converted to density, momenta and total energy before constructing the arithmetic mean\. A clean\-room periodic Rusanov finite\-volume map with density\-weighted viscous fluxes and SSPRK2 time integration defined the comparison operator; finite differences supplied Jacobian actions\. Because this operator does not replay the dataset generator exactly, these results characterize an operator\-mismatched regime\. The partial path usedα∈\{0,0\.25,0\.5,0\.75,1,1\.25\}\\alpha\\in\\\{0,0\.25,0\.5,0\.75,1,1\.25\\\}with reference\-free admissibility and defect checks\. Additional medoid results are given in Supplementary Section S14\.

### Numerical references used for evaluation

Reference losses were computed from independent high\-resolution solutions\. For the eight Burgers and reaction–diffusion libraries, Burgers references used integrating\-factor RK4 on independent 768\- and 1,536\-point Fourier grids; reaction–diffusion used 511\- and 767\-point interior grids\. The 16\-library comparison increased these pairs to 1,536/3,072 and 767/1,023, respectively\. Reaction–diffusion references used damped Newton iteration with requested relative and absolute tolerances10−1110^\{\-11\}and10−1310^\{\-13\}, and a preconditioned linear solve with requested relative tolerance10−1210^\{\-12\}\. Sine–Gordon references used spectral velocity\-Verlet at 384 spatial points with 6,144 time steps and at 768 points with 12,288 steps\. Refinement, residual, replay, and Jacobian\-closure checks are detailed in Supplementary Section S6\. For PDEBench, terminal states in the released numerical trajectories provided the reference states\.

These realizations share the decision object but use different discretizations, anchor rules, and evaluation metrics\.

### Panel\-level model selection

For an unlabeled selection panelAA, aggregate the same per\-input score,

S¯i​\(A\)=1\|A\|​∑a∈ASi​\(a\),i^​\(A\)=arg⁡mini​S¯i​\(A\)\.\\bar\{S\}\_\{i\}\(A\)=\\frac\{1\}\{\|A\|\}\\sum\_\{a\\in A\}S\_\{i\}\(a\),\\qquad\\widehat\{i\}\(A\)=\\arg\\min\_\{i\}\\bar\{S\}\_\{i\}\(A\)\.\(42\)For squared\-Hilbert tasks, the instance\-wise comparison identity averages exactly over the panel\. WithGi​j,atG\_\{ij,a\}^\{t\}andGi​j,aqG\_\{ij,a\}^\{q\}denoting the truth and diagnostic gaps for inputaa,

1\|A\|​∑a∈A\(Gi​j,at−Gi​j,aq\)=2\|A\|​∑a∈A⟨ta−qa,yi​\(a\)−yj​\(a\)⟩M\.\\frac\{1\}\{\|A\|\}\\sum\_\{a\\in A\}\(G\_\{ij,a\}^\{t\}\-G\_\{ij,a\}^\{q\}\)=\\frac\{2\}\{\|A\|\}\\sum\_\{a\\in A\}\\left\\langle t\_\{a\}\-q\_\{a\},y\_\{i\}\(a\)\-y\_\{j\}\(a\)\\right\\rangle\_\{M\}\.\(43\)Thus panel\-level selection averages the same decision coordinates over the inputs\. On an evaluation setDD, onlyℳi^​\(A\)\\mathcal\{M\}\_\{\\widehat\{i\}\(A\)\}is run\. Its mean task lossR⁡\(i^\)R\(\\widehat\{i\}\)is compared with the best fixed checkpointi⋆=arg⁡mini⁡R⁡\(i\)i^\{\\star\}=\\arg\\min\_\{i\}R\(i\)through absolute regret and

NRegret⁡\(i^\)=R⁡\(i^\)−R⁡\(i⋆\)max⁡\{R⁡\(iworst\)−R⁡\(i⋆\),10−8\}\.\\operatorname\{NRegret\}\(\\widehat\{i\}\)=\\frac\{R\(\\widehat\{i\}\)\-R\(i^\{\\star\}\)\}\{\\max\\\{R\(i^\{\\mathrm\{worst\}\}\)\-R\(i^\{\\star\}\),10^\{\-8\}\\\}\}\.\(44\)Uncertainty estimates treat inputs as the sampling units and preserve the dependence among candidate pairs from the same input\.

### Use of generative language tools

Generative language tools assisted manuscript organization, prose revision and code review\. The authors verified the reported mathematics and numerical results and take responsibility for the final manuscript\. No experimental data were generated by these tools\.

## Data availability

The numerical data underlying Figs\. 2–6 and all quantitative Supplementary figures and tables are provided in the accompanying Source Data file\. The processed data and supporting files used in the analyses are available at[https://github\.com/ToughClimb/OpPropRank\-open\-source](https://github.com/ToughClimb/OpPropRank-open-source); the version used in this study is[commitdafb2a0](https://github.com/ToughClimb/OpPropRank-open-source/commit/dafb2a06d0815a6678174fc9e1aadc9055b7e69b)\. The larger archive of candidate predictions, reference solutions and selection results is available in the[data\-v1\.0\.0release](https://github.com/ToughClimb/OpPropRank-open-source/releases/tag/data-v1.0.0)\. The PDEBench data and pretrained models used in this study are available under CC BY 4\.0 from[DaRUS\-2986](https://doi.org/10.18419/DARUS-2986)and[DaRUS\-2987](https://doi.org/10.18419/DARUS-2987), respectively\.

## Code availability

The code and tests needed to reproduce the numerical analyses are available at[https://github\.com/ToughClimb/OpPropRank\-open\-source](https://github.com/ToughClimb/OpPropRank-open-source)\(commit[dafb2a0](https://github.com/ToughClimb/OpPropRank-open-source/commit/dafb2a06d0815a6678174fc9e1aadc9055b7e69b)\)\. The accompanying source package also contains the scripts used to generate all six main figures\.

## Author contributions

H\.L\. conceived the study, developed the theory and methodology, implemented the software, designed and performed the computational experiments, analysed the data, prepared the figures and drafted the manuscript\. F\.L\. supervised the research, provided project guidance and reviewed and edited the manuscript\. Both authors reviewed and approved the final manuscript\.

## Funding

This work was supported by the Science and Technology Development Program of Jilin Province under Grant No\. 20250102032JC\.

## Ethics declarations

No ethical approval was required because this study used computational experiments and publicly available or generated data and did not involve human participants or animals\.

## Competing interests

The authors declare no competing interests\.

## References

- Li et al\. \[2021\]Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar\.Fourier neural operator for parametric partial differential equations\.In*International Conference on Learning Representations*, 2021\.URL[https://openreview\.net/forum?id=c8P9NQVtmnO](https://openreview.net/forum?id=c8P9NQVtmnO)\.
- Lu et al\. \[2021\]Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis\.Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators\.*Nature Machine Intelligence*, 3:218–229, 2021\.doi:[10\.1038/s42256\-021\-00302\-5](https://doi.org/10.1038/s42256-021-00302-5)\.
- Kovachki et al\. \[2023\]Nikola Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar\.Neural operator: Learning maps between function spaces with applications to PDEs\.*Journal of Machine Learning Research*, 24\(89\):1–97, 2023\.URL[https://www\.jmlr\.org/papers/v24/21\-1524\.html](https://www.jmlr.org/papers/v24/21-1524.html)\.
- Raonić et al\. \[2023\]Bogdan Raonić, Roberto Molinaro, Tim De Ryck, Tobias Rohner, Francesca Bartolucci, Rima Alaifari, Siddhartha Mishra, and Emmanuel de Bézenac\.Convolutional neural operators for robust and accurate learning of PDEs\.In*Advances in Neural Information Processing Systems*, volume 36, 2023\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/hash/f3c1951b34f7f55ffaecada7fde6bd5a\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper_files/paper/2023/hash/f3c1951b34f7f55ffaecada7fde6bd5a-Abstract-Conference.html)\.
- Takamoto et al\. \[2022\]Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Dan MacKinlay, Francesco Alesiani, Dirk Pflüger, and Mathias Niepert\.PDEBench: An extensive benchmark for scientific machine learning\.In*36th Conference on Neural Information Processing Systems Datasets and Benchmarks Track*, 2022\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/hash/0a9747136d411fb83f0cf81820d44afb\-Abstract\-Datasets\_and\_Benchmarks\.html](https://proceedings.neurips.cc/paper_files/paper/2022/hash/0a9747136d411fb83f0cf81820d44afb-Abstract-Datasets_and_Benchmarks.html)\.
- Gupta and Brandstetter \[2023\]Jayesh K\. Gupta and Johannes Brandstetter\.Towards multi\-spatiotemporal\-scale generalized PDE modeling\.*Transactions on Machine Learning Research*, 2023\.URL[https://openreview\.net/forum?id=dPSTDbGtBY](https://openreview.net/forum?id=dPSTDbGtBY)\.
- Ohana et al\. \[2024\]Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina J\. Agocs, Miguel Beneitez, Marsha Berger, Blakesley Burkhart, Stuart B\. Dalziel, Drummond B\. Fielding, Daniel Fortunato, Jared A\. Goldberg, Keiya Hirashima, Yan\-Fei Jiang, Rich R\. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S\. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo\-Saint Blancard, François Rozet, Liam H\. Parker, Miles Cranmer, and Shirley Ho\.The well: a large\-scale collection of diverse physics simulations for machine learning\.In*Advances in Neural Information Processing Systems*, volume 37, pages 44989–45037\. Curran Associates, Inc\., 2024\.doi:[10\.52202/079017\-1430](https://doi.org/10.52202/079017-1430)\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/hash/4f9a5acd91ac76569f2fe291b1f4772b\-Abstract\-Datasets\_and\_Benchmarks\_Track\.html](https://proceedings.neurips.cc/paper_files/paper/2024/hash/4f9a5acd91ac76569f2fe291b1f4772b-Abstract-Datasets_and_Benchmarks_Track.html)\.Datasets and Benchmarks Track\.
- Wei et al\. \[2023\]Zhao Wei, Jian Cheng Wong, Nicholas Sung, Abhishek Gupta, Chin Chun Ooi, Pao\-Hsiung Chiu, My Ha Dao, and Yew Soon Ong\.How to select physics\-informed neural networks in the absence of ground truth: A pareto front\-based strategy\.In*ICML 2023 Workshop on the Synergy of Scientific and Machine Learning Modelling*, 2023\.URL[https://icml\.cc/virtual/2023/28612](https://icml.cc/virtual/2023/28612)\.Workshop paper\.
- Garg et al\. \[2022\]Saurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur, and Hanie Sedghi\.Leveraging unlabeled data to predict out\-of\-distribution performance\.In*International Conference on Learning Representations*, 2022\.URL[https://openreview\.net/forum?id=o\_HsiMPYh\_x](https://openreview.net/forum?id=o_HsiMPYh_x)\.
- Guillory et al\. \[2021\]Devin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell, and Ludwig Schmidt\.Predicting with confidence on unseen distributions\.In*Proceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 1114–1124, 2021\.doi:[10\.1109/ICCV48922\.2021\.00117](https://doi.org/10.1109/ICCV48922.2021.00117)\.
- Raissi et al\. \[2019\]M\. Raissi, P\. Perdikaris, and George E\. Karniadakis\.Physics\-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.*Journal of Computational Physics*, 378:686–707, 2019\.doi:[10\.1016/j\.jcp\.2018\.10\.045](https://doi.org/10.1016/j.jcp.2018.10.045)\.
- Karniadakis et al\. \[2021\]George Em Karniadakis, Ioannis G\. Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang\.Physics\-informed machine learning\.*Nature Reviews Physics*, 3:422–440, 2021\.doi:[10\.1038/s42254\-021\-00314\-5](https://doi.org/10.1038/s42254-021-00314-5)\.
- Li et al\. \[2024\]Zongyi Li, Hongkai Zheng, Nikola Kovachki, David Jin, Haoxuan Chen, Burigede Liu, Kamyar Azizzadenesheli, and Anima Anandkumar\.Physics\-informed neural operator for learning partial differential equations\.*ACM/IMS Journal of Data Science*, 1\(3\):1–27, 2024\.doi:[10\.1145/3648506](https://doi.org/10.1145/3648506)\.
- Cao et al\. \[2023\]Lianghao Cao, Thomas O’Leary\-Roseberry, Prashant K\. Jha, J\. Tinsley Oden, and Omar Ghattas\.Residual\-based error correction for neural operator accelerated infinite\-dimensional bayesian inverse problems\.*Journal of Computational Physics*, 486:112104, 2023\.doi:[10\.1016/j\.jcp\.2023\.112104](https://doi.org/10.1016/j.jcp.2023.112104)\.
- Jha \[2024\]Prashant K\. Jha\.Residual\-based error corrector operator to enhance accuracy and reliability of neural operator surrogates of nonlinear variational boundary\-value problems\.*Computer Methods in Applied Mechanics and Engineering*, 419:116595, 2024\.doi:[10\.1016/j\.cma\.2023\.116595](https://doi.org/10.1016/j.cma.2023.116595)\.
- Huang and Perdikaris \[2026\]Xinquan Huang and Paris Perdikaris\.PhysicsCorrect: A training\-free approach for stable neural PDE simulations\.*Proceedings of the AAAI Conference on Artificial Intelligence*, 40\(26\):22057–22065, 2026\.doi:[10\.1609/aaai\.v40i26\.39360](https://doi.org/10.1609/aaai.v40i26.39360)\.
- Prudhomme and Oden \[1999\]Serge Prudhomme and J\. Tinsley Oden\.On goal\-oriented error estimation for elliptic problems: Application to the control of pointwise errors\.*Computer Methods in Applied Mechanics and Engineering*, 176\(1–4\):313–331, 1999\.doi:[10\.1016/S0045\-7825\(98\)00343\-0](https://doi.org/10.1016/S0045-7825(98)00343-0)\.
- Oden and Prudhomme \[2001\]J\. Tinsley Oden and Serge Prudhomme\.Goal\-oriented error estimation and adaptivity for the finite element method\.*Computers & Mathematics with Applications*, 41\(5–6\):735–756, 2001\.doi:[10\.1016/S0898\-1221\(00\)00317\-5](https://doi.org/10.1016/S0898-1221(00)00317-5)\.
- Becker and Rannacher \[2001\]Roland Becker and Rolf Rannacher\.An optimal control approach to a posteriori error estimation in finite element methods\.*Acta Numerica*, 10:1–102, 2001\.doi:[10\.1017/S0962492901000010](https://doi.org/10.1017/S0962492901000010)\.
- Giles and Süli \[2002\]Michael B\. Giles and Endre Süli\.Adjoint methods for PDEs: A posteriori error analysis and postprocessing by duality\.*Acta Numerica*, 11:145–236, 2002\.doi:[10\.1017/S096249290200003X](https://doi.org/10.1017/S096249290200003X)\.
- Hesthaven et al\. \[2016\]Jan S\. Hesthaven, Gianluigi Rozza, and Benjamin Stamm\.*Certified Reduced Basis Methods for Parametrized Partial Differential Equations*\.SpringerBriefs in Mathematics\. Springer, 2016\.doi:[10\.1007/978\-3\-319\-22470\-1](https://doi.org/10.1007/978-3-319-22470-1)\.
- Fanaskov et al\. \[2024\]Vladimir Fanaskov, Alexander Rudikov, and Ivan Oseledets\.Neural functional a posteriori error estimates\.*arXiv preprint arXiv:2402\.05585*, 2024\.doi:[10\.48550/arXiv\.2402\.05585](https://doi.org/10.48550/arXiv.2402.05585)\.
- Eiras et al\. \[2024\]Francisco Eiras, Adel Bibi, Rudy R\. Bunel, Krishnamurthy Dvijotham, Philip Torr, and M\. Pawan Kumar\.Efficient error certification for physics\-informed neural networks\.In*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pages 12318–12347\. PMLR, 2024\.URL[https://proceedings\.mlr\.press/v235/eiras24a\.html](https://proceedings.mlr.press/v235/eiras24a.html)\.
- Ronneberger et al\. \[2015\]Olaf Ronneberger, Philipp Fischer, and Thomas Brox\.U\-Net: Convolutional networks for biomedical image segmentation\.In*Medical Image Computing and Computer\-Assisted Intervention – MICCAI 2015*, volume 9351 of*Lecture Notes in Computer Science*, pages 234–241\. Springer, 2015\.doi:[10\.1007/978\-3\-319\-24574\-4\_28](https://doi.org/10.1007/978-3-319-24574-4_28)\.
- Trefethen \[2000\]Lloyd N\. Trefethen\.*Spectral Methods in MATLAB*, volume 10 of*Software, Environments, and Tools*\.Society for Industrial and Applied Mathematics, Philadelphia, PA, 2000\.doi:[10\.1137/1\.9780898719598](https://doi.org/10.1137/1.9780898719598)\.
- Hairer et al\. \[2006\]Ernst Hairer, Christian Lubich, and Gerhard Wanner\.*Geometric Numerical Integration: Structure\-Preserving Algorithms for Ordinary Differential Equations*, volume 31 of*Springer Series in Computational Mathematics*\.Springer, Berlin, Heidelberg, 2 edition, 2006\.doi:[10\.1007/3\-540\-30666\-8](https://doi.org/10.1007/3-540-30666-8)\.

Similar Articles

@AnimaAnandkumar: This is something I have been emphasizing since we started our work on Neural Operators. We very quickly went from simp…

X AI KOLs Following

Anima Anandkumar highlights that neural operators, despite simple benchmarks, have achieved massive speedups (10,000–million times) in hard real-world problems like high-resolution AI weather modeling (FourCastNet) and nuclear fusion turbulence, referencing a new paper showing learned solvers become more cost-effective as PDE tasks get harder.

Sequential Physics-Constrained Neural Operator Forward Modeling for the $\textit{Norne}$ Reservoir System

arXiv cs.LG

This paper presents a comprehensive mathematical framework for sequential surrogate modeling of three-phase black-oil reservoir dynamics using Fourier Neural Operators (FNO) and physics-informed variants (PINO), applied to the Norne benchmark reservoir. Theoretical contributions include functional-analytic formulation, covariate shift analysis, physics-constrained spectral stability, and truncated backpropagation gradient analysis.

Physics-Unrolled Neural Operator for Wireless Field Modeling

arXiv cs.LG

The paper introduces Physics-Unrolled Hybrid Neural Operator (PU-HNO), a model that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs by capturing propagation effects like reflection, diffraction, and scattering, outperforming existing baselines.