STCO: Conditional Neural Operators for Time-Dependent PDEs

arXiv cs.AI Papers

Summary

The paper presents STCO, a conditional neural operator for prescribed-condition operator learning in time-dependent PDEs, which improves prediction accuracy by incorporating various condition fields and shows significant error reductions across multiple architectures.

arXiv:2608.20477v1 Announce Type: new Abstract: Neural operators have emerged as efficient surrogates for time-dependent physical systems governed by partial differential equations (PDEs), but their future-state predictions are often conditioned only on observed states and static problem descriptors. For control or optimization, however, body motion, inflow, or forcing are prescribed for the query without being determined solely by the observed state. We introduce the Spatiotemporal Conditional Operator (STCO) for prescribed-condition operator learning (PCOL), a common interface that supplies prescribed target-time condition fields to heterogeneous backbone architectures while retaining their architecture-specific core computation and context pathways. Its condition interface combines Flow-Aware Graph Leaf (FAGL) with Dual-Site Feature-wise Linear Modulation (DSFiLM). Non-learned FAGL uses vorticity from the final observed frame to construct a fixed-cardinality adaptive partition, then co-locates the observed history and target-time condition fields at its regional coordinates. DSFiLM injects separate motion, inflow, and force routes before and after operator computation through current-feature-driven slot- and channel-wise gates. We evaluate twelve matched backbone architectures with different existing physical and temporal inputs. The immersed-boundary computational fluid dynamics (CFD) benchmark spans prescribed motion, inflow disturbances, body-force actuation, and morphology. Across twelve matched backbones, three regimes, and two lead ranges, STCO yields mean paired reductions of 31.1% in relative-L2 field error and 24.7% in normalized pressure-derived load error. It also lowers longer-lead field error for 11 backbones, while interventions on individual condition groups produce measurable prediction changes for every group evaluated.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:10 AM

# STCO: Conditional Neural Operators for Time-Dependent PDEs
Source: [https://arxiv.org/html/2608.20477](https://arxiv.org/html/2608.20477)
###### Abstract

Neural operators have emerged as efficient surrogates for time\-dependent physical systems governed by partial differential equations \(PDEs\), but their future\-state predictions are often conditioned only on observed states and static problem descriptors\. For control or optimization, however, body motion, inflow, or forcing are prescribed for the query without being determined solely by the observed state\. We introduce the Spatiotemporal Conditional Operator \(STCO\) for prescribed\-condition operator learning \(PCOL\), a common interface that supplies prescribed target\-time condition fields to heterogeneous backbone architectures while retaining their architecture\-specific core computation and context pathways\. Its condition interface combines Flow\-Aware Graph Leaf \(FAGL\) with Dual\-Site Feature\-wise Linear Modulation \(DSFiLM\)\. Non\-learned FAGL uses vorticity from the final observed frame to construct a fixed\-cardinality adaptive partition, then co\-locates the observed history and target\-time condition fields at its regional coordinates\. DSFiLM injects separate motion, inflow, and force routes before and after operator computation through current\-feature\-driven slot\- and channel\-wise gates\. We evaluate twelve matched backbone architectures with different existing physical and temporal inputs\. The immersed\-boundary computational fluid dynamics \(CFD\) benchmark spans prescribed motion, inflow disturbances, body\-force actuation, and morphology\. Across twelve matched backbones, three regimes, and two lead ranges, STCO yields mean paired reductions of 31\.1% in relative\-L2L\_\{2\}field error and 24\.7% in normalized pressure\-derived load error\. It also lowers longer\-lead field error for 11 backbones, while interventions on individual condition groups produce measurable prediction changes for every group evaluated\.

## Introduction

Neural operators learn mappings between function spaces and provide efficient surrogates for parametric partial differential equations \(PDEs\)\([Li et al\. 2021](https://arxiv.org/html/2608.20477#bib.bib20);[Lu et al\. 2021](https://arxiv.org/html/2608.20477#bib.bib24);[Li et al\. 2023b](https://arxiv.org/html/2608.20477#bib.bib22)\)\. For time\-dependent PDEs, they commonly forecast a future field from an initial state or observed trajectory, with lead time or auxiliary trajectories supplied as context\([Alkin et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib1);[Herde et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib10);[Mousavi et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib28);[Kassaï Koupaï et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib19)\)\. This formulation is effective when the requested future is determined by the observed dynamics and fixed problem descriptors\.

Control and optimization often require a different query\. The system response must be evaluated under a future scenario specified for the query, such as a gust, target body geometry, or prescribed boundary input\([Bartels 2013](https://arxiv.org/html/2608.20477#bib.bib3);[Sedky et al\. 2022](https://arxiv.org/html/2608.20477#bib.bib35);[Han et al\. 2021](https://arxiv.org/html/2608.20477#bib.bib8);[Hu and Liu 2025](https://arxiv.org/html/2608.20477#bib.bib12)\)\. The history describes the current dynamics, while the prescribed condition distinguishes the future to be evaluated\. As shown in Figure[1](https://arxiv.org/html/2608.20477#Sx1.F1), we formulate this prediction problem as*prescribed\-condition operator learning*\(PCOL\)\. In PCOL, a conditional neural operator maps an observed history, a lead time, and physical condition fields specified at the target time to the corresponding future response\. This query structure is also shared by predictive world models that infer future states from past observations and a supplied action or intervention\([NVIDIA et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib29);[Hafner et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib7)\)\. PCOL resolves the response as a PDE solution field that can be evaluated in downstream control or optimization\.

![Refer to caption](https://arxiv.org/html/2608.20477v1/figures/submission/prescribed_condition_pde.png)Figure 1:Prescribed\-condition operator learning in moving\-body flow\. The dashed frame marks the PCOL map from observed history and prescribed condition to future response\. Examples include a vortical encounter, body motion, a convecting gust, and body\-force actuation\. Response panels show vorticity and pressure\-derived load histories from CFD ground truth \(GT\)\. The predicted velocity–pressure field supports downstream control and optimization\.We examine PCOL through two\-dimensional incompressible moving\-body flow\. Geometry and motion alter the boundary, inflow disturbances modify inlet data, and distributed forces enter the momentum equation, allowing distinct prescribed mechanisms to be varied within one Navier–Stokes system\([Mittal and Iaccarino 2005](https://arxiv.org/html/2608.20477#bib.bib27);[Sotiropoulos and Yang 2014](https://arxiv.org/html/2608.20477#bib.bib36)\)\. This setting exposes three linked challenges\.1\) Representation alignment\.Observed and prescribed fields differ in channels, spatial support, and physical meaning\.2\) Selective condition modulation\.The model must preserve the roles of geometry, inflow, and forcing while adapting their influence to the current features\.3\) Sparse\-event supervision\.Brief force and inflow events provide few condition\-revealing pairs under uniform temporal sampling\.

We introduce the*Spatiotemporal Conditional Operator*\(STCO\), shown in Figure[2](https://arxiv.org/html/2608.20477#Sx3.F2), as a common condition interface for heterogeneous backbone architectures\. Its*Flow\-Aware Graph Leaf*\(FAGL\) representation forms a fixed\-cardinality, vorticity\-aware partition and co\-locates historical and prescribed fields at shared regional slots\.*Dual\-Site Feature\-wise Linear Modulation*\(DSFiLM\) assigns framewise motion, inflow, and force to separate routes, with signed\-distance geometry localizing the motion route\. Its independently parameterized IN\-DSFiLM and OUT\-DSFiLM modules act before and after the backbone core, respectively, while current features gate each route by slot and channel\. External\-activity sampling increases exposure to onset crossings and active intervals\. STCO adds target\-time conditioning while retaining each backbone’s core computation and existing context pathways\.

We evaluate this interface through twelve matched pairs of baseline \(Base\) and STCO configurations\. Within each pair, FAGL, external\-activity sampling, data, backbone, decoder, and training protocol are fixed\. The Base models retain their original inputs\. Eleven receive observed\-frame conditions, and five also encode lead time\. STCO preserves these inputs and activates DSFiLM, which combines target\-time condition routes with lead\-time modulation\. The paired differences measure the predictive value of this complete interface under an otherwise matched representation and training protocol\. STCO yields a positive mean field gain across the six regime–lead strata for all 12 backbones, lowers the field error beyond the training lead range for 11, and reduces pressure\-derived load error overall\.*Modulation counterfactual*\(MCF\) interventions establish prediction sensitivity to every evaluated spatial condition group\.

Our contributions are summarized below\.

- •We formulate PCOL as predicting a future response from observed history and physical conditions prescribed for the query\. We address this problem with STCO, a common condition interface that retains the core computation of heterogeneous backbone architectures\.
- •We introduce vorticity\-aware FAGL for adaptive spatial alignment of observed fields and prescribed conditions\. DSFiLM modulates separate motion, inflow, and force routes before and after backbone computation, with their influence gated by the current features\.
- •We construct a moving\-body\-flow benchmark and evaluate STCO in matched comparisons across twelve backbones\. External\-activity sampling increases exposure to sparse force and inflow events\. MCF measures prediction sensitivity to each evaluated spatial condition group\.

## Related Work

Prescribed response and conditional operators\.Moving\-boundary predictors infer the next flow field from recent flow fields and boundary\-position histories\([Han et al\. 2021](https://arxiv.org/html/2608.20477#bib.bib8)\)\. A closely related rigid\-body fluid–structure interaction study predicts the terminal fluid state from an initial state and a prescribed motion sequence over the full prediction interval\([Zhong and Meidani 2025](https://arxiv.org/html/2608.20477#bib.bib39)\)\. It compares direct concatenation, temporal pooling, and sequential coupling\. Boundary\-control operators map temporal boundary inputs to boundary\-output trajectories\([Hu and Liu 2025](https://arxiv.org/html/2608.20477#bib.bib12)\)\. Other neural operators evaluate candidate controls within an outer optimization loop or predict controller gains and feedback inputs\([Hwang et al\. 2022](https://arxiv.org/html/2608.20477#bib.bib14);[Bhan, Shi, and Krstić 2024](https://arxiv.org/html/2608.20477#bib.bib4)\)\. PCOL addresses forward response from observed history and target\-time condition fields, leaving control selection downstream\. MIONet represents products of function spaces\([Jin, Meng, and Lu 2022](https://arxiv.org/html/2608.20477#bib.bib15)\), while GNOT, GINO, and Unisolver process heterogeneous inputs, geometry, or structured PDE descriptors\([Hao et al\. 2023](https://arxiv.org/html/2608.20477#bib.bib9);[Li et al\. 2023b](https://arxiv.org/html/2608.20477#bib.bib22);[Zhou et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib40)\)\. GEPS learns environment conditioning, and time\-dependent operators incorporate lead time or context trajectories\([Kassaï Koupaï et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib18);[Herde et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib10);[Mousavi et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib28);[Kassaï Koupaï et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib19)\)\. STCO organizes multiple target\-time physical fields through one interface rather than introducing another backbone architecture\.

Adaptive representation and modulation\.Irregular\-domain operators use coordinate deformation, graph–grid transfer, regional graphs, or latent tokens\([Li et al\. 2023a](https://arxiv.org/html/2608.20477#bib.bib21);[Li et al\. 2023b](https://arxiv.org/html/2608.20477#bib.bib22);[Mousavi et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib28);[Alkin et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib1)\)\. UPT also conditions Transformer and Perceiver blocks on time or velocity context through feature modulation\([Alkin et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib1)\)\. FAGL draws on flow\-guided refinement\([Kamkar et al\. 2011](https://arxiv.org/html/2608.20477#bib.bib16)\)and graph\-based component merging\([Felzenszwalb and Huttenlocher 2004](https://arxiv.org/html/2608.20477#bib.bib5)\)to allocate a fixed slot budget and align fields at shared indices\. FiLM applies condition\-dependent affine transformations, SPADE makes them spatially varying, and squeeze–excitation derives channel weights from current features\([Perez et al\. 2018](https://arxiv.org/html/2608.20477#bib.bib32);[Park et al\. 2019](https://arxiv.org/html/2608.20477#bib.bib30);[Hu, Shen, and Sun 2018](https://arxiv.org/html/2608.20477#bib.bib11)\)\. DSFiLM combines these principles through separate spatial condition routes, local state\-dependent gates, and independently learned IN\-DSFiLM and OUT\-DSFiLM modules\.

## Methodology

### Problem Formulation

Setup\.We instantiate PCOL for time\-dependent PDEs in incompressible moving\-body flow\. Let𝒙∈Ω⁡\(t\)⊂ℝ2\\bm\{x\}\\in\\Omega\(t\)\\subset\\mathbb\{R\}^\{2\}denote a point in the time\-dependent fluid domain and∂Ωb​\(t\)\\partial\\Omega\_\{b\}\(t\)the moving body boundary\. Using body chordLL, reference inflow speedU∞U\_\{\\infty\}, and constant densityρ\\rhoas the length, velocity, and density scales, we scale time, pressure, and acceleration byL/U∞L/U\_\{\\infty\},ρ​U∞2/2\\rho U\_\{\\infty\}^\{2\}/2, andU∞2/LU\_\{\\infty\}^\{2\}/L, respectively\. After dropping dimensionless superscripts, the velocity𝒖=\(u,v\)\\bm\{u\}=\(u,v\)and pressure coefficientppsatisfy

∂t𝒖\+\(𝒖⋅∇\)𝒖\\displaystyle\\partial\_\{t\}\\bm\{u\}\+\(\\bm\{u\}\\cdot\\nabla\)\\bm\{u\}=−12∇p\+Re−1∇2𝒖\+𝒈,\\displaystyle=\-\\tfrac\{1\}\{2\}\\nabla p\+\\mathrm\{Re\}^\{\-1\}\\nabla^\{2\}\\bm\{u\}\+\\bm\{g\},\(1\)∇⋅𝒖\\displaystyle\\nabla\\cdot\\bm\{u\}=0\.\\displaystyle=0\.Here,Re=U∞​L/ν\\mathrm\{Re\}=U\_\{\\infty\}L/\\nu,ν\\nuis the kinematic viscosity, and𝒈⁡\(𝒙,t\)\\bm\{g\}\(\\bm\{x\},t\)is a prescribed body\-force acceleration\. The prescribed boundary conditions are𝒖=𝒖b\\bm\{u\}=\\bm\{u\}\_\{b\}on∂Ωb​\(t\)\\partial\\Omega\_\{b\}\(t\)and𝒖=𝒆x\+𝒖bc\\bm\{u\}=\\bm\{e\}\_\{x\}\+\\bm\{u\}\_\{\\mathrm\{bc\}\}on the inletΓin\\Gamma\_\{\\mathrm\{in\}\}\. Here,𝒖b\\bm\{u\}\_\{b\}is the prescribed body velocity,𝒆x=\(1,0\)⊤\\bm\{e\}\_\{x\}=\(1,0\)^\{\\top\}is the nondimensional base inflow, and𝒖bc\\bm\{u\}\_\{\\mathrm\{bc\}\}is the prescribed inflow perturbation\. Body kinematics define the signed\-distance fieldψ⁡\(𝒙,t\)\\psi\(\\bm\{x\},t\), whose zero level set is∂Ωb​\(t\)\\partial\\Omega\_\{b\}\(t\)\. We sampleTTstored frames at times\{tn\}n=0T−1\\\{t\_\{n\}\\\}\_\{n=0\}^\{T\-1\}\. At framenn, the domain isΩn=Ω⁡\(tn\)\\Omega\_\{n\}=\\Omega\(t\_\{n\}\), and for𝒙∈Ωn\\bm\{x\}\\in\\Omega\_\{n\}we define the response and prescribed\-condition fields

𝒚n​\(𝒙\)\\displaystyle\\bm\{y\}\_\{n\}\(\\bm\{x\}\)=\[un​\(𝒙\),vn​\(𝒙\),pn​\(𝒙\)\]⊤,\\displaystyle=\[u\_\{n\}\(\\bm\{x\}\),v\_\{n\}\(\\bm\{x\}\),p\_\{n\}\(\\bm\{x\}\)\]^\{\\top\},\(2\)𝒄n​\(𝒙\)\\displaystyle\\bm\{c\}\_\{n\}\(\\bm\{x\}\)=\[ψn​\(𝒙\),Δ​ψn​\(𝒙\),𝒈n​\(𝒙\)⊤,𝒖bc,n​\(𝒙\)⊤\]⊤\.\\displaystyle=\[\\psi\_\{n\}\(\\bm\{x\}\),\\Delta\\psi\_\{n\}\(\\bm\{x\}\),\\bm\{g\}\_\{n\}\(\\bm\{x\}\)^\{\\top\},\\bm\{u\}\_\{\\mathrm\{bc\},n\}\(\\bm\{x\}\)^\{\\top\}\]^\{\\top\}\.The framewise signed\-distance change isΔ​ψn=ψn−ψn−1\\Delta\\psi\_\{n\}=\\psi\_\{n\}\-\\psi\_\{n\-1\}forn≥1n\\geq 1, withΔ​ψ0=0\\Delta\\psi\_\{0\}=0\. The six condition channels encode geometry and framewise motion throughψn\\psi\_\{n\}andΔ​ψn\\Delta\\psi\_\{n\}, distributed forcing through𝒈n\\bm\{g\}\_\{n\}, and inflow perturbation through𝒖bc,n\\bm\{u\}\_\{\\mathrm\{bc\},n\}\.

Objective\.Letnrn\_\{r\}be the final observed reference index andnq\>nrn\_\{q\}\>n\_\{r\}the queried future index\. Subscriptsrrandqqhenceforth denote evaluation atnrn\_\{r\}andnqn\_\{q\}\. The observed history isℋr=\{\(𝒚nk,𝒄nk\)\}k=1K\\mathcal\{H\}\_\{r\}=\\\{\(\\bm\{y\}\_\{n\_\{k\}\},\\bm\{c\}\_\{n\_\{k\}\}\)\\\}\_\{k=1\}^\{K\}, whereKKis the number of observed frames andn1<⋯<nK=nrn\_\{1\}<\\cdots<n\_\{K\}=n\_\{r\}\. The query lead isΔ​n=nq−nr\\Delta n=n\_\{q\}\-n\_\{r\}, represented by the normalized scalarτ=Δ​n/\(T−1\)\\tau=\\Delta n/\(T\-1\)\. PCOL seeks an operator that maps the observed history and the condition prescribed for this query to its future response

𝒚^q​\(𝒙\)=𝒢θ​\(ℋr,𝒄q,𝒙,τ\)≈𝒚q​\(𝒙\),𝒙∈Ωq,\\widehat\{\\bm\{y\}\}\_\{q\}\(\\bm\{x\}\)=\\mathcal\{G\}\_\{\\theta\}\\bigl\(\\mathcal\{H\}\_\{r\},\\bm\{c\}\_\{q\},\\bm\{x\},\\tau\\bigr\)\\approx\\bm\{y\}\_\{q\}\(\\bm\{x\}\),\\qquad\\bm\{x\}\\in\\Omega\_\{q\},\(3\)where𝒚q=𝒚nq\\bm\{y\}\_\{q\}=\\bm\{y\}\_\{n\_\{q\}\},𝒄q=𝒄nq\\bm\{c\}\_\{q\}=\\bm\{c\}\_\{n\_\{q\}\},Ωq=Ωnq\\Omega\_\{q\}=\\Omega\_\{n\_\{q\}\}, andθ\\thetacontains the learned parameters\.

### STCO Architecture

![Refer to caption](https://arxiv.org/html/2608.20477v1/figures/submission/STCO_AAAI2027_framework_observed_support_600dpi.jpg)Figure 2:STCO architecture\. \(a\) Overall pipeline from observed history and prescribed target\-time conditions to the velocity–pressure response on𝑿q\\bm\{X\}\_\{q\}\. FAGL provides regional coordinates𝑿r\\bm\{X\}\_\{r\}, while IN\-DSFiLM and OUT\-DSFiLM bracket a selected backbone core\. \(b\) FAGL constructs an observed\-vorticity\-aware partition and aligns historical aggregates and target\-time conditions on𝑿r\\bm\{X\}\_\{r\}by four\-neighbor inverse\-distance weighting\. \(c\) DSFiLM applies lead\-time scaling and state\-gated motion, inflow, and force shifts, with signed distance localizing motion near the body\.STCO realizes the PCOL map through adaptive spatial alignment and conditional modulation, as shown in Figure[2](https://arxiv.org/html/2608.20477#Sx3.F2)\(a\)\. FAGL co\-locates the observed history and target\-time condition fields at shared regional slots\. Backbone\-specific adapters connect IN\-DSFiLM and OUT\-DSFiLM to the selected backbone core while preserving its computation\.

Let𝑿r=\[𝒙r,ℓ\]ℓ=1Nℓ∈ℝNℓ×2\\bm\{X\}\_\{r\}=\[\\bm\{x\}\_\{r,\\ell\}\]\_\{\\ell=1\}^\{N\_\{\\ell\}\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times 2\}denote theNℓN\_\{\\ell\}final\-observed\-frame slot coordinates, and let𝑿q=\[𝒙q,j\]j=1Nq∈ℝNq×2\\bm\{X\}\_\{q\}=\[\\bm\{x\}\_\{q,j\}\]\_\{j=1\}^\{N\_\{q\}\}\\in\\mathbb\{R\}^\{N\_\{q\}\\times 2\}contain theNqN\_\{q\}requested coordinates, with𝒙q,j∈Ωq\\bm\{x\}\_\{q,j\}\\in\\Omega\_\{q\}\. Let𝒁¯r∈ℝK×Nℓ×9\\overline\{\\bm\{Z\}\}\_\{r\}\\in\\mathbb\{R\}^\{K\\times N\_\{\\ell\}\\times 9\}contain the aligned history features of the three response and six condition channels, and let𝑪¯q∈ℝNℓ×6\\overline\{\\bm\{C\}\}\_\{q\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times 6\}contain the target\-time condition sampled at the same coordinates\. Itsℓ\\ellth slot is𝑪¯q,ℓ=\[ψ¯q,ℓ,Δ​ψ¯q,ℓ,𝒈¯q,ℓ⊤,𝒖¯bc,q,ℓ⊤\]⊤\\overline\{\\bm\{C\}\}\_\{q,\\ell\}=\[\\overline\{\\psi\}\_\{q,\\ell\},\\overline\{\\Delta\\psi\}\_\{q,\\ell\},\\overline\{\\bm\{g\}\}\_\{q,\\ell\}^\{\\top\},\\overline\{\\bm\{u\}\}\_\{\\mathrm\{bc\},q,\\ell\}^\{\\top\}\]^\{\\top\}, Here, overbars mark aligned values before standardization\. Vorticity determines the observed\-frame partition but does not enter either tensor\. Fixed channelwise maps standardize both quantities after aggregation and alignment, giving𝒁r=𝒩z​\(𝒁¯r\)\\bm\{Z\}\_\{r\}=\\mathcal\{N\}\_\{z\}\(\\overline\{\\bm\{Z\}\}\_\{r\}\)and𝑪q=𝒩c​\(𝑪¯q\)\\bm\{C\}\_\{q\}=\\mathcal\{N\}\_\{c\}\(\\overline\{\\bm\{C\}\}\_\{q\}\)\. The unstandardized nondimensional signed\-distance component𝝍¯q∈ℝNℓ\\overline\{\\bm\{\\psi\}\}\_\{q\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\}is retained for boundary localization\. The query leadτ\\tauremains a separate scalar input\. Evaluating the prescribed fieldψq\\psi\_\{q\}on𝑿q\\bm\{X\}\_\{q\}defines the query\-frame fluid mask used for training and evaluation\. For backbonebb, letℰb\\mathcal\{E\}\_\{b\},𝒪b\\mathcal\{O\}\_\{b\}, and𝒟b\\mathcal\{D\}\_\{b\}denote its input adapter, core computation, and dense readout, and letℳbin\\mathcal\{M\}\_\{b\}^\{\\mathrm\{in\}\}andℳbout\\mathcal\{M\}\_\{b\}^\{\\mathrm\{out\}\}denote its IN\-DSFiLM and OUT\-DSFiLM maps\. STCO computes

𝒉0\\displaystyle\\bm\{h\}\_\{0\}=ℰb​\(𝒁r,𝑿r,τ\),\\displaystyle=\\mathcal\{E\}\_\{b\}\(\\bm\{Z\}\_\{r\},\\bm\{X\}\_\{r\},\\tau\),𝒉1\\displaystyle\\bm\{h\}\_\{1\}=ℳbin​\(𝒉0,𝑪q,𝝍¯q,τ\),\\displaystyle=\\mathcal\{M\}\_\{b\}^\{\\mathrm\{in\}\}\(\\bm\{h\}\_\{0\},\\bm\{C\}\_\{q\},\\overline\{\\bm\{\\psi\}\}\_\{q\},\\tau\),\(4\)𝒉2\\displaystyle\\bm\{h\}\_\{2\}=𝒪b​\(𝒉1,𝑿r,τ\),\\displaystyle=\\mathcal\{O\}\_\{b\}\(\\bm\{h\}\_\{1\},\\bm\{X\}\_\{r\},\\tau\),𝒉3\\displaystyle\\bm\{h\}\_\{3\}=ℳbout​\(𝒉2,𝑪q,𝝍¯q,τ\),\\displaystyle=\\mathcal\{M\}\_\{b\}^\{\\mathrm\{out\}\}\(\\bm\{h\}\_\{2\},\\bm\{C\}\_\{q\},\\overline\{\\bm\{\\psi\}\}\_\{q\},\\tau\),𝒚^q​\(𝑿q\)\\displaystyle\\widehat\{\\bm\{y\}\}\_\{q\}\(\\bm\{X\}\_\{q\}\)=𝒟b​\(𝒉3,𝑿r,𝑿q\)\.\\displaystyle=\\mathcal\{D\}\_\{b\}\(\\bm\{h\}\_\{3\},\\bm\{X\}\_\{r\},\\bm\{X\}\_\{q\}\)\.Here,𝒉0,𝒉1∈ℝNℓ×Din\\bm\{h\}\_\{0\},\\bm\{h\}\_\{1\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times D\_\{\\mathrm\{in\}\}\}and𝒉2,𝒉3∈ℝNℓ×Dout\\bm\{h\}\_\{2\},\\bm\{h\}\_\{3\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times D\_\{\\mathrm\{out\}\}\}, whereDinD\_\{\\mathrm\{in\}\}andDoutD\_\{\\mathrm\{out\}\}are the respective interface widths\. For the shared DSFiLM form,\(𝒉in,𝒉out\)=\(𝒉0,𝒉1\)\(\\bm\{h\}\_\{\\mathrm\{in\}\},\\bm\{h\}\_\{\\mathrm\{out\}\}\)=\(\\bm\{h\}\_\{0\},\\bm\{h\}\_\{1\}\)for IN\-DSFiLM and\(𝒉2,𝒉3\)\(\\bm\{h\}\_\{2\},\\bm\{h\}\_\{3\}\)for OUT\-DSFiLM\. The two modulation maps have independent parameters\.

#### Flow\-Aware Graph Leaf\.

FAGL maps each observed grid toNℓN\_\{\\ell\}adaptive regional slots through graph construction, partition, and remapping, as shown in Figure[2](https://arxiv.org/html/2608.20477#Sx3.F2)\(b\)\.Graph\.At observed framenn, letωn​\(𝒙\)=∂xvn​\(𝒙\)−∂yun​\(𝒙\)\\omega\_\{n\}\(\\bm\{x\}\)=\\partial\_\{x\}v\_\{n\}\(\\bm\{x\}\)\-\\partial\_\{y\}u\_\{n\}\(\\bm\{x\}\)denote the scalar vorticity andμn​\(𝒙\)=\|ωn​\(𝒙\)\|\\mu\_\{n\}\(\\bm\{x\}\)=\|\\omega\_\{n\}\(\\bm\{x\}\)\|its magnitude\. An eight\-neighbor grid graph assigns edge dissimilarity from spatial separation and contrasts inμn\\mu\_\{n\}and‖∇μn‖2\\\|\\nabla\\mu\_\{n\}\\\|\_\{2\}\. Processing edges in increasing dissimilarity merges adjacent components when the connecting edge is small relative to their internal variation\([Felzenszwalb and Huttenlocher 2004](https://arxiv.org/html/2608.20477#bib.bib5)\)\. This retains sharp vorticity changes for refinement\.Partition\.Each eligible componentCCin refinement batchℬ\\mathcal\{B\}is ranked by

𝒫n​\(C\)=\|C\|2\+∑𝒙i∈Cμn​\(𝒙i\)μ¯ℬ\+εμ\\mathcal\{P\}\_\{n\}\(C\)=\\frac\{\|C\|\}\{2\}\+\\frac\{\\sum\_\{\\bm\{x\}\_\{i\}\\in C\}\\mu\_\{n\}\(\\bm\{x\}\_\{i\}\)\}\{\\bar\{\\mu\}\_\{\\mathcal\{B\}\}\+\\varepsilon\_\{\\mu\}\}\(5\)where\|C\|\|C\|is its number of grid points,μ¯ℬ\\bar\{\\mu\}\_\{\\mathcal\{B\}\}is the mean ofμn\\mu\_\{n\}overℬ\\mathcal\{B\}, andεμ=10−6\\varepsilon\_\{\\mu\}=10^\{\-6\}stabilizes the denominator\. The two terms preserve spatial coverage and direct more slots to integrated vorticity activity\. Inspired by feature\-driven refinement for vortex\-dominated flows\([Kamkar et al\. 2011](https://arxiv.org/html/2608.20477#bib.bib16)\), selected components are median\-split along the best ofxx,yy, and signedωn\\omega\_\{n\}; the signed candidate separates opposing rotations when selected\. Refinement ends atNℓN\_\{\\ell\}or when no eligible split remains, after which the collection is normalized to exactlyNℓN\_\{\\ell\}slots\.Remap\.For the grid points𝒮n,ℓ\\mathcal\{S\}\_\{n,\\ell\}in slotℓ\\ell, let𝒜n,ℓ​\[𝒇\]\\mathcal\{A\}\_\{n,\\ell\}\[\\bm\{f\}\]be the mean of field𝒇\\bm\{f\}and𝒙n,ℓ\\bm\{x\}\_\{n,\\ell\}their centroid\. Historical aggregates and the target\-time condition are aligned to𝑿r\\bm\{X\}\_\{r\}by

𝒛¯k,ℓ\\displaystyle\\overline\{\\bm\{z\}\}\_\{k,\\ell\}=IDW4⁡\(\{𝒙nk,j,𝒜nk,j​\[𝒚nk,𝒄nk\]\}j=1Nℓ,𝒙r,ℓ\),\\displaystyle=\\operatorname\{IDW\}\_\{4\}\\\!\\left\(\\left\\\{\\bm\{x\}\_\{n\_\{k\},j\},\\mathcal\{A\}\_\{n\_\{k\},j\}\[\\bm\{y\}\_\{n\_\{k\}\},\\bm\{c\}\_\{n\_\{k\}\}\]\\right\\\}\_\{j=1\}^\{N\_\{\\ell\}\},\\bm\{x\}\_\{r,\\ell\}\\right\),\(6\)𝑪¯q,ℓ\\displaystyle\\overline\{\\bm\{C\}\}\_\{q,\\ell\}=IDW4⁡\[𝒄q\]​\(𝒙r,ℓ\)\.\\displaystyle=\\operatorname\{IDW\}\_\{4\}\[\\bm\{c\}\_\{q\}\]\(\\bm\{x\}\_\{r,\\ell\}\)\.Here,IDW4\\operatorname\{IDW\}\_\{4\}uses normalized inverse\-distance weights over four nearest sources; the second line interpolates directly from the target\-time condition mesh\. Stacking𝒛¯k,ℓ\\overline\{\\bm\{z\}\}\_\{k,\\ell\}gives𝒁¯r\\overline\{\\bm\{Z\}\}\_\{r\}\. The supplementary material specifies the merge, refinement, interpolation, and slot\-normalization rules\.

#### Dual\-Site Feature\-wise Linear Modulation\.

Spatial alignment places prescribed mechanisms on common slots but leaves their influence on current features unresolved\. As detailed in Figure[2](https://arxiv.org/html/2608.20477#Sx3.F2)\(c\), DSFiLM combines feature\-wise affine conditioning in FiLM\([Perez et al\. 2018](https://arxiv.org/html/2608.20477#bib.bib32)\), spatially varying modulation in SPADE\([Park et al\. 2019](https://arxiv.org/html/2608.20477#bib.bib30)\), and feature\-dependent channel gating in squeeze–excitation\([Hu, Shen, and Sun 2018](https://arxiv.org/html/2608.20477#bib.bib11)\)\. IN\-DSFiLM acts before the backbone core and OUT\-DSFiLM before dense readout\. The two sites have independent parameters\.

Consider a slot at either site and omit site and slot indices\. Let𝒉in∈ℝD\\bm\{h\}\_\{\\mathrm\{in\}\}\\in\\mathbb\{R\}^\{D\}be the feature entering that module and𝒉out\\bm\{h\}\_\{\\mathrm\{out\}\}its modulated output\. Three independent two\-layer networks encode standardized localΔ​ψq\\Delta\\psi\_\{q\},𝒖bc,q\\bm\{u\}\_\{\\mathrm\{bc\},q\}, and𝒈q\\bm\{g\}\_\{q\}channels as the shifts𝜷mot\\bm\{\\beta\}\_\{\\mathrm\{mot\}\},𝜷inflow\\bm\{\\beta\}\_\{\\mathrm\{inflow\}\}, and𝜷force∈ℝD\\bm\{\\beta\}\_\{\\mathrm\{force\}\}\\in\\mathbb\{R\}^\{D\}\. Motion acts near the body, so its shift is multiplied byw=exp\[−ψ¯q2/\(2δ2\)\]w=\\exp\[\-\\overline\{\\psi\}\_\{q\}^\{\\,2\}/\(2\\delta^\{2\}\)\]\.ψ¯q\\overline\{\\psi\}\_\{q\}denotes the unstandardized signed distance at that slot andδ\>0\\delta\>0is a learned bandwidth\. A fourth two\-layer network maps the sinusoidal encoding ofτ\\tauto a lead\-time scale𝜸⁡\(τ\)∈ℝD\\bm\{\\gamma\}\(\\tau\)\\in\\mathbb\{R\}^\{D\}\. This scale is shared across slots\. A sigmoid affine projection of𝒉in\\bm\{h\}\_\{\\mathrm\{in\}\}produces three channel\-wise gates𝜶mot\\bm\{\\alpha\}\_\{\\mathrm\{mot\}\},𝜶inflow\\bm\{\\alpha\}\_\{\\mathrm\{inflow\}\}, and𝜶force∈\(0,1\)D\\bm\{\\alpha\}\_\{\\mathrm\{force\}\}\\in\(0,1\)^\{D\}\. The modulation is

𝒉out\\displaystyle\\bm\{h\}\_\{\\mathrm\{out\}\}=𝒉in⊙\[𝟏\+𝜸⁡\(τ\)\]\+𝜶mot⊙\(w​𝜷mot\)\\displaystyle=\\bm\{h\}\_\{\\mathrm\{in\}\}\\odot\[\\bm\{1\}\+\\bm\{\\gamma\}\(\\tau\)\]\+\\bm\{\\alpha\}\_\{\\mathrm\{mot\}\}\\odot\(w\\bm\{\\beta\}\_\{\\mathrm\{mot\}\}\)\(7\)\+𝜶inflow⊙𝜷inflow\+𝜶force⊙𝜷force,\\displaystyle\+\\bm\{\\alpha\}\_\{\\mathrm\{inflow\}\}\\odot\\bm\{\\beta\}\_\{\\mathrm\{inflow\}\}\+\\bm\{\\alpha\}\_\{\\mathrm\{force\}\}\\odot\\bm\{\\beta\}\_\{\\mathrm\{force\}\},where𝟏∈ℝD\\bm\{1\}\\in\\mathbb\{R\}^\{D\}is the all\-ones vector and⊙\\odotdenotes elementwise multiplication\. Inspired by the identity\-preserving zero\-initialization principle of adaLN\-Zero\([Peebles and Xie 2023](https://arxiv.org/html/2608.20477#bib.bib31)\), the final scale and shift layers are initialized at zero, so Equation[7](https://arxiv.org/html/2608.20477#Sx3.E7)begins with𝒉out=𝒉in\\bm\{h\}\_\{\\mathrm\{out\}\}=\\bm\{h\}\_\{\\mathrm\{in\}\}\.

### External\-Activity Sampling

Brief body\-force and inflow events yield few condition\-revealing pairs under uniform reference–query sampling\. For simulationmm, letam​\(n\)∈\[0,1\]a\_\{m\}\(n\)\\in\[0,1\]be the spatial mean of‖𝒈m,n‖22\+‖𝒖bc,m,n‖22\\\|\\bm\{g\}\_\{m,n\}\\\|\_\{2\}^\{2\}\+\\\|\\bm\{u\}\_\{\\mathrm\{bc\},m,n\}\\\|\_\{2\}^\{2\}over grid points that are fluid in at least one stored frame, normalized by its maximum over that simulation\. For simulations without external force or inflow,am​\(n\)=0a\_\{m\}\(n\)=0for everynn\. Within the selected simulation, an admissible reference–query pair is sampled according to

Pr⁡\(nr,Δ​n∣m\)∝ηm​\(nr,Δ​n\)​\[ε\+∑j=0Δ​nam​\(nr\+j\)\]\.\\Pr\(n\_\{r\},\\Delta n\\mid m\)\\propto\\eta\_\{m\}\(n\_\{r\},\\Delta n\)\\left\[\\varepsilon\+\\sum\_\{j=0\}^\{\\Delta n\}a\_\{m\}\(n\_\{r\}\+j\)\\right\]\.\(8\)The sum accumulates external activity from the reference frame through the query frame\. The floorε=0\.1\\varepsilon=0\.1gives every admissible pair nonzero probability\. The onset factorηm​\(nr,Δ​n\)=10\\eta\_\{m\}\(n\_\{r\},\\Delta n\)=10when an interval crosses the first frame where any force or inflow component exceeds10−310^\{\-3\}in magnitude\. It is11otherwise\. Motion remains part of𝒄q\\bm\{c\}\_\{q\}but is excluded fromama\_\{m\}, so the activity weighting targets force and inflow events\. Every matched Base–STCO pair receives identical temporal and spatial samples\. The supplementary material gives the complete rule\.

## Experiments

### Experimental Setup

Datasets\.The benchmark contains142142two\-dimensional simulations of incompressible moving\-body flow generated withWaterLily\.jl\([Weymouth and Font 2025](https://arxiv.org/html/2608.20477#bib.bib38)\)\. Each simulation contains256256stored frames, giving36,35236\{,\}352frames in total\. All simulations useRe=5000\\mathrm\{Re\}=5000\. The benchmark spans five swimmer morphologies, an undisturbed baseline, and six prescribed\-condition families\. Three families use analytic instantiations of Lamb–Oseen vortex encounters, convecting finite\-width velocity gusts, and localized transverse gusts\. Their physical settings are motivated by prior experimental and computational studies\([Hufstedler and McKeon 2019](https://arxiv.org/html/2608.20477#bib.bib13);[Bartels 2013](https://arxiv.org/html/2608.20477#bib.bib3);[Sedky et al\. 2022](https://arxiv.org/html/2608.20477#bib.bib35)\)\. The remaining families comprise localized Gaussian body\-force actuation, prescribed pose or undulation changes, and compound motion–disturbance cases\. The held\-out simulations vary in amplitude, onset, placement, duration, waveform, and morphology\. The training, validation, and test sets contain9292,88, and4242simulations, respectively, with no simulation shared across splits\. The reference swimming kinematics, analytic condition profiles, waveform variants, parameter values, and case allocation are given in the supplementary material\.

Backbones\.We evaluate the same STCO interface on twelve architecture cores adapted to a common FAGL\-to\-query contract: MGN\([Pfaff et al\. 2021](https://arxiv.org/html/2608.20477#bib.bib33)\), RIGNO\([Mousavi et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib28)\), PINO\([Li et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib23)\), Poseidon\([Herde et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib10)\), GAOT\([Wen et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib37)\), GINO\([Li et al\. 2023b](https://arxiv.org/html/2608.20477#bib.bib22)\), Transolver\+\+\([Luo et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib25)\), Unisolver\([Zhou et al\. 2025](https://arxiv.org/html/2608.20477#bib.bib40)\), MPP\([McCabe et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib26)\), CALM\-PDE\([Hagnberger, Musekamp, and Niepert 2025](https://arxiv.org/html/2608.20477#bib.bib6)\), UPT\([Alkin et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib1)\), and GEPS\([Kassaï Koupaï et al\. 2024](https://arxiv.org/html/2608.20477#bib.bib18)\)\. They span graph, spectral, geometry\-aware, transformer, and latent operator architectures\. The labels denote adapted architecture cores rather than reproductions of the original training recipes\. All configurations are trained from scratch with the common data objective\. PINO uses its FNO core without the physics\-informed loss, and Poseidon uses its ScOT core\. MPP denotes a Transformer core without multiple\-physics pretraining, while GEPS uses low\-rank context layers without inference\-time adaptation\.

Implementation\.Vorticity\-aware FAGL represents each grid with12,00012\{,\}000adaptive slots, approximately9\.4%9\.4\\%of the mean fluid\-grid cardinality\. Across five representative simulations, mean channelwise reconstruction errors range from1\.5%1\.5\\%to11\.8%11\.8\\%\. With one observed frame, each epoch samples10,00010\{,\}000reference–query pairs with replacement atΔ​n=1\\Delta n=1–2020\. Each pair resamples15,00015\{,\}000supervision coordinates, nominally comprising80%80\\%FAGL\-region representatives and20%20\\%additional nonrepresentative locations\. After framewise outlier removal and fixed standardization, all models minimize the fluid\-masked channel\-mean relative\-L2L\_\{2\}loss over standardizeduu,vv, andpp\. Training runs for8080epochs on one H200 GPU using AdamW with an initial learning rate of5×10−45\\times 10^\{\-4\}and batch size3232\. In paired experiments, Base and STCO share the split, samples, FAGL, retained backbone context, adapters, decoder, query geometry, and optimization\. Base bypasses both DSFiLM sites, whereas STCO activates the complete interface\.

Table 1:Regime\-balanced accuracy and MCF sensitivity across twelve matched backbones\. ID/OOD denoteΔ​n=1\\Delta n=1–2020/2121–4040; errors macro\-average three regimes\. PositiveΔ\\Deltafavors STCO, and bold marks the lower paired error\. The Average row macro\-averages backbone\-level entries, including paired reductions\. MeanΔy\\Delta\_\{y\}averages six field gains; MCF reports lead\-time and mean spatial\-condition sensitivity\.![[Uncaptioned image]](https://arxiv.org/html/2608.20477v1/figures/submission/stco_performance_triptych.png)

Figure 3:Cross\-backbone accuracy, condition dependence, and inference cost\.\(a\)H200 median dense\-query latency atNq=264,196N\_\{q\}=264\{,\}196and batch size one versus six\-stratum regime\-balancedEyE\_\{y\}; arrows connect matched Base and STCO configurations\.\(b\)Cross\-backbone median and interquartile range ofEyE\_\{y\}at each lead\. Annotations average the relative reductions of exact\-lead medians within ID\-Lead and OOD\-Lead, whereas MeanΔy\\Delta\_\{y\}in Table[1](https://arxiv.org/html/2608.20477#Sx4.T1)averages six stratum\-level gains per backbone\.\(c\)Cross\-backbone means of ID/OODEyE\_\{y\},EFpE\_\{F\_\{p\}\}, MCF, and latency under shared normalization; MCF equally weights lead, moving\-boundary, force, and inflow responses\. Accuracy and MCF use the same fixed population of5,0045\{,\}004reference–query pairs, with MCF applying controlled input interventions to each pair\. Outward denotes lower error or latency and stronger MCF; polygon area is not an aggregate score\.
![Refer to caption](https://arxiv.org/html/2608.20477v1/figures/submission/stco_spatial_comparison.png)Figure 4:Prescribed\-condition velocity–pressure responses and pressure\-derived loads\.\(a\)CFD GT and matched Base/STCO absolute errors for Poseidon, PINO, and GEPS on a vortical encounter atnr=128n\_\{r\}=128,Δ​n=10\\Delta n=10; rows showuu,vv, andpp\.\(b,c\)CFD GT, GAOT Base/STCO predictions, absolutevverrors, and independent\-query loads for prescribed yaw \(nr=128n\_\{r\}=128\) and transverse gust \(nr=126n\_\{r\}=126\) atΔ​n=20\\Delta n=20\. Backgrounds mark ID/OOD\-Lead, and vertical lines mark displayed fields\. Annotations give relative\-L2L\_\{2\}errors; color limits are shared within comparisons\.Evaluation\.A predefined set of5,0045\{,\}004unique reference–query pairs from all4242test simulations is shared across configurations\. Its six834834\-pair strata combine three regimes \(pre\-event, motion transition, and external onset\) with ID\-Lead \(Δ​n=1\\Delta n=1–2020\) or OOD\-Lead \(Δ​n=21\\Delta n=21–4040\)\. All predictions are scored on target\-frame fluid points of the complete514×514514\\times 514mesh\. Within each stratum,EyE\_\{y\}pools queries and fluid points per eligible simulation to compute channelwise relative\-L2L\_\{2\}errors, then macro\-averages overuu,vv,pp, and simulations\. The ID\-Lead, OOD\-Lead, and overall scores average three, three, and six strata, respectively\.EFpE\_\{F\_\{p\}\}macro\-averages per\-simulation joint pressure\-derived drag–lift RMSE normalized by fixed training\-set RMS scales\. Both metrics use inverse\-standardized predictions\. Latency is the batch\-one H200 forward\-pass time with dense decoding but without FAGL preprocessing\.

Modulation counterfactual \(MCF\) sensitivity measures functional dependence onr∈\{τ,ψ,Δ​ψ,𝒈,𝒖bc\}r\\in\\\{\\tau,\\psi,\\Delta\\psi,\\bm\{g\},\\bm\{u\}\_\{\\mathrm\{bc\}\}\\\}\. It is evaluated on the same fixed setℐ\\mathcal\{I\}of5,0045\{,\}004pairs shared by all model cells\. Let𝒀^i,j\\widehat\{\\bm\{Y\}\}\_\{i,j\}and𝒀^i,j\(r\)\\widehat\{\\bm\{Y\}\}^\{\(r\)\}\_\{i,j\}denote the original and intervened inverse\-standardized full\-mesh predictions for channeljj\. The reported score is

MCFr=13​\|ℐ\|​∑i∈ℐ∑j∈\{u,v,p\}‖𝑴i⊙\(𝒀^i,j−𝒀^i,j\(r\)\)‖2max⁡\{‖𝑴i⊙𝒀^i,j‖2,10−6\}\.\\operatorname\{MCF\}\_\{r\}=\\frac\{1\}\{3\|\\mathcal\{I\}\|\}\\sum\_\{i\\in\\mathcal\{I\}\}\\sum\_\{j\\in\\\{u,v,p\\\}\}\\frac\{\\left\\\|\\bm\{M\}\_\{i\}\\odot\\left\(\\widehat\{\\bm\{Y\}\}\_\{i,j\}\-\\widehat\{\\bm\{Y\}\}^\{\(r\)\}\_\{i,j\}\\right\)\\right\\\|\_\{2\}\}\{\\max\\\!\\left\\\{\\left\\\|\\bm\{M\}\_\{i\}\\odot\\widehat\{\\bm\{Y\}\}\_\{i,j\}\\right\\\|\_\{2\},10^\{\-6\}\\right\\\}\}\.\(9\)Here,𝑴i\\bm\{M\}\_\{i\}is the fixed fluid mask\. Each intervention holds the observed history, other condition groups, slot and query coordinates, fluid mask, and random state fixed\. MCF measures functional input dependence, complementingEyE\_\{y\}andEFpE\_\{F\_\{p\}\}, which quantify accuracy against paired CFD ground truth\. The supplementary material specifies the load metric, complete aggregation, intervention distributions, and deterministic seeding\.

### Main Results

STCO improves accuracy for PCOL across twelve heterogeneous backbones \(Table[1](https://arxiv.org/html/2608.20477#Sx4.T1)\)\. The matched gains evaluate the complete condition interface, rather than an individual route or modulation site\.

STCO lowersEyE\_\{y\}in 68/72 regime–lead comparisons \(mean: 31\.1%\) and yields a positive six\-stratum mean field gain for all 12 backbones\. This includes all 11 backbones retaining observed\-frame conditions and all 5 with existing lead\-time inputs\. The complete interface therefore adds predictive value alongside the retained observed\-condition and lead\-time pathways\. Across the same comparisons,EFpE\_\{F\_\{p\}\}decreases by 24\.7% on average and improves in 63 cases\. The cross\-backbone field gain is40\.4%40\.4\\%in ID\-Lead and23\.1%23\.1\\%in OOD\-Lead, while load reductions remain comparable at19\.3%19\.3\\%and20\.4%20\.4\\%\.

#### Resolved physical responses\.

Figure[4](https://arxiv.org/html/2608.20477#Sx4.F4)links the aggregate gains to resolved fields and pressure\-derived loads\. On the common vortical query, STCO reduces eight of nine displayed channel errors, with lower residuals around the body, shear layers, and wake\. For prescribed yaw and transverse gust, the transverse\-velocity error decreases by62\.8%62\.8\\%and70\.4%70\.4\\%, respectively, atΔ​n=20\\Delta n=20\. Across the 40 direct queries in each case, lift and drag RMSE reductions span33\.6%33\.6\\%–75\.1%75\.1\\%\. Together, the motion and inflow examples connect improved condition\-specific fields to pressure\-derived responses relevant to downstream evaluation\.

#### Longer leads, condition dependence, and cost\.

11 backbones retain lower OOD\-LeadEyE\_\{y\}beyond the training range\. In Figure[3](https://arxiv.org/html/2608.20477#Sx4.F3)\(b\), the cross\-backbone median remains below Base at all forty leads, with mean reductions of 53\.4% in ID\-Lead and 45\.6% in OOD\-Lead\. Although both medians rise overall across OOD\-Lead, the advantage persists\. MCF is nonzero for every evaluated spatial condition group and STCO backbone, establishing condition dependence alongside the paired CFD accuracy metrics\. Spatial MCF peaks for GINO, while field gain peaks for GAOT, separating sensitivity from accuracy\. Figure[3](https://arxiv.org/html/2608.20477#Sx4.F3)\(a\) shows that these accuracy gains are achieved at comparable inference latency across matched pairs\.

#### Limitations\.

The evaluation uses one seed and one observed frame in a two\-dimensional moving\-body Navier–Stokes system from one CFD solver\.EFpE\_\{F\_\{p\}\}excludes viscous shear, and the largest yaw excursions remain underestimated\. The load curves assemble independent target\-time queries\. Complete condition trajectories, cross\-time\-consistent prediction, broader PDEs, and closed\-loop action selection remain future work\.

## Conclusion

PCOL formulates future\-response prediction from observed history and prescribed target\-time fields\. STCO couples vorticity\-aware FAGL, DSFiLM, and activity sampling across heterogeneous backbones\. Across twelve matched pairs, it reduces mean field and load error within and beyond the training lead range, with condition dependence at comparable inference cost\.

## Supplementary Material

This supplement details the benchmark, FAGL implementation, training, evaluation, and extended results\.

## Appendix ABenchmark

Each simulation containsT=256T=256stored frames generated with the immersed\-boundary computational fluid dynamics \(CFD\) solverWaterLily\.jl\([Weymouth and Font 2025](https://arxiv.org/html/2608.20477#bib.bib38)\)\. The CFD response fields serve as ground truth \(GT\)\. At framenn, the response and prescribed\-condition fields are𝒚n=\[un,vn,pn\]⊤\\bm\{y\}\_\{n\}=\[u\_\{n\},v\_\{n\},p\_\{n\}\]^\{\\top\}and𝒄n=\[ψn,Δ​ψn,𝒈n⊤,𝒖bc,n⊤\]⊤\\bm\{c\}\_\{n\}=\[\\psi\_\{n\},\\Delta\\psi\_\{n\},\\bm\{g\}\_\{n\}^\{\\top\},\\bm\{u\}\_\{\\mathrm\{bc\},n\}^\{\\top\}\]^\{\\top\}\. The reference and query indices satisfynq=nr\+Δ​nn\_\{q\}=n\_\{r\}\+\\Delta n, with normalized leadτ=Δ​n/\(T−1\)\\tau=\\Delta n/\(T\-1\)\. All experiments use the one\-frame historyℋr=\{\(𝒚r,𝒄r\)\}\\mathcal\{H\}\_\{r\}=\\\{\(\\bm\{y\}\_\{r\},\\bm\{c\}\_\{r\}\)\\\}\.

### Simulation Families and Splits

Figure[5](https://arxiv.org/html/2608.20477#A1.F5)illustrates the undisturbed baseline and six prescribed\-condition families through representative condition fields, CFD responses, and load histories\. Table[2](https://arxiv.org/html/2608.20477#A1.T2)gives the92/8/4292/8/42simulation\-level train/validation/test allocation\.

Figure 5:Representative benchmark simulations\. Rows correspond to the undisturbed, convecting\-gust, transverse\-gust, vortical\-encounter, actuation, motion, and compound families\. The first column shows prescribed condition fields, with motion represented by the one\-frame signed\-distance changeΔ​ψ\\Delta\\psi\. The remaining columns show CFD GT fields and CFD load histories\. Dashed markers identify the displayed frames\. The learned target is the fluid\-region velocity–pressure field\(u,v,p\)\(u,v,p\)\.
### Flow, Kinematics, and Prescribed Conditions

All simulations use the same base inflow,Re=5000\\mathrm\{Re\}=5000, NACA 0012 section, and solver configuration\. The514×514514\\times 514grid resolves one chord withLpix=128L\_\{\\mathrm\{pix\}\}=128lattice units, and the nondimensional variables setL=U∞=ρ=1L=U\_\{\\infty\}=\\rho=1\. For body\-frame chordwise coordinatexx, defines=clip⁡\(x/L,0,1\)s=\\operatorname\{clip\}\(x/L,0,1\)\. The reference centerline is

yb​\(s,t\)\\displaystyle y\_\{b\}\(s,t\)=L​Ab​\(s\)​cos⁡\(2​π​sλ\+−ωkin​t\+ϕ\),\\displaystyle=LA\_\{b\}\(s\)\\cos\\\!\\left\(\\frac\{2\\pi s\}\{\\lambda^\{\+\}\}\-\\omega\_\{\\mathrm\{kin\}\}t\+\\phi\\right\),\(10\)Ab​\(s\)\\displaystyle A\_\{b\}\(s\)=A0,b\+A1,b​s\+A2,b​s2,\\displaystyle=A\_\{0,b\}\+A\_\{1,b\}s\+A\_\{2,b\}s^\{2\},ωkin\\displaystyle\\omega\_\{\\mathrm\{kin\}\}=2​π​St​U∞2​Amax​L\.\\displaystyle=\\frac\{2\\pi\\mathrm\{St\}\\,U\_\{\\infty\}\}\{2A\_\{\\max\}L\}\.Here,bbindexes morphology,Ab​\(s\)A\_\{b\}\(s\)is its dimensionless amplitude envelope,λ\+\\lambda^\{\+\}is the nondimensional body wavelength, andϕ\\phiis the phase\. Table[3](https://arxiv.org/html/2608.20477#A1.T3)lists five amplitude envelopes with tail amplitudeAmax=Ab​\(1\)=0\.1A\_\{\\max\}=A\_\{b\}\(1\)=0\.1\. All cases use Strouhal numberSt=0\.4\\mathrm\{St\}=0\.4,λ\+=1\\lambda^\{\+\}=1, andϕ=0\\phi=0\. The kinematic period isTcyc=2​π/ωkinT\_\{\\mathrm\{cyc\}\}=2\\pi/\\omega\_\{\\mathrm\{kin\}\}\.

vgust​\(x,t\)=Ag​F​\(ξ\),ξ=U∞​\(t−t0\)−\(x−x0\)w\.v\_\{\\mathrm\{gust\}\}\(x,t\)=A\_\{g\}F\(\\xi\),\\qquad\\xi=\\frac\{U\_\{\\infty\}\(t\-t\_\{0\}\)\-\(x\-x\_\{0\}\)\}\{w\}\.\(11\)Here,vgustv\_\{\\mathrm\{gust\}\}is the prescribed transverse\-velocity perturbation,t0t\_\{0\}is the onset time,x0x\_\{0\}andwware the streamwise reference position and profile scale, andAgA\_\{g\}is the peak amplitude\. Convecting gusts use Gaussian, sine, or one\-minus\-cosine profiles, and transverse gusts use top\-hat, sine\-squared, or trapezoidal profiles\([Bartels 2013](https://arxiv.org/html/2608.20477#bib.bib3);[Andreu\-Angulo et al\. 2020](https://arxiv.org/html/2608.20477#bib.bib2);[Sedky, Biler, and Jones 2022](https://arxiv.org/html/2608.20477#bib.bib34)\); the Gaussian profile isF⁡\(ξ\)=e−4​ln⁡2​\(ξ−1/2\)2F\(\\xi\)=e^\{\-4\\ln 2\(\\xi\-1/2\)^\{2\}\}\.

The Lamb–Oseen vortical inflow is

𝒖vrt\(𝒙,t\)=Γ2​π​r2\(1−e−αr2/rc2\)𝒓⟂\.\\bm\{u\}\_\{\\mathrm\{vrt\}\}\(\\bm\{x\},t\)=\\frac\{\\Gamma\}\{2\\pi r^\{2\}\}\\left\(1\-e^\{\-\\alpha r^\{2\}/r\_\{c\}^\{2\}\}\\right\)\\bm\{r\}^\{\\perp\}\.\(12\)Here,𝒙=\(x,y\)\\bm\{x\}=\(x,y\),Γ\\Gammais the circulation,rcr\_\{c\}is the core radius, and\(x0,y0\)\(x\_\{0\},y\_\{0\}\)is the vortex center att=t0t=t\_\{0\}\. With𝒙c​\(t\)=\(x0\+U∞​\(t−t0\),y0\)\\bm\{x\}\_\{c\}\(t\)=\(x\_\{0\}\+U\_\{\\infty\}\(t\-t\_\{0\}\),y\_\{0\}\), define𝒓=𝒙−𝒙c​\(t\)\\bm\{r\}=\\bm\{x\}\-\\bm\{x\}\_\{c\}\(t\),r=‖𝒓‖2r=\\\|\\bm\{r\}\\\|\_\{2\}, and𝒓⟂=\(−ry,rx\)\\bm\{r\}^\{\\perp\}=\(\-r\_\{y\},r\_\{x\}\)\. The vortex is zero fort<t0t<t\_\{0\}, andα=1\.25643\\alpha=1\.25643places the peak tangential velocity atr=rcr=r\_\{c\}\.

Body\-force actuation uses

𝒈\(𝒙,t\)=Af\(t\)e−∥𝒙−𝒙f∥22/\(2σ2\)𝒆y\.\\bm\{g\}\(\\bm\{x\},t\)=A\_\{f\}\(t\)e^\{\-\\\|\\bm\{x\}\-\\bm\{x\}\_\{f\}\\\|\_\{2\}^\{2\}/\(2\\sigma^\{2\}\)\}\\bm\{e\}\_\{y\}\.\(13\)Here,𝒙f\\bm\{x\}\_\{f\}andσ\\sigmaare the forcing center and spatial scale, and𝒆y\\bm\{e\}\_\{y\}is the transverse unit vector\. The one\-cosine amplitude isAf​\(t\)=12​A​\[1−cos⁡\(2​π​\(t−t0\)/d\)\]A\_\{f\}\(t\)=\\tfrac\{1\}\{2\}A\[1\-\\cos\(2\\pi\(t\-t\_\{0\}\)/d\)\]on\[t0,t0\+d\]\[t\_\{0\},t\_\{0\}\+d\]and zero otherwise, whereAAandddare its peak amplitude and duration\.

Prescribed motion uses

χ⁡\(t\)=χ∞​ℛ​\(t−t0d\)\.\\chi\(t\)=\\chi\_\{\\infty\}\\mathcal\{R\}\\\!\\left\(\\frac\{t\-t\_\{0\}\}\{d\}\\right\)\.\(14\)Here,χ⁡\(t\)\\chi\(t\)is the displacement or yaw angle,χ∞\\chi\_\{\\infty\}is its terminal value, andddis the ramp duration\.ℛ⁡\(ϑ\)\\mathcal\{R\}\(\\vartheta\)equals00forϑ≤0\\vartheta\\leq 0,12​\(1−cos⁡π​ϑ\)\\tfrac\{1\}\{2\}\(1\-\\cos\\pi\\vartheta\)for0<ϑ<10<\\vartheta<1, and11forϑ≥1\\vartheta\\geq 1\. At target frameqq, the gust and vortex fields define𝒖bc,q\\bm\{u\}\_\{\\mathrm\{bc\},q\}, actuation defines𝒈q\\bm\{g\}\_\{q\}, and prescribed motion determinesψq\\psi\_\{q\}andΔ​ψq=ψq−ψq−1\\Delta\\psi\_\{q\}=\\psi\_\{q\}\-\\psi\_\{q\-1\}\. Table[4](https://arxiv.org/html/2608.20477#A1.T4)summarizes the varied factors and profiles\.

Table 2:Composition of the moving\-body\-flow benchmark\. Training, validation, and test contain9292,88, and4242simulations\. The test set covers all six prescribed\-condition families\.Table 3:Centerline\-amplitude coefficients in Equation[10](https://arxiv.org/html/2608.20477#A1.E10)\. All five profiles satisfyAb​\(1\)=Amax=0\.1A\_\{b\}\(1\)=A\_\{\\max\}=0\.1while varying the chordwise amplitude envelope\.Table 4:Prescribed\-condition families and varied benchmark factors\. Compound cases pair one motion command with one disturbance\.
### Frames, Candidate Pairs, and Preprocessing

With256256frames per simulation, the benchmark contains142×256=36,352142\\times 256=36\{,\}352frames, including92×256=23,55292\\times 256=23\{,\}552for training\. The216216reference indicesnr=20,…,235n\_\{r\}=20,\\ldots,235and2020training leadsΔ​n=1,…,20\\Delta n=1,\\ldots,20give92×216×20=397,44092\\times 216\\times 20=397\{,\}440admissible training pairs\. Each epoch samples10,00010\{,\}000pairs with replacement\.

After framewise response outlier removal, fixed channelwise maps standardize the aligned response and prescribed\-condition inputs\. Metrics use inverse\-standardized response fields\.

## Appendix BFAGL Implementation

FAGL derives the observed vorticityωr=∂xvr−∂yur\\omega\_\{r\}=\\partial\_\{x\}v\_\{r\}\-\\partial\_\{y\}u\_\{r\}and its magnitudeμr=\|ωr\|\\mu\_\{r\}=\|\\omega\_\{r\}\|from𝒚r\\bm\{y\}\_\{r\}\. For simulationmm, let𝒱m\\mathcal\{V\}\_\{m\}denote the time\-union fluid grid determined by its prescribed geometry sequence\. The observed\-frame regional slots are

\{𝒮r,ℓ\}ℓ=1Nℓ=ΠNℓ​\(μr,𝒱m\),Nℓ=12,000,\\\{\\mathcal\{S\}\_\{r,\\ell\}\\\}\_\{\\ell=1\}^\{N\_\{\\ell\}\}=\\Pi\_\{N\_\{\\ell\}\}\(\\mu\_\{r\},\\mathcal\{V\}\_\{m\}\),\\qquad N\_\{\\ell\}=12\{,\}000,\(15\)whereΠNℓ\\Pi\_\{N\_\{\\ell\}\}denotes the FAGL partitioner and𝒮r,ℓ\\mathcal\{S\}\_\{r,\\ell\}contains the grid points assigned to slotℓ\\ell\. Observed response and condition fields are averaged within each slot\. The prescribed target\-time fields\(ψq,Δ​ψq,𝒈q,𝒖bc,q\)\(\\psi\_\{q\},\\Delta\\psi\_\{q\},\\bm\{g\}\_\{q\},\\bm\{u\}\_\{\\mathrm\{bc\},q\}\)are aligned to the slot centroids by four\-neighbor inverse\-distance interpolation\.

FAGL combines graph merging with bounded refinement to form a fixed\-cardinality representation\. On the eight\-neighbor grid graph, an edge\(i,j\)\(i,j\)receives dissimilaritydi​jd\_\{ij\}from grid distance, the contrast in observed vorticity magnitudeμr\\mu\_\{r\}, and the contrast in‖∇μr‖2\\\|\\nabla\\mu\_\{r\}\\\|\_\{2\}\. The three cues are normalized within each frame\.

Edges are scanned once in ascendingdi​jd\_\{ij\}\. ComponentsAAandBBmerge when

di​j≤min⁡\{Int⁡\(A\)\+κ\|A\|,Int⁡\(B\)\+κ\|B\|\},d\_\{ij\}\\leq\\min\\\!\\left\\\{\\operatorname\{Int\}\(A\)\+\\frac\{\\kappa\}\{\|A\|\},\\operatorname\{Int\}\(B\)\+\\frac\{\\kappa\}\{\|B\|\}\\right\\\},\(16\)whereInt⁡\(C\)\\operatorname\{Int\}\(C\)is the largest accepted edge withinCCandκ\\kappais the frame\-adaptive merge threshold\.

If the merge returns fewer thanNℓN\_\{\\ell\}regions, refinement prioritizes

𝒫r​\(C\)=\|C\|2\+∑𝒙i∈Cμr​\(𝒙i\)μ¯ℬ\+10−6\.\\mathcal\{P\}\_\{r\}\(C\)=\\frac\{\|C\|\}\{2\}\+\\frac\{\\sum\_\{\\bm\{x\}\_\{i\}\\in C\}\\mu\_\{r\}\(\\bm\{x\}\_\{i\}\)\}\{\\bar\{\\mu\}\_\{\\mathcal\{B\}\}\+10^\{\-6\}\}\.\(17\)Here,\|C\|\|C\|is the number of points in regionCC, andμ¯ℬ\\bar\{\\mu\}\_\{\\mathcal\{B\}\}is the pointwise mean ofμr\\mu\_\{r\}over refinement batchℬ\\mathcal\{B\}\. Regions with at least four points are median\-split along the best ofxx,yy, and signedωr\\omega\_\{r\}, selected by balance and vorticity contrast\. Refinement stops atNℓN\_\{\\ell\}regions or when no eligible split remains, after which the collection is normalized to exactlyNℓN\_\{\\ell\}slots\. Counts aboveNℓN\_\{\\ell\}are reduced by retaining the largest regions; counts belowNℓN\_\{\\ell\}are padded with empty observed slots centered on the grid before target conditions are aligned\.

Each nonempty slot stores the mean coordinate and observed fields of its member points\.

All interpolation and dense decoding use four\-neighbor inverse\-distance weighting,

IDW4⁡\[𝒇\]​\(𝒙\)\\displaystyle\\operatorname\{IDW\}\_\{4\}\[\\bm\{f\}\]\(\\bm\{x\}\)=∑j∈𝒩4​\(𝒙\)wj​\(𝒙\)​𝒇j,\\displaystyle=\\sum\_\{j\\in\\mathcal\{N\}\_\{4\}\(\\bm\{x\}\)\}w\_\{j\}\(\\bm\{x\}\)\\bm\{f\}\_\{j\},\(18\)wj​\(𝒙\)\\displaystyle w\_\{j\}\(\\bm\{x\}\)=ρj​\(𝒙\)−1∑k∈𝒩4​\(𝒙\)ρk​\(𝒙\)−1,\\displaystyle=\\frac\{\\rho\_\{j\}\(\\bm\{x\}\)^\{\-1\}\}\{\\sum\_\{k\\in\\mathcal\{N\}\_\{4\}\(\\bm\{x\}\)\}\\rho\_\{k\}\(\\bm\{x\}\)^\{\-1\}\},ρj​\(𝒙\)\\displaystyle\\rho\_\{j\}\(\\bm\{x\}\)=max⁡\{‖𝒙−𝒙j‖2,10−8\}\.\\displaystyle=\\max\\\{\\\|\\bm\{x\}\-\\bm\{x\}\_\{j\}\\\|\_\{2\},10^\{\-8\}\\\}\.Here,𝒩4​\(𝒙\)\\mathcal\{N\}\_\{4\}\(\\bm\{x\}\)indexes the four nearest source coordinates and𝒇j\\bm\{f\}\_\{j\}is the source value at𝒙j\\bm\{x\}\_\{j\}\. FAGL usesωr\\omega\_\{r\}only to construct the observed\-frame partition\. The learned operator receives regional means of\(u,v,p,ψ,Δ​ψ,𝒈,𝒖bc\)\(u,v,p,\\psi,\\Delta\\psi,\\bm\{g\},\\bm\{u\}\_\{\\mathrm\{bc\}\}\)\. Figure[6](https://arxiv.org/html/2608.20477#A2.F6)reports reconstruction across slot counts\.

Figure 6:FAGL reconstruction across slot counts\. In panel \(a\), parenthetical labels giveNℓN\_\{\\ell\}as a percentage of the complete514×514514\\times 514mesh\. \(a\) Fluid\-region relative\-L2L\_\{2\}error for constant\-region reconstructions\. Curves and bands give the mean and one standard deviation over five simulations at frame 128\. AtNℓ=12,000N\_\{\\ell\}=12\{,\}000, the mean errors are1\.5%1\.5\\%,11\.8%11\.8\\%,6\.9%6\.9\\%, and7\.6%7\.6\\%foruu,vv,pp, andωr\\omega\_\{r\}\. \(b\) GT vorticity and representative regional reconstructions with region boundaries\. \(c\) Signed residualsω^r−ωr\\widehat\{\\omega\}\_\{r\}\-\\omega\_\{r\}\. FAGL usesωr\\omega\_\{r\}for partitioning, while the operator predicts\(u,v,p\)\(u,v,p\)\.
## Appendix CTraining and Evaluation

To isolate the predictive contribution of the target\-time condition interface, each matched Base–STCO pair uses the same split, temporal and spatial samples, FAGL representation, retained backbone context, adapters, decoder, query geometry, and optimization\. Base bypasses both DSFiLM sites, whereas STCO activates the complete interface with𝒄q\\bm\{c\}\_\{q\}andτ\\tau\.

### External\-Activity Sampling and Optimization

External force and inflow disturbances are active over only part of each trajectory, so uniform pair sampling would draw many intervals without external activity\. We therefore sample each training example in two stages: first a simulation, and then a reference–query pair within that simulation\.

For simulationmm, let𝒱m\\mathcal\{V\}\_\{m\}denote its time\-union fluid grid,\|𝒱m\|\|\\mathcal\{V\}\_\{m\}\|its number of points, andhha grid\-point index\. The fields𝒈m,n,h\\bm\{g\}\_\{m,n,h\}and𝒖bc,m,n,h\\bm\{u\}\_\{\\mathrm\{bc\},m,n,h\}are the prescribed body\-force acceleration and inflow perturbation at framennand pointhh\. We define the case\-normalized external activity at framennas

a~m​\(n\)\\displaystyle\\widetilde\{a\}\_\{m\}\(n\)=1\|𝒱m\|​∑h∈𝒱m\(‖𝒈m,n,h‖22\+‖𝒖bc,m,n,h‖22\),\\displaystyle=\\frac\{1\}\{\|\\mathcal\{V\}\_\{m\}\|\}\\sum\_\{h\\in\\mathcal\{V\}\_\{m\}\}\\left\(\\\|\\bm\{g\}\_\{m,n,h\}\\\|\_\{2\}^\{2\}\+\\\|\\bm\{u\}\_\{\\mathrm\{bc\},m,n,h\}\\\|\_\{2\}^\{2\}\\right\),\(19\)am⋆\\displaystyle a\_\{m\}^\{\\star\}=maxn′⁡a~m​\(n′\),\\displaystyle=\\max\_\{n^\{\\prime\}\}\\widetilde\{a\}\_\{m\}\(n^\{\\prime\}\),am​\(n\)\\displaystyle a\_\{m\}\(n\)=\{a~m​\(n\)/am⋆,am⋆\>10−12,0,otherwise\.\\displaystyle=\\begin\{cases\}\\widetilde\{a\}\_\{m\}\(n\)/a\_\{m\}^\{\\star\},&a\_\{m\}^\{\\star\}\>10^\{\-12\},\\\\ 0,&\\text\{otherwise\}\.\\end\{cases\}Here,a~m​\(n\)\\widetilde\{a\}\_\{m\}\(n\)is the spatially averaged raw activity,am⋆a\_\{m\}^\{\\star\}is its maximum over framesn′n^\{\\prime\}, andam​\(n\)∈\[0,1\]a\_\{m\}\(n\)\\in\[0,1\]is the normalized activity trace\. The10−1210^\{\-12\}threshold avoids division by a numerically zero maximum\. Prescribed body motion is excluded from this sampling score and remains represented byψ\\psiandΔ​ψ\\Delta\\psiin the model input\.

At the first stage, undisturbed and zero\-external\-activity simulations receive weightbm=0\.3b\_\{m\}=0\.3, while all others receivebm=1b\_\{m\}=1\. We samplemmwithPr⁡\(m\)=bm/∑m′∈ℳtrainbm′\\Pr\(m\)=b\_\{m\}/\\sum\_\{m^\{\\prime\}\\in\\mathcal\{M\}\_\{\\mathrm\{train\}\}\}b\_\{m^\{\\prime\}\}, whereℳtrain\\mathcal\{M\}\_\{\\mathrm\{train\}\}is the training simulation set andm′m^\{\\prime\}is its summation index\.

At the second stage, the admissible pairs are𝒜m=\{\(nr,Δn\):nr∈\{20,…,235\},Δn∈\{1,…,20\}\}\\mathcal\{A\}\_\{m\}=\\\{\(n\_\{r\},\\Delta n\):n\_\{r\}\\in\\\{20,\\ldots,235\\\},\\ \\Delta n\\in\\\{1,\\ldots,20\\\}\\\}\.nrn\_\{r\}is the reference frame andΔ​n\\Delta nis the forecast lead in stored frames\. If simulationmmcontains an external\-activity onset, letnon,mn\_\{\\mathrm\{on\},m\}be the first frame at which the absolute value of any component of𝒈\\bm\{g\}or𝒖bc\\bm\{u\}\_\{\\mathrm\{bc\}\}exceeds10−310^\{\-3\}anywhere on the grid, and define

ηm​\(nr,Δ​n\)\\displaystyle\\eta\_\{m\}\(n\_\{r\},\\Delta n\)=\{10,nr<non,m≤nr\+Δ​n,1,otherwise,\\displaystyle=\\begin\{cases\}10,&n\_\{r\}<n\_\{\\mathrm\{on\},m\}\\leq n\_\{r\}\+\\Delta n,\\\\ 1,&\\text\{otherwise\},\\end\{cases\}\(20\)wm​\(nr,Δ​n\)\\displaystyle w\_\{m\}\(n\_\{r\},\\Delta n\)=ηm​\(nr,Δ​n\)​\[0\.1\+∑j=0Δ​nam​\(nr\+j\)\],\\displaystyle=\\eta\_\{m\}\(n\_\{r\},\\Delta n\)\\left\[0\.1\+\\sum\_\{j=0\}^\{\\Delta n\}a\_\{m\}\(n\_\{r\}\+j\)\\right\],Pr⁡\(nr,Δ​n∣m\)\\displaystyle\\Pr\(n\_\{r\},\\Delta n\\mid m\)=wm​\(nr,Δ​n\)∑\(nr′,Δ​n′\)∈𝒜mwm​\(nr′,Δ​n′\)\.\\displaystyle=\\frac\{w\_\{m\}\(n\_\{r\},\\Delta n\)\}\{\\sum\_\{\(n^\{\\prime\}\_\{r\},\\Delta n^\{\\prime\}\)\\in\\mathcal\{A\}\_\{m\}\}w\_\{m\}\(n^\{\\prime\}\_\{r\},\\Delta n^\{\\prime\}\)\}\.Here,ηm\\eta\_\{m\}is the onset\-crossing multiplier,wmw\_\{m\}is the unnormalized pair weight,jjindexes frames afternrn\_\{r\}, and\(nr′,Δ​n′\)\(n^\{\\prime\}\_\{r\},\\Delta n^\{\\prime\}\)indexes candidate pairs in the normalizing sum\. For simulations without an external\-activity onset,ηm=1\\eta\_\{m\}=1for every pair\. The activity sum covers the complete reference\-to\-query interval; the constant0\.10\.1retains nonzero probability for quiescent intervals, andηm\\eta\_\{m\}increases the probability of intervals that cross the first onset\. The complete sampling distribution isPr⁡\(m,nr,Δ​n\)=Pr⁡\(m\)​Pr⁡\(nr,Δ​n∣m\)\\Pr\(m,n\_\{r\},\\Delta n\)=\\Pr\(m\)\\Pr\(n\_\{r\},\\Delta n\\mid m\)\. Figure[7](https://arxiv.org/html/2608.20477#A3.F7)illustrates the resulting activity and onset weighting\.

Figure 7:External\-activity sampling\. \(a\) Case\-normalized traces for the undisturbed baseline, convecting gust \(CVG\), transverse gust \(TG\), Lamb–Oseen vortex \(VRT\), and body\-force actuation \(ACT\)\. \(b,c\) Four thousand pairs sampled from one transverse\-gust simulation\. Activity weighting favors active intervals, and the onset factor further favors onset\-crossing pairs\.Each sampled reference–query pair draws15,00015\{,\}000supervised coordinates\. The sampler targets80%80\\%unique representatives of observed\-frame FAGL regions\. The remaining nominal budget comprises12%12\\%nonrepresentative locations weighted by observed\|ωr\|\|\\omega\_\{r\}\|and8%8\\%sampled uniformly\. All evaluation metrics use the full mesh\.

Training uses batch size3232,8080epochs, and AdamW with learning rate5×10−45\\times 10^\{\-4\}, weight decay10−310^\{\-3\},\(β1,β2\)=\(0\.9,0\.95\)\(\\beta\_\{1\},\\beta\_\{2\}\)=\(0\.9,0\.95\), gradient clipping at11, and cosine decay to10−510^\{\-5\}\. Let𝒚~j,𝒚~^j∈ℝB×Nq\\widetilde\{\\bm\{y\}\}\_\{j\},\\widehat\{\\widetilde\{\\bm\{y\}\}\}\_\{j\}\\in\\mathbb\{R\}^\{B\\times N\_\{q\}\}collect the standardized target and prediction for channeljjover a batch, and let𝑴∈\{0,1\}B×Nq\\bm\{M\}\\in\\\{0,1\\\}^\{B\\times N\_\{q\}\}be the corresponding query\-frame fluid mask\. Here,BBis the batch size,NqN\_\{q\}is the number of supervised coordinates, andθ\\thetadenotes the trainable parameters\. The loss is

ℒ⁡\(θ\)=13​∑j∈\{u,v,p\}‖𝑴⊙\(𝒚~^j−𝒚~j\)‖2‖𝑴⊙𝒚~j‖2\.\\mathcal\{L\}\(\\theta\)=\\frac\{1\}\{3\}\\sum\_\{j\\in\\\{u,v,p\\\}\}\\frac\{\\\|\\bm\{M\}\\odot\(\\widehat\{\\widetilde\{\\bm\{y\}\}\}\_\{j\}\-\\widetilde\{\\bm\{y\}\}\_\{j\}\)\\\|\_\{2\}\}\{\\\|\\bm\{M\}\\odot\\widetilde\{\\bm\{y\}\}\_\{j\}\\\|\_\{2\}\}\.\(21\)With incomplete batches dropped, the10,00010\{,\}000training pairs give312312optimization steps per epoch\. Each matched Base–STCO pair uses the same training and validation pair\-sampling seeds\. Every five epochs, validation evaluates1,0001\{,\}000sampled pairs from the eight validation simulations atΔ​n=1\\Delta n=1–2020, using15,00015\{,\}000coordinates per pair\. Each result uses the checkpoint with the lowest validation relative\-L2L\_\{2\}error\. Each of the 24 configurations is trained from scratch on a single NVIDIA H200 GPU\.

### Predefined Evaluation Set

Evaluation scores all target\-frame fluid points on the complete514×514514\\times 514mesh\. The predefined evaluation set contains5,0045\{,\}004unique reference–query pairs from all4242test simulations, divided into six834834\-pair regime–lead strata and shared across all configurations\. ID\-Lead isΔ​n=1\\Delta n=1–2020, and OOD\-Lead isΔ​n=21\\Delta n=21–4040\. Pairs are assigned to the pre\-event regime whennr\+Δ​n<neventn\_\{r\}\+\\Delta n<n\_\{\\mathrm\{event\}\}, to motion transition whennstart≤nr<nr\+Δ​n≤nendn\_\{\\mathrm\{start\}\}\\leq n\_\{r\}<n\_\{r\}\+\\Delta n\\leq n\_\{\\mathrm\{end\}\}, and to external onset whennr<nevent≤nr\+Δ​nn\_\{r\}<n\_\{\\mathrm\{event\}\}\\leq n\_\{r\}\+\\Delta n\. Here,neventn\_\{\\mathrm\{event\}\}is the annotated disturbance onset, whilenstartn\_\{\\mathrm\{start\}\}andnendn\_\{\\mathrm\{end\}\}delimit the prescribed motion transition\. Table[5](https://arxiv.org/html/2608.20477#A3.T5)gives the stratum composition\.

Table 5:Composition of the predefined5,0045\{,\}004\-pair evaluation set\. Each regime–lead stratum contains834834pairs\. Metrics macro\-average the contributing simulations within each stratum\. A simulation may contribute pairs to more than one regime\.
### Field and Pressure\-Derived Load Metrics

Letς\\varsigmaindex a regime–lead stratum,𝒞ς\\mathcal\{C\}\_\{\\varsigma\}its contributing simulations, and𝒬m,ς\\mathcal\{Q\}\_\{m,\\varsigma\}the selected pair indices from simulationmm\. For channelj∈\{u,v,p\}j\\in\\\{u,v,p\\\}and mesh pointhh, letym,i,h,jy\_\{m,i,h,j\}andy^m,i,h,j\\widehat\{y\}\_\{m,i,h,j\}be the GT and predicted target values, and letMm,i,hM\_\{m,i,h\}be the query\-frame fluid mask\. For valueszm,i,hz\_\{m,i,h\}on these pairs, define the masked norm‖z‖m,ς2=∑i∈𝒬m,ς∑hMm,i,h​\|zm,i,h\|2\\\|z\\\|\_\{m,\\varsigma\}^\{2\}=\\sum\_\{i\\in\\mathcal\{Q\}\_\{m,\\varsigma\}\}\\sum\_\{h\}M\_\{m,i,h\}\|z\_\{m,i,h\}\|^\{2\}\. The field metric is

Em,ς,j\\displaystyle E\_\{m,\\varsigma,j\}=‖y^m,j−ym,j‖m,ςmax⁡\{‖ym,j‖m,ς,10−6\},\\displaystyle=\\frac\{\\\|\\widehat\{y\}\_\{m,j\}\-y\_\{m,j\}\\\|\_\{m,\\varsigma\}\}\{\\max\\\{\\\|y\_\{m,j\}\\\|\_\{m,\\varsigma\},10^\{\-6\}\\\}\},\(22\)Ey\(ς\)\\displaystyle E\_\{y\}^\{\(\\varsigma\)\}=13​\|𝒞ς\|​∑m∈𝒞ς∑j∈\{u,v,p\}Em,ς,j\.\\displaystyle=\\frac\{1\}\{3\|\\mathcal\{C\}\_\{\\varsigma\}\|\}\\sum\_\{m\\in\\mathcal\{C\}\_\{\\varsigma\}\}\\sum\_\{j\\in\\\{u,v,p\\\}\}E\_\{m,\\varsigma,j\}\.
The pressure\-derived drag and lift coefficients are reconstructed as

𝑪p​\(t\)\\displaystyle\\bm\{C\}\_\{p\}\(t\)=\[CD,p​\(t\)CL,p​\(t\)\]⊤\\displaystyle=\\begin\{bmatrix\}C\_\{D,p\}\(t\)&C\_\{L,p\}\(t\)\\end\{bmatrix\}^\{\\top\}=−∫∂Ωb​\(t\)p\(𝒙,t\)𝒏\(𝒙,t\)ds\.\\displaystyle=\-\\int\_\{\\partial\\Omega\_\{b\}\(t\)\}p\(\\bm\{x\},t\)\\bm\{n\}\(\\bm\{x\},t\)\\,\\mathrm\{d\}s\.Here,∂Ωb​\(t\)\\partial\\Omega\_\{b\}\(t\)is the prescribed body boundary,𝒏\\bm\{n\}its outward unit normal, andd​s\\mathrm\{d\}sthe boundary arc\-length element\. GT and predicted loads use the same query\-frame signed\-distance contour and quadrature\. The stored field is the pressure coefficientppdefined in the main paper\. Fork∈\{D,L\}k\\in\\\{D,L\\\}, the normalized pressure\-derived load error is

Em,ς,Fp2\\displaystyle E\_\{m,\\varsigma,F\_\{p\}\}^\{\\,2\}=12​\|𝒬m,ς\|∑i∈𝒬m,ς∑k∈\{D,L\}\\displaystyle=\\frac\{1\}\{2\|\\mathcal\{Q\}\_\{m,\\varsigma\}\|\}\\sum\_\{i\\in\\mathcal\{Q\}\_\{m,\\varsigma\}\}\\sum\_\{k\\in\\\{D,L\\\}\}\(23\)\(C^m,i,k,p−Cm,i,k,psk\)2,\\displaystyle\\left\(\\frac\{\\widehat\{C\}\_\{m,i,k,p\}\-C\_\{m,i,k,p\}\}\{s\_\{k\}\}\\right\)^\{2\},EFp\(ς\)\\displaystyle E\_\{F\_\{p\}\}^\{\(\\varsigma\)\}=1\|𝒞ς\|​∑m∈𝒞ςEm,ς,Fp,\\displaystyle=\\frac\{1\}\{\|\\mathcal\{C\}\_\{\\varsigma\}\|\}\\sum\_\{m\\in\\mathcal\{C\}\_\{\\varsigma\}\}E\_\{m,\\varsigma,F\_\{p\}\},whereCm,i,k,pC\_\{m,i,k,p\}is pressure\-derived componentkkfor pairiifrom simulationmm, andsD=0\.1187s\_\{D\}=0\.1187andsL=0\.7959s\_\{L\}=0\.7959are RMS scales over all397,440397\{,\}440admissible training pairs\.

For eitherE∈\{Ey,EFp\}E\\in\\\{E\_\{y\},E\_\{F\_\{p\}\}\\\}, band and overall errors are

EID\\displaystyle E^\{\\mathrm\{ID\}\}=13​∑ς∈𝒮IDE\(ς\),\\displaystyle=\\frac\{1\}\{3\}\\sum\_\{\\varsigma\\in\\mathcal\{S\}\_\{\\mathrm\{ID\}\}\}E^\{\(\\varsigma\)\},EOOD\\displaystyle E^\{\\mathrm\{OOD\}\}=13​∑ς∈𝒮OODE\(ς\),\\displaystyle=\\frac\{1\}\{3\}\\sum\_\{\\varsigma\\in\\mathcal\{S\}\_\{\\mathrm\{OOD\}\}\}E^\{\(\\varsigma\)\},\(24\)Eall\\displaystyle E^\{\\mathrm\{all\}\}=16​∑ςE\(ς\)\.\\displaystyle=\\frac\{1\}\{6\}\\sum\_\{\\varsigma\}E^\{\(\\varsigma\)\}\.Here,𝒮ID\\mathcal\{S\}\_\{\\mathrm\{ID\}\}and𝒮OOD\\mathcal\{S\}\_\{\\mathrm\{OOD\}\}contain the three regime strata in their respective lead bands\. The aggregation gives equal weight to contributing simulations and regime strata, whileEyE\_\{y\}also weights the three response channels equally\. Paired gain is100​\(EBase−ESTCO\)/EBase100\(E\_\{\\mathrm\{Base\}\}\-E\_\{\\mathrm\{STCO\}\}\)/E\_\{\\mathrm\{Base\}\}\.

## Appendix DExtended Results

Table 6:Stratified paired gains and route\-level MCF on the predefined5,0045\{,\}004\-pair evaluation set\. Positive gains indicate lower STCO error, bold marks negative gains, and “External” denotes external onset\. STCO improves68/7268/72field cells and63/7263/72pressure\-derived load cells, with mean gains of31\.1%31\.1\\%and24\.7%24\.7\\%\. For MCF,τ\\tauis replaced by another same\-band lead,ψ\\psiandΔ​ψ\\Delta\\psiare drawn from fixed channelwise Gaussian marginals, and𝒈\\bm\{g\}and𝒖bc\\bm\{u\}\_\{\\mathrm\{bc\}\}use active benchmark fields\. Deterministic draws are shared across models; unchanged\-input repeats define the numerical floor\. GAOT has the largest mean field gain, and GINO the largest spatial MCF\.Figure 8:Learning and lead\-time dynamics across twelve matched Base–STCO pairs\. \(a\) Training objective and \(b\) validation relative\-L2L\_\{2\}error over8080epochs\. \(c\) Velocity–pressure field errorEyE\_\{y\}and \(d\) pressure\-derived load errorEFpE\_\{F\_\{p\}\}over the predefined5,0045\{,\}004\-pair evaluation set\. Thin curves show individual backbones; thick curves and bands show the cross\-backbone median and interquartile range\. Each lead is an independent target\-time query\. Gray and red backgrounds denote ID\-Lead and OOD\-Lead\. Before smoothing, STCO lowers field error in466/480466/480and load error in454/480454/480backbone–lead comparisons\.Figure 9:Prescribed\-yaw response across twelve matched backbones\. Atnr=128n\_\{r\}=128andΔ​n=20\\Delta n=20, rows show GT, Base/STCO transverse velocityvv, absolute errors, and fluid\-region relative\-L2L\_\{2\}errorEvE\_\{v\}, followed by pressure\-derived lift \(solid\) and drag \(dashed\) over independent target\-time queries atΔ​n=1\\Delta n=1–4040\. Field, error, and load\-axis scales are shared across rows\. Gray/red backgrounds mark ID/OOD\-Lead\. STCO lowersEvE\_\{v\}for9/129/12backbones, with a mean reduction of18\.9%18\.9\\%\. GAOT and UPT show the largest reductions,62\.8%62\.8\\%and57\.5%57\.5\\%\.Figure 10:Transverse\-gust response across twelve matched backbones atnr=126n\_\{r\}=126andΔ​n=20\\Delta n=20\. Rows use the field and pressure\-derived load layout of Figure[9](https://arxiv.org/html/2608.20477#A4.F9)\. The load curves span independent target\-time queries atΔ​n=1\\Delta n=1–4040\. STCO lowersEvE\_\{v\}for all twelve backbones, with a median reduction of58\.5%58\.5\\%\. It lowers the pressure\-derived load error of Equation[23](https://arxiv.org/html/2608.20477#A3.E23)over the forty leads for11/1211/12backbones, with a median reduction of34\.3%34\.3\\%\.Figure 11:Prescribed\-yaw transition across twelve STCO configurations\. \(a\) Transverse velocityvvat the observed framenr=44n\_\{r\}=44, target\-time signed\-distance changeΔ​ψq\\Delta\\psi\_\{q\}, and GTvvatΔ​n=20\\Delta n=20\. \(b\) Pointwise absolute errors\|v^−v\|\|\\widehat\{v\}\-v\|\. Labels report full\-mesh fluid\-region relative\-L2L\_\{2\}errorEvE\_\{v\}\. Observed and GTvvshare one color limit, and all error panels share another\. The medianEvE\_\{v\}across the twelve models is0\.2800\.280\. Errors concentrate near the moving body and wake\.Figure 12:Compound vortical\-inflow and motion response across twelve STCO configurations\. \(a\) Transverse velocityvvat the observed framenr=22n\_\{r\}=22, target\-timevbc,qv\_\{\\mathrm\{bc\},q\}overlaid withΔ​ψq\\Delta\\psi\_\{q\}, and GTvvatΔ​n=20\\Delta n=20\. The two condition fields share the range\[−0\.18,0\.18\]\[\-0\.18,0\.18\]\. \(b\) Pointwise absolute errors\|v^−v\|\|\\widehat\{v\}\-v\|\. Labels report full\-mesh fluid\-region relative\-L2L\_\{2\}errorEvE\_\{v\}\. The medianEvE\_\{v\}across the twelve models is0\.2100\.210\. Errors concentrate near the moving body and disturbed wake\.

## Acknowledgments

This research is funded by Engineering Start\-up Grant of King’s College London, and the Daiwa Anglo\-Japanese Foundation through Daiwa Foundation Awards \(14465/15310\)\. The authors acknowledge the use of King’s Computational Research, Engineering and Technology Environment \(CREATE\) in conducting this research\([King’s College London 2026](https://arxiv.org/html/2608.20477#bib.bib17)\)\.

## References

- Alkin et al\. \(2024\)Alkin, B\.; Fürst, A\.; Schmid, S\.; Gruber, L\.; Holzleitner, M\.; and Brandstetter, J\. 2024\.Universal physics transformers: A framework for efficiently scaling neural operators\.In*Advances in Neural Information Processing Systems*, volume 37\.
- Andreu\-Angulo et al\. \(2020\)Andreu\-Angulo, I\.; Babinsky, H\.; Biler, H\.; Sedky, G\.; and Jones, A\. R\. 2020\.Effect of transverse gust velocity profiles\.*AIAA Journal*, 58\(12\):5123–5133\.
- Bartels \(2013\)Bartels, R\. E\. 2013\.Developing an accurate CFD based gust model for the truss braced wing aircraft\.In*31st AIAA Applied Aerodynamics Conference*, AIAA Paper 2013\-3044\.
- Bhan, Shi, and Krstić \(2024\)Bhan, L\.; Shi, Y\.; and Krstić, M\. 2024\.Neural operators for bypassing gain and control computations in PDE backstepping\.*IEEE Transactions on Automatic Control*, 69\(8\):5310–5325\.
- Felzenszwalb and Huttenlocher \(2004\)Felzenszwalb, P\. F\.; and Huttenlocher, D\. P\. 2004\.Efficient graph\-based image segmentation\.*International Journal of Computer Vision*, 59\(2\):167–181\.
- Hagnberger, Musekamp, and Niepert \(2025\)Hagnberger, J\.; Musekamp, D\.; and Niepert, M\. 2025\.CALM\-PDE: Continuous and adaptive convolutions for latent space modeling of time\-dependent PDEs\.In*Advances in Neural Information Processing Systems*, volume 38\.
- Hafner et al\. \(2025\)Hafner, D\.; Pasukonis, J\.; Ba, J\.; and Lillicrap, T\. 2025\.Mastering diverse control tasks through world models\.*Nature*, 640:647–653\.
- Han et al\. \(2021\)Han, R\.; Zhang, Z\.; Wang, Y\.; Liu, Z\.; Zhang, Y\.; and Chen, G\. 2021\.Hybrid deep neural network based prediction method for unsteady flows with moving boundary\.*Acta Mechanica Sinica*, 37\(10\):1557–1566\.
- Hao et al\. \(2023\)Hao, Z\.; Wang, Z\.; Su, H\.; Ying, C\.; Dong, Y\.; Liu, S\.; Cheng, Z\.; Song, J\.; and Zhu, J\. 2023\.GNOT: A general neural operator transformer for operator learning\.In*Proceedings of the 40th International Conference on Machine Learning*, volume 202 of*Proceedings of Machine Learning Research*, 12556–12569\.
- Herde et al\. \(2024\)Herde, M\.; Raonić, B\.; Rohner, T\.; Käppeli, R\.; Molinaro, R\.; de Bézenac, E\.; and Mishra, S\. 2024\.Poseidon: Efficient foundation models for PDEs\.In*Advances in Neural Information Processing Systems*, volume 37\.
- Hu, Shen, and Sun \(2018\)Hu, J\.; Shen, L\.; and Sun, G\. 2018\.Squeeze\-and\-excitation networks\.In*Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition*, 7132–7141\.
- Hu and Liu \(2025\)Hu, H\.; and Liu, C\. 2025\.Safe PDE boundary control with neural operators\.In*Proceedings of the 7th Annual Learning for Dynamics & Control Conference*, volume 283 of*Proceedings of Machine Learning Research*, 513–526\.
- Hufstedler and McKeon \(2019\)Hufstedler, E\. A\. L\.; and McKeon, B\. J\. 2019\.Vortical gusts: Experimental generation and interaction with wing\.*AIAA Journal*, 57\(3\):921–931\.
- Hwang et al\. \(2022\)Hwang, R\.; Lee, J\. Y\.; Shin, J\. Y\.; and Hwang, H\. J\. 2022\.Solving PDE\-constrained control problems using operator learning\.*Proceedings of the AAAI Conference on Artificial Intelligence*, 36\(4\):4504–4512\.
- Jin, Meng, and Lu \(2022\)Jin, P\.; Meng, S\.; and Lu, L\. 2022\.MIONet: Learning multiple\-input operators via tensor product\.*SIAM Journal on Scientific Computing*, 44\(6\):A3490–A3514\.
- Kamkar et al\. \(2011\)Kamkar, S\. J\.; Wissink, A\. M\.; Sankaran, V\.; and Jameson, A\. 2011\.Feature\-driven Cartesian adaptive mesh refinement for vortex\-dominated flows\.*Journal of Computational Physics*, 230\(16\):6271–6298\.
- King’s College London \(2026\)King’s College London\. 2026\.King’s Computational Research, Engineering and Technology Environment \(CREATE\)\.Retrieved August 20, 2026, fromhttps://doi\.org/10\.18742/rnvf\-m076\.
- Kassaï Koupaï et al\. \(2024\)Kassaï Koupaï, A\.; Mifsut Benet, J\.; Yin, Y\.; Vittaut, J\.\-N\.; and Gallinari, P\. 2024\.GEPS: Boosting generalization in parametric PDE neural solvers through adaptive conditioning\.In*Advances in Neural Information Processing Systems*, volume 37\.
- Kassaï Koupaï et al\. \(2025\)Kassaï Koupaï, A\.; Le Boudec, L\.; Serrano, L\.; and Gallinari, P\. 2025\.ENMA: Tokenwise autoregression for continuous neural PDE operators\.In*Advances in Neural Information Processing Systems*, volume 38\.
- Li et al\. \(2021\)Li, Z\.; Kovachki, N\.; Azizzadenesheli, K\.; et al\. 2021\.Fourier neural operator for parametric partial differential equations\.In*International Conference on Learning Representations*\.
- Li et al\. \(2023a\)Li, Z\.; Huang, D\. Z\.; Liu, B\.; and Anandkumar, A\. 2023\.Fourier neural operator with learned deformations for PDEs on general geometries\.*Journal of Machine Learning Research*, 24\(388\):1–26\.
- Li et al\. \(2023b\)Li, Z\.; Kovachki, N\.; Choy, C\.; et al\. 2023\.Geometry\-informed neural operator for large\-scale 3D PDEs\.In*Advances in Neural Information Processing Systems*, volume 36\.
- Li et al\. \(2024\)Li, Z\.; Zheng, H\.; Kovachki, N\.; Jin, D\.; Chen, H\.; Liu, B\.; Azizzadenesheli, K\.; and Anandkumar, A\. 2024\.Physics\-informed neural operator for learning partial differential equations\.*ACM/IMS Journal of Data Science*, 1\(3\):1–27\.
- Lu et al\. \(2021\)Lu, L\.; Jin, P\.; Pang, G\.; Zhang, Z\.; and Karniadakis, G\. E\. 2021\.Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators\.*Nature Machine Intelligence*, 3:218–229\.
- Luo et al\. \(2025\)Luo, H\.; Wu, H\.; Zhou, H\.; Xing, L\.; Di, Y\.; Wang, J\.; and Long, M\. 2025\.Transolver\+\+: An accurate neural solver for PDEs on million\-scale geometries\.In*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, 41432–41449\.
- McCabe et al\. \(2024\)McCabe, M\.; Régaldo\-Saint Blancard, B\.; Parker, L\.; Ohana, R\.; Cranmer, M\.; et al\. 2024\.Multiple physics pretraining for spatiotemporal surrogate models\.In*Advances in Neural Information Processing Systems*, volume 37\.
- Mittal and Iaccarino \(2005\)Mittal, R\.; and Iaccarino, G\. 2005\.Immersed boundary methods\.*Annual Review of Fluid Mechanics*, 37:239–261\.
- Mousavi et al\. \(2025\)Mousavi, S\.; Wen, S\.; Lingsch, L\.; Herde, M\.; Raonić, B\.; and Mishra, S\. 2025\.RIGNO: A graph\-based framework for robust and accurate operator learning for PDEs on arbitrary domains\.In*Advances in Neural Information Processing Systems*, volume 38\.
- NVIDIA et al\. \(2025\)NVIDIA; Agarwal, N\.; Ali, A\.; Bala, M\.; et al\. 2025\.Cosmos world foundation model platform for physical AI\.arXiv:2501\.03575\.
- Park et al\. \(2019\)Park, T\.; Liu, M\.\-Y\.; Wang, T\.\-C\.; and Zhu, J\.\-Y\. 2019\.Semantic image synthesis with spatially\-adaptive normalization\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, 2337–2346\.
- Peebles and Xie \(2023\)Peebles, W\.; and Xie, S\. 2023\.Scalable diffusion models with transformers\.In*Proceedings of the IEEE/CVF International Conference on Computer Vision*, 4195–4205\.
- Perez et al\. \(2018\)Perez, E\.; Strub, F\.; de Vries, H\.; Dumoulin, V\.; and Courville, A\. 2018\.FiLM: Visual reasoning with a general conditioning layer\.*Proceedings of the AAAI Conference on Artificial Intelligence*, 32\(1\):3942–3951\.
- Pfaff et al\. \(2021\)Pfaff, T\.; Fortunato, M\.; Sánchez\-González, A\.; and Battaglia, P\. 2021\.Learning mesh\-based simulation with graph networks\.In*International Conference on Learning Representations*\.
- Sedky, Biler, and Jones \(2022\)Sedky, G\.; Biler, H\.; and Jones, A\. R\. 2022\.Experimental comparison of a sinusoidal and trapezoidal transverse gust\.*AIAA Journal*, 60\(5\):3347–3351\.
- Sedky et al\. \(2022\)Sedky, G\.; Gementzopoulos, A\.; Andreu\-Angulo, I\.; Lagor, F\. D\.; and Jones, A\. R\. 2022\.Physics of gust response mitigation in open\-loop pitching manoeuvres\.*Journal of Fluid Mechanics*, 944:A38\.
- Sotiropoulos and Yang \(2014\)Sotiropoulos, F\.; and Yang, X\. 2014\.Immersed boundary methods for simulating fluid–structure interaction\.*Progress in Aerospace Sciences*, 65:1–21\.
- Wen et al\. \(2025\)Wen, S\.; Kumbhat, A\.; Lingsch, L\.; Mousavi, S\.; Zhao, Y\.; Chandrashekar, P\.; and Mishra, S\. 2025\.Geometry\-aware operator transformer as an efficient and accurate neural surrogate for PDEs on arbitrary domains\.In*Advances in Neural Information Processing Systems*, volume 38\.
- Weymouth and Font \(2025\)Weymouth, G\. D\.; and Font, B\. 2025\.WaterLily\.jl: A differentiable and backend\-agnostic Julia solver for incompressible viscous flow around dynamic bodies\.*Computer Physics Communications*, 315:109748\.
- Zhong and Meidani \(2025\)Zhong, W\.; and Meidani, H\. 2025\.Benchmarking of neural operators for rigid\-body fluid–structure interaction\.In*Machine Learning and the Physical Sciences Workshop at NeurIPS*\.
- Zhou et al\. \(2025\)Zhou, H\.; Ma, Y\.; Wu, H\.; Wang, H\.; and Long, M\. 2025\.Unisolver: PDE\-conditional transformers towards universal neural PDE solvers\.In*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, 79061–79088\.

Similar Articles

Operator Learning for Cubic Nonlinear Schr\"odinger Equation on Periodic Domains

arXiv cs.LG

This paper presents a geometry-conditioned Fourier Neural Operator (FNO) to learn the solution operator for the cubic nonlinear Schrödinger equation on periodic domains with varying aspect ratios. Numerical experiments show the model captures distinct Sobolev norm behaviors on rational and irrational tori, demonstrating geometry-aware neural operators for dispersive PDEs.

Joint discovery of governing partial differential equations from multi-source datasets by competitive optimization

arXiv cs.LG

This paper presents MCO-PDE, a competitive optimization framework that discovers shared partial differential equations from multiple observational datasets by combining neural surrogates, soft-competitive weighting, and genetic algorithms for structure search. It demonstrates high accuracy in recovering canonical equations from limited data and handles complex geometries and real-world experiments.

Sequential Physics-Constrained Neural Operator Forward Modeling for the $\textit{Norne}$ Reservoir System

arXiv cs.LG

This paper presents a comprehensive mathematical framework for sequential surrogate modeling of three-phase black-oil reservoir dynamics using Fourier Neural Operators (FNO) and physics-informed variants (PINO), applied to the Norne benchmark reservoir. Theoretical contributions include functional-analytic formulation, covariate shift analysis, physics-constrained spectral stability, and truncated backpropagation gradient analysis.