VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows

arXiv cs.LG Papers

Summary

Introduces VATO, a vortex-force-aware transformer operator that improves predictions of unsteady separated aerodynamic flows by reducing errors in velocity, pressure, and vorticity through force-aware neural operator learning.

arXiv:2609.00507v1 Announce Type: new Abstract: Accurate prediction of unsteady separated flows is challenging because the aerodynamic loads depend on nonlinear separation and vortex-shedding dynamics. Although high-fidelity CFD resolves these mechanisms, its cost limits repeated use in design and control. Standard field-level surrogate training, however, does not distinguish the flow regions that contribute most strongly to the aerodynamic loads. We introduce VATO (Vortex-Force-Aware Transformer Operator), which couples the Vortex Force Map (VFM) method to a geometry-aware neural operator through two complementary mechanisms. VATO-S adds training-only supervision of the local VFM force-contribution field, with no increase in model size or inference cost. VATO-A uses VFM contribution and sensitivity fields to prioritise force-relevant source locations for residual cross attention. The methods are evaluated on unsteady CFD data for double-edged-plate aerofoils over 54 trajectories from nine geometries. Over lead times of 1-20~ms, VATO-S reduces velocity, pressure, and vorticity errors by 10.4\%, 1.0\%, and 15.6\%, respectively, while VATO-A achieves reductions of 15.8\%, 7.5\%, and 31.2\%. VATO-S gives the lowest VFM-derived drag error, whereas VATO-A gives the lowest pressure-derived lift and drag errors. Over lead times extending 50\% beyond the training range, VATO-A retains a 26.9\% reduction in vorticity error and larger improvements in all four force readouts, despite reduced gains in velocity and pressure. These results show that force-aware operator learning can improve both flow-field prediction and aerodynamic functional accuracy in unsteady separated flows.
Original Article
View Cached Full Text

Cached at: 09/02/26, 06:15 AM

# A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows
Source: [https://arxiv.org/html/2609.00507](https://arxiv.org/html/2609.00507)
Xingxin YangCorresponding author:These authors contributed equally to this work\.Affiliation:Department of Engineering, King’s College London, London, WC2R 2LS, UKZhan ZhangCorresponding author:These authors contributed equally to this work\.Affiliation:Department of Engineering, King’s College London, London, WC2R 2LS, UKJuan LiEmail:[juan\.li@kcl\.ac\.uk](mailto:[email protected])Corresponding author:Corresponding author\.Affiliation:Department of Engineering, King’s College London, London, WC2R 2LS, UK

###### Abstract

Accurate prediction of unsteady separated flows is challenging because the aerodynamic loads depend on nonlinear separation and vortex\-shedding dynamics\. Although high\-fidelity CFD resolves these mechanisms, its cost limits repeated use in design and control\. Standard field\-level surrogate training, however, does not distinguish the flow regions that contribute most strongly to the aerodynamic loads\. We introduce VATO \(Vortex\-Force\-Aware Transformer Operator\), which couples the Vortex Force Map \(VFM\) method to a geometry\-aware neural operator through two complementary mechanisms\. VATO\-S adds training\-only supervision of the local VFM force\-contribution field, with no increase in model size or inference cost\. VATO\-A uses VFM contribution and sensitivity fields to prioritise force\-relevant source locations for residual cross attention\. The methods are evaluated on unsteady CFD data for double\-edged\-plate aerofoils over 54 trajectories from nine geometries\. Over lead times of 1–20 ms, VATO\-S reduces velocity, pressure, and vorticity errors by 10\.4%, 1\.0%, and 15\.6%, respectively, while VATO\-A achieves reductions of 15\.8%, 7\.5%, and 31\.2%\. VATO\-S gives the lowest VFM\-derived drag error, whereas VATO\-A gives the lowest pressure\-derived lift and drag errors\. Over lead times extending 50% beyond the training range, VATO\-A retains a 26\.9% reduction in vorticity error and larger improvements in all four force readouts, despite reduced gains in velocity and pressure\. These results show that force\-aware operator learning can improve both flow\-field prediction and aerodynamic functional accuracy in unsteady separated flows\.

###### Keywords:

neural operator , vortex\-aware operator learning , residual cross attention , unsteady separated flow

## 1Introduction

Accurate prediction of unsteady aerodynamic loads is essential for the design, optimisation, and control of vehicles operating in separated flow\. In such regimes, lift and drag are governed by flow separation, vortex formation, and vortex shedding, whose strongly nonlinear and history\-dependent dynamics are difficult to represent with linearised aerodynamic models\. High\-fidelity computational fluid dynamics \(CFD\) can resolve these mechanisms, but the cost of repeated simulations remains prohibitive for iterative design and control\. This has motivated the development of data\-driven surrogate models that approximate the evolution of the flow and its aerodynamic response at substantially reduced computational cost\.

Machine learning has been widely applied across fluid mechanics over the past decade\[[1](https://arxiv.org/html/2609.00507#bib.bib1),[2](https://arxiv.org/html/2609.00507#bib.bib2)\], while data\-driven modelling of unsteady aerodynamics has developed into a substantial research area of its own\[[3](https://arxiv.org/html/2609.00507#bib.bib3)\]\. Existing approaches can broadly be separated according to whether they predict aerodynamic quantities directly or reconstruct the underlying flow field\. At the level of aerodynamic coefficients, Yao et al\.\[[4](https://arxiv.org/html/2609.00507#bib.bib4)\]optimised the endurance coefficient of a tail\-sitter UAV using a multi\-objective genetic algorithm, Lou et al\.\[[5](https://arxiv.org/html/2609.00507#bib.bib5)\]modified aerofoil geometry with a double deep Q\-network using a neural\-network reward, Zhu et al\.\[[6](https://arxiv.org/html/2609.00507#bib.bib6)\]learned a closure between turbulent eddy viscosity and mean flow variables, and Zhao et al\.\[[7](https://arxiv.org/html/2609.00507#bib.bib7)\]predicted the lift\-to\-drag characteristics of Mars helicopter aerofoils\. Liu et al\.\[[8](https://arxiv.org/html/2609.00507#bib.bib8)\]review this broader class of surrogate\-based aerodynamic shape modelling and optimisation\. These approaches demonstrate that learned surrogates can replace repeated solver evaluations when integral aerodynamic quantities are the primary outputs, but models based on steady or time\-averaged quantities do not directly resolve the unsteady flow structures responsible for transient load histories\. Field\-level surrogates instead seek to reconstruct the spatial flow state from which aerodynamic quantities can subsequently be evaluated\. Convolutional and fully connected architectures have reproduced aerodynamic fields and coefficients across varied geometries\[[9](https://arxiv.org/html/2609.00507#bib.bib9),[10](https://arxiv.org/html/2609.00507#bib.bib10),[11](https://arxiv.org/html/2609.00507#bib.bib11)\], while recurrent architectures have been developed for unsteady modelling\[[12](https://arxiv.org/html/2609.00507#bib.bib12),[13](https://arxiv.org/html/2609.00507#bib.bib13),[14](https://arxiv.org/html/2609.00507#bib.bib14)\]and high\-incidence prediction\[[15](https://arxiv.org/html/2609.00507#bib.bib15)\]\. Recursive temporal prediction, however, remains susceptible to accumulated error over long horizons\. Physics\-informed formulations constrain the learned field by the governing equations\[[16](https://arxiv.org/html/2609.00507#bib.bib16)\]\. Neural operators learn mappings between function spaces that transfer across discretisations, through branch\-trunk architectures built on the universal approximation theorem for operators\[[17](https://arxiv.org/html/2609.00507#bib.bib17)\], spectral parameterisations of the integral kernel\[[18](https://arxiv.org/html/2609.00507#bib.bib18)\], and formulations that impose boundary conditions for aerofoil flows\[[19](https://arxiv.org/html/2609.00507#bib.bib19)\]\. Physics\-guided reconstruction has recovered high\-fidelity fields from sparse or low\-resolution inputs with diffusion models\[[20](https://arxiv.org/html/2609.00507#bib.bib20),[21](https://arxiv.org/html/2609.00507#bib.bib21)\]and regularised sparse compressible\-flow reconstruction with discrete conservation residuals\[[22](https://arxiv.org/html/2609.00507#bib.bib22)\]\.Geometry\-aware operator transformers further extend operator learning to unstructured meshes and non\-uniform domains\[[23](https://arxiv.org/html/2609.00507#bib.bib23)\], enabling a single surrogate to operate across families of geometries\.

Despite these advances, the training signal in field\-level surrogates is applied uniformly over the computational domain\. A relative error on velocity, pressure, or vorticity assigns every mesh point the same weight, so the wake and far field, which occupy most of the mesh, dominate the objective, while the aerodynamic load originates in a small fraction of the domain\. The objective therefore contains no mechanism that identifies the flow structures on which the downstream design or control task depends\. The Vortex Force Map \(VFM\) method\[[24](https://arxiv.org/html/2609.00507#bib.bib24),[25](https://arxiv.org/html/2609.00507#bib.bib25),[26](https://arxiv.org/html/2609.00507#bib.bib26)\]provides a mechanics\-based means of making this distinction\. VFM solves an auxiliary potential problem determined by the body geometry and force direction, producing a vortex\-force factor field whose interaction with the local velocity and vorticity identifies the spatial contribution of the flow to lift or drag\. The auxiliary field requires no additional flow solution and can therefore be precomputed for each geometry and incidence before operator training\. VFM was introduced to estimate unsteady forces from measured or computed snapshots\[[24](https://arxiv.org/html/2609.00507#bib.bib24),[25](https://arxiv.org/html/2609.00507#bib.bib25)\]and has subsequently been extended to three\-dimensional \(3\-D\) unsteady flow\[[26](https://arxiv.org/html/2609.00507#bib.bib26)\]\. He et al\.\[[33](https://arxiv.org/html/2609.00507#bib.bib33)\]combined vortex\-force information with a graph convolution attention network to infer velocity and vortex\-force contributions from incomplete flow measurements and thereby recover force coefficients for flow around a circular cylinder, demonstrating that vortex\-force information can guide reconstruction from sparse observations\. However, the question remains: whether vortex\-force awareness can be introduced to shape the learning and information\-routing mechanisms of a neural operator for full\-field unsteady prediction\.

To address this question, we introduce VATO \(Vortex\-Force\-Aware Transformer Operator\), a family of neural operators that couples the VFM method to a geometry\-aware transformer backbone\[[23](https://arxiv.org/html/2609.00507#bib.bib23)\]at two interfaces\. VATO\-S supervises the per\-point VFM contribution field during training, leaving the architecture, the parameter count, and the inference\-time operator of the backbone unchanged, and therefore suits deployments in which inference cost is fixed\. VATO\-A uses the current flow vorticity together with the VFM contribution and sensitivity fields to identify regions relevant to Lift and Drag, and admits the ordinary flow state from those regions through residual cross attention without embedding the VFM values, providing larger field improvements where additional inference cost is acceptable\. Both configurations are trained with a flow\-aware sampler, and a sampling\-matched control is retained alongside the retrained backbone reference so that the effect of each coupling is separated from the effect of the sampler\. Neither coupling is tied to a particular section: the vortex\-force fields follow from the shape and the incidence alone, and the backbone accepts arbitrary unstructured meshes\.

The double\-edged\-plate \(DEP\) aerofoil proposed for Martian rotorcraft provides a demanding test case for this purpose\. The thin Martian atmosphere places rotor blades at low Reynolds numbers of order10410^\{4\}\-10510^\{5\}, where leading\-edge separation bubbles form readily\[[27](https://arxiv.org/html/2609.00507#bib.bib27)\], may burst into fully separated states\[[28](https://arxiv.org/html/2609.00507#bib.bib28)\], and the separation location in this regime materially affects drag\[[29](https://arxiv.org/html/2609.00507#bib.bib29)\]\. Sharp\-edged sections address this sensitivity by imposing the separation point geometrically\[[30](https://arxiv.org/html/2609.00507#bib.bib30)\], so that the load is governed by the vortex structures rather than by the onset of separation\. Traub and Coffman\[[31](https://arxiv.org/html/2609.00507#bib.bib31)\]showed experimentally that leading\- and trailing\-edge folds fix the location at which separation begins and improve the efficiency of thin plate aerofoils at low Reynolds number, and Koning et al\.\[[32](https://arxiv.org/html/2609.00507#bib.bib32)\]evaluated and optimised the resulting DEP family for Martian rotor applications\. Varying the two fold angles produces a family of geometries with the same separation mechanism but different vortex structures and aerodynamic responses\. This provides a suitable test case for force\-aware operator learning\.

The main contributions of this work are as follows\.

1. 1\.A training\-only coupling of the VFM method to a neural operator, VATO\-S, which supervises where the aerodynamic load originates and improves field prediction at the parameter count and inference cost of the backbone\.
2. 2\.An architectural coupling, VATO\-A, in which VFM contribution and sensitivity fields identify Lift\- and Drag\-relevant regions whose flow state enters residual cross attention, yielding the strongest field accuracy of the family, most prominently in vorticity\.
3. 3\.An attribution protocol built on a retrained reference and a sampling\-matched control, which separates the effect of the vortex\-force couplings from that of the training sampler and shows the two act on different quantities\.
4. 4\.An evaluation across held\-out incidences and a 50% lead\-time extension, in which \(CLC\_\{L\}\) and \(CDC\_\{D\}\) are recovered from the predicted fields through two independent readouts, a pressure\-surface integral and the VFM volume integral\.

The remainder of this paper is organised as follows\. Section[2](https://arxiv.org/html/2609.00507#S2)describes the DEP dataset and its CFD configuration\. Section[3](https://arxiv.org/html/2609.00507#S3)presents the backbone, the two couplings, and the training protocol\. Section[4](https://arxiv.org/html/2609.00507#S4)reports the benchmark, the force functionals, and the predicted fields\. Section[5](https://arxiv.org/html/2609.00507#S5)discusses the findings and their limitations, and Section[6](https://arxiv.org/html/2609.00507#S6)concludes\.

## 2Dataset

The dataset consists of two\-dimensional unsteady CFD solutions for a family of DEP aerofoils, spanning fourteen geometries and thirteen angles of attack\.

### 2\.1Geometry and CFD configuration

The DEP aerofoil is parameterised by independent leading\- and trailing\-edge fold angles,θ1\\theta\_\{1\}andθ2\\theta\_\{2\}, measured as the interior angle at each fold vertex, so that180∘180^\{\\circ\}denotes an unfolded edge and a smaller angle denotes a larger deflection of the surface at that vertex\. Fourteen geometries span a continuous envelope from the most folded shape \(θ1=156\.2∘\\theta\_\{1\}=156\.2^\{\\circ\},θ2=116\.6∘\\theta\_\{2\}=116\.6^\{\\circ\}\) to the flat\-plate reference \(θ1=θ2=180∘\\theta\_\{1\}=\\theta\_\{2\}=180^\{\\circ\}\), shown in Figure[1](https://arxiv.org/html/2609.00507#S2.F1)a,b\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/dataset_geometry_cfd_validation.png)Figure 1:Dataset geometry and CFD cross\-check\. \(a\) Envelope in\(θ1,θ2\)\(\\theta\_\{1\},\\theta\_\{2\}\)interior\-fold\-angle space; the dashed diagonal denotesθ1=θ2\\theta\_\{1\}=\\theta\_\{2\}\. \(b\) Chord\-normalised surface contours of the fourteen DEP geometries, coloured by geometry index\. \(c\) Time\-averaged sectionalCLC\_\{L\}from the present STAR\-CCM\+ simulations, compared with the transition\-model and fully laminar results computed by Koning et al\.\[[32](https://arxiv.org/html/2609.00507#bib.bib32)\]on the matched configuration\.Each geometry and angle\-of\-attack pair was solved as a two\-dimensional unsteady case in STAR\-CCM\+ on an unstructured mesh of 39,552 to 44,600 cells, corresponding to 24,404 to 29,178 mesh points\. The Martian atmospheric conditions are a density of0\.01459​kg​m−30\.01459\\,\\mathrm\{kg\\,m^\{\-3\}\}and a dynamic viscosity of1\.22×10−5​Pa​s1\.22\\times 10^\{\-5\}\\,\\mathrm\{Pa\\,s\}, and the freestream speed of69\.6​m​s−169\.6\\,\\mathrm\{m\\,s^\{\-1\}\}follows Koning et al\.\[[32](https://arxiv.org/html/2609.00507#bib.bib32)\], giving a chord Reynolds number of 10,117\. Turbulence is treated with the SSTkk–ω\\omegamodel\. The computational domain extends five chords upstream and eight chords downstream of the section, and time integration uses a physical time step of5×10−5​s5\\times 10^\{\-5\}\\,\\mathrm\{s\}\.

Sampling began at initialisation and continued until theCLC\_\{L\}andCDC\_\{D\}settled into a regular periodic oscillation, which required between 4,000 and 10,000 time steps depending on the geometry and the angle of attack\. Recording every 20 steps gives trajectories of 200 to 500 frames at intervals of 1 ms, each containing the initial transient followed by the established shedding\. Velocity and pressure were exported at every sampled step\. The angle of attack ranges from1∘1^\{\\circ\}to13∘13^\{\\circ\}\.

The CFD setup was cross\-checked against the independent Mars\-rotorcraft baseline of Koning et al\.\[[32](https://arxiv.org/html/2609.00507#bib.bib32)\]on a matched flat\-plate and DEP configuration\. Their computations solve the compressible RANS equations in OVERFLOW 2\.2o with the SA\-neg\-1a one\-equation model coupled to the AFT2017b transition model, and a second set of cases is run fully laminar\. The two baselines agree closely with each other over the operating envelope of interest, so at this Reynolds number the sectional lift is set by the geometrically fixed separation rather than by the turbulence closure\. The present time\-averagedCLC\_\{L\}agree with both baselines across that envelope \(Figure[1](https://arxiv.org/html/2609.00507#S2.F1)c\)\.

### 2\.2Dataset training settings

The dataset spans fourteen geometries at thirteen angles of attack: nine geometries were simulated across the full incidence range and five at1∘1^\{\\circ\},6∘6^\{\\circ\}, and12∘12^\{\\circ\}only, giving 132 trajectories\. Seven incidences provide 68 training trajectories, on which the lead time is sampled at random from 1 to 20 ms, and four further incidences provide 36 validation trajectories evaluated at a fixed lead time of 1 ms to monitor the field objective during training; all reported comparisons use the final epoch\-300 checkpoint, so no checkpoint was selected on the validation set\. The test set contains 54 trajectories from the nine fully simulated geometries at six incidences, none of which appears in training, and of these6∘6^\{\\circ\}and12∘12^\{\\circ\}appear in neither the training nor the validation set and therefore test interpolation to an unseen angle of attack\. Evaluation covers every integer lead time from 1 to 30 ms, extending 50% beyond the longest lead time covered in training and testing extrapolation in lead time \(Table[1](https://arxiv.org/html/2609.00507#S2.T1)\)\.

Table 1:Data split by angle of attack\. Lead times are sampled at random within the stated range during training and evaluated at every integer value within it at test time\. The validation set serves only to monitor the field objective and is evaluated at a single lead time\.

## 3Method

### 3\.1GAOT backbone and VATO variants

All four reported configurations use the GAOT backbone\[[23](https://arxiv.org/html/2609.00507#bib.bib23)\]\. The backbone consists of three core modules: a multiscale attentional graph neural operator \(MAGNO\) encoder, which maps unstructured physical coordinates onto a regular latent grid; a vision transformer, which divides the latent grid into patch tokens and processes them with global attention; and a MAGNO decoder, which maps the processed latent features back onto arbitrary query coordinates\. In the validated configuration, the latent grid size is128×128128\\times 128, the lifting channel dimension is 96, the number of transformer layers is six, the number of attention heads is eight, the MAGNO neighbourhood radius is 0\.03 with multiscale factors\[0\.5,1\.0\]\[0\.5,1\.0\], and at most 512 neighbours are aggregated per query point\.

Two common adaptations are used for the present dataset\. First, the encoder and decoder operate on different point sets, so that the model accepts a sampled source state and predicts on the native CFD mesh\. Second, geometry and time are provided through a signed\-distance field and temporal scalars, as defined in Section[3\.2](https://arxiv.org/html/2609.00507#S3.SS2)\. The vortex\-force information is then coupled to this backbone along two routes, which define the two variants \(Figure[2](https://arxiv.org/html/2609.00507#S3.F2)\)\. VATO\-S retains the architecture above and changes only the training objective, through the contribution field supervision defined in Section 3\.6\. VATO\-A instead inserts three zero\-initialised residual attention paths\. After transformer processing, each latent patch reads 256 ordinary source tokens selected from the full current frame using the VFM contribution and sensitivity fields\. Within the MAGNO decoder, each output query reads nearby ordinary source points through gated local attention\. After the preliminary prediction, each native\-mesh query reads the same 256 selected tokens again through a global cross attention residual\. The VFM values determine which full\-frame locations are prioritised but are not embedded in the selected token content\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/vato_architecture.png)Figure 2:VFM coupling to the shared GAOT backbone\. A parameter\-free prioritisation rule forms 256 ordinary source tokens from the current input frame: 64 indices retain geometric coverage, while the remaining 192 are balanced over signed Lift and Drag contribution groups and VFM sensitivity\. VFM values determine the selected locations but are not embedded in the token content\. VATO\-A augments the backbone through three zero\-initialised residual paths\. Processed latent patches read the selected tokens through an eight\-head, width\-192 cross attention; each output query reads up to 32 nearby ordinary source points through a gated four\-head, width\-96 local attention with radius 0\.04; and each preliminary output query reads the selected tokens through a second eight\-head, width\-192 cross attention\. The common field objective combines physical\-space UVP and mesh\-curl signed\-log\-vorticity losses\. VATO\-S retains the GAOT architecture and additionally applies the training\-only VFM contribution field loss, whereas VATO\-A uses the field objective without that VFM loss\.
### 3\.2Prediction task

The prediction target is the future full\-mesh field𝐲⁡\(tin\+τ,𝐗\)\\mathbf\{y\}\(t\_\{\\mathrm\{in\}\}\+\\tau,\\mathbf\{X\}\)\. Each model writes the output as

𝐲^​\(tin\+τ,𝐗\)=Gθ​\(𝐲⁡\(tin,𝐗s\),𝐜⁡\(𝐗s\),tin,τ,𝐗\),\\hat\{\\mathbf\{y\}\}\(t\_\{\\mathrm\{in\}\}\+\\tau,\\,\\mathbf\{X\}\)=G\_\{\\theta\}\\\!\\left\(\\mathbf\{y\}\(t\_\{\\mathrm\{in\}\},\\mathbf\{X\}\_\{s\}\),\\,\\mathbf\{c\}\(\\mathbf\{X\}\_\{s\}\),\\,t\_\{\\mathrm\{in\}\},\\,\\tau,\\,\\mathbf\{X\}\\right\),\(1\)where𝐗=\{𝐱i\}i=1N\\mathbf\{X\}=\\\{\\mathbf\{x\}\_\{i\}\\\}\_\{i=1\}^\{N\}is the native CFD mesh,𝐗s⊂𝐗\\mathbf\{X\}\_\{s\}\\subset\\mathbf\{X\}is the 12,000\-point input set used during training,tint\_\{\\mathrm\{in\}\}is the source time, andτ=ttarget−tin\\tau=t\_\{\\mathrm\{target\}\}\-t\_\{\\mathrm\{in\}\}is the requested scalar lead time\. At each mesh point, the state vector is

𝐲i​\(t\)=\[ui​\(t\),vi​\(t\),pi​\(t\)\]T\.\\mathbf\{y\}\_\{i\}\(t\)=\[u\_\{i\}\(t\),v\_\{i\}\(t\),p\_\{i\}\(t\)\]^\{T\}\.\(2\)Hereuiu\_\{i\}andviv\_\{i\}are velocity components in m s\-1andpip\_\{i\}is pressure in Pa\. The spatial condition field𝐜\\mathbf\{c\}contains the signed distance to the aerofoil surface,

The input at each source point is\[u,v,p\]T\[u,v,p\]^\{T\}together withccand the temporal scalarstint\_\{\\mathrm\{in\}\}andτ\\tau\(Figure[2](https://arxiv.org/html/2609.00507#S3.F2)\)\.

Because the CFD mesh is strongly refined near the body, with roughly 72% of points lying within a near\-wall band of signed distance≤0\.1​c\\leq 0\.1c, an unstratified draw of encoder points would under\-represent the wake\. The sampler therefore reserves 6% of the encoder budget for body\-wall points, caps the near\-wall band at 50%, and assigns the remaining 44% to the far field, while in the reported benchmark every configuration uses𝐗s=𝐗\\mathbf\{X\}\_\{s\}=\\mathbf\{X\}as its encoder source\. Training uses direct single\-step prediction, mapping a source state to a single requested lead time rather than through a recursive rollout\.

### 3\.3Pressure surface integral

CLC\_\{L\}andCDC\_\{D\}are recovered from a predicted field through two readouts, The first integrates the pressure over the body wall\. LetΩb\\Omega\_\{b\}denote the solid aerofoil region and∂Ωb\\partial\\Omega\_\{b\}its boundary\. With𝐧\\mathbf\{n\}the unit normal directed from the body into the fluid, the pressure force per unit span is

𝐅press=−∫∂Ωbp𝐧ds\.\\mathbf\{F\}^\{\\mathrm\{press\}\}=\-\\int\_\{\\partial\\Omega\_\{b\}\}p\\,\\mathbf\{n\}\\,ds\.\(3\)The Cartesian components are projected onto the wind\-axis drag and lift directions using the angle of attackα\\alpha:

FDpress\\displaystyle F\_\{D\}^\{\\mathrm\{press\}\}=Fxpress​cos⁡α\+Fypress​sin⁡α,\\displaystyle=F\_\{x\}^\{\\mathrm\{press\}\}\\cos\\alpha\+F\_\{y\}^\{\\mathrm\{press\}\}\\sin\\alpha,\(4\)FLpress\\displaystyle F\_\{L\}^\{\\mathrm\{press\}\}=−Fxpress​sin⁡α\+Fypress​cos⁡α\.\\displaystyle=\-F\_\{x\}^\{\\mathrm\{press\}\}\\sin\\alpha\+F\_\{y\}^\{\\mathrm\{press\}\}\\cos\\alpha\.\(5\)With freestream densityρ∞\\rho\_\{\\infty\}, speedU∞U\_\{\\infty\}, and reference chordcc, the corresponding pressure\-induced force coefficient is

Ckpress=Fkpress12​ρ∞​U∞2​c,k∈\{L,D\}\.C\_\{k\}^\{\\mathrm\{press\}\}=\\frac\{F\_\{k\}^\{\\mathrm\{press\}\}\}\{\\frac\{1\}\{2\}\\rho\_\{\\infty\}U\_\{\\infty\}^\{2\}c\},\\qquad k\\in\\\{L,D\\\}\.\(6\)In the implementation, Eq\. \([3](https://arxiv.org/html/2609.00507#S3.E3)\) is evaluated by quadrature over the body\-boundary edges of each sample’s native mesh\. For the pressure readout,ppis given by either the target or model\-predicted pressure field\.

### 3\.4VFM volume integral

The VFM volume integral provides a second, geometrically distinct path to lift and drag, supported on the wake and separation region rather than the body wall\. LetΩ\\Omegabe the fluid domain and𝐞k\\mathbf\{e\}\_\{k\}the wind\-axis unit vector fork∈\{L,D\}k\\in\\\{L,D\\\}\. VFM introduces a geometry\- and wind\-axis\-dependent auxiliary potential

\{∇2ϕk=0,𝐱∈Ω,∇ϕk⋅𝐧=𝐞k⋅𝐧,𝐱∈∂Ωb,∇ϕk→0,\|𝐱\|→∞,\\begin\{cases\}\\nabla^\{2\}\\phi\_\{k\}=0,&\\mathbf\{x\}\\in\\Omega,\\\\ \\nabla\\phi\_\{k\}\\cdot\\mathbf\{n\}=\\mathbf\{e\}\_\{k\}\\cdot\\mathbf\{n\},&\\mathbf\{x\}\\in\\partial\\Omega\_\{b\},\\\\ \\nabla\\phi\_\{k\}\\rightarrow 0,&\|\\mathbf\{x\}\|\\rightarrow\\infty,\\end\{cases\}\(7\)from which the vortex\-force factor is obtained as

𝚲k=∇⟂ϕk=\(ϕk,y,−ϕk,x\)T≡\(Pk,Qk\)T\.\\bm\{\\Lambda\}\_\{k\}=\\nabla^\{\\perp\}\\phi\_\{k\}=\(\\phi\_\{k,y\},\-\\phi\_\{k,x\}\)^\{T\}\\equiv\(P\_\{k\},Q\_\{k\}\)^\{T\}\.\(8\)VATO uses the vortex\-pressure contribution of this formulation\. Taking the constant fluid density asρ=ρ∞\\rho=\\rho\_\{\\infty\}, the force per unit span is

FkVFM=ρ​∫Ω\(𝚲k⋅𝐮\)​ωz​𝑑A\.F\_\{k\}^\{\\mathrm\{VFM\}\}=\\rho\\int\_\{\\Omega\}\\left\(\\bm\{\\Lambda\}\_\{k\}\\cdot\\mathbf\{u\}\\right\)\\omega\_\{z\}\\,dA\.\(9\)With the freestream dynamic\-pressure scale, the corresponding force coefficient is

CkVFM=FkVFM12​ρ​U∞2​c=112​U∞2​c​∫Ω\(𝚲k⋅𝐮\)​ωz​𝑑A,k∈\{L,D\}\.C\_\{k\}^\{\\mathrm\{VFM\}\}=\\frac\{F\_\{k\}^\{\\mathrm\{VFM\}\}\}\{\\frac\{1\}\{2\}\\rho U\_\{\\infty\}^\{2\}c\}=\\frac\{1\}\{\\frac\{1\}\{2\}U\_\{\\infty\}^\{2\}c\}\\int\_\{\\Omega\}\\left\(\\bm\{\\Lambda\}\_\{k\}\\cdot\\mathbf\{u\}\\right\)\\omega\_\{z\}\\,dA,\\qquad k\\in\\\{L,D\\\}\.\(10\)The fieldsPL​\(x\)P\_\{L\}\(x\),QL​\(x\)Q\_\{L\}\(x\),PD​\(x\)P\_\{D\}\(x\), andQD​\(x\)Q\_\{D\}\(x\)are precomputed for each geometry and angle of attack and sampled at the corresponding mesh points \(Figure[3](https://arxiv.org/html/2609.00507#S3.F3)a–d\)\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/vfm_physics_basis.png)Figure 3:Physical basis for the VFM coupling, shown for\(θ1,θ2\)=\(174\.1∘,155\.0∘\)\(\\theta\_\{1\},\\theta\_\{2\}\)=\(174\.1^\{\\circ\},155\.0^\{\\circ\}\)at AoA=12∘=12^\{\\circ\}\. \(a–d\) Panel\-method coefficient fieldsPL,QL,PD,QDP\_\{L\},Q\_\{L\},P\_\{D\},Q\_\{D\}\. \(e\) Lift and drag recovered on paired left/right ordinates from the target force record, the pressure integral on the target pressure field, and the VFM integral evaluated on the target velocity and mesh curl; the VFM mean absolute relative errors are 2\.09% forCLC\_\{L\}and 2\.22% forCDC\_\{D\}\. \(f\) Cumulative absolute force contribution after ranking fluid points by local VFM magnitude\. At the displayed input frame \(t=0\.073t=0\.073s\) the top 10% contributes approximately 79% of lift and 77% of drag, and the top 20% approximately 90% of both\.For physics validation and predicted\-field post\-processing, Eq\. \([10](https://arxiv.org/html/2609.00507#S3.E10)\) is evaluated in discrete form on the query mesh\. The spanwise vorticity is reconstructed asw≡ωz=∂v/∂x−∂u/∂yw\\equiv\\omega\_\{z\}=\\partial v/\\partial x\-\\partial u/\\partial yusing the native mesh derivatives\. Let𝒱\\mathcal\{V\}be the valid fluid\-point set andaia\_\{i\}the nodal quadrature weight; in the sampled training path,aia\_\{i\}also contains the stratified\-sampling correction\. The discrete coefficient is

CkVFM=112​U∞2​c​∑i∈𝒱\[Pk,i​ui\+Qk,i​vi\]​ωi​ai,k∈\{L,D\}\.C\_\{k\}^\{\\mathrm\{VFM\}\}=\\frac\{1\}\{\\frac\{1\}\{2\}U\_\{\\infty\}^\{2\}c\}\\sum\_\{i\\in\\mathcal\{V\}\}\\bigl\[P\_\{k,i\}u\_\{i\}\+Q\_\{k,i\}v\_\{i\}\\bigr\]\\,\\omega\_\{i\}\\,a\_\{i\},\\qquad k\\in\\\{L,D\\\}\.\(11\)For predicted\-field readout,\(ui,vi,ωi\)\(u\_\{i\},v\_\{i\},\\omega\_\{i\}\)in Eq\. \([11](https://arxiv.org/html/2609.00507#S3.E11)\) are replaced by\(u^i,v^i,ω^i\)\(\\hat\{u\}\_\{i\},\\hat\{v\}\_\{i\},\\hat\{\\omega\}\_\{i\}\), withω^i\\hat\{\\omega\}\_\{i\}reconstructed from the predicted velocity by the same mesh\-curl operator\.

Both operators are validated on the target fields before being applied to predicted fields\. The validation uses a representative test trajectory at an incidence excluded from training, with\(θ1,θ2\)=\(174\.1∘,155\.0∘\)\(\\theta\_\{1\},\\theta\_\{2\}\)=\(174\.1^\{\\circ\},155\.0^\{\\circ\}\)and AoA=12∘=12^\{\\circ\}over a0\.180\.18s physical time, and takes theCLC\_\{L\}andCDC\_\{D\}recorded by the solver as the reference\. The viscous contribution to these coefficients is negligible on this case, so the reference is governed by the pressure force\. The coefficients recovered through Eqs\. \([3](https://arxiv.org/html/2609.00507#S3.E3)\) agree closely with the target values, with correlation coefficientsr=1\.0000r=1\.0000forCLC\_\{L\}andr=0\.9996r=0\.9996forCDC\_\{D\}\. The VFM integral, computed fromuu,vv, andω\\omegawithout any pressure information, givesr=0\.9721r=0\.9721andr=0\.9695r=0\.9695and matches the target to mean absolute relative errors of 2\.09% and 2\.22% \(Figure[3](https://arxiv.org/html/2609.00507#S3.F3)e\)\. Both operators recover accurate coefficients from different information\. The agreement establishes both the accuracy of the VFM formulation on this flow and the reliability of the two operators used to recover lift and drag from predicted fields in Section[4](https://arxiv.org/html/2609.00507#S4)\.

The same integrand, evaluated pointwise on the target field by combining the four panel\-method coefficients with the reconstructed vorticity and quadrature weight, gives the pointwise force\-contribution mapsΦL,i\\Phi\_\{L,i\}andΦD,i\\Phi\_\{D,i\}\. At the displayed input\-frame snapshot \(t=0\.073t=0\.073s\), the top 10% of fluid points account for approximately 79% of the absolute lift contribution and 77% of the absolute drag contribution; the corresponding top\-20% shares are approximately 90% \(Figure[3](https://arxiv.org/html/2609.00507#S3.F3)f\)\. This concentration motivates the fixed source budget in Section[3\.7](https://arxiv.org/html/2609.00507#S3.SS7)\.

### 3\.5Data preprocessing

All CFD fields are first represented as point data on the native mesh\. For each geometry–incidence trajectory, an inlet\-background state𝐛g\\mathbf\{b\}\_\{g\}is estimated by averaging\[u,v,p\]T\[u,v,p\]^\{T\}over the upstream bandΩup\\Omega\_\{\\mathrm\{up\}\}of the first frame\. The state field used for learning is the corresponding perturbation field,

𝐲~i​\(t\)=𝐲i​\(t\)−𝐛g\.\\tilde\{\\mathbf\{y\}\}\_\{i\}\(t\)=\\mathbf\{y\}\_\{i\}\(t\)\-\\mathbf\{b\}\_\{g\}\.\(12\)The residual state field and the spatial condition field are then standardised using training\-set statistics,

𝐲i∗​\(t\)=𝐲~i​\(t\)−𝝁y𝝈y,ci∗=ci−μcσc\.\\mathbf\{y\}\_\{i\}^\{\*\}\(t\)=\\frac\{\\tilde\{\\mathbf\{y\}\}\_\{i\}\(t\)\-\\bm\{\\mu\}\_\{y\}\}\{\\bm\{\\sigma\}\_\{y\}\},\\qquad c\_\{i\}^\{\*\}=\\frac\{c\_\{i\}\-\\mu\_\{c\}\}\{\\sigma\_\{c\}\}\.\(13\)Here𝐲~i​\(t\)\\tilde\{\\mathbf\{y\}\}\_\{i\}\(t\)is the perturbation state at pointii,𝐲i∗​\(t\)\\mathbf\{y\}^\{\*\}\_\{i\}\(t\)andci∗c^\{\*\}\_\{i\}are its standardised counterpart and the standardised condition field, andμy\\mu\_\{y\},σy\\sigma\_\{y\},μc\\mu\_\{c\},σc\\sigma\_\{c\}are per\-channel means and standard deviations computed over the training set\.

The temporal scalars are scaled separately before being appended to the model input\. Points inside the aerofoil, identified from the signed\-distance field, are masked in the state and appended temporal channels, while signed distance remains available as the geometry condition\. Vorticity is not used as a predicted state channel\.

### 3\.6VFM contribution field supervision

VATO\-S preserves the GAOT architecture and adds supervision on the two per\-point contribution fields whose sums give the discrete VFMCLC\_\{L\}andCDC\_\{D\}\. Fork∈\{L,D\}k\\in\\\{L,D\\\}, the VFM contribution field at mesh pointiiis

Φk,i=\(Pk,i​ui\+Qk,i​vi\)​ωi​ai12​U∞2​c\.\\Phi\_\{k,i\}=\\frac\{\(P\_\{k,i\}u\_\{i\}\+Q\_\{k,i\}v\_\{i\}\)\\,\\omega\_\{i\}a\_\{i\}\}\{\\frac\{1\}\{2\}U\_\{\\infty\}^\{2\}c\}\.\(14\)The auxiliary objective compares the predicted and targetCLC\_\{L\}/CDC\_\{D\}maps as one vector\-valued field,

LVFM​\-​contrib=∑i∈𝒱∑k∈\{L,D\}\(Φ^k,i−Φk,i\)2max⁡\(∑i∈𝒱∑k∈\{L,D\}Φk,i2,ε\)\.L^\{\\mathrm\{VFM\\mbox\{\-\}contrib\}\}=\\frac\{\\sum\_\{i\\in\\mathcal\{V\}\}\\sum\_\{k\\in\\\{L,D\\\}\}\\left\(\\hat\{\\Phi\}\_\{k,i\}\-\\Phi\_\{k,i\}\\right\)^\{2\}\}\{\\max\\\!\\left\(\\sum\_\{i\\in\\mathcal\{V\}\}\\sum\_\{k\\in\\\{L,D\\\}\}\\Phi\_\{k,i\}^\{2\},\\varepsilon\\right\)\}\.\(15\)The predicted contribution fieldΦ^\\hat\{\\Phi\}uses predicted velocity and curl\-derived vorticity; pressure does not enter Eq\. \([14](https://arxiv.org/html/2609.00507#S3.E14)\)\. The target fieldΦ\\Phiis detached from the gradient graph\. Equation \([15](https://arxiv.org/html/2609.00507#S3.E15)\) is evaluated for each sample withε=10−8\\varepsilon=10^\{\-8\}and then averaged over the valid samples in the batch\. The target contribution field is available only during training\. At inference, VATO\-S receives the same state, geometry, and time inputs as the matched GAOT control and therefore adds no model parameters or inference\-time operator\.

### 3\.7Source prioritisation and residual cross attention

VATO\-A uses current flow vorticity together with VFM contribution and sensitivity fields to identify regions relevant to Lift and Drag\. This information prioritises source locations but is neither embedded as an input field nor used as an auxiliary loss\. From the full native input frame, the model combines the mesh curl of\(u,v\)\(u,v\), the precomputed VFM basis, nodal quadrature weight, and dynamic pressure to form signed local contributions and their normalised positive and negative groups:

Φk,i\\displaystyle\\Phi\_\{k,i\}=\(Pk,i​ui\+Qk,i​vi\)​ωi​ai12​U∞2​c,\\displaystyle=\\frac\{\(P\_\{k,i\}u\_\{i\}\+Q\_\{k,i\}v\_\{i\}\)\\,\\omega\_\{i\}a\_\{i\}\}\{\\frac\{1\}\{2\}U\_\{\\infty\}^\{2\}c\},k∈\{L,D\},\\displaystyle k\\in\\\{L,D\\\},\(16\)σk\\displaystyle\\sigma\_\{k\}=\(1\|𝒱\|​∑j∈𝒱Φk,j2\)1/2,\\displaystyle=\\left\(\\frac\{1\}\{\|\\mathcal\{V\}\|\}\\sum\_\{j\\in\\mathcal\{V\}\}\\Phi\_\{k,j\}^\{2\}\\right\)^\{1/2\},Φ^k,i\\displaystyle\\widehat\{\\Phi\}\_\{k,i\}=clip⁡\(Φk,imax⁡\(σk,εr\),−γ,γ\),\\displaystyle=\\operatorname\{clip\}\\\!\\left\(\\frac\{\\Phi\_\{k,i\}\}\{\\max\(\\sigma\_\{k\},\\varepsilon\_\{r\}\)\},\-\\gamma,\\gamma\\right\),rk,i\+\\displaystyle r\_\{k,i\}^\{\+\}=max⁡\(Φ^k,i,0\),\\displaystyle=\\max\\\!\\left\(\\widehat\{\\Phi\}\_\{k,i\},0\\right\),rk,i−\\displaystyle r\_\{k,i\}^\{\-\}=max⁡\(−Φ^k,i,0\)\.\\displaystyle=\\max\\\!\\left\(\-\\widehat\{\\Phi\}\_\{k,i\},0\\right\)\.Here𝒱\\mathcal\{V\}is the valid fluid\-point set,σk\\sigma\_\{k\}is the framewise root mean square \(RMS\) contribution,εr=10−12\\varepsilon\_\{r\}=10^\{\-12\}prevents division by zero, andγ=6\\gamma=6is the fixed clipping limit\. The four local VFM sensitivity factors, defined with vorticity held fixed, are

sk,u,i=Pk,i​ωi​ai12​U∞2​c,sk,v,i=Qk,i​ωi​ai12​U∞2​c,k∈\{L,D\}\.s\_\{k,u,i\}=\\frac\{P\_\{k,i\}\\omega\_\{i\}a\_\{i\}\}\{\\frac\{1\}\{2\}U\_\{\\infty\}^\{2\}c\},\\qquad s\_\{k,v,i\}=\\frac\{Q\_\{k,i\}\\omega\_\{i\}a\_\{i\}\}\{\\frac\{1\}\{2\}U\_\{\\infty\}^\{2\}c\},\\qquad k\\in\\\{L,D\\\}\.\(17\)Form∈\{u,v\}m\\in\\\{u,v\\\}, these factors are normalised and combined as

ηk\\displaystyle\\eta\_\{k\}=\[12​\|𝒱\|​∑j∈𝒱\(sk,u,j2\+sk,v,j2\)\]1/2,\\displaystyle=\\left\[\\frac\{1\}\{2\|\\mathcal\{V\}\|\}\\sum\_\{j\\in\\mathcal\{V\}\}\\left\(s\_\{k,u,j\}^\{2\}\+s\_\{k,v,j\}^\{2\}\\right\)\\right\]^\{1/2\},\(18\)s^k,m,i\\displaystyle\\widehat\{s\}\_\{k,m,i\}=clip⁡\(sk,m,imax⁡\(ηk,εr\),−γ,γ\),\\displaystyle=\\operatorname\{clip\}\\\!\\left\(\\frac\{s\_\{k,m,i\}\}\{\\max\(\\eta\_\{k\},\\varepsilon\_\{r\}\)\},\-\\gamma,\\gamma\\right\),ris\\displaystyle r\_\{i\}^\{s\}=\[14​∑k∈\{L,D\}∑m∈\{u,v\}s^k,m,i2\]1/2,\\displaystyle=\\left\[\\frac\{1\}\{4\}\\sum\_\{k\\in\\\{L,D\\\}\}\\sum\_\{m\\in\\\{u,v\\\}\}\\widehat\{s\}\_\{k,m,i\}^\{2\}\\right\]^\{1/2\},riall\\displaystyle r\_\{i\}^\{\\mathrm\{all\}\}=\[\(ris\)2\+Φ^L,i2\+Φ^D,i2\]1/2\.\\displaystyle=\\left\[\(r\_\{i\}^\{s\}\)^\{2\}\+\\widehat\{\\Phi\}\_\{L,i\}^\{2\}\+\\widehat\{\\Phi\}\_\{D,i\}^\{2\}\\right\]^\{1/2\}\.The scorerisr\_\{i\}^\{s\}forms the fifth prioritisation group, whileriallr\_\{i\}^\{\\mathrm\{all\}\}fills any unassigned quota\. All prioritisation values and selected indices are detached, and only input\-frame quantities are used\. Exactly 64 tokens are reserved for spatial coverage over the valid fluid domain\. The remaining 192\-token budget is balanced overrL\+r\_\{L\}^\{\+\},rL−r\_\{L\}^\{\-\},rD\+r\_\{D\}^\{\+\},rD−r\_\{D\}^\{\-\}, and combined VFM sensitivity, with unfilled quotas assigned by the combined VFM contribution–sensitivity score\.

Each selected token carries the baseline source features\[u,v,p,c,tin,τ\]\[u,v,p,c,t\_\{\\mathrm\{in\}\},\\tau\]and coordinates\(x,y\)\(x,y\), not the VFM values used for prioritisation\. Let𝐙F\\mathbf\{Z\}\_\{F\}denote these 256 full\-frame tokens,𝐒𝒩⁡\(q\)\\mathbf\{S\}\_\{\\mathcal\{N\}\(q\)\}the ordinary source tokens in the local neighbourhood of queryqq,𝐇\\mathbf\{H\}the processed latent patches,𝐡q\\mathbf\{h\}\_\{q\}the MAGNO\-decoded feature at queryqq, and𝐲^q0\\hat\{\\mathbf\{y\}\}\_\{q\}^\{\\,0\}the preliminary UVP prediction\. VATO\-A applies

𝐇\+\\displaystyle\\mathbf\{H\}^\{\+\}=𝐇\+𝒜patch​\(𝐇,𝐙F\),\\displaystyle=\\mathbf\{H\}\+\\mathcal\{A\}\_\{\\mathrm\{patch\}\}\(\\mathbf\{H\},\\mathbf\{Z\}\_\{F\}\),\(19\)𝐡q\+\\displaystyle\\mathbf\{h\}\_\{q\}^\{\+\}=𝐡q\+𝐠q⊙𝒜local​\(𝐡q,𝐒𝒩⁡\(q\)\),\\displaystyle=\\mathbf\{h\}\_\{q\}\+\\mathbf\{g\}\_\{q\}\\odot\\mathcal\{A\}\_\{\\mathrm\{local\}\}\(\\mathbf\{h\}\_\{q\},\\mathbf\{S\}\_\{\\mathcal\{N\}\(q\)\}\),𝐲^q\\displaystyle\\hat\{\\mathbf\{y\}\}\_\{q\}=𝐲^q0\+𝒜query​\(\[𝐲^q0,𝐱q\],𝐙F\)\.\\displaystyle=\\hat\{\\mathbf\{y\}\}\_\{q\}^\{\\,0\}\+\\mathcal\{A\}\_\{\\mathrm\{query\}\}\(\[\\hat\{\\mathbf\{y\}\}\_\{q\}^\{\\,0\},\\mathbf\{x\}\_\{q\}\],\\mathbf\{Z\}\_\{F\}\)\.Here each𝒜\\mathcal\{A\}is multi\-head scaled dot\-product attention, with its first argument providing queries and its second providing keys and values;⊙\\odotdenotes channel\-wise multiplication\. The operators𝒜patch\\mathcal\{A\}\_\{\\mathrm\{patch\}\}and𝒜query\\mathcal\{A\}\_\{\\mathrm\{query\}\}are eight\-head, width\-192 residual cross attention modules\. The former is applied after Transformer processing and before MAGNO decoding; the latter is applied to the preliminary UVP output in query chunks of 2,048\. The local path uses four\-head, width\-96 attention over the neighbourhood𝒩⁡\(q\)\\mathcal\{N\}\(q\)of at most 32 ordinary source points within radius 0\.04, including relative displacement and distance\. Its learned sigmoid gate𝐠q∈\(0,1\)96\\mathbf\{g\}\_\{q\}\\in\(0,1\)^\{96\}modulates the local residual before the final decoder projection\. All three output projections are zero\-initialised\. Figure[4](https://arxiv.org/html/2609.00507#S3.F4)visualises the prioritisation rule at the trained 256\-token budget\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/v317_goal_focus_allocation.png)Figure 4:Source prioritisation for the input frame att=0\.073t=0\.073s with\(θ1,θ2\)=\(174\.1∘,155\.0∘\)\(\\theta\_\{1\},\\theta\_\{2\}\)=\(174\.1^\{\\circ\},155\.0^\{\\circ\}\), at \(a\) AoA=3∘=3^\{\\circ\}and \(b\)12∘12^\{\\circ\}\. The stacked left maps show the RMS\-normalised signed VFM Lift and Drag contributions, with markers identifying the indices retained from their positive and negative groups; display values are clipped to\[−0\.5,0\.5\]\[\-0\.5,0\.5\]without changing the indices\. The right maps show the 192 VFM\-prioritised indices over the input\-frame vorticity, with the 64 spatial\-coverage indices omitted for clarity\. The displayed allocation uses the trained 256\-token budget\. Prioritisation uses only detached input\-frame quantities, and the selected tokens carry ordinary state and coordinates rather than VFM values\. The vorticity background uses the common\[−1000,1000\]\[\-1000,1000\]s\-1scale on the native mesh\.
### 3\.8Training objective and protocol

The GAOT reference, matched GAOT control, and both VATO variants optimise physical\-space, squared relativeL2L\_\{2\}errors for velocity and pressure together with a mesh\-curl vorticity term:

Ltotal\\displaystyle L^\{\\mathrm\{total\}\}=∑f∈\{u,v,p\}αf​‖f^−f‖2,𝒱2max⁡\(‖f‖2,𝒱2,εℓ\)\+αw​‖w^log−wlog‖2,𝒱2max⁡\(‖wlog‖2,𝒱2,εℓ\),\\displaystyle=\\sum\_\{f\\in\\\{u,v,p\\\}\}\\alpha\_\{f\}\\frac\{\\\|\\hat\{f\}\-f\\\|\_\{2,\\mathcal\{V\}\}^\{2\}\}\{\\max\(\\\|f\\\|\_\{2,\\mathcal\{V\}\}^\{2\},\\varepsilon\_\{\\ell\}\)\}\+\\alpha\_\{w\}\\frac\{\\\|\\hat\{w\}\_\{\\log\}\-w\_\{\\log\}\\\|\_\{2,\\mathcal\{V\}\}^\{2\}\}\{\\max\(\\\|w\_\{\\log\}\\\|\_\{2,\\mathcal\{V\}\}^\{2\},\\varepsilon\_\{\\ell\}\)\},\(20\)wlog\\displaystyle w\_\{\\log\}=sign⁡\(ω\)​log⁡\(1\+\|ω\|\),\\displaystyle=\\operatorname\{sign\}\(\\omega\)\\log\(1\+\|\\omega\|\),w^log\\displaystyle\\hat\{w\}\_\{\\log\}=sign⁡\(ω^\)​log⁡\(1\+\|ω^\|\)\.\\displaystyle=\\operatorname\{sign\}\(\\hat\{\\omega\}\)\\log\(1\+\|\\hat\{\\omega\}\|\)\.whereω^\\hat\{\\omega\}is reconstructed from the mesh curl of the predicted velocity andεℓ=10−8\\varepsilon\_\{\\ell\}=10^\{\-8\}\. Each ratio is evaluated per sample over valid fluid points, excluding body\-interior and padded points, and then averaged over the batch\. Velocity and pressure are denormalised and the inlet background is restored before this physical\-space objective is evaluated\. All four channel weights are one\. VATO\-S usesLtotal\+LVFM​\-​contribL^\{\\mathrm\{total\}\}\+L^\{\\mathrm\{VFM\\mbox\{\-\}contrib\}\}with unit weight on the second term\. The GAOT reference, matched GAOT control, and VATO\-A use onlyLtotalL^\{\\mathrm\{total\}\}\. Pressure\-integral, scalar force\-coefficient, and VFM\-integral losses have zero weight in all four configurations and are used only as diagnostic readouts\.

Optimisation uses AdamW with weight decay1×10−41\\times 10^\{\-4\}, drop\-path 0\.1, attention dropout 0\.1, and a cosine\-annealed learning rate from1×10−41\\times 10^\{\-4\}to5×10−55\\times 10^\{\-5\}\. Training uses 8,192 sampled pairs per epoch, batch size 8 per rank, and four\-rank distributed data parallelism for 300 epochs\. All reported comparisons use the final epoch\-300 checkpoint\.

The four configurations are summarised in Table[2](https://arxiv.org/html/2609.00507#S3.T2)\. The GAOT reference uses uniform sampling, meaning the default training sampler without trajectory reweighting\. The matched GAOT control, VATO\-S, and VATO\-A instead share flow\-aware trajectory sampling\. Each trajectory receives a fixed flow\-variation scoreDcD\_\{c\}, formed from the spatial standard deviation of the increment between source and target frames and weighted 0\.20, 0\.45, 0\.25, and 0\.10 overuu,vv,pp, and the variation of that quantity between pairs, soDcD\_\{c\}is governed by the transverse velocity and carries no vorticity contribution\. Source–target pairs are drawn with replacement in proportion tomax⁡\(0\.05,Dc1/2\)\\max\(0\.05,\\,D\_\{c\}^\{1/2\}\), which every pair of a trajectory shares, so the number of pairs a trajectory offers enters its sampling mass alongsideDcD\_\{c\}\. The scores are computed once before training, are not updated from model errors, and are absent at inference\. Because the matched control carries this sampler but no vortex\-force coupling, it isolates the effect of the coupling from that of the sampler\.

Table 2:The four configurations\. GAOT is the published architecture retrained under the common task protocol, so the values listed come from this retraining rather than from the cited paper\. Uniform denotes the default training sampler without trajectory reweighting; flow\-aware denotes fixed training\-only trajectory weighting derived from the variation between source and target frames\. GAOT differs from GAOT \(matched\) only in the sampler, and GAOT \(matched\) differs from VATO\-S only in the contribution field loss, so each row isolates one change from the row above\. Inference is the median preprocessing\-plus\-forward time per sample for batch size 8 on an NVIDIA H200, excluding data loading and force post\-processing\.
### 3\.9Evaluation metrics

Field errors are evaluated separately for the velocity vector, pressure, and vorticity over the valid fluid\-point set𝒱\\mathcal\{V\}:

r​L2U​V\\displaystyle rL\_\{2\}^\{UV\}=\[∑i∈𝒱\(\(u^i−ui\)2\+\(v^i−vi\)2\)∑i∈𝒱\(ui2\+vi2\)\]1/2,\\displaystyle=\\left\[\\frac\{\\sum\_\{i\\in\\mathcal\{V\}\}\\bigl\(\(\\hat\{u\}\_\{i\}\-u\_\{i\}\)^\{2\}\+\(\\hat\{v\}\_\{i\}\-v\_\{i\}\)^\{2\}\\bigr\)\}\{\\sum\_\{i\\in\\mathcal\{V\}\}\(u\_\{i\}^\{2\}\+v\_\{i\}^\{2\}\)\}\\right\]^\{1/2\},\(21\)r​L2f\\displaystyle rL\_\{2\}^\{f\}=\[∑i∈𝒱\(f^i−fi\)2∑i∈𝒱fi2\]1/2,f∈\{p,w\}\.\\displaystyle=\\left\[\\frac\{\\sum\_\{i\\in\\mathcal\{V\}\}\(\\hat\{f\}\_\{i\}\-f\_\{i\}\)^\{2\}\}\{\\sum\_\{i\\in\\mathcal\{V\}\}f\_\{i\}^\{2\}\}\\right\]^\{1/2\},\\qquad f\\in\\\{p,w\\\}\.\(22\)Bothwwandw^\\hat\{w\}are reconstructed from the corresponding velocity field with the same native mesh\-curl operator\. Tables and figures report100​r​L2100\\,rL\_\{2\}in percent\. For operatoro∈\{press,VFM\}o\\in\\\{\\mathrm\{press\},\\mathrm\{VFM\}\\\}and force directionk∈\{L,D\}k\\in\\\{L,D\\\}, the sequence MAE for one source anchor and lead\-time block𝒯\\mathcal\{T\}is

MAEko=1\|𝒯\|​∑τ∈𝒯\|C^ko​\(τ\)−Cko​\(τ\)\|\.\\operatorname\{MAE\}\_\{k\}^\{o\}=\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\}\\left\|\\hat\{C\}\_\{k\}^\{o\}\(\\tau\)\-C\_\{k\}^\{o\}\(\\tau\)\\right\|\.\(23\)HereC^ko​\(τ\)\\hat\{C\}\_\{k\}^\{o\}\(\\tau\)andCko​\(τ\)C\_\{k\}^\{o\}\(\\tau\)apply the same operatorooto the predicted and target fields, respectively, at lead timeτ\\tau\. For each reported lead\-time window, pair\-level field errors and anchor\-level force MAEs are averaged over anchors within a trajectory, over the six incidences within a geometry, and then with equal weight over the nine geometries\. Lead\-time\-resolved figures retainτ\\taubefore applying the same anchor–incidence–geometry hierarchy\. Body\-interior points are excluded from all field metrics\.

## 4Results

All four configurations reduce both the training and validation field objectives under the common 300\-epoch protocol \(Figure[5](https://arxiv.org/html/2609.00507#S4.F5)\)\. The GAOT reference ends at a validation field loss of 0\.3741 and the matched GAOT control at 0\.3569\. VATO\-S ends at 0\.3602, marginally above the matched control, while the benchmark below shows it to be substantially more accurate over the evaluated horizon; the validation view is restricted to a one\-frame lead time, so the contribution field target carries no measurable advantage on single\-step pointwise accuracy\. VATO\-A ends at 0\.2325, a 34\.9% reduction relative to the matched control\. These curves describe one completed run per configuration\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_training_dynamics.png)Figure 5:Training dynamics under the common 300\-epoch protocol\. \(a\) Training and \(b\) validation values of the field objective shared by all four configurations, defined by Eq\. \([20](https://arxiv.org/html/2609.00507#S3.E20)\) with unit channel weights\. The auxiliary contribution field term of VATO\-S is excluded so that all four configurations share the same ordinate\. Each point is the exact aggregate logged at a five\-epoch interval, with no smoothing or interpolation\.### 4\.1Comprehensive direct\-field benchmark

The final 300\-epoch checkpoints of the four configurations were applied to the test set defined in Section[2\.2](https://arxiv.org/html/2609.00507#S2.SS2)\. Each of the 54 trajectories supplies 28 source anchors and every integer lead time from 1 to 30 ms, giving 45,360 source–target combinations per configuration\. Errors follow the aggregation of Section[3\.9](https://arxiv.org/html/2609.00507#S3.SS9)and are reported in Table[3](https://arxiv.org/html/2609.00507#S4.T3)\.

Table 3:Field errors and force diagnostics on the 54\-trajectory benchmark, with 45,360 source–target combinations per configuration\. \(a\) Fluid\-only relativeL2L\_\{2\}error in percent, with velocity pooled overuuandvvthrough their joint vector energy\. \(b\) Sequence mean absolute error in theCLC\_\{L\}andCDC\_\{D\}coefficients, obtained by applying the pressure\-surface and vortex\-force\-volume operators to the predicted and the target fields\. Aggregation follows Section[3\.9](https://arxiv.org/html/2609.00507#S3.SS9)\. Parentheses give the reduction against the GAOT reference in the same column, with positive values favouring the named configuration\. Lead times 1–20 ms match the training horizon and 21–30 ms extend it by 50%\. Values are shown to one decimal place, and column minima are bold\.\(a\) Field prediction: relativeL2L\_\{2\}error \(%\)

\(b\)CLC\_\{L\}andCDC\_\{D\}MAE from predicted fields

The two GAOT configurations separate the headline comparison from attribution to the training protocol\. The GAOT reference uses uniform sampling and is the denominator for all displayed percentages\. Relative to it, flow\-aware sampling in the matched GAOT control increases velocity, pressure, and vorticity error by 0\.8%, 9\.4%, and 0\.7% over lead times 1–20 ms and by 2\.6%, 7\.1%, and 2\.0% over lead times 21–30 ms\. It nevertheless reduces every force MAE, with gains ranging from 2\.6% to 19\.5%\. Flow\-aware sampling therefore lowers the force\-diagnostic errors at the cost of broad field accuracy\. Both VATO configurations retain this sampling regime\.

Relative to the GAOT reference, VATO\-S reduces velocity and vorticity error by 10\.4% and 15\.6% over lead times 1–20 ms and by 3\.8% and 13\.7% over lead times 21–30 ms; its pressure changes are\+1\.0%\+1\.0\\%and−1\.6%\-1\.6\\%\. Measured against the matched GAOT control, which shares its sampling regime, the pressure reductions are 9\.5% and 5\.2%, so the auxiliary term recovers the pressure accuracy that flow\-aware sampling gives up and, over lead times 1–20 ms, returns it to the level of the reference\. VATO\-A gives reductions of 15\.8%, 7\.5%, and 31\.2% over lead times 1–20 ms and 6\.5%, 0\.04%, and 26\.9% over lead times 21–30 ms\. It is therefore the minimum\-error configuration in all six field cells, although its extrapolated\-pressure advantage over the GAOT reference is negligible\. Paired geometry–incidence bootstrap intervals exclude zero for five of the six VATO\-A reductions; extrapolated pressure is the exception\.

The force readouts separate the two configurations differently\. VATO\-A gives the lowest pressure\-derivedCLC\_\{L\}andCDC\_\{D\}MAE in both windows, reducing them by 13\.9% and 17\.1% over lead times 1–20 ms and by 17\.2% and 21\.7% over lead times 21–30 ms\. It also reduces VFM\-derivedCLC\_\{L\}MAE by 28\.0% and 29\.4%\. VATO\-S instead gives the lowest VFM\-derivedCDC\_\{D\}MAE, with reductions of 26\.7% and 33\.6%\. The two complete configurations therefore have distinct functional profiles: VATO\-A gives the stronger overall field and pressure\-derived force result, whereas VATO\-S is most favourable for the VFM\-derived Drag readout\. Because VATO\-A changes source prioritisation, three residual attention paths, and trainable capacity together, this contrast does not isolate the effect of the VFM interface alone\.

The matched GAOT control and VATO\-S each contain 19\.09 million parameters and require 34\.8 ms per sample\. VATO\-A contains 20\.08 million parameters and requires 56\.9 ms per sample, about 63\.4% above the matched control\. The additional cost arises from its latent\-patch, local\-query, and output\-query residual attention paths\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_lag30_absolute_profiles.png)Figure 6:Lead\-time\-resolved absolute errors on the 54\-trajectory population\. At each lead time, errors are averaged over the 28 anchors within each trajectory, over the six evaluated incidences within each geometry, and then with equal weight over the nine geometries\. \(a\) Equal\-channel UVP relativeL2L\_\{2\}error for GAOT, GAOT \(matched\), VATO\-S, and VATO\-A\. \(b\) Equal mean of pressure\-derivedCLC\_\{L\}andCDC\_\{D\}MAE at each lead time for all four configurations, with the VFM\-derived VATO\-A result shown as an additional dashed curve\. In\-distribution \(ID\) training range \(τ=1\\tau=1–2020ms\) and the temporal extrapolation beyond the maximum training lead\-time, out\-of\-distribution \(OOD\) range \(τ=21\\tau=21–3030ms\)\.Resolving the same errors by lead time shows the ordering of the four configurations to be stable across the horizon rather than produced by a particular window \(Figure[6](https://arxiv.org/html/2609.00507#S4.F6)\)\. Figure[7](https://arxiv.org/html/2609.00507#S4.F7)resolves the improvement of VATO\-A over the reference by lead time\. Panel \(a\) reports the equal\-channel UVP relativeL2L\_\{2\}improvement, which reaches close to 24% at the shortest lead times and decreases as the horizon grows, giving window\-level values of 11\.8% over the trained range and 3\.0% over the extrapolation window\. Panel \(b\) reports the improvement in the pressure\-derivedCLC\_\{L\}andCDC\_\{D\}mean absolute error on the same population, which holds between roughly 12% and 20% across the trained range\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_lag30_improvement_profiles.png)Figure 7:Lead\-time\-resolved improvement of VATO\-A over the GAOT reference on the same reported 54\-trajectory population\. \(a\) Equal\-channel UVP relativeL2L\_\{2\}improvement\. \(b\) Pressure\-derivedCLC\_\{L\}andCDC\_\{D\}MAE improvement\. Improvement is defined as100​\(EGAOT−EVATO​\-​A\)/EGAOT100\(E\_\{\\mathrm\{GAOT\}\}\-E\_\{\\mathrm\{VATO\\text\{\-\}A\}\}\)/E\_\{\\mathrm\{GAOT\}\}, whereEEdenotes UVP relativeL2L\_\{2\}error in \(a\) and the correspondingCLC\_\{L\}orCDC\_\{D\}MAE in \(b\); positive values favour VATO\-A\. The dashed horizontal segments in \(a\) report window\-level improvement after averaging each model’s error over the ID and OOD\.Beyond 20 ms the two force curves rise together rather than following the field metric downward, reaching close to 20% forCLC\_\{L\}and close to 25% forCDC\_\{D\}near 25 ms before falling back toward 30 ms\.CLC\_\{L\}andCDC\_\{D\}are generated by the shed vortices, so these readouts respond to how well the coherent structures are preserved rather than to the pointwise accuracy of the field as a whole\. The rise indicates that the reference reproduces those structures markedly less well once the requested lead time leaves the range seen in training, while VATO\-A sustains them over part of the extrapolation window, and the subsequent fall toward 30 ms follows as the prediction of both configurations degrades further\.

### 4\.2Aerodynamic functional consistency

The force MAEs of Table[3](https://arxiv.org/html/2609.00507#S4.T3)pool over geometry and incidence\. Figure[8](https://arxiv.org/html/2609.00507#S4.F8)instead resolves the VFM\-derivedCLC\_\{L\}andCDC\_\{D\}by incidence for four geometries spanning the thickness range of the reported population, with each point a benchmark\-window mean over the same 28 source anchors and direct lead times 1–30 ms\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_lag30_force_polars.png)Figure 8:Fixed\-geometry VATO predictions of VFM\-derived meanCLC\_\{L\}andCDC\_\{D\}against matched targets\. \(a–d\) The thinnest, median, second\-thickest, and thickest geometries of the nine geometry sweep, with\(θ1,θ2\)\(\\theta\_\{1\},\\theta\_\{2\}\)equal to\(178\.32∘,172\.41∘\)\(178\.32^\{\\circ\},172\.41^\{\\circ\}\),\(174\.95∘,158\.19∘\)\(174\.95^\{\\circ\},158\.19^\{\\circ\}\),\(172\.46∘,149\.04∘\)\(172\.46^\{\\circ\},149\.04^\{\\circ\}\), and\(171\.63∘,146\.31∘\)\(171\.63^\{\\circ\},146\.31^\{\\circ\}\)\. Within each geometry and AoA, each point is the mean force coefficient over all direct predictions at lead times 1–30 ms\. Circles are the target values; diamonds and crosses are the mean force coefficients predicted by VATO\-S and VATO\-A\. Open markers denoteCLC\_\{L\}and filled markers denoteCDC\_\{D\}\. The overlay in each panel shows a target\-vorticity snapshot at AoA=12∘=12^\{\\circ\}for the same geometry, rendered on the native mesh with the commonw∈\[−1000,1000\]w\\in\[\-1000,1000\]s\-1range\.The four fixed\-geometry views expose incidence\-dependent behaviour that an across\-geometry coefficient average would hide\. Both VATO variants reproduce the overallCLC\_\{L\}rise and the lower\-amplitudeCDC\_\{D\}trend, including the high\-incidence turning of the thicker geometries\. Agreement remains geometry\-dependent: VATO\-A is generally closer for the thinnest geometry, whereas VATO\-S more closely follows the AoA=12∘=12^\{\\circ\}response of the two thickest geometries\.

### 4\.3Representative and diagnostic field predictions

Figure[9](https://arxiv.org/html/2609.00507#S4.F9)examines the output\-level residual path on a single frame with the geometry, AoA, source time, and lead time fixed in advance\. Panel \(a\) shows the prioritisation rule of Figure[4](https://arxiv.org/html/2609.00507#S3.F4)realised at the trained budget of 256 tokens, with the selected locations lying on the vortices above the section and in the near wake\. In panel \(b\), the RMS magnitude of the output correction is distributed along the upper surface and coincides with the vortices convected downstream from the sharp leading edge, which generate the suction that carries the load\. Bypassing this path on the same checkpoint, with the latent residual and the local attention left active, raises the vorticity relativeL2L\_\{2\}error from 58\.4% to 64\.2%, and the resulting difference in panel \(e\) attains its maximum along the same surface, averaging 368\.1 s\-1within the displayed crop\. These values apply to this intervention on this frame\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/vato_main_correction_diagnostic.png)Figure 9:Output\-level residual cross\-attention correction in VATO\-A for one input from the test population, with the geometry, AoA, source time, and lead time fixed in advance\. The geometry is\(θ1,θ2\)=\(174\.1∘,155\.0∘\)\(\\theta\_\{1\},\\theta\_\{2\}\)=\(174\.1^\{\\circ\},155\.0^\{\\circ\}\)at AoA=12∘=12^\{\\circ\}, with source time 0\.081 s and lead timeτ=10\\tau=10ms\. \(a\) Current\-input vorticitywwwith the realised VFM\-prioritised source locations at the trained budget of 256 tokens; the spatial\-coverage locations are omitted from the display\. \(b\) RMS magnitude of the output\-level residual correction across the normalised\[u,v,p\]\[u,v,p\]channels, linearly stretched between its fluid\-point 5th and 99th percentiles\. \(c\) VATO\-A prediction with the output correction\. \(d\) Prediction from the same checkpoint after bypassing only the selected\-token global query\-readout residual, with the latent residual and the local geometric attention left active\. \(e\) Pointwise absolute vorticity difference between \(c\) and \(d\)\. Vorticity panels use\[−1000,1000\]\[\-1000,1000\]s\-1and the difference panel\[0,500\]\[0,500\]s\-1; clipping affects the display only\.The three models are compared at12∘12^\{\\circ\}, an incidence held out from both training and validation, at fixed geometry \(Figure[10](https://arxiv.org/html/2609.00507#S4.F10)\)\. The frame att=0\.081t=0\.081s serves as the source, and predictions are made 5, 10, 20, 25, and 30 ms ahead of it, with the pressure field in the upper row and the vorticity field in the lower row of each block\. The baseline shown alongside the two VATO configurations is GAOT rather than the matched control, GAOT being the more accurate of the two on every field metric of Table[3](https://arxiv.org/html/2609.00507#S4.T3)\. All three models predict the shed structures at a lead time of 5 ms\. From a lead time of 10 ms onwards the GAOT prediction degrades visibly: adjacent cores on the upper surface are no longer resolved as separate structures and merge, so the predicted vorticity distribution departs from the target in shape while its overall position is retained\. The vortex shed downstream of the trailing edge is still predicted by GAOT at 20 ms, although with a markedly different shape, and is absent at 25 ms, whereas both VATO configurations continue to resolve it at 25 and 30 ms\. The vorticity errors follow the same ordering at every displayed lead time: VATO\-S is below GAOT throughout, and VATO\-A is the most accurate of the three, with an error of 0\.334 at 30 ms against 0\.392 for VATO\-S and 0\.437 for GAOT\. In pressure the ordering is less uniform, VATO\-A being the most accurate at 20 and 30 ms and VATO\-S at the three remaining lead times\.

Figure[11](https://arxiv.org/html/2609.00507#S4.F11)shows the same comparison across the six test incidences at a fixed lead time of 20 ms\. As the flow develops from attached to separated and the prediction becomes more demanding, VATO\-A is more accurate than GAOT in both fields at every incidence, by 29% to 43% in vorticity and by up to 71% in pressure at9∘9^\{\\circ\}\. VATO\-S also resolves the vortex structures that GAOT fails to recover, improving on it in vorticity at every incidence and in pressure once the flow separates\. VATO\-A, the most accurate of the three on the vorticity field, is examined against the target over five randomly selected geometries and four incidences \(Figure[12](https://arxiv.org/html/2609.00507#S4.F12)\), which extends the comparison from single cases to a broader sample of the population\. The twenty cells cover the flow states the family spans: a thin attached shear layer along the surface, a regular vortex street in the wake downstream of the trailing edge, a single large separated region over the upper surface, and discrete cores arranged along the upper surface\. VATO\-A predicts each of them, together with the transitions between them that occur as incidence increases within a geometry and as the fold angles change at fixed incidence\. The relativeL2L\_\{2\}errors range from 0\.151 to 0\.411, the largest values arising where several discrete cores must be positioned individually\.

![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_fixed_geometry_leadtime_pw.png)Figure 10:Predictions of the pressure and vorticity fields by the three models for the same case,\(θ1,θ2\)=\(174\.1∘,155\.0∘\)\(\\theta\_\{1\},\\theta\_\{2\}\)=\(174\.1^\{\\circ\},155\.0^\{\\circ\}\)at AoA=12∘=12^\{\\circ\}, with a source time of0\.0810\.081s\. Rows show lead times of 5, 10, 20, 25, and 30 ms; columns compare the target with the three predictions\. Each cell places the pressure field above the vorticity field, and the numbers on the prediction tiles give the unclipped relativeL2L\_\{2\}error over the valid fluid points\. All panels usex/c∈\[−0\.30,3\.50\]x/c\\in\[\-0\.30,3\.50\]andy/c∈\[−0\.73,0\.77\]y/c\\in\[\-0\.73,0\.77\], with colour rangesp∈\[−80,40\]p\\in\[\-80,40\]Pa andw∈\[−1000,1000\]w\\in\[\-1000,1000\]s\-1\.![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_fixed_geometry_aoa_pw.png)Figure 11:Predictions by the three models for the same geometry,\(θ1,θ2\)=\(176\.6∘,165\.1∘\)\(\\theta\_\{1\},\\theta\_\{2\}\)=\(176\.6^\{\\circ\},165\.1^\{\\circ\}\), at a lead time of 20 ms\. Rows span the six test incidences; columns compare the target with the three predictions\. Each cell places the pressure field above the vorticity field, and the numbers on the prediction tiles give the unclipped relativeL2L\_\{2\}error over the valid fluid points\. All panels sharex/c∈\[−0\.30,3\.50\]x/c\\in\[\-0\.30,3\.50\]andy/c∈\[−0\.73,0\.77\]y/c\\in\[\-0\.73,0\.77\], with colour rangesp∈\[−80,40\]p\\in\[\-80,40\]Pa andω∈\[−1000,1000\]\\omega\\in\[\-1000,1000\]s\-1\.![Refer to caption](https://arxiv.org/html/2609.00507v1/figures/benchmark_geometry_aoa_vorticity_atlas.png)Figure 12:VATO\-A predictions of the vorticity field across geometry and incidence at a lead time of 20 ms\. Rows show five randomly selected test geometries indexed by the interior fold angles\(θ1,θ2\)\(\\theta\_\{1\},\\theta\_\{2\}\); columns show four incidences\. Each cell places the target vorticity above the VATO\-A prediction, and the number on each prediction gives the unclipped relativeL2L\_\{2\}error over the valid fluid points\. All panels sharex/c∈\[−0\.30,3\.50\]x/c\\in\[\-0\.30,3\.50\],y/c∈\[−0\.73,0\.77\]y/c\\in\[\-0\.73,0\.77\], andw∈\[−1000,1000\]w\\in\[\-1000,1000\]s\-1\.

## 5Discussion and Limitations

The training sampler and the vortex\-force coupling act on different quantities\. Relative to the uniformly sampled reference, the matched control raises every field error and lowers every force MAE, so flow\-aware sampling changes the balance between the two families of metric rather than giving a uniformly stronger configuration\. Both VATO configurations improve on that control in velocity, pressure, and vorticity across both lead\-time windows, which attributes the field improvement to the coupling rather than to the sampler\.

Each interface improves the functional it is aligned with\. VATO\-S is trained on the contribution field and gives the lowest VFM\-derived Drag error, at the parameter count and inference cost of the backbone\. VATO\-A is trained on neither functional and improves the predicted field broadly, including pressure, so both readouts follow and it gives the lowest pressure\-derivedCLC\_\{L\}andCDC\_\{D\}errors, at about 63% more measured inference time\. The choice between them follows from which functional is required and how much inference time is available: VATO\-S adds nothing at inference, while VATO\-A spends additional communication and capacity on the predicted field\.

Outside the trained lead\-time range the two families of metric separate again\. Over the 50% temporal extension, the pointwise margins of VATO\-A fall, its vorticity margin is retained at 26\.9%, and all four force margins increase\. Lift and drag are generated by the shed vortices, so the quantities that hold their advantage there are those that depend on the coherent structures rather than on the pointwise accuracy of the field as a whole\. A surrogate placed in a design or control loop is queried at horizons its training set did not fix, and the readouts it is asked to supply are the force functionals\.

Where each interface enters bounds what can be inferred from it\. The contribution field constrains force\-relevant combinations of velocity and vorticity, does not uniquely determine the underlying state, and contains no pressure term, so its improvements do not establish that supervising it alone recovers every force\-bearing flow feature\. In VATO\-A the VFM values determine only which locations are prioritised and are not embedded in learned query, key, or value content, but the reported comparison changes the prioritisation rule, the residual attention paths, and the trainable capacity together\. Its result is an effect of the complete configuration, and isolating the rule would require a geometry\-only prioritisation control matched in path and capacity\.

The reported population contains 54 trajectories from nine geometries represented during training at other incidences, so the evidence covers unseen incidences on known geometries rather than unseen shapes\. Thirty\-six of these trajectories also provided the 1\-ms validation view, and all comparisons use frozen final checkpoints, so no selection was performed on them\. Every configuration is represented by one training seed, and the hierarchical bootstrap quantifies variation over geometry and incidence rather than over retraining\. Training encodes 12,000 sampled source points whereas the benchmark supplies the full native mesh to every configuration\. The force results are matched\-operator diagnostics, with the same operator applied to predicted and target fields, and both operators are inviscid in construction and recover the pressure\-derived component of the load\. Multiple seeds, unseen geometries, cross\-solver tests, calibrated force validation, and autoregressive evaluation would be needed to support stronger claims\.

## 6Conclusion

We introduced VATO, a family of neural operators that couples the Vortex Force Map method to a geometry\-aware transformer backbone at two interfaces\. The coupling draws on a geometry\-only auxiliary potential problem and requires no additional flow solution\. It was evaluated on double\-edged\-plate aerofoils at incidences excluded from training, against a retrained backbone reference and a sampling\-matched control\.

VATO\-S supervises the per\-point contribution field during training, leaving the architecture, the parameter count, and the inference cost of the backbone unchanged\. Relative to the reference it reduces velocity and vorticity error by 10\.4% and 15\.6% over the trained lead\-time range, and it gives the lowest VFM\-derived Drag error of the family\. VATO\-A uses the same information to prioritise 256 source locations for residual cross attention, reduces velocity, pressure, and vorticity error by 15\.8%, 7\.5%, and 31\.2% over the same range, and gives the lowest pressure\-derivedCLC\_\{L\}andCDC\_\{D\}errors, at about 63% more measured inference time\.

Over a horizon 50% longer than the longest lag seen in training, the pointwise margins of VATO\-A fall while its vorticity margin is retained at 26\.9% and all four force margins increase\. The advantage outside the trained range therefore lies in the quantities that carry the load\. The sampling\-matched control separates a second effect along the same line, in which flow\-aware sampling on its own lowers force error and raises pointwise field error\.

The two interfaces cover a range of deployment conditions, from a training\-only intervention that leaves inference unchanged to one that spends additional inference time on the predicted field\. Separating the contribution of the prioritisation rule from that of the additional attention paths, and testing on unseen geometries, are the next steps\.

## Acknowledgements

This research was funded by the Engineering Start\-up Grant from King’s College London and by the Daiwa Anglo\-Japanese Foundation through Daiwa Foundation Awards \(14465/15310\)\. The authors acknowledge the use of King’s Computational Research, Engineering and Technology Environment \(CREATE\) in conducting this research\. The authors appreciate financial support from the China Scholarship Council Program \(202508440025\)

## Declaration of generative AI and AI\-assisted technologies in the writing process

During the preparation of this work the authors used generative AI tools to assist with language editing and manuscript structuring\. The authors reviewed and edited the content as needed and take full responsibility for the content of the publication\.

## References

- \[1\]S\.L\. Brunton, B\.R\. Noack, P\. Koumoutsakos, Machine learning for fluid mechanics, Annu\. Rev\. Fluid Mech\. 52 \(2020\) 477–508\.[https://doi\.org/10\.1146/annurev\-fluid\-010719\-060214](https://doi.org/10.1146/annurev-fluid-010719-060214)\.
- \[2\]R\. Vinuesa, S\.L\. Brunton, Enhancing computational fluid dynamics with machine learning, Nat\. Comput\. Sci\. 2 \(6\) \(2022\) 358–366\.[https://doi\.org/10\.1038/s43588\-022\-00264\-7](https://doi.org/10.1038/s43588-022-00264-7)\.
- \[3\]J\. Kou, W\. Zhang, Data\-driven modeling for unsteady aerodynamics and aeroelasticity, Prog\. Aerosp\. Sci\. 125 \(2021\) 100725\.[https://doi\.org/10\.1016/j\.paerosci\.2021\.100725](https://doi.org/10.1016/j.paerosci.2021.100725)\.
- \[4\]X\. Yao, W\. Liu, W\. Han, G\. Li, Q\. Ma, Development of response surface model of endurance time and structural parameter optimization for a tailsitter UAV, Sensors 20 \(6\) \(2020\) 1766\.[https://doi\.org/10\.3390/s20061766](https://doi.org/10.3390/s20061766)\.
- \[5\]J\. Lou, R\. Chen, J\. Liu, Y\. Bao, Y\. You, Z\. Chen, Aerodynamic optimization of airfoil based on deep reinforcement learning, Phys\. Fluids 35 \(3\) \(2023\) 037128\.[https://doi\.org/10\.1063/5\.0137002](https://doi.org/10.1063/5.0137002)\.
- \[6\]L\. Zhu, W\. Zhang, J\. Kou, Y\. Liu, Machine learning methods for turbulence modeling in subsonic flows around airfoils, Phys\. Fluids 31 \(1\) \(2019\) 015105\.[https://doi\.org/10\.1063/1\.5061693](https://doi.org/10.1063/1.5061693)\.
- \[7\]P\. Zhao, X\. Gao, B\. Zhao, H\. Liu, J\. Wu, Z\. Deng, Machine learning assisted prediction of airfoil lift\-to\-drag characteristics for Mars helicopter, Aerospace 10 \(7\) \(2023\) 614\.[https://doi\.org/10\.3390/aerospace10070614](https://doi.org/10.3390/aerospace10070614)\.
- \[8\]X\. Liu, S\. Yang, H\. Sun, Z\. Wang, X\. Guan, Y\. Gu, Y\. Wang, Review of deep learning\-based aerodynamic shape surrogate models and optimization for airfoils and blade profiles, Phys\. Fluids 37 \(4\) \(2025\) 041304\.[https://doi\.org/10\.1063/5\.0268466](https://doi.org/10.1063/5.0268466)\.
- \[9\]C\. White, D\.M\. Ushizima, C\. Farhat, Fast neural network predictions from constrained aerodynamics datasets, in: AIAA SciTech 2020 Forum, AIAA Paper 2020\-0364, 2020\.[https://doi\.org/10\.2514/6\.2020\-0364](https://doi.org/10.2514/6.2020-0364)\.
- \[10\]Y\. Zhang, W\.J\. Sung, D\.N\. Mavris, Application of convolutional neural network to predict airfoil lift coefficient, in: 2018 AIAA/ASCE/AHS/ASC Structures, Structural Dynamics, and Materials Conference, AIAA Paper 2018\-1903, 2018\.[https://doi\.org/10\.2514/6\.2018\-1903](https://doi.org/10.2514/6.2018-1903)\.
- \[11\]J\. Zhang, X\. Zhao, Machine\-learning\-based surrogate modeling of aerodynamic flow around distributed structures, AIAA J\. 59 \(3\) \(2021\) 868–879\.[https://doi\.org/10\.2514/1\.J059877](https://doi.org/10.2514/1.J059877)\.
- \[12\]H\.E\. Tekaslan, Y\. Demiroglu, M\. Nikbay, Surrogate unsteady aerodynamic modeling with autoencoders and LSTM networks, in: AIAA SciTech 2022 Forum, AIAA Paper 2022\-0508, 2022\.[https://doi\.org/10\.2514/6\.2022\-0508](https://doi.org/10.2514/6.2022-0508)\.
- \[13\]K\. Li, J\. Kou, W\. Zhang, Unsteady aerodynamic reduced\-order modeling based on machine learning across multiple airfoils, Aerosp\. Sci\. Technol\. 119 \(2021\) 107173\.[https://doi\.org/10\.1016/j\.ast\.2021\.107173](https://doi.org/10.1016/j.ast.2021.107173)\.
- \[14\]X\. Ding, C\. Gong, H\. Su, C\. Li, W\. Li, X\. Jia, High\-fidelity modeling of unsteady aerodynamic loads under structural vibration using dual modal spaces and LSTM networks, Aerosp\. Sci\. Technol\. 176 \(Part A\) \(2026\) 111927\.[https://doi\.org/10\.1016/j\.ast\.2026\.111927](https://doi.org/10.1016/j.ast.2026.111927)\.
- \[15\]W\. Dong, X\. Wang, Q\. Lin, C\. Cheng, L\. Zhu, A weighted feature fusion model for unsteady aerodynamic modeling at high angles of attack, Aerospace 11 \(5\) \(2024\) 339\.[https://doi\.org/10\.3390/aerospace11050339](https://doi.org/10.3390/aerospace11050339)\.
- \[16\]J\.H\. Harmening, F\. Pioch, L\. Fuhrig, F\.\-J\. Peitzmann, D\. Schramm, O\. el Moctar, Data\-assisted training of a physics\-informed neural network to predict the separated Reynolds\-averaged turbulent flow field around an airfoil under variable angles of attack, Neural Comput\. Appl\. 36 \(25\) \(2024\) 15353–15371\.[https://doi\.org/10\.1007/s00521\-024\-09883\-9](https://doi.org/10.1007/s00521-024-09883-9)\.
- \[17\]L\. Lu, P\. Jin, G\. Pang, Z\. Zhang, G\.E\. Karniadakis, Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Nat\. Mach\. Intell\. 3 \(3\) \(2021\) 218–229\.[https://doi\.org/10\.1038/s42256\-021\-00302\-5](https://doi.org/10.1038/s42256-021-00302-5)\.
- \[18\]Z\. Li, N\. Kovachki, K\. Azizzadenesheli, B\. Liu, K\. Bhattacharya, A\. Stuart, A\. Anandkumar, Fourier neural operator for parametric partial differential equations, in: International Conference on Learning Representations \(ICLR\), 2021\.[https://openreview\.net/forum?id=c8P9NQVtmnO](https://openreview.net/forum?id=c8P9NQVtmnO)\.
- \[19\]Y\. Dai, Y\. An, Z\. Li, J\. Zhang, C\. Yu, Fourier neural operator with boundary conditions for efficient prediction of steady airfoil flows, Appl\. Math\. Mech\. \(Engl\. Ed\.\) 44 \(11\) \(2023\) 2019–2038\.[https://doi\.org/10\.1007/s10483\-023\-3050\-9](https://doi.org/10.1007/s10483-023-3050-9)\.
- \[20\]D\. Shu, Z\. Li, A\. Barati Farimani, A physics\-informed diffusion model for high\-fidelity flow field reconstruction, J\. Comput\. Phys\. 478 \(2023\) 111972\.[https://doi\.org/10\.1016/j\.jcp\.2023\.111972](https://doi.org/10.1016/j.jcp.2023.111972)\.
- \[21\]Y\. Chen, L\. Huang, W\. Xu, F\. Yang, Efficient high\-fidelity three\-dimensional super\-resolution reconstruction of swirling flame via spaced physically consistent diffusion model, J\. Comput\. Phys\. 549 \(2026\) 114631\.[https://doi\.org/10\.1016/j\.jcp\.2025\.114631](https://doi.org/10.1016/j.jcp.2025.114631)\.
- \[22\]T\. Zhu, B\. Si, L\. Fu, Y\. Lu, SFVnet: Finite\-volume informed U\-net for compressible flow prediction with sparse data under ill\-conditions, J\. Comput\. Phys\. 552 \(2026\) 114696\.[https://doi\.org/10\.1016/j\.jcp\.2026\.114696](https://doi.org/10.1016/j.jcp.2026.114696)\.
- \[23\]S\. Wen, A\. Kumbhat, L\. Lingsch, S\. Mousavi, Y\. Zhao, P\. Chandrashekar, S\. Mishra, Geometry aware operator transformer as an efficient and accurate neural surrogate for PDEs on arbitrary domains, in: Advances in Neural Information Processing Systems, Vol\. 38, 2025\.[https://openreview\.net/forum?id=HXFvNkNt0n](https://openreview.net/forum?id=HXFvNkNt0n)\.
- \[24\]S\. Ōtomo, P\. Gehlert, H\. Babinsky, J\. Li, Vortex force map method to estimate unsteady forces from snapshot flowfield measurements, Exp\. Fluids 66 \(3\) \(2025\) 64\.[https://doi\.org/10\.1007/s00348\-025\-03962\-w](https://doi.org/10.1007/s00348-025-03962-w)\.
- \[25\]J\. Li, Z\.\-N\. Wu, Vortex force map method for viscous flows of general airfoils, J\. Fluid Mech\. 836 \(2018\) 145–166\.[https://doi\.org/10\.1017/jfm\.2017\.783](https://doi.org/10.1017/jfm.2017.783)\.
- \[26\]J\. Li, X\. Zhao, M\. Graham, Vortex force maps for three\-dimensional unsteady flows with application to a delta wing, J\. Fluid Mech\. 900 \(2020\) A36\.[https://doi\.org/10\.1017/jfm\.2020\.515](https://doi.org/10.1017/jfm.2020.515)\.
- \[27\]A\. Aprovitola, L\. Iuspa, G\. Pezzella, A\. Viviani, Optimization procedure for wings flying in the Martian atmosphere, Acta Astronaut\. 235 \(2025\) 339–354\.[https://doi\.org/10\.1016/j\.actaastro\.2025\.05\.023](https://doi.org/10.1016/j.actaastro.2025.05.023)\.
- \[28\]M\. Carreño Ruiz, D\. D’Ambrosio, Implicit and explicit large eddy simulations in Martian aerodynamics, in: AIAA AVIATION Forum and ASCEND 2025, AIAA Paper 2025\-3585, 2025\.[https://doi\.org/10\.2514/6\.2025\-3585](https://doi.org/10.2514/6.2025-3585)\.
- \[29\]G\.R\. Spedding, J\. McArthur, Span efficiencies of wings at low Reynolds numbers, J\. Aircr\. 47 \(1\) \(2010\) 120–128\.[https://doi\.org/10\.2514/1\.44247](https://doi.org/10.2514/1.44247)\.
- \[30\]A\. Rizzi, Separated and vortical flow in aircraft aerodynamics: a CFD perspective, Aeronaut\. J\. 127 \(1313\) \(2023\) 1065–1103\.[https://doi\.org/10\.1017/aer\.2023\.39](https://doi.org/10.1017/aer.2023.39)\.
- \[31\]L\.W\. Traub, C\. Coffman, Efficient low\-Reynolds\-number airfoils, J\. Aircr\. 56 \(5\) \(2019\) 1987–2003\.[https://doi\.org/10\.2514/1\.C035515](https://doi.org/10.2514/1.C035515)\.
- \[32\]W\.J\. Koning, E\.A\. Romander, W\. Johnson, Optimization of low Reynolds number airfoils for Martian rotor applications using an evolutionary algorithm, in: AIAA SciTech 2020 Forum, AIAA Paper 2020\-0084, 2020\.[https://doi\.org/10\.2514/6\.2020\-0084](https://doi.org/10.2514/6.2020-0084)\.
- \[33\]X\. He, Y\. Wang, J\. Li, Flow completion network: Inferring the fluid dynamics from incomplete flow information using graph neural networks,Physics of Fluids34 \(8\) \(2022\) 087114\.[https://doi\.org/10\.1063/5\.0097688](https://doi.org/10.1063/5.0097688)

Similar Articles

A Variational Optimal Transport Operator on Incompressible Flow

arXiv cs.LG

The paper presents VIOT, a variational incompressible optimal transport operator that uses a Fourier Neural Operator to predict divergence-free velocity fields for efficient incompressible density transport, achieving orders of magnitude speedup over traditional optimization methods.

LLT: Local Linear Transformer for PDE Operator Learning

arXiv cs.LG

Introduces LLT, a transformer-based neural operator that combines linear global attention with local spatial mixing for PDE learning. It achieves competitive accuracy and faster training compared to baselines on multiple PDE problems.