Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains
Summary
The paper proposes GeoLAMP, a geometry-aware latent autoregressive generative model for solving multiphysics partial differential equations in complex geometries, using a dual-encoder architecture and causal self-attention transformer with flow matching for stable and scalable predictions.
View Cached Full Text
Cached at: 09/02/26, 06:11 AM
# Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains
Source: [https://arxiv.org/html/2609.00297](https://arxiv.org/html/2609.00297)
Minghui Xu11footnotemark:1Tapan MukerjiAffiliation:Department of Energy Science and EngineeringAffiliation:Stanford UniversityAffiliation:Stanford, CA 94305Affiliation:\{ziwang3, minghuix, mukerji\}@stanford\.edu
###### Abstract
Solving multiphysics partial differential equations \(PDEs\) remains a major challenge in scientific computing, especially for highly complexμ\\mum\-scale tortuous geometries critical to energy and chemical engineering\. We address this challenge by proposing aGeometry\-awareLatentAutoregressive generativeModel forPDEs \(GeoLAMP\) for solving physics within highly irregular and tortuous structures\. GeoLAMP introduces a dual\-encoder architecture on graph representations to jointly capture global topology and fine\-scale geometric features, enabling an effective transition from real\-space fields to compact latent representations\. In the latent space, we propose a causal self\-attention transformer with flow matching to model temporal dynamics, allowing stable and scalable block\-wise autoregressive prediction\. A flexible decoder reconstructs high\-resolution physical fields on arbitrary points\. We establish three multiphysics benchmark datasets in complex geometries, covering reactive flow, heat convection, and elasticity\. GeoLAMP consistently achieves the most stable autoregression performance on these datasets, maintaining low errors throughout the entire rollout horizon\. Our results provide a systematic study of geometry\-aware learning for PDEs inμ\\mum\-scale complex geometries and offer new insights into block\-wise time marching of latent autoregressive PDE modeling via a flow matching framework\.
## 1Introduction
Scientific machine learning for partial differential equations \(PDEs\) has gained substantial attention and has multiple methodological directions, ranging from physics\-informed fitting\[[1](https://arxiv.org/html/2609.00297#bib.bib1),[2](https://arxiv.org/html/2609.00297#bib.bib5),[3](https://arxiv.org/html/2609.00297#bib.bib8)\]to operator learning\[[4](https://arxiv.org/html/2609.00297#bib.bib10),[5](https://arxiv.org/html/2609.00297#bib.bib9),[6](https://arxiv.org/html/2609.00297#bib.bib35),[7](https://arxiv.org/html/2609.00297#bib.bib11)\]and latent modeling\[[8](https://arxiv.org/html/2609.00297#bib.bib31),[9](https://arxiv.org/html/2609.00297#bib.bib32),[10](https://arxiv.org/html/2609.00297#bib.bib40)\]\. It shows strong potential to reshape conventional computational workflows in many scientific and engineering scenarios, including weather forecasting\[[11](https://arxiv.org/html/2609.00297#bib.bib12)\], turbulence reconstruction\[[12](https://arxiv.org/html/2609.00297#bib.bib14)\], and geological engineering\[[13](https://arxiv.org/html/2609.00297#bib.bib15)\]\. However, compared with the widely studied physical problems, multiphysics processes occurring inside complex meso\-scale structures \(μ\\mum\) have received far less attention\[[14](https://arxiv.org/html/2609.00297#bib.bib30)\]\. More importantly, meso\-scale physics often governs macroscopic scientific and engineering behavior, such as natural hydrogen\[[15](https://arxiv.org/html/2609.00297#bib.bib2)\], chip cooling\[[16](https://arxiv.org/html/2609.00297#bib.bib4)\]and battery design\[[17](https://arxiv.org/html/2609.00297#bib.bib3)\]\. This cross\-scale dependence poses significant computational challenges\. Therefore, it is necessary to develop and apply deep learning methods specifically tailored to physical systems defined on meso\-scale complex geometries\[[18](https://arxiv.org/html/2609.00297#bib.bib24),[19](https://arxiv.org/html/2609.00297#bib.bib23)\]\.
Figure 1:Schematic of GeoLAMP architectureTo handle complex geometries, several machine\-learning\-based studies have proposed efficient solutions by incorporating geometric information into model architectures\[[20](https://arxiv.org/html/2609.00297#bib.bib44),[21](https://arxiv.org/html/2609.00297#bib.bib43),[22](https://arxiv.org/html/2609.00297#bib.bib37)\]\. Point\-cloud methods represent complex geometries as unordered sets of spatial points, thereby avoiding the need for structured grids\[[23](https://arxiv.org/html/2609.00297#bib.bib45),[6](https://arxiv.org/html/2609.00297#bib.bib35),[24](https://arxiv.org/html/2609.00297#bib.bib36),[25](https://arxiv.org/html/2609.00297#bib.bib13),[26](https://arxiv.org/html/2609.00297#bib.bib28),[27](https://arxiv.org/html/2609.00297#bib.bib46)\]\. Graph\-based methods further augment these representations with explicit neighborhood connectivity, enabling message passing or graph\-kernel operators to model local physical interactions and long\-range geometry\-dependent transport\[[28](https://arxiv.org/html/2609.00297#bib.bib33),[29](https://arxiv.org/html/2609.00297#bib.bib34),[30](https://arxiv.org/html/2609.00297#bib.bib41),[31](https://arxiv.org/html/2609.00297#bib.bib42),[32](https://arxiv.org/html/2609.00297#bib.bib16)\]\. More recently, GeoPT\[[33](https://arxiv.org/html/2609.00297#bib.bib25)\]introduces lifted geometric pre\-training, which augments large\-scale geometry data with synthetic dynamics to bridge the gap between static shape representation and physics simulation\. However, real\-space geometric complexity increases dramatically with domain size and problem scale, especially for meso\-scale structures represented by millions of voxels\.
To address this difficulty, latent\-space methods\[[34](https://arxiv.org/html/2609.00297#bib.bib18),[35](https://arxiv.org/html/2609.00297#bib.bib19),[8](https://arxiv.org/html/2609.00297#bib.bib31),[10](https://arxiv.org/html/2609.00297#bib.bib40)\]encode real\-space data into compact latent representations and model temporal evolution in latent space, leading to substantially improved efficiency and flexibility through compression\. For complex domains, latent modeling is more appealing if the geometry\-aware backbones can compress irregular point clouds or meshes into regular and compact latent representations while preserving real\-space information\[[27](https://arxiv.org/html/2609.00297#bib.bib46),[28](https://arxiv.org/html/2609.00297#bib.bib33)\]\. Meso\-scale geometries characterized by highly tortuous and heterogeneous structures calls for more reffned geometry\-aware encoder–decoder designs that can compress complex domains while preserving geometry\-sensitive physical information\.
Once the dynamics are represented in a latent space, diffusion\-based generative modeling\[[36](https://arxiv.org/html/2609.00297#bib.bib50),[37](https://arxiv.org/html/2609.00297#bib.bib55),[38](https://arxiv.org/html/2609.00297#bib.bib26),[39](https://arxiv.org/html/2609.00297#bib.bib57),[40](https://arxiv.org/html/2609.00297#bib.bib51),[41](https://arxiv.org/html/2609.00297#bib.bib52),[42](https://arxiv.org/html/2609.00297#bib.bib53),[43](https://arxiv.org/html/2609.00297#bib.bib54)\]provides a principled paradigm for accurate and stable temporal prediction\. At scale, latent tokenization and transformer backbones have further shown that high\-dimensional generation becomes more efficient when the model operates in a compact latent space\[[44](https://arxiv.org/html/2609.00297#bib.bib47),[45](https://arxiv.org/html/2609.00297#bib.bib48),[46](https://arxiv.org/html/2609.00297#bib.bib49),[47](https://arxiv.org/html/2609.00297#bib.bib58),[48](https://arxiv.org/html/2609.00297#bib.bib27)\]\. These ideas are increasingly relevant to scientific surrogate modeling\[[49](https://arxiv.org/html/2609.00297#bib.bib39),[50](https://arxiv.org/html/2609.00297#bib.bib38),[35](https://arxiv.org/html/2609.00297#bib.bib19)\]\.
In this study, we propose aGeometry\-awareLatentAutoregressive generativeModel forPDEs \(GeoLAMP\) to learn time\-dependent multiphysics behavior in complex meso\-scale geometries through latent representations, as shown in Fig\.[1](https://arxiv.org/html/2609.00297#S1.F1)\. Specifically, GeoLAMP consists of three main components:\(i\)\(i\)Global and Local Encoders \(GE & LE\) uses multiple graph neural operator \(GNO\) layers to encode global/local information from graph pairs of query points sampled by the farthest\-point algorithm \(FPS\) and curvature\-based algorithm \(CBS\), separately;\(ii\)\(ii\)Causal Self\-attention Transformer performs block\-wise temporal autoregression by applying masked self\-attention over timestep tokens; and\(iii\)\(iii\)Arbitrary Decoder \(AD\) decodes generated latent representations back to arbitrary locations\. GeoLAMP first encodes real\-space information into a latent space using GE and LE\. The temporal dynamics are then autoregressively generated from the initial latent state\. Finally, AD decodes the generated latent representations back to real space at arbitrary spatial locations and time points\. We constructed three meso\-scale physics datasets\. GeoLAMP achieves the most stable autoregression behavior on them, maintaining low errors throughout the entire rollout horizon\.
#### Contributions
GeoLAMP offers a promising direction for connecting real\-space physics in complex tortuous geometries with latent representations by capturing global structure while preserving local geometric details\. The causal self\-attention transformer behavior suggests that block\-wise rollout goes beyond simple regression of future states and instead captures deeper temporal dependencies in the underlying physical dynamics\. In addition to the above contributions, we establish three meso\-scale physics datasets, including reactive flow, heat convection, and elasticity\.
## 2Problem setup in complex geometry
We focus on solving multiphysics PDEs withKKsolution variables, denoted bysK\(x,t\)s^\{K\}\(x,t\), in a complex domainΩ⊆ℝd\\Omega\\subseteq\\mathbb\{R\}^\{d\}over timet∈\[0,T\]t\\in\[0,T\], subject to boundary conditionsBC\(x,t\)BC\(x,t\)and initial conditionsIC\(x\)IC\(x\), as defined in Eq\.[1](https://arxiv.org/html/2609.00297#S2.E1)\. Herein,ℒa\\mathcal\{L\}\{a\}denotes a physics\-dependent differential operator parameterized by coefficients or source termsaa\. We construct three datasets covering different physical processes and meso\-scale structures: flow\-reactive transport in circular\-packed porous structures, convective heat transfer in stochastic\-field\-generated structures, and elasticity in foam structures\. All datasets are generated using the finite element method \(FEM\) in COMSOL\[[51](https://arxiv.org/html/2609.00297#bib.bib7)\], with detailed descriptions provided in Appendix[A](https://arxiv.org/html/2609.00297#A1)\. Compared with large\-scale benchmark datasets, which typically focus on single\-physics Navier–Stokes equations and turbulent flows, our datasets emphasize the coupled effects of complex geometries and source terms in nonlinear multiphysics transport problems\. Accurately capturing both global physical behavior and local heterogeneous variations poses significant challenges\.
ℒa∘sK\(x,t\)=0,x∈ΩsK\(x,t\)=BC\(x,t\),x∈∂Ω,t∈\[0,T\],sK\(x,0\)=IC\(x\),x∈Ω\\begin\{array\}\[\]\{ll\}\\mathcal\{L\}\_\{a\}\\circ s^\{K\}\(x,t\)=0,&x\\in\\Omega\\\\\[3\.0pt\] s^\{K\}\(x,t\)=BC\(x,t\),&x\\in\\partial\\Omega,\\;t\\in\[0,T\],\\\\\[3\.0pt\] s^\{K\}\(x,0\)=IC\(x\),&x\\in\\Omega\\end\{array\}\(1\)
## 3Geometry\-aware latent generative model
### 3\.1Encoder and decoder
The physical information in real space is organized from FEM mesh outputs as point sets, denoted byxi,siKi=1N\{x\_\{i\},s\_\{i\}^\{K\}\}\{i=1\}^\{N\}at timett\. Since the full point set is typically dense and computationally expensive, directly mapping all real\-space points to latent representations is inefficient\. To address this, we design global and local encoders \(GE & LE\) that compress the full field through tailored point selection while preserving as much geometry\-sensitive physical information as possible\. The arbitrary decoder maps the generated latent representations back to real\-space points at arbitrary resolutions\. All of them are built upon graph neural operators \(GNOs\)\[[28](https://arxiv.org/html/2609.00297#bib.bib33)\], which aggregate information from points within a query ballBBof radiusrbr\_\{b\}and map it onto a regular latent gridyj,gjDj=1M\{y\_\{j\},g\_\{j\}^\{D\}\}\_\{j=1\}^\{M\}at timettthrough Eq\.[2](https://arxiv.org/html/2609.00297#S3.E2), following a Green’s function approximation\.
gj\(y0\)≈∑x∈Bκ\(x,yj0\)s\(x\)μ\(x\)g\_\{j\}\(y^\{0\}\)\\approx\\sum\_\{x\\in B\}\\kappa\(x,y^\{0\}\_\{j\}\)\\,s\(x\)\\,\\mu\(x\)\(2\)Then, followed with multiple GNO blocks, the real space information is compressed to latent representations, where in each layer, the approximation follows
gj\(yl\)≈∑yjl−1∈Blκ\(yl,yjl−1\)s\(yjl−1\)μ\(yjl−1\)g\_\{j\}\(y^\{l\}\)\\approx\\sum\_\{y^\{l\-1\}\_\{j\}\\in B^\{l\}\}\\kappa\(y^\{l\},y^\{l\-1\}\_\{j\}\)\\,s\(y\_\{j\}^\{l\-1\}\)\\,\\mu\(y\_\{j\}^\{l\-1\}\)\(3\)In order to comprehensively capture the global and local information, GE and LE takes different embedded points from real space, and they are trained together as a variational auto\-encoders \(VAE\)\.
#### Global encoder
Given the full input point cloud\{xi\}i=1N\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}, wherexi∈ℝdx\_\{i\}\\in\\mathbb\{R\}^\{d\}, GE automatically selectsngn\_\{g\}representative points using the farthest point sampling \(FPS\)\[[52](https://arxiv.org/html/2609.00297#bib.bib17)\]as the input\. Starting from an initial pointxi1x\_\{i\_\{1\}\}, the sampled set at stepttis denoted asSt=\{xi1,xi2,…,xit\}S\_\{t\}=\\\{x\_\{i\_\{1\}\},x\_\{i\_\{2\}\},\\ldots,x\_\{i\_\{t\}\}\\\}, at each iteration, the next point is selected by maximizing its minimum Euclidean distance to the existing sampled set:
xit\+1=argmaxx∈P∖St\(minxj∈St‖x−xj‖2\)\.x\_\{i\_\{t\+1\}\}=\\arg\\max\_\{x\\in P\\setminus S\_\{t\}\}\\left\(\\min\_\{x\_\{j\}\\in S\_\{t\}\}\\\|x\-x\_\{j\}\\\|\_\{2\}\\right\)\.\(4\)This greedy procedure approximately solves the max–min sampling objective:
maxS⊂P,\|S\|=ngminx∈Pminxj∈S‖x−xj‖2\.\\max\_\{S\\subset P,\\ \|S\|=n\_\{g\}\}\\;\\min\_\{x\\in P\}\\min\_\{x\_\{j\}\\in S\}\\\|x\-x\_\{j\}\\\|\_\{2\}\.\(5\)Therefore, GE intends to cover the full spatial domain and avoids local clustering, providing a robust geometric representation for global information, including the whole trend from inlet to outlet and the boundary nodes\.
#### Local encoder
Taking the remaining points after GE sampling, the local encoder \(LE\) further selectsnln\_\{l\}points that represent the geometric signatures in regions with sharp structural variations\. We introduce a curvature\-based sampling \(CBS\) strategy, which aims to find points located inside local pores and highly curved local regions\. Given a point set\{xi\}i=1N\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}, wherexi∈ℝdx\_\{i\}\\in\\mathbb\{R\}^\{d\}, for each pointxix\_\{i\}, itskk\-nearest neighbors are first constructed as
𝒩k\(xi\)=\{xi1,xi2,…,xik\}\.\\mathcal\{N\}\_\{k\}\(x\_\{i\}\)=\\\{x\_\{i\_\{1\}\},x\_\{i\_\{2\}\},\\ldots,x\_\{i\_\{k\}\}\\\}\.\(6\)
The local geometric variation aroundxix\_\{i\}is characterized by the covariance matrix of neighbor points:
𝐂i=1k∑xj∈𝒩k\(xi\)\(xj−x¯i\)\(xj−x¯i\)⊤,\\mathbf\{C\}\_\{i\}=\\frac\{1\}\{k\}\\sum\_\{x\_\{j\}\\in\\mathcal\{N\}\_\{k\}\(x\_\{i\}\)\}\(x\_\{j\}\-\\bar\{x\}\_\{i\}\)\(x\_\{j\}\-\\bar\{x\}\_\{i\}\)^\{\\top\},\(7\)where the local centroid is defined as
x¯i=1k∑xj∈𝒩k\(xi\)xj\.\\bar\{x\}\_\{i\}=\\frac\{1\}\{k\}\\sum\_\{x\_\{j\}\\in\\mathcal\{N\}\_\{k\}\(x\_\{i\}\)\}x\_\{j\}\.\(8\)
By performing eigenvalue decomposition on𝐂i\\mathbf\{C\}\_\{i\}, we obtainλ1≤λ2≤⋯≤λd\\lambda\_\{1\}\\leq\\lambda\_\{2\}\\leq\\cdots\\leq\\lambda\_\{d\}, and then the local curvature score is defined as
κi=λ1∑m=1dλm,\\kappa\_\{i\}=\\frac\{\\lambda\_\{1\}\}\{\\sum\_\{m=1\}^\{d\}\\lambda\_\{m\}\},\(9\)
In addition to the global information provided by the GE, the local point distributions inside flow channels and curved pores are captured by the LE by selecting a largerλ1\\lambda\_\{1\}\. Finally, thenln\_\{l\}points with the largest curvature scores are selected\. LE allows model to focus on local geometric details that are most informative for preserving fine\-scale behaviors and heterogeneous multiphysics interactions, complementing the global point distribution of GE\.
#### Arbitrary decoder
The arbitrary decoder \(AD\) is designed to decode the generated latent representations back to real\-space points at arbitrary locations and time steps\. To handle meshes with variable numbers of nodes, we introduce a simple masking strategy into the GNO block, where information is propagated only through valid graph nodes while empty nodes are masked out\. Moreover, the full field can be reconstructed at any desired resolution through iterative decoding with AD\.
### 3\.2Flow matching and autoregressive generation
After mapping the original physical\-space states into a compact latent token space, we apply flow matching to learn a transport from a Gaussian distribution to the latent data distribution, enabling autoregressive generation in the latent space\. Specifically, we develop two prediction methods by modifying the scalable interpolant transformer\[[38](https://arxiv.org/html/2609.00297#bib.bib26)\]\(SiT\) backbone, as illustrated in Fig\.[2](https://arxiv.org/html/2609.00297#S3.F2)\. In the first method, the previous two latent states are patchified into tokens and used as conditional context through the cross\-attention mechanism of the transformer block for next\-step prediction\. In addition, optional auxiliary information, such as velocity fields or structural features, can be incorporated into the conditional context, providing the model with richer physics\-relevant guidance\. We refer to this method as GeoLAMP\-S in Fig\.[2](https://arxiv.org/html/2609.00297#S3.F2)\(a\), where “S” denotes single\-step prediction\. In the second method, the previous latent tokens are directly incorporated into a causal self\-attention block, enabling block\-wise prediction in anMM\-to\-MMmanner\. We refer to this method as GeoLAMP\-B, where “B” denotes blockwise prediction\. In this work, we setM=8M=8and perform non\-overlapping blockwise prediction\. Under this setting, all conditional tokens precede the target query tokens in the causal attention sequence, so each query token can attend to all conditional key tokens while the causal mask is preserved\. In Fig\.[2](https://arxiv.org/html/2609.00297#S3.F2)\(b\), the orange entries denote allowed attention, and the black square highlights that a noisy token at stepLLattends to all preceding conditional key tokens and itself\. This autoregressive formulation of spatiotemporal dynamics enables the model to benefit from Transformer scalability, as a single forward pass provides prediction errors for multiple target locations\.
Figure 2:Illustration of latent\-space autoregressive generative paradigmsFor both prediction paradigms, we train the latent generative model using a linear flow\-matching objective\. Let𝐳\\mathbf\{z\}denote the clean target latent tokens and letϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)denote standard Gaussian noise with the same shape as𝐳\\mathbf\{z\}\. Givent∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\), we construct the interpolated latent state
𝐱t=\(1−t\)ϵ\+t𝐳,\\mathbf\{x\}\_\{t\}=\(1\-t\)\\boldsymbol\{\\epsilon\}\+t\\mathbf\{z\},\(10\)which transports samples from the Gaussian distribution att=0t=0to the latent data distribution att=1t=1\. The corresponding target velocity is
𝐮t=d𝐱tdt=𝐳−ϵ\.\\mathbf\{u\}\_\{t\}=\\frac\{d\\mathbf\{x\}\_\{t\}\}\{dt\}=\\mathbf\{z\}\-\\boldsymbol\{\\epsilon\}\.\(11\)
The neural network is trained to predict this velocity field from the noisy latent state, the flow\-matching time, and the conditional context tokens𝐜\\mathbf\{c\}\. The training objective is
ℒFM=𝔼𝐳,ϵ,t\[‖𝐯θ\(𝐱t,t,𝐜\)−𝐮t‖22\]\.\\mathcal\{L\}\_\{\\mathrm\{FM\}\}=\\mathbb\{E\}\_\{\\mathbf\{z\},\\boldsymbol\{\\epsilon\},t\}\\left\[\\left\\\|\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{c\}\)\-\\mathbf\{u\}\_\{t\}\\right\\\|\_\{2\}^\{2\}\\right\]\.\(12\)
For block\-wise prediction, the clean latent variable𝐳\\mathbf\{z\}corresponds only to the target block\. Given a conditioning window\[𝐳s,…,𝐳s\+M−1\]\[\\mathbf\{z\}\_\{s\},\\ldots,\\mathbf\{z\}\_\{s\+M\-1\}\], the model predicts the target block\[𝐳s\+M,…,𝐳s\+2M−1\]\[\\mathbf\{z\}\_\{s\+M\},\\ldots,\\mathbf\{z\}\_\{s\+2M\-1\}\]\. Flow matching is applied only to this target block, while the conditioning block is provided through𝐜\\mathbf\{c\}\. Equivalently, the block\-wise objective is an average over target latent positions:
ℒFMblock=𝔼𝐳,ϵ,t\[1\|Ωtar\|∑i∈Ωtar‖𝐯θ\(𝐱t,t,𝐜\)i−𝐮t,i‖22\],\\mathcal\{L\}\_\{\\mathrm\{FM\}\}^\{\\mathrm\{block\}\}=\\mathbb\{E\}\_\{\\mathbf\{z\},\\boldsymbol\{\\epsilon\},t\}\\left\[\\frac\{1\}\{\|\\Omega\_\{\\mathrm\{tar\}\}\|\}\\sum\_\{i\\in\\Omega\_\{\\mathrm\{tar\}\}\}\\left\\\|\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{c\}\)\_\{i\}\-\\mathbf\{u\}\_\{t,i\}\\right\\\|\_\{2\}^\{2\}\\right\],\(13\)whereΩtar\\Omega\_\{\\mathrm\{tar\}\}denotes the set of target\-block latent positions\. In implementation, this target selection is performed implicitly by constructing𝐱t\\mathbf\{x\}\_\{t\}only over the target block and by extracting velocity predictions only from the noisy target\-token positions\.
During inference, samples are generated by solving the learned probability\-flow ODE
d𝐱tdt=𝐯θ\(𝐱t,t,𝐜\),𝐱0∼𝒩\(𝟎,𝐈\),t∈\[0,1\]\.\\frac\{d\\mathbf\{x\}\_\{t\}\}\{dt\}=\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t,\\mathbf\{c\}\),\\qquad\\mathbf\{x\}\_\{0\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),\\qquad t\\in\[0,1\]\.\(14\)The terminal latent state𝐱1\\mathbf\{x\}\_\{1\}gives the predicted target latent block\. For autoregressive rollout, the predicted latent states are appended to the conditioning window and the procedure is repeated\. The terminal latents are then passed through the decoder to reconstruct the corresponding physical\-space states\. For the reported test rollouts, we use Euler integration with 50 discretization steps\.
## 4Experiment
We first evaluate GeoLAMP on three datasets, where accuracy is measured on the global–local \(G&L\) point sets consistent with the graph pairs used by the GE and LE modules\. We then investigate how the proposed dual\-encoder design affects autoregressive prediction by comparing it with random sampling strategies\. In addition, we evaluate high\-resolution reconstruction by decoding the autoregressively generated latent states back to real space through the arbitrary\-point decoder \(AD\)\. Finally, we analyze the inference scheme and attention patterns of the block\-wise generative module to better understand how temporal dependencies are learned\.
### 4\.1Model accuracy
We conduct autoregressive experiments on three datasets using 4096 global–local \(G&L\) points sampled from the method in Section\.[3\.1](https://arxiv.org/html/2609.00297#S3.SS1)\. The first two datasets are nonlinear flow–transport problems with source terms on solid surfaces\. The third dataset is a linear elasticity problem with nonlinear boundary conditions, where we intend to test whether model can capture the dynamic boundary presence\. We compare GeoLAMP with recent regression\-based methods and evaluate performance using the relativeL2L\_\{2\}error over the full autoregressive rollout as well as at the final prediction step, as reported in Table[1](https://arxiv.org/html/2609.00297#S4.T1)\. GeoLAMP achieves the best last\-step prediction accuracy across all three datasets, demonstrating its effectiveness in long\-horizon forecasting\. Although Transolver\[[24](https://arxiv.org/html/2609.00297#bib.bib36)\]attains a lower average rollout error than GeoLAMP on the elasticity dataset, GeoLAMP exhibits similar average and last\-step errors, indicating more stable long\-term autoregressive performance\.
We additionally include two geometry\-aware operator baselines and one matched deterministic control\. Geo\-FNO\[[53](https://arxiv.org/html/2609.00297#bib.bib59)\]learns a diffeomorphic deformation from the irregular domain to a regular grid and applies an FNO on that grid\. NUNO\[[54](https://arxiv.org/html/2609.00297#bib.bib60)\]decomposes the non\-uniform domain into subdomains, interpolates each onto a local regular grid, and applies an FNO per subdomain\. Both are trained under the same blockwiseMM\-to\-MM\(M=8M=8\) autoregression protocol, matched model size, and identical training and inference settings as GeoLAMP\-B, and are evaluated on the same 4096 G&L points\. Latent\-Det\. is a matched deterministic control that shares the frozen VAE, tokens, causal block mask,M=8M=8protocol and commit stride, matched depth and width, and identical conditioning with GeoLAMP\-B, but replaces the flow\-matching objective and its iterative sampler with direct MSE regression of the same latent block, using one forward pass per block\. Among the geometry\-aware baselines, GeoLAMP\-B attains the lowest mean and last\-step error on all three datasets\. The deterministic control closes much of the gap on the two transport datasets, but degrades sharply on long\-horizon elasticity \(last\-step 1\.9324 vs\. 0\.3823\), indicating that the flow\-matching formulation contributes most where the rollout horizon is most demanding\.
Fig\.[3](https://arxiv.org/html/2609.00297#S4.F3)shows the rollout error histories for three representative methods, including the two prediction paradigms of GeoLAMP and Transolver\[[24](https://arxiv.org/html/2609.00297#bib.bib36)\]\. Although all methods exhibit some degree of error accumulation during rollout, GeoLAMP maintain more controlled error growth, with no significant increase over time in both cases\. Overall, GeoLAMP demonstrates high autoregressive accuracy and strong flexibility in bridging real\-space resolution and latent representations\.
Table 1:RelativeL2L\_\{2\}errors for autoregressive mean and last\-step predictions on three datasets\. Best/second\-best results arebold/underlined\. Suffixes \-S/\-B denote single\-step/blockwise prediction\. Rows above the horizontal rule are end\-to\-end geometry\-aware baselines; rows below share the same frozen dual\-encoder latent pipeline and rollout protocol, so they isolate the generative formulation\.Figure 3:Rollout relativeL2L\_\{2\}error on Heat convection, Reactive flow, and Elasticity\. Curves show the mean over test trajectories, with shaded regions indicating the 10th\-90th percentile range\.
### 4\.2Effects of geometry representation
We further examine the contribution of the proposed GE & LE architecture to autoregressive prediction and evaluate the model’s ability to reconstruct physical fields on super\-resolved point sets\.
#### Designed global & local encoder
GeoLAMP employs a dual\-encoder architecture on graph pairs formulated from FPS and CBS points, enabling effective extraction of global and local information from complex real\-space geometries to latent representations\. To demonstrate the effectiveness of the proposed global & local dual\-encoder design, we compare GeoLAMP with a variant that encodes purely randomly sampled 4096 points \(RS\)\. The autoregressive accuracy on the three datasets is reported in Table[2](https://arxiv.org/html/2609.00297#S4.T2)\. The results show that the proposed GE & LE dual encoder improves accuracy by more than 20% on the two flow–transport datasets while maintaining comparable performance on the elasticity dataset\. This demonstrates that carefully designed point selection is essential for bridging real and latent space, and is crucial for accurate latent\-space prediction\.
#### Super\-resolution by arbitrary decoder
Originating from the flexibility of the GNO\-based decoder, the AD enables the generated latent states to be decoded back to real\-space fields on meshes of arbitrary resolution\. We further evaluate the super\-resolution \(SR\) performance using both randomly sampled \(RS\) points and the proposed global–local \(G&L\) point sets\. The super\-resolved autoregression results from G&L points for three datasets are shown in Fig\.[4](https://arxiv.org/html/2609.00297#S4.F4)\. The accuracy is given in Table[2](https://arxiv.org/html/2609.00297#S4.T2), the G&L point selection not only achieves higher accuracy at its own sampled locations, but also leads to better global super\-resolution performance\. Especially for the elasticity dataset, the SR result from G&L points gains about 33% improvement\. The whole comparison of results with RS and G&L points still highlight a key distinction in scientific PDE learning: spatial representation directly influence solution accuracy\. This effect is further amplified in complex geometries, where physical behavior can vary sharply across subregions due to tortuous pathways and localized geometric features\.
Table 2:RelativeL2L\_\{2\}error comparison across different sampling method and corresponding super resolution\. We report the autoregression mean and the last\-step error\. Best results are inbold\.
### 4\.3Rollout commit step and block length flexibility
#### Commit step influence
The blockwise prediction paradigm provides the flexibility to commit different numbers of predicted steps during inference\. Specifically, after predicting a block of future latent states, the model can commit only the first few steps and then re\-predict the next block, or commit more steps at once before the next autoregressive update\. We investigate this effect on three datasets by varying the number of committed stepskk, as shown in Fig\.[5](https://arxiv.org/html/2609.00297#S4.F5)\. Across all three datasets, committing more steps generally leads to lower rollout error, especially at later rollout timesteps\. This trend suggests that, for the blockwise model, frequent re\-prediction with smallkkcan accumulate errors more rapidly, while largerkkreduces the number of autoregressive updates and therefore improves rollout stability\. The effect is particularly evident for Elasticity, where small committed steps lead to a sharp late\-stage error increase, whereas larger committed steps substantially suppress this error growth\. In addition, increasingkkreduces the number of autoregressive update steps, leading to a linear reduction in time\-marching cost\.
Figure 4:GeoLAMP results \(bottom\) vs\. ground truth \(top\)\. \(a\) Heat convection in stochastic\-field structure\. \(b\) Reactive flow in sphere\-packed structure\. \(c\) Elasticity in foam structure\.Figure 5:RelativeL2L\_\{2\}rollout error for different committed steps during blockwise prediction\.
#### Flexible block length at inference
Beyond varying the commit stride, GeoLAMP\-B can also vary the temporal block itself at inference\. The conditioning lengthCCcan be matched to the available history, while increasing the output lengthPPreduces sequential rollout calls by generating more frames per block\. This follows from the frame\-group representation: each conditioning or noisy target frame forms a group of spatial tokens, and the embedder, self\-attention blocks, and token\-wise output head are shared across frames\. Thus, neitherCCnorPPdetermines any learned tensor shape\. To generatePPframes, we appendPPnoisy frame groups and rebuild the parameter\-free block mask, such that each target group attends to allCCconditioning groups and to itself, but not to the other targets\.
For each dataset, we vary bothCCandPPusing the same frozen pretrained model and evaluate on 200 held\-out trajectories \(Table[3](https://arxiv.org/html/2609.00297#S4.T3)\)\. AtC=8C=8, increasing the output length toP=10P=10or1212does not increase error by more than1%1\\%relative to the nativeC=P=8C=P=8rollout, and sometimes reduces it\. AtP=16P=16, the number of model evaluations and wall\-clock rollout time decrease by4343–50%50\\%and1414–25%25\\%, respectively, while peak batch memory increases by1212–16%16\\%under identical settings on one NVIDIA A100 GPU \(40 GB\)\. This efficiency gain is dataset dependent in accuracy: the normalized error remains near the native setting for heat convection \(1\.011\.01\) and elasticity \(0\.9770\.977\), but rises to1\.421\.42for reactive flow\. At fixedP=8P=8, using only one conditioning frame increases the normalized error to2\.742\.74and2\.042\.04for heat convection and reactive flow, respectively\. The maximum error ratio across datasets decreases to1\.061\.06atC=6C=6and1\.021\.02atC=7C=7, indicating that most of the useful history is captured by six to seven frames\.
Table 3:Inference\-time block\-length flexibility of pretrained GeoLAMP\-B models\. Values are horizon\-averaged physical relativeL2L\_\{2\}errors normalized by the corresponding nativeC=P=8C=P=8rollout \(lower is better;1\.001\.00is native\)\. Left: output length atC=8C=8, with commit stridePP\. Right: conditioning length atP=8P=8\. Each value averages 200 held\-out trajectories\.
### 4\.4Attention variations along the temporal axis
We analyze temporal changes by isolating the attention from noisy target queriesNjN\_\{j\}to conditioning framesCiC\_\{i\}\. For each query, the learned attention scores are visualized as heatmaps in Fig\.[6](https://arxiv.org/html/2609.00297#S4.F6)\(a\)–\(c\), where white regions indicate masked entries\. In addition, the attention weights renormalized overC1C\_\{1\}–C8C\_\{8\}are shown as line plots in Fig\.[6](https://arxiv.org/html/2609.00297#S4.F6)\(d\)–\(f\), measuring the relative preference among conditioning frames\. Heat convection and elasticity exhibit similar recency preferences\. This is likely due to the monotonic spatiotemporal behavior of these systems, where temperature and stress increase continuously across the entire domain throughout the process, as shown in Fig\.[4](https://arxiv.org/html/2609.00297#S4.F4)\. As a result, nearby states remain spatiotemporally consistent and continuous, leading the model to focus more strongly on the most recent frames, which contain more correlated information\.
In contrast, reactive flow displays a different pattern in our setting\. For near\-future prediction, the initial conditioning frames remain nearly as important as the most recent frames\. For later target frames, the attention gradually shifts toward recent inputs, while the contribution from the earliest frames decreases substantially\. This behavior likely originates from the underlying physics\. As shown in Fig\.[4](https://arxiv.org/html/2609.00297#S4.F4), the concentration dynamics form a wormhole pattern, driven by the competition between convective transport and surface reaction consumption\[[55](https://arxiv.org/html/2609.00297#bib.bib29)\]\. During this process, reaction fronts preferentially migrate along paths of lower resistance\. As a result, different local regions evolve asynchronously in time: reaction fronts along different transport regions alternately advance toward the outlet during different time intervals\. The resulting attention weights therefore exhibit curved structures, reflecting the need to capture time\-lagged front dynamics, namely the alternative advancing behavior of the reaction front across different regions\. Together, these patterns suggest that capturing long\-term dependencies, rather than relying solely on the most recent frames, is beneficial for scientific PDE learning\.
Figure 6:Attention strength along the temporal axis for the three dynamics\.
## 5Related Work
Recent advances in diffusion\-based generative modeling\[[36](https://arxiv.org/html/2609.00297#bib.bib50),[41](https://arxiv.org/html/2609.00297#bib.bib52),[37](https://arxiv.org/html/2609.00297#bib.bib55),[56](https://arxiv.org/html/2609.00297#bib.bib56),[38](https://arxiv.org/html/2609.00297#bib.bib26),[39](https://arxiv.org/html/2609.00297#bib.bib57)\]have motivated a new class of PDE surrogates that model time\-dependent dynamics through denoising\-based generation rather than direct deterministic regression, offering long\-horizon stability\. However, applying diffusion generative models directly in real space can be expensive for high\-resolution PDEs\. Latent generative modeling addresses this issue by learning temporal evolution in the lower\-dimensional space\.
DiffusionPDE\[[57](https://arxiv.org/html/2609.00297#bib.bib21)\]models the joint distribution of PDE coefficients and solutions with a diffusion prior, enabling forward and inverse PDE solving\. DPOT\[[50](https://arxiv.org/html/2609.00297#bib.bib38)\]scales denoising\-based operator learning through large\-scale PDE pretraining, showing that generative\-style objectives can improve data efficiency and generalization across diverse physical systems\. M2PDE\[[58](https://arxiv.org/html/2609.00297#bib.bib22)\]uses diffusion models for compositional multiphysics\. Lola\[[59](https://arxiv.org/html/2609.00297#bib.bib20)\]provides an empirical study of latent diffusion models for dynamical systems, showing that latent\-space emulation can remain accurate under high compression rates while diffusion\-based emulators outperform non\-generative baselines\. These methods are attractive for PDE dynamics because they can reduce rollout drift beyond standard next\-step regression\. However, how to design rollout schemes that balance sampling cost and error accumulation remains underexplored\[[50](https://arxiv.org/html/2609.00297#bib.bib38)\]\. In addition, constructing latent representations that preserve geometry\-sensitive physics in complex structures remains challenging\[[59](https://arxiv.org/html/2609.00297#bib.bib20)\]\.
The PDE problems in complex geometries usually present in unstructured data formats rather than pixel\-wise format\. The key challenge is not only to represent irregular domains, but also to compress them into latent variables without losing geometry\-sensitive physical information\. Point\-cloud encoders provide a direct way to process unstructured spatial samples by aggregating features from local neighborhoods and progressively enlarging the receptive field\[[23](https://arxiv.org/html/2609.00297#bib.bib45),[27](https://arxiv.org/html/2609.00297#bib.bib46),[25](https://arxiv.org/html/2609.00297#bib.bib13)\]\. Graph\-based encoders offer a complementary strategy by explicitly using connectivity between spatial samples\. In mesh\- or graph\-based PDE learning, nodes represent physical states at sampled locations, while edges encode local neighborhoods, mesh adjacency, or radius\-based interactions\[[30](https://arxiv.org/html/2609.00297#bib.bib41),[31](https://arxiv.org/html/2609.00297#bib.bib42),[21](https://arxiv.org/html/2609.00297#bib.bib43)\]\. Furthermore, graph neural operators provide a more operator\-oriented encoder–decoder mechanism for mapping between irregular physical fields and latent representations\[[28](https://arxiv.org/html/2609.00297#bib.bib33)\]\.
## 6Conclusions
In this work, we propose GeoLAMP for multiphysics PDE learning in meso\-scale complex 2D geometries\. The specially designed dual encoder enables the model to capture both global physical behavior and local geometric features\. We propose a casual self\-attention transformer in a flow matching framework, and achieve block\-wise autoregressive prediction\. We construct three complex\-geometry datasets, including heat convection, reactive flow, and elasticity\. GeoLAMP achieves higher accuracy in long\-time autoregression than regression\-based models, while demonstrating strong flexibility in geometry representation and resolution reconstruction for tortuous structures\. In addition, block\-wise attention analysis shows that the model captures physically meaningful temporal dependencies beyond simple one\-step correlations\. Several directions remain for future work\. First, although the dual\-encoder architecture improves geometry\-aware latent representation, the model still faces challenges in reconstructing high\-quality physical fields in highly complex geometries\. Second, while block\-wise inference improves inference efficiency, the training cost increases with the block size\. We plan to address this by introducing more efficient attention mechanisms\. Moreover, meso\-scale multiphysics encompasses a broad range of scientific and engineering problems in 3D, which remain to be further explored\. Overall, GeoLAMP demonstrates promising potential for providing fast surrogate models in chemical and energy engineering applications\.
## References
- \[1\]M\. Raissi, P\. Perdikaris, and G\. E\. Karniadakis\(2019\)Physics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational Physics378,pp\. 686–707\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2018.10.045),[Link](https://doi.org/10.1016/j.jcp.2018.10.045)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[2\]G\. E\. Karniadakis, I\. G\. Kevrekidis, L\. Lu, P\. Perdikaris, S\. Wang, and L\. Yang\(2021\)Physics\-informed machine learning\.Nature Reviews Physics3\(6\),pp\. 422–440\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[3\]G\. Pang, L\. Lu, and G\. E\. Karniadakis\(2019\)FPINNs: fractional physics\-informed neural networks\.SIAM Journal on Scientific Computing41\(4\),pp\. A2603–A2626\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[4\]Z\. Li, N\. Kovachki, K\. Azizzadenesheli, B\. Liu, K\. Bhattacharya, A\. Stuart, and A\. Anandkumar\(2021\)Fourier neural operator for parametric partial differential equations\.arXiv preprint arXiv:2010\.08895\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2010.08895),[Link](https://arxiv.org/abs/2010.08895)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[5\]L\. Lu, P\. Jin, G\. Pang, Z\. Zhang, and G\. E\. Karniadakis\(2021\)Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators\.Nature Machine Intelligence3,pp\. 218–229\.External Links:[Document](https://dx.doi.org/10.1038/s42256-021-00302-5),[Link](https://doi.org/10.1038/s42256-021-00302-5)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[6\]Z\. Li, K\. Meidani, and A\. B\. Farimani\(2023\)Transformer for partial differential equations’ operator learning\.arXiv preprint arXiv:2205\.13671\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2205.13671),[Link](https://arxiv.org/abs/2205.13671)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1),[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[Table 1](https://arxiv.org/html/2609.00297#S4.T1.11.3.1.1)\.
- \[7\]K\. Azizzadenesheli, N\. Kovachki, Z\. Li, M\. Liu\-Schiaffini, J\. Kossaifi, and A\. Anandkumar\(2024\)Neural operators for accelerating scientific simulations and design\.Nature Reviews Physics6,pp\. 320–328\.External Links:[Document](https://dx.doi.org/10.1038/s42254-024-00712-5),[Link](https://doi.org/10.1038/s42254-024-00712-5)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[8\]T\. Wu, T\. Maruyama, and J\. Leskovec\(2022\)Learning to accelerate partial differential equations via latent global evolution\.Advances in Neural Information Processing Systems35,pp\. 2240–2253\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1),[§1](https://arxiv.org/html/2609.00297#S1.p3.1)\.
- \[9\]V\. Iakovlev, M\. Heinonen, and H\. Lähdesmäki\(2023\)Learning space\-time continuous latent neural pdes from partially observed states\.Advances in Neural Information Processing Systems36,pp\. 26372–26395\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[10\]K\. Kontolati, S\. Goswami, G\. E\. Karniadakis, and M\. D\. Shields\(2024\)Learning nonlinear operators in latent spaces for real\-time predictions of complex dynamics in physical systems\.Nature Communications15,pp\. 5101\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-49411-w),[Link](https://doi.org/10.1038/s41467-024-49411-w)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1),[§1](https://arxiv.org/html/2609.00297#S1.p3.1)\.
- \[11\]L\. Chen, X\. Zhong, H\. Li, J\. Wu, B\. Lu, D\. Chen, S\. Xie, L\. Wu, Q\. Chao, C\. Lin,et al\.\(2024\)A machine learning model that outperforms conventional global subseasonal forecast models\.Nature Communications15\(1\),pp\. 6425\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[12\]R\. Wang, K\. Kashinath, M\. Mustafa, A\. Albert, and R\. Yu\(2020\)Towards physics\-informed deep learning for turbulent flow prediction\.InProceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining,pp\. 1457–1466\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[13\]G\. Wen, Z\. Li, K\. Azizzadenesheli, A\. Anandkumar, and S\. M\. Benson\(2022\)U\-fno—an enhanced fourier neural operator\-based deep\-learning model for multiphase flow\.Advances in Water Resources163,pp\. 104180\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[14\]L\. Chen, A\. He, J\. Zhao, Q\. Kang, Z\. Li, J\. Carmeliet, N\. Shikazono, and W\. Tao\(2022\)Pore\-scale modeling of complex transport phenomena in porous media\.Progress in Energy and Combustion Science88,pp\. 100968\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[15\]Y\. Mathur and T\. Mukerji\(2025\)Changes in physical properties of rocks during serpentinization and implications for natural hydrogen exploration\.The Leading Edge44\(6\),pp\. 479–487\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[16\]R\. Van Erp, R\. Soleimanzadeh, L\. Nela, G\. Kampitsis, and E\. Matioli\(2020\)Co\-designing electronics with microfluidics for more sustainable cooling\.Nature585\(7824\),pp\. 211–216\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[17\]X\. Lu, A\. Bertei, D\. P\. Finegan, C\. Tan, S\. R\. Daemi, J\. S\. Weaving, K\. B\. O’Regan, T\. M\. Heenan, G\. Hinds, E\. Kendrick,et al\.\(2020\)3D microstructure design of lithium\-ion battery electrodes assisted by x\-ray nano\-computed tomography and modelling\.Nature communications11\(1\),pp\. 2079\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[18\]J\. Li, W\. Ge, W\. Wang, N\. Yang, X\. Liu, L\. Wang, X\. He, X\. Wang, J\. Wang, and M\. Kwauk\(2013\)From multiscale modeling to meso\-science\.Springer\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[19\]K\. Agrawal, P\. N\. Loezos, M\. Syamlal, and S\. Sundaresan\(2001\)The role of meso\-scale structures in rapid gas–solid flows\.Journal of Fluid Mechanics445,pp\. 151–185\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p1.1)\.
- \[20\]M\. M\. Bronstein, J\. Bruna, T\. Cohen, and P\. Veličković\(2021\)Geometric deep learning: grids, groups, graphs, geodesics, and gauges\.arXiv preprint arXiv:2104\.13478\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2104.13478),[Link](https://arxiv.org/abs/2104.13478)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1)\.
- \[21\]J\. Brandstetter, D\. E\. Worrall, and M\. Welling\(2022\)Message passing neural pde solvers\.arXiv preprint arXiv:2202\.03376\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2202.03376),[Link](https://arxiv.org/abs/2202.03376)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[22\]M\. Takamoto, T\. Praditia, R\. Leiteritz, D\. MacKinlay, F\. Alesiani, D\. Pflüger, and M\. Niepert\(2022\)PDEBench: an extensive benchmark for scientific machine learning\.arXiv preprint arXiv:2210\.07182\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2210.07182),[Link](https://arxiv.org/abs/2210.07182)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1)\.
- \[23\]C\. R\. Qi, L\. Yi, H\. Su, and L\. J\. Guibas\(2017\)PointNet\+\+: deep hierarchical feature learning on point sets in a metric space\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Link](https://papers.nips.cc/paper/7095-pointnet-deep-hierarchical-feature-learning-on-point-sets-in-a-metric-space)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[24\]H\. Wu, H\. Luo, H\. Wang, J\. Wang, and M\. Long\(2024\)Transolver: a fast transformer solver for pdes on general geometries\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 53681–53705\.External Links:[Link](https://proceedings.mlr.press/v235/wu24r.html),[Document](https://dx.doi.org/10.48550/arXiv.2402.02366)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.00297#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.00297#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.00297#S4.T1.11.5.1.1)\.
- \[25\]A\. Kashefi and T\. Mukerji\(2021\)Point\-cloud deep learning of porous media for permeability prediction\.Physics of Fluids33\(9\)\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[26\]J\. Liang and H\. Zhao\(2013\)Solving partial differential equations on point clouds\.SIAM Journal on Scientific Computing35\(3\),pp\. A1461–A1486\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1)\.
- \[27\]Y\. Wang, Y\. Sun, Z\. Liu, S\. E\. Sarma, M\. M\. Bronstein, and J\. M\. Solomon\(2019\)Dynamic graph cnn for learning on point clouds\.ACM Transactions on Graphics38\(5\),pp\. 146:1–146:12\.External Links:[Document](https://dx.doi.org/10.1145/3326362),[Link](https://doi.org/10.1145/3326362)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§1](https://arxiv.org/html/2609.00297#S1.p3.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[28\]Z\. Li, N\. Kovachki, K\. Azizzadenesheli, B\. Liu, K\. Bhattacharya, A\. Stuart, and A\. Anandkumar\(2020\)Neural operator: graph kernel network for partial differential equations\.arXiv preprint arXiv:2003\.03485\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2003.03485),[Link](https://arxiv.org/abs/2003.03485)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§1](https://arxiv.org/html/2609.00297#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.00297#S3.SS1.p1.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[29\]Z\. Li, N\. B\. Kovachki, C\. Choy, B\. Li, J\. Kossaifi, S\. P\. Otta, M\. A\. Nabian, M\. Stadler, C\. Hundt, K\. Azizzadenesheli, and A\. Anandkumar\(2023\)Geometry\-informed neural operator for large\-scale 3d pdes\.arXiv preprint arXiv:2309\.00583\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2309.00583),[Link](https://arxiv.org/abs/2309.00583)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[Table 1](https://arxiv.org/html/2609.00297#S4.T1.11.4.1.1)\.
- \[30\]T\. Pfaff, M\. Fortunato, A\. Sanchez\-Gonzalez, and P\. W\. Battaglia\(2021\)Learning mesh\-based simulation with graph networks\.arXiv preprint arXiv:2010\.03409\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2010.03409),[Link](https://arxiv.org/abs/2010.03409)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[31\]A\. Sanchez\-Gonzalez, J\. Godwin, T\. Pfaff, R\. Ying, J\. Leskovec, and P\. W\. Battaglia\(2020\)Learning to simulate complex physics with graph networks\.arXiv preprint arXiv:2002\.09405\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2002.09405),[Link](https://arxiv.org/abs/2002.09405)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1),[§5](https://arxiv.org/html/2609.00297#S5.p3.1)\.
- \[32\]J\. Chung, R\. Ahmad, W\. Sun, W\. Cai, and T\. Mukerji\(2024\)Prediction of effective elastic moduli of rocks using graph neural networks\.Computer Methods in Applied Mechanics and Engineering421,pp\. 116780\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1)\.
- \[33\]H\. Wu, M\. Guo, Z\. Li, Z\. Dou, M\. Long, K\. He, and W\. Matusik\(2026\)GeoPT: scaling physics simulation via lifted geometric pre\-training\.arXiv preprint arXiv:2602\.20399\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p2.1)\.
- \[34\]T\. Gao, H\. Zheng, W\. Deng, H\. Feng, T\. Zhang, R\. Feng, Q\. Chen, and T\. Wu\(2026\)GenCP: towards generative modeling paradigm of coupled physics\.arXiv preprint arXiv:2601\.19541\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p3.1)\.
- \[35\]Z\. Li, S\. Patil, F\. Ogoke, D\. Shu, W\. Zhen, M\. Schneier, J\. R\. Buchanan Jr, and A\. B\. Farimani\(2025\)Latent neural pde solver: a reduced\-order modeling framework for partial differential equations\.Journal of Computational Physics524,pp\. 113705\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p3.1),[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[36\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.arXiv preprint arXiv:2006\.11239\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2006.11239),[Link](https://arxiv.org/abs/2006.11239)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1),[§5](https://arxiv.org/html/2609.00297#S5.p1.1)\.
- \[37\]Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le\(2023\)Flow matching for generative modeling\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PqvMRDCJT9t),[Document](https://dx.doi.org/10.48550/arXiv.2210.02747)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1),[§5](https://arxiv.org/html/2609.00297#S5.p1.1)\.
- \[38\]W\. Peebles and S\. Xie\(2023\)Scalable diffusion models with transformers\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 4195–4205\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1),[§3\.2](https://arxiv.org/html/2609.00297#S3.SS2.p1.1),[§5](https://arxiv.org/html/2609.00297#S5.p1.1)\.
- \[39\]X\. Liu, C\. Gong, and Q\. Liu\(2023\)Flow straight and fast: learning to generate and transfer data with rectified flow\.arXiv preprint arXiv:2209\.03003\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2209.03003),[Link](https://arxiv.org/abs/2209.03003)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1),[§5](https://arxiv.org/html/2609.00297#S5.p1.1)\.
- \[40\]A\. Q\. Nichol and P\. Dhariwal\(2021\)Improved denoising diffusion probabilistic models\.arXiv preprint arXiv:2102\.09672\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2102.09672),[Link](https://arxiv.org/abs/2102.09672)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[41\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.arXiv preprint arXiv:2011\.13456\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2011.13456),[Link](https://arxiv.org/abs/2011.13456)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1),[§5](https://arxiv.org/html/2609.00297#S5.p1.1)\.
- \[42\]J\. Song, C\. Meng, and S\. Ermon\(2021\)Denoising diffusion implicit models\.arXiv preprint arXiv:2010\.02502\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2010.02502),[Link](https://arxiv.org/abs/2010.02502)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[43\]T\. Karras, M\. Aittala, T\. Aila, and S\. Laine\(2022\)Elucidating the design space of diffusion\-based generative models\.arXiv preprint arXiv:2206\.00364\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2206.00364),[Link](https://arxiv.org/abs/2206.00364)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[44\]A\. van den Oord, O\. Vinyals, and K\. Kavukcuoglu\(2017\)Neural discrete representation learning\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Link](https://papers.nips.cc/paper/7210-neural-discrete-representation-learning)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[45\]P\. Esser, R\. Rombach, and B\. Ommer\(2021\)Taming transformers for high\-resolution image synthesis\.arXiv preprint arXiv:2012\.09841\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2012.09841),[Link](https://arxiv.org/abs/2012.09841)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[46\]R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer\(2022\)High\-resolution image synthesis with latent diffusion models\.arXiv preprint arXiv:2112\.10752\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2112.10752),[Link](https://arxiv.org/abs/2112.10752)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[47\]W\. Peebles and S\. Xie\(2023\)Scalable diffusion models with transformers\.arXiv preprint arXiv:2212\.09748\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2212.09748),[Link](https://arxiv.org/abs/2212.09748)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[48\]N\. Ma, M\. Goldstein, M\. S\. Albergo, N\. M\. Boffi, E\. Vanden\-Eijnden, and S\. Xie\(2024\)Sit: exploring flow and diffusion\-based generative models with scalable interpolant transformers\.InEuropean Conference on Computer Vision,pp\. 23–40\.Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[49\]P\. Lippe, B\. S\. Veeling, P\. Perdikaris, R\. E\. Turner, and J\. Brandstetter\(2023\)PDE\-refiner: achieving accurate long rollouts with neural pde solvers\.arXiv preprint arXiv:2308\.05732\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2308.05732),[Link](https://arxiv.org/abs/2308.05732)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1)\.
- \[50\]Z\. Hao, C\. Su, S\. Liu, J\. Berner, C\. Ying, H\. Su, A\. Anandkumar, J\. Song, and J\. Zhu\(2024\)DPOT: auto\-regressive denoising operator transformer for large\-scale pde pre\-training\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 17616–17635\.External Links:[Link](https://proceedings.mlr.press/v235/hao24d.html),[Document](https://dx.doi.org/10.48550/arXiv.2403.03542)Cited by:[§1](https://arxiv.org/html/2609.00297#S1.p4.1),[§5](https://arxiv.org/html/2609.00297#S5.p2.1)\.
- \[51\]C\. Multiphysics\(1998\)Introduction to comsol multiphysics®\.COMSOL Multiphysics, Burlington, MA, accessed Feb9\(2018\),pp\. 32\.Cited by:[Appendix A](https://arxiv.org/html/2609.00297#A1.p1.1),[§2](https://arxiv.org/html/2609.00297#S2.p1.1)\.
- \[52\]C\. Moenning and N\. A\. Dodgson\(2003\)Fast marching farthest point sampling\.Technical reportUniversity of Cambridge, Computer Laboratory\.Cited by:[§3\.1](https://arxiv.org/html/2609.00297#S3.SS1.SSS0.Px1.p1.1)\.
- \[53\]Z\. Li, D\. Z\. Huang, B\. Liu, and A\. Anandkumar\(2023\)Fourier neural operator with learned deformations for pdes on general geometries\.Journal of Machine Learning Research24\(388\),pp\. 1–26\.External Links:[Link](https://jmlr.org/papers/v24/23-0064.html),[Document](https://dx.doi.org/10.48550/arXiv.2207.05209)Cited by:[§4\.1](https://arxiv.org/html/2609.00297#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2609.00297#S4.T1.11.6.1.1)\.
- \[54\]S\. Liu, Z\. Hao, C\. Ying, H\. Su, Z\. Cheng, and J\. Zhu\(2023\)Nuno: a general framework for learning parametric pdes with non\-uniform data\.InInternational Conference on Machine Learning,pp\. 21658–21671\.Cited by:[§4\.1](https://arxiv.org/html/2609.00297#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2609.00297#S4.T1.11.7.1.1)\.
- \[55\]Z\. Wang, L\. Chen, H\. Wei, Z\. Dai, Q\. Kang, and W\. Tao\(2022\)Pore\-scale study of mineral dissolution in heterogeneous structures and deep learning prediction of permeability\.Physics of Fluids34\(11\)\.Cited by:[Appendix A](https://arxiv.org/html/2609.00297#A1.SS0.SSS0.Px2.p1.1),[§4\.4](https://arxiv.org/html/2609.00297#S4.SS4.p2.1)\.
- \[56\]M\. S\. Albergo, N\. M\. Boffi, and E\. Vanden\-Eijnden\(2023\)Stochastic interpolants: a unifying framework for flows and diffusions\.arXiv preprint arXiv:2303\.08797\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2303.08797),[Link](https://arxiv.org/abs/2303.08797)Cited by:[§5](https://arxiv.org/html/2609.00297#S5.p1.1)\.
- \[57\]J\. Huang, G\. Yang, Z\. Wang, and J\. J\. Park\(2024\)DiffusionPDE: generative pde\-solving under partial observation\.Advances in Neural Information Processing Systems37,pp\. 130291–130323\.Cited by:[§5](https://arxiv.org/html/2609.00297#S5.p2.1)\.
- \[58\]T\. Zhang, Z\. Liu, F\. Qi, Y\. Jiao, and T\. Wu\(2024\)M2PDE: compositional generative multiphysics and multi\-component pde simulation\.arXiv preprint arXiv:2412\.04134\.Cited by:[§5](https://arxiv.org/html/2609.00297#S5.p2.1)\.
- \[59\]F\. Rozet, R\. Ohana, M\. McCabe, G\. Louppe, F\. Lanusse, and S\. Ho\(2025\)Lost in latent space: an empirical study of latent diffusion models for physics emulation\.arXiv preprint arXiv:2507\.02608\.Cited by:[§5](https://arxiv.org/html/2609.00297#S5.p2.1)\.
- \[60\]J\. Ohser and K\. Schladitz\(2009\)3D images of materials structures: processing and analysis\.John Wiley & Sons\.Cited by:[Appendix A](https://arxiv.org/html/2609.00297#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2609.00297#A1.SS0.SSS0.Px3.p1.1)\.
## Appendix ADataset Generation
Overall, we reconstruct three distinct meso\-scale geometries for three different physical problems, as shown in Fig\.[7](https://arxiv.org/html/2609.00297#A1.F7)\. For each dataset, 2,000 structures are generated, and the corresponding physical PDEs are solved to obtain 2,000 spatio\-temporal solution sequences\. Each dataset is then randomly split into training, validation, and test sets with 1,500, 300, and 200 sequences, respectively\. All simulation is implemented in COMSOL\[[51](https://arxiv.org/html/2609.00297#bib.bib7)\]\. We show one sequences of each dataset in Fig\.[8](https://arxiv.org/html/2609.00297#A1.F8)\.
Figure 7:Meso\-scale structure reconstruction#### Heat convection in stochastic\-field structure\.
The geometry is generated using a stochastic field method implemented in GeoDict\[[60](https://arxiv.org/html/2609.00297#bib.bib6)\], as shown in Fig\.[7](https://arxiv.org/html/2609.00297#A1.F7)\(a\), yielding a highly heterogeneous porous structure within a24,000×24,000,μm224\{,\}000\\times 24\{,\}000,\\mu\\mathrm\{m\}^\{2\}domain\. The flow is governed by the incompressible Navier–Stokes equations coupled with an energy transport equation, where heat exchange occurs at the solid–fluid interface\. Water is used as the working fluid\. Periodic boundary conditions are imposed along the flow direction with a pressure drop of3,Pa3,\\mathrm\{Pa\}, allowing the fluid to exit through the outlet and re\-enter from the inlet while continuously exchanging heat with the solid phase under a prescribed heat flux of1000,W/m21000,\\mathrm\{W\}/\\mathrm\{m\}^\{2\}\. This setup represents pore\-scale heat\-transfer processes commonly encountered in battery systems and low\-heat\-flux thermal management applications\.
ρ\(∂𝐮∂t\+𝐮⋅∇𝐮\)=−∇p\+μ∇2𝐮,x∈Ωf∇⋅𝐮=0,x∈Ωf∂T∂t\+𝐮⋅∇T=α∇2T,x∈Ωf−kT∇T⋅𝐧=q,x∈∂Ωs\\begin\{array\}\[\]\{ll\}\\rho\\left\(\\frac\{\\partial\\mathbf\{u\}\}\{\\partial t\}\+\\mathbf\{u\}\\cdot\\nabla\\mathbf\{u\}\\right\)=\-\\nabla p\+\\mu\\nabla^\{2\}\\mathbf\{u\},&x\\in\\Omega\_\{f\}\\\\\[3\.0pt\] \\nabla\\cdot\\mathbf\{u\}=0,&x\\in\\Omega\_\{f\}\\\\\[3\.0pt\] \\frac\{\\partial T\}\{\\partial t\}\+\\mathbf\{u\}\\cdot\\nabla T=\\alpha\\nabla^\{2\}T,&x\\in\\Omega\_\{f\}\\\\\[3\.0pt\] \-k\_\{T\}\\nabla T\\cdot\\mathbf\{n\}=q,&x\\in\\partial\\Omega\_\{s\}\\end\{array\}\(15\)
Figure 8:Meso\-scale physics datasets
#### Reactive flow in sphere\-packed structure\.
The complex geometry is constructed using a stochastic Monte Carlo sphere\-packing process\[[55](https://arxiv.org/html/2609.00297#bib.bib29)\], as shown in Fig\.[7](https://arxiv.org/html/2609.00297#A1.F7)\(b\)\. Specifically, multiple spheres with radius14,μm14,\\mu\\mathrm\{m\}are randomly initialized and iteratively adjusted within a240×240,μm2240\\times 240,\\mu\\mathrm\{m\}^\{2\}computational domain\. In this domain, the incompressible Navier–Stokes equations are coupled with a solute transport equation, where reactants are advected from the left inlet to the right outlet and consumed at the solid surfaces through a surface reaction term\. The inlet velocity, inlet concentration, and surface reaction rate are set to7\.2×10−4,m/s7\.2\\times 10^\{\-4\},\\mathrm\{m\}/\\mathrm\{s\},0\.1,mol/m30\.1,\\mathrm\{mol\}/\\mathrm\{m\}^\{3\}, and104,m/s10^\{4\},\\mathrm\{m\}/\\mathrm\{s\}, respectively\. This setup captures key mechanisms in subsurface reactive transport and is representative of applications such as CO2sequestration and natural hydrogen generation\.
ρ\(∂𝐮∂t\+𝐮⋅∇𝐮\)=−∇p\+μ∇2𝐮,x∈Ωf∇⋅𝐮=0,x∈Ωf∂c∂t\+𝐮⋅∇c=D∇2c,x∈Ωf−D∇c⋅𝐧=kc,x∈∂Ωs\\begin\{array\}\[\]\{ll\}\\rho\\left\(\\frac\{\\partial\\mathbf\{u\}\}\{\\partial t\}\+\\mathbf\{u\}\\cdot\\nabla\\mathbf\{u\}\\right\)=\-\\nabla p\+\\mu\\nabla^\{2\}\\mathbf\{u\},&x\\in\\Omega\_\{f\}\\\\\[3\.0pt\] \\nabla\\cdot\\mathbf\{u\}=0,&x\\in\\Omega\_\{f\}\\\\\[3\.0pt\] \\frac\{\\partial c\}\{\\partial t\}\+\\mathbf\{u\}\\cdot\\nabla c=D\\nabla^\{2\}c,&x\\in\\Omega\_\{f\}\\\\\[3\.0pt\] \-D\\nabla c\\cdot\\mathbf\{n\}=kc,&x\\in\\partial\\Omega\_\{s\}\\end\{array\}\(16\)
#### Elasticity in foam structure\.
The porous geometry is generated using a foam reconstruction method implemented in GeoDict\[[60](https://arxiv.org/html/2609.00297#bib.bib6)\], as shown in Fig\.[7](https://arxiv.org/html/2609.00297#A1.F7)\(c\), within a2,400×2,400,μm22\{,\}400\\times 2\{,\}400,\\mu\\mathrm\{m\}^\{2\}domain\. The mechanical response is modeled by infinitesimal linear elasticity under dynamic loading conditions\. Copper is used as the solid material, reflecting its common use in metallic foams\. Although the governing equations are quasi\-static, time\-dependent boundary displacements are applied to mimic realistic loading scenarios\. This dataset evaluates the model’s ability to capture deformation responses under temporally varying boundary conditions in complex porous media\.
∇⋅σ=0,x∈Ωσ=λ\(∇⋅u\)I\+2με,x∈Ωε=12\(∇u\+\(∇u\)⊤\),x∈Ωu\(x,t\)=−0\.001H⋅1−cos\(2πt\)2,x∈Γtop\\begin\{array\}\[\]\{ll\}\\nabla\\cdot\\sigma=0,&x\\in\\Omega\\\\\[3\.0pt\] \\sigma=\\lambda\(\\nabla\\cdot u\)I\+2\\mu\\varepsilon,&x\\in\\Omega\\\\\[3\.0pt\] \\varepsilon=\\frac\{1\}\{2\}\(\\nabla u\+\(\\nabla u\)^\{\\top\}\),&x\\in\\Omega\\\\\[3\.0pt\] u\(x,t\)=\-0\.001H\\cdot\\frac\{1\-\\cos\(2\\pi t\)\}\{2\},&x\\in\\Gamma\_\{top\}\\end\{array\}\(17\)
## Appendix BModel hyperparameter and implementation details
### B\.1Model hyperparameter
#### Global & Local VAE\.
The autoencoder compresses each per\-timestep scalar field on the unstructured mesh into a64×8×864\\times 8\\times 8latent\. Two independent GNO encoders use radius\-graph integral transforms with radius0\.150\.15and MLP layers\[80,80,80\]\[80,80,80\], projecting the input points onto a shared32×3232\\times 32grid with 128 channels\. The two encoded grids are summed element\-wise\. The CNN encoder uses a base hidden width of 128, channel multipliers\(1,2,4\)\(1,2,4\), two ResNet blocks per resolution level, and self\-attention at resolutions16×1616\\times 16and8×88\\times 8, producing an8×8×648\\times 8\\times 64bottleneck with mean and log\-variance\. The KL weight is set to10−610^\{\-6\}\. The decoder mirrors the CNN encoder and reconstructs a32×32×12832\\times 32\\times 128grid, followed by a GNO decoder with MLP layers\[512,256\]\[512,256\]and radius0\.150\.15\.
#### Super\-resolution Decoder\.
The super\-resolution decoder shares the CNN\-decoder and GNO\-decoder backend with the VAE\. It takes a predicted8×8×648\\times 8\\times 64latent and upsamples it through three resolution levels with base hidden width 128, channel multipliers\(1,2,4\)\(1,2,4\), two ResNet blocks per level, and self\-attention at resolutions8×88\\times 8and16×1616\\times 16\. The CNN decoder outputs a32×32×12832\\times 32\\times 128feature grid\. A GNO decoder with MLP layers\[512,256\]\[512,256\]and radius0\.150\.15maps the grid features to the full mesh with approximately1212–21×10321\\times 10^\{3\}valid vertices and produces a one\-channel field prediction\.
#### GeoLAMP\-S\.
GeoLAMP\-S denoises a64×8×864\\times 8\\times 8target latent under conditioning\. With𝚙𝚊𝚝𝚌𝚑\_𝚜𝚒𝚣𝚎=1\\mathtt\{patch\\\_size\}=1, each latent is patch\-embedded into 64 tokens with hidden width 768 and sinusoidal 2\-D positional embeddings\. The flow\-matching time and autoregressive step index are embedded using sinusoidal\-MLP timestep embedders and combined as the AdaLN\-Zero conditioning vector\. The transformer uses 12 blocks, hidden width 768, 12 attention heads, andmlp\_ratio=4\\mathrm\{mlp\\\_ratio\}=4\. Each block contains AdaLN\-modulated self\-attention, cross\-attention to context tokens, and an AdaLN\-modulated GeGLU feed\-forward layer\. The final AdaLN\-modulated linear layer maps the tokens back to 64\-channel patches\.
#### GeoLAMP\-B\.
GeoLAMP\-B uses the same backbone and hyperparameter settings as GeoLAMP\-S, but performs non\-overlapping block\-wise prediction\. It predictsM=8M=8future latent velocity fields in one flow\-matching evaluation\. Conditional and noisy target latents are patch\-embedded into a causal token sequence, and the noisy target token groups are unpatchified and concatenated to produce the eight predicted velocity fields\.
#### GINO\.
GINO uses the same encoder\-decoder structure as the Global & Local VAE \(changed to AE\), with the variational bottleneck replaced by a deterministic autoencoder, and are trained together\. The FNO operates on the8×88\\times 8latent grid\. A1×11\\times 1lifting convolution maps the input to a 752\-channel hidden representation, using 192 input channels for reaction flow and heat convection and 128 input channels for elasticity\. The trunk contains 10 spectral blocks with\(4,4\)\(4,4\)Fourier modes, complex\-valued mode\-wise weights,1×11\\times 1pointwise convolutional residual bypasses, GELU activations, and outer residual connections\. A final1×11\\times 1projection convolution maps the hidden tensor to 64 output channels\.
#### OFormer\.
OFormer operates directly on 4096 mesh nodes sampled by uniform and curvature\-biased strategies\. The encoder uses aConv2d\(5→128\)\\mathrm\{Conv2d\}\(5\\to 128\)stem with kernel size\(2,1\)\(2,1\), learned node\-type embeddings, and relative positional encodings\. The encoder contains six Galerkin linear\-attention blocks with hidden dimension 128, head dimension 128, one attention head, and feed\-forward hidden dimension 512\. The decoder uses Gaussian Fourier coordinate features, a two\-layer MLP with hidden dimension 128, one Galerkin cross\-attention block, one linear\-attention block, and a three\-layer MLP prediction head\.
#### Transolver\.
Transolver operates directly on 4096 real\-space query points from the concatenated uniform and curvature point clouds\. The input feature dimension is 6 for reaction flow and heat convection and 4 for elasticity\. A preprocessing MLP maps the input to width 256, followed by a learned placeholder vector and a sinusoidal physical\-time embedding processed by a two\-layer SiLU MLP\. The trunk contains 6 Transolver blocks with hidden size 256, 8 heads, 32 slice tokens per head, dropout 0, andmlp\_ratio=2\\mathrm\{mlp\\\_ratio\}=2\. A final LayerNorm and linear head predict one scalar field value per point\.
#### Geo\-FNO\.
Geo\-FNO learns a deformation from the irregular input domain to a regular computational grid and applies an FNO trunk on that grid, mapping back to the physical points through the inverse deformation\. It is trained on all three datasets under the same blockwiseMM\-to\-MM\(M=8M=8\) autoregression protocol, with matched model size, identical training strategy, and identical inference protocol as GeoLAMP\-B, and is evaluated on the same 4096 G&L points of the test set\.
#### NUNO\.
NUNO partitions the non\-uniform point set into subdomains, interpolates each subdomain onto a local regular grid, and applies an FNO per subdomain before scattering the predictions back to the original points\. It follows the same blockwiseMM\-to\-MM\(M=8M=8\) protocol, matched model size, training strategy, inference protocol, and 4096 G&L evaluation points as Geo\-FNO and GeoLAMP\-B\.
#### Latent\-Det\.
Latent\-Det\. is a matched deterministic control for the generative formulation\. It uses the same frozen Global & Local VAE, the same latent tokens, the same causal block mask, the sameM=8M=8blockwise protocol and commit stride, matched transformer depth and width, and identical conditioning as GeoLAMP\-B\. The flow\-matching objective and its iterative ODE sampler are replaced by direct MSE regression of the same target latent block, so inference uses a single forward pass per block rather than 50 Euler steps\.
#### DiT\-S\.
DiT\-S denoises a64×8×864\\times 8\\times 8target latent using DDPM noise prediction\. With𝚙𝚊𝚝𝚌𝚑\_𝚜𝚒𝚣𝚎=1\\mathtt\{patch\\\_size\}=1, the noisy target latent is embedded into 64 tokens with hidden width 768 and fixed sinusoidal 2\-D positional embeddings\. The diffusion timestep is embedded using a sinusoidal\-MLP timestep embedder for AdaLN\-Zero modulation\. The transformer uses 12 blocks, hidden width 768, 12 attention heads, andmlp\_ratio=4\\mathrm\{mlp\\\_ratio\}=4\. Each block contains AdaLN\-modulated self\-attention, cross\-attention to context tokens, and an AdaLN\-modulated GeGLU feed\-forward layer\. The final AdaLN\-modulated linear layer maps the tokens back to 64\-channel patches\.
### B\.2Training settings
The training configurations are summarized in Table[4](https://arxiv.org/html/2609.00297#A2.T4)\. For the heat convection and reactive flow datasets, the input includes the velocity field as an optional context\. For the elasticity dataset, the model takes only the von Mises stress field as input\. The VAE and latent dynamics model are trained separately\. The arbitrary decoder is also trained independently, using precomputed latent vectors from partial points as input and full\-mesh data as the target\. All training tasks are conducted on a single NVIDIA A100 GPU with 80 GB of memory\.
Table 4:Training hyperparameters for baseline models, GeoLAMP variants, the VAE, and the arbitrary decoder\.Similar Articles
Latent PDE mapping for efficient physics-informed learning across geometries with limited data
Introduces latent PDE mapping, a physics-informed learning technique that enables efficient geometric generalization with sparse training data by pulling back PDE residuals to a latent geometry. Demonstrates significant error reduction on cardiac electrophysiology PDE benchmarks using only 15 training samples.
Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
Presents GeoIncNO, a geometry-aware incremental neural operator that improves long-horizon PDE prediction via residual latent increments and mean-fluctuation decoupled reconstruction, achieving better stability and spectral fidelity on 1D/2D/3D benchmarks.
Modeling Unknown Nonlocal PDE Systems via Flow Map Learning
This paper presents a flow-map learning framework for modeling unknown nonlocal PDEs directly from solution data, avoiding explicit nonlocal operator evaluation. The method learns finite-time evolution operators in modal or nodal space and demonstrates accurate long-time prediction for fractional diffusion and wave equations.
AeroJEPA: Learning Semantic Latent Representations for Scalable 3D Aerodynamic Field Modeling
This paper introduces AeroJEPA, a Joint-Embedding Predictive Architecture for scalable 3D aerodynamic field modeling. It addresses limitations in current surrogate models by predicting semantic latent representations of flow fields, enabling efficient high-fidelity analysis and design optimization.
LLT: Local Linear Transformer for PDE Operator Learning
Introduces LLT, a transformer-based neural operator that combines linear global attention with local spatial mixing for PDE learning. It achieves competitive accuracy and faster training compared to baselines on multiple PDE problems.