StationPDE: 面向站点的地表PDE学习用于多站多变量天气预测

arXiv cs.LG 论文

摘要

StationPDE是一个面向站点的地表PDE学习模型,用于多站多变量天气预测。它构建地形感知的连续场并建模物理动力学,以超越最先进的基准,实现9.6%的MSE降低。

arXiv:2609.22123v1 Announce Type: new Abstract: Multi-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution. Meanwhile, PDE-based weather models provide interpretable physical dynamics, yet require continuous fields and upper-air variables unavailable in surface station data. To bridge this gap, we propose StationPDE, a station-oriented surface PDE learning model. StationPDE constructs a terrain-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper-air inference. Surface wind transport explicitly evolves observable weather variables, while upper-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper-air variables. A parallel data-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station-level multivariate forecasting. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state-of-the-art baselines, reducing MSE by about $9.6\%$ on average compared with the strongest baseline. Code and implementation details are available at https://github.com/hnu-vis/StationPDE.
查看原文
查看缓存全文

缓存时间: 2026/09/22 09:12

# StationPDE: Station-Oriented Surface PDE Learning for Multi-Station Multivariate Weather Forecasting
Source: [https://arxiv.org/html/2609.22123](https://arxiv.org/html/2609.22123)
Changjian ChenAffiliation:Hunan UniversityCorrespondence to:[changjianchen@hnu\.edu\.cn](mailto:[email protected])Rongwen LiAffiliation:Hunan UniversityHongwu LiuKun FangZhuo TangAffiliation:Hunan UniversityCorrespondence to:[ztang@hnu\.edu\.cn](mailto:[email protected])

###### Abstract

Multi\-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations\. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution\. Meanwhile, PDE\-based weather models provide interpretable physical dynamics, yet require continuous fields and upper\-air variables unavailable in surface station data\. To bridge this gap, we propose StationPDE, a station\-oriented surface PDE learning model\. StationPDE constructs a terrain\-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper\-air inference\. Surface wind transport explicitly evolves observable weather variables, while upper\-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper\-air variables\. A parallel data\-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station\-level multivariate forecasting\. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state\-of\-the\-art baselines, reducing MSE by about9\.6%9\.6\\%on average compared with the strongest baseline\. Code and implementation details are available at[https://github\.com/hnu\-vis/StationPDE](https://github.com/hnu-vis/StationPDE)\.

\\icml@noticeprintedtrue††footnotetext:1Hunan University\. Correspondence to: Xiao Wang <wang0l1@hnu\.edu\.cn\>, Changjian Chen <changjianchen@hnu\.edu\.cn\>, Zhuo Tang <ztang@hnu\.edu\.cn\>\.

## 1Introduction

Multi\-station multivariate weather forecasting forecasts future weather variables at multiple weather stations from their historical observations\. It is widely used in weather\-sensitive applications, such as local weather services, risk management, and renewable energy scheduling\([19](https://arxiv.org/html/2609.22123#bib.bib5)\)\. Existing multi\-station weather forecasting models typically adopt data\-driven strategies to model inter\-station dependencies\([1](https://arxiv.org/html/2609.22123#bib.bib3);[2](https://arxiv.org/html/2609.22123#bib.bib8);[14](https://arxiv.org/html/2609.22123#bib.bib7)\)\. While these models can capture statistical correlations among observed stations, they usually characterize weather evolution through discrete station\-to\-station relationships, which may fail to reflect the continuous evolution of the underlying weather process\. A recent model, CDPNet\([15](https://arxiv.org/html/2609.22123#bib.bib17)\), moves beyond discrete station modeling by constructing a spatially continuous field from station observations and estimating diffusive differences over time\. However, their evolution process remains largely data\-driven and lacks explicit modeling of multivariate physical coupling\.

![Refer to caption](https://arxiv.org/html/2609.22123v1/intro.png)Figure 1:Motivation of station\-oriented surface PDE modeling\. \(a\) Wind\-aligned redistribution of surface relative humidity observed from MeteoNet stations\. \(b\) Surface PDE approximation under unavailable upper\-air dynamics, where surface wind transport and upper\-air influence are modeled on the station\-derived surface field\.Figure[1](https://arxiv.org/html/2609.22123#S1.F1)\(a\) illustrates why such multivariate physical coupling is important in multi\-station multivariate weather forecasting\. In a real case from the MeteoNet dataset\([4](https://arxiv.org/html/2609.22123#bib.bib2)\)in northwestern France, future surface relative humidity is not solely determined by each station’s own history\. Instead, the dominant surface wind att0t\_\{0\}\(Fig\.[1](https://arxiv.org/html/2609.22123#S1.F1)A\) is followed by a clear redistribution of relative humidity after six hours \(Fig\.[1](https://arxiv.org/html/2609.22123#S1.F1)B\), suggesting a wind\-driven transport process across stations\. This observation indicates that weather forecasting requires more than learning which stations are statistically related; it also requires modeling how weather variables are transported over space\.

A natural way to introduce such physical evolution is to embed atmospheric partial differential equation \(PDE\) processes into weather forecasting models\. For example, WeatherGFT\([16](https://arxiv.org/html/2609.22123#bib.bib13)\)incorporates PDE\-based processes into the forward pass and demonstrates the value of explicit physical dynamics for weather forecasting\. However, existing PDE\-based weather models are mainly designed to evolve continuous states on regular continuous fields, whereas surface station observations and available only at discrete locations\([20](https://arxiv.org/html/2609.22123#bib.bib9);[4](https://arxiv.org/html/2609.22123#bib.bib2)\)\. Moreover, standard atmospheric PDE formulations often rely on multi\-pressure\-level upper\-air variables, such as geopotential and vertical velocity, while surface station datasets only have surface observations \(*e\.g*\., Fig\.[1](https://arxiv.org/html/2609.22123#S1.F1)C\)\. This leaves a clear gap:multi\-station weather forecasting models lack explicit physical evolution, while existing PDE\-based weather models cannot be directly applied when only surface station observations are available\.

To bridge this gap, we propose StationPDE, a station\-oriented surface PDE learning model for multi\-station multivariate weather forecasting\. StationPDE first constructs a continuous surface field from discrete station observations through terrain\-aware surface field initialization\. To address the incomplete specification of physical evolution caused by missing upper\-air variables, StationPDE decomposes the evolution process into two complementary parts:*surface wind transport*and*upper\-air inference*\. For*surface wind transport*, since the related surface variables are available in station observations, it directly incorporates these variables, including thermodynamic state, pressure transport, humidity evolution, and wind dynamics, in the PDE\. For*upper\-air inference*, since the related variables are not available, we use learnable horizontal diffusion to approximate their missing influence on the surface field, enabling local weather changes to propagate across neighboring regions \(*e\.g*\., Fig\.[1](https://arxiv.org/html/2609.22123#S1.F1)D\)\. Together, surface wind transport and horizontal diffusion form a complete PDE system for modeling the physical evolution of the surface field\. A parallel data\-driven diffusion branch further captures residual motion patterns, and an adaptive router integrates the two forecasts and maps the fused continuous field back to station\-level multivariate forecasts\.

In summary, our contributions are as follows\.

- •A station\-oriented surface PDE learning model, which introduces explicit physical evolution into multi\-station multivariate weather forecasting using only surface station observations\.
- •A surface PDE evolution design, which decomposes surface field evolution into surface wind transport and upper\-air inference, making explicit physical evolution feasible from surface station observations\.
- •Extensive experiments on real\-world datasets, which show that our model outperforms state\-of\-the\-art baselines on Weather2K and MeteoNet, reducing MSE by about9\.6%9\.6\\%on average compared with the strongest baseline\.

## 2Related Work

### 2\.1Data\-Driven Multi\-Station Forecasting Models

Data\-driven multi\-station forecasting models mainly learn temporal patterns and inter\-station dependencies from historical observations\. For example, Autoformer\([13](https://arxiv.org/html/2609.22123#bib.bib4)\)replaces standard self\-attention with an Auto\-Correlation mechanism and seasonal\-trend decomposition for long\-term forecasting\. Building on Autoformer, Corrformer\([14](https://arxiv.org/html/2609.22123#bib.bib7)\)introduces a multi\-correlation mechanism that unifies spatial cross\-correlation and temporal auto\-correlation for regional or global forecasting\. Formulating multi\-station weather forecasting as a graph learning problem, MasterGNN\+\([2](https://arxiv.org/html/2609.22123#bib.bib8)\)learns multi\-view station graphs from geographic distance and environmental context\.

Recent general multivariate time series models complement multi\-station models by improving variable\-wise dependency modeling\. iTransformer\([5](https://arxiv.org/html/2609.22123#bib.bib14)\)applies Transformer components on inverted dimensions, embedding each variate’s historical series as a token to capture multivariate correlations\. TimeXer\([12](https://arxiv.org/html/2609.22123#bib.bib15)\)models forecasting with exogenous variables, reconciling endogenous and exogenous information through patch\-wise self\-attention and variate\-wise cross\-attention\.

However, these models mainly operate on observed sequences rather than constructing continuous weather fields from discrete stations\. Recently, CDPNet\([15](https://arxiv.org/html/2609.22123#bib.bib17)\)moves beyond purely discrete station modeling by lifting discrete station observations into continuous fields\. Nevertheless, its evolution process remains largely data\-driven and does not explicitly characterize the physical coupling among weather variables\. In contrast, StationPDE constructs a station\-derived surface continuous field and evolves it with station\-oriented surface PDE processes\.

![Refer to caption](https://arxiv.org/html/2609.22123v1/pipline.png)Figure 2:The overview of the StationPDE structure\.
### 2\.2Physical\-Evolution Weather Forecasting Models

Physical knowledge has been increasingly introduced into weather forecasting models to improve interpretability and physical consistency\. For example, DGFormer\([17](https://arxiv.org/html/2609.22123#bib.bib11)\)is a physics\-guided station\-level weather forecasting model that uses a continuously varying graph topology and inserts domain knowledge into spatial\-temporal graph learning\. For discrete station observations, physics\-guided weather reconstruction\([9](https://arxiv.org/html/2609.22123#bib.bib12)\)reconstructs dense wind and pressure fields from discrete station measurements under Navier\-Stokes regularization\. PhyDL\-NWP\([6](https://arxiv.org/html/2609.22123#bib.bib18)\)computes physical terms by automatic differentiation and uses physics\-informed losses with latent force parameterization to guide continuous weather modeling\. However, these models mainly use physical knowledge as graph guidance, loss regularization, or parameterized correction, rather than explicit station\-oriented PDE forward evolution over weather variables\.

Recently, WeatherGFT\([16](https://arxiv.org/html/2609.22123#bib.bib13)\)embeds physical evolution directly into the model forward process, using a PDE kernel for fine\-grained physical evolution and a parallel neural branch for adaptive correction\. However, such PDE\-based weather models are usually designed for regular continuous fields and rely on upper\-air variables unavailable in surface station observations\. In contrast, StationPDE performs surface PDE evolution over observable weather variables on a station\-derived continuous field, enabling PDE\-guided multi\-station forecasting from surface station data\.

## 3Problem Formulation

We study multi\-station multivariate weather forecasting from surface station observations\. ConsiderNNweather stations distributed in a region\. Each station has fixed geographic information, including longitude, latitude, and altitude\. We denote the station geographic information as𝐋station∈ℝN×3\\mathbf\{L\}^\{\\mathrm\{station\}\}\\in\\mathbb\{R\}^\{N\\times 3\}\.

At each time steptt, the observable weather variables at all stations are denoted by𝐒t∈ℝN×C\\mathbf\{S\}\_\{t\}\\in\\mathbb\{R\}^\{N\\times C\}\. In this work, we useC=5C=5observable variables ordered as𝐒t=\[u,v,p,T,R​H\]\\mathbf\{S\}\_\{t\}=\[u,v,p,T,RH\], wherepp,TT, andR​HRHdenote air pressure, air temperature, and relative humidity, respectively, and wind speed and direction are converted into horizontal componentsuuandvvalong the eastward and northward directions\. Given the historical station observations over the pastTinT\_\{\\mathrm\{in\}\}time steps,𝐒t−Tin:t−1∈ℝTin×N×C\\mathbf\{S\}\_\{t\-T\_\{\\mathrm\{in\}\}:t\-1\}\\in\\mathbb\{R\}^\{T\_\{\\mathrm\{in\}\}\\times N\\times C\}, the goal is to generate future station\-level weather forecasts for the nextToutT\_\{\\mathrm\{out\}\}time steps:

𝐒t−Tin:t−1,𝐋⟶𝐒^t:t\+Tout−1,\\mathbf\{S\}\_\{t\-T\_\{\\mathrm\{in\}\}:t\-1\},\\mathbf\{L\}\\longrightarrow\\hat\{\\mathbf\{S\}\}\_\{t:t\+T\_\{\\mathrm\{out\}\}\-1\},\(1\)where𝐒^t:t\+Tout−1∈ℝTout×N×C\\hat\{\\mathbf\{S\}\}\_\{t:t\+T\_\{\\mathrm\{out\}\}\-1\}\\in\\mathbb\{R\}^\{T\_\{\\mathrm\{out\}\}\\times N\\times C\}\.

In this work, we focus on introducing explicit physical evolution into this station\-based forecasting setting\. StationPDE lifts the station observations𝐒\\mathbf\{S\}into a station\-derived surface continuous field, evolves the field with station\-oriented surface PDE processes, and maps the evolved field back to station\-level forecasts\. To represent this continuous field, we discretize the study region into a regularH×WH\\times Wgrid, whereHHandWWdenote the numbers of grid cells along the latitude and longitude directions, respectively, andM=H​WM=HW\. We denote the geographic information of all grid cells by𝐋grid∈ℝM×3\\mathbf\{L\}^\{\\mathrm\{grid\}\}\\in\\mathbb\{R\}^\{M\\times 3\}\.

## 4The StationPDE Framework

In this section, we introduce the proposed StationPDE model for multi\-station multivariate weather forecasting from surface station observations\. As shown in Fig\.[2](https://arxiv.org/html/2609.22123#S2.F2), StationPDE consists of four key components: \(1\) aTerrain\-Aware Surface Field Initializationmodule that lifts discrete station observations into a station\-derived surface continuous field; \(2\) aStation\-Oriented PDE Branchthat models surface field evolution by combining surface wind transport with upper\-air inference over observable weather variables; \(3\) aData\-Driven Diffusion Branchthat captures complementary motion and correction patterns beyond the simplified surface PDEs; and \(4\) anAdaptive Routerthat integrates the PDE\-based and data\-driven forecasts and maps the fused field back to station\-level multivariate forecasts\. The following subsections describe these components in detail\.

### 4\.1Terrain\-Aware Surface Field Initialization

The first step of StationPDE is to construct a surface continuous field from discrete and irregular station observations\. Before spatial lifting, we convert observable variables𝐒t=\[u,v,p,T,R​H\]\\mathbf\{S\}\_\{t\}=\[u,v,p,T,RH\]into physical state channelsΨ⁡\(𝐒t\)=\[u,v,p,θ,q\]\\Psi\(\\mathbf\{S\}\_\{t\}\)=\[u,v,p,\\theta,q\], whereθ\\thetaandqqare potential temperature and specific humidity\. The inverse transformationΨ−1​\(⋅\)\\Psi^\{\-1\}\(\\cdot\)maps the physical state back to observable variables, with details provided in the appendix\.

CDPNet\([15](https://arxiv.org/html/2609.22123#bib.bib17)\)shows that interpolating station observations over latitude and longitude can construct a continuous spatial field from discrete stations\. However, such location\-only interpolation ignores terrain effects, which is problematic for surface variables such as temperature and pressure\. Therefore, StationPDE performs terrain\-aware surface field initialization with inverse distance weighting \(IDW\) interpolation, in which both horizontal location and altitude are used to compute interpolation weights\. Let𝐋stationi,:\\mathbf\{L\}^\{\\mathrm\{station\}\}\_\{i,:\}and𝐋gridm,:\\mathbf\{L\}^\{\\mathrm\{grid\}\}\_\{m,:\}denote the geographic information of stationiiand grid cellmm, respectively\. Forc∈\{lon,lat,alt\}c\\in\\\{\\mathrm\{lon\},\\mathrm\{lat\},\\mathrm\{alt\}\\\}, we define the coordinate difference asΔ​lm​ic=Lm,cgrid−Li,cstation\\Delta l^\{c\}\_\{mi\}=L^\{\\mathrm\{grid\}\}\_\{m,c\}\-L^\{\\mathrm\{station\}\}\_\{i,c\}\. The grid altitude is first estimated from station altitudes by horizontal IDW\. For each grid cellmm, we select a neighboring station set𝒩⁡\(m\)\\mathcal\{N\}\(m\)and compute the terrain\-aware distance between stationiiand grid cellmmas

dm​i=\[\\displaystyle d\_\{mi\}=\\Big\[\(R​cos⁡\(l¯lat\)​Δ​lm​ilon\)2\+\(R​Δ​lm​ilat\)2\\displaystyle\\big\(R\\cos\(\\bar\{l\}\_\{\\mathrm\{lat\}\}\)\\Delta l^\{\\mathrm\{lon\}\}\_\{mi\}\\big\)^\{2\}\+\\big\(R\\Delta l^\{\\mathrm\{lat\}\}\_\{mi\}\\big\)^\{2\}\(2\)\+\(waltΔlaltm​i\)2\]1/2,i∈𝒩\(m\),\\displaystyle\+\\big\(w\_\{\\mathrm\{alt\}\}\\Delta l^\{\\mathrm\{alt\}\}\_\{mi\}\\big\)^\{2\}\\Big\]^\{1/2\},\\quad i\\in\\mathcal\{N\}\(m\),whereRRis the Earth radius,l¯lat\\bar\{l\}\_\{\\mathrm\{lat\}\}is the mean latitude of the study region, andwalt=5w\_\{\\mathrm\{alt\}\}=5controls the contribution of altitude to the interpolation distance\.

Based on this distance, the terrain\-aware IDW coefficient from stationiito grid cellmmis defined as

αm​i=\(dm​i\+ϵ\)−1∑j∈𝒩⁡\(m\)\(dm​j\+ϵ\)−1,i∈𝒩⁡\(m\),\\alpha\_\{mi\}=\\frac\{\(d\_\{mi\}\+\\epsilon\)^\{\-1\}\}\{\\sum\_\{j\\in\\mathcal\{N\}\(m\)\}\(d\_\{mj\}\+\\epsilon\)^\{\-1\}\},\\quad i\\in\\mathcal\{N\}\(m\),\(3\)wherejjindexes neighboring stations for normalization andϵ\\epsilonis a small constant for numerical stability\. Each historical station state is then lifted to grid cellmmby

𝐆τ,m=∑i∈𝒩⁡\(m\)αm​i\[Ψ\(𝐒τ\)\]i,τ=t−Tin,…,t−1\.\\mathbf\{G\}\_\{\\tau,m\}=\\sum\_\{i\\in\\mathcal\{N\}\(m\)\}\\alpha\_\{mi\}\[\\Psi\(\\mathbf\{S\}\_\{\\tau\}\)\]\_\{i\},\\quad\\tau=t\-T\_\{\\mathrm\{in\}\},\\ldots,t\-1\.\(4\)This produces a historical surface continuous field𝐆t−Tin:t−1∈ℝTin×C×H×W\\mathbf\{G\}\_\{t\-T\_\{\\mathrm\{in\}\}:t\-1\}\\in\\mathbb\{R\}^\{T\_\{\\mathrm\{in\}\}\\times C\\times H\\times W\}\.

PDE evolution is driven by an initial spatial state\. If we directly use the latest interpolated field𝐆t−1\\mathbf\{G\}\_\{t\-1\}, the temporal information in the historical window may be underused\. We therefore introduce a history\-conditioned encoder\. The historical tensor is concatenated with initialization condition fields𝐔t\\mathbf\{U\}\_\{t\}, including location and historical time features, to generate a candidate initial field:

𝐗~t=ℰϕ\(\[𝐆t−Tin:t−1,𝐔t\]\),\\widetilde\{\\mathbf\{X\}\}\_\{t\}=\\mathcal\{E\}\_\{\\phi\}\\big\(\[\\mathbf\{G\}\_\{t\-T\_\{\\mathrm\{in\}\}:t\-1\},\\mathbf\{U\}\_\{t\}\]\\big\),\(5\)whereℰϕ\\mathcal\{E\}\_\{\\phi\}denotes the history\-conditioned encoder, implemented with lightweight convolutional layers to aggregate the historical sequence and condition fields\. The initial field is then obtained by a gated correction:

𝐗t=𝐆t−1\+𝐌tinit⊙\(𝐗~t−𝐆t−1\),\\mathbf\{X\}\_\{t\}=\\mathbf\{G\}\_\{t\-1\}\+\\mathbf\{M\}\_\{t\}^\{\\mathrm\{init\}\}\\odot\(\\widetilde\{\\mathbf\{X\}\}\_\{t\}\-\\mathbf\{G\}\_\{t\-1\}\),\(6\)where𝐌tinit\\mathbf\{M\}\_\{t\}^\{\\mathrm\{init\}\}is a sigmoid initialization gate generated from the latest continuous field, encoded history features, and initialization condition fields\. The resulting𝐗t∈ℝC×H×W\\mathbf\{X\}\_\{t\}\\in\\mathbb\{R\}^\{C\\times H\\times W\}serves as the initial surface field for both the station\-oriented PDE branch and the data\-driven diffusion branch\.

### 4\.2Station\-Oriented PDE Branch

The station\-derived surface continuous field𝐗t\\mathbf\{X\}\_\{t\}provides a spatial state for PDE evolution, but it only contains surface observable variables rather than full atmospheric states\. Thus, standard atmospheric PDE formulations that rely on upper\-air variables, such as vertical velocity and geopotential, cannot be directly applied\. To address the incomplete specification of physical evolution caused by missing upper\-air variables, StationPDE decomposes surface field evolution into two complementary parts:*surface wind transport*and*upper\-air inference*\.

*Surface wind transport*\. Let𝐕=\(u,v\)\\mathbf\{V\}=\(u,v\)denote the surface horizontal wind field, and let∇=\(∂x,∂y\)\\nabla=\(\\partial\_\{x\},\\partial\_\{y\}\)denote the horizontal gradient operator\. Following standard atmospheric PDE formulations,*surface wind transport*is modeled by the surface computable term−𝐕⋅∇a\-\\mathbf\{V\}\\cdot\\nabla afor each state variablea∈\{u,v,p,θ,q\}a\\in\\\{u,v,p,\\theta,q\\\}\. This term describes how observable weather variables are transported over the continuous surface field by horizontal wind\.

*Upper\-air inference*\. We consider the missing influence of vertical momentum transport and geopotential gradient forcing, both of which require upper\-air variables unavailable in surface station observations\. Without these terms, the surface field lacks part of the large\-scale spatial forcing needed to propagate local weather changes across neighboring regions\. We therefore introduce learnable horizontal diffusion for each state variable\. Its Laplacian form spreads local anomalies over the surface field, providing a surface\-level approximation to the missing upper\-air influence\. This design is inspired by the explicit spatial diffusion treatment in the WRF\-ARW technical note\([8](https://arxiv.org/html/2609.22123#bib.bib1)\), where second\-order horizontal diffusion is introduced for model variables, with corresponding forms for both momentum and scalar variables\.

Accordingly, for each state variableaa, we use the following transport\-diffusion form:

∂a∂t=−𝐕⋅∇a\+κa∇2a\+ℛa,\\frac\{\\partial a\}\{\\partial t\}=\-\\mathbf\{V\}\\cdot\\nabla a\+\\kappa\_\{a\}\\nabla^\{2\}a\+\\mathcal\{R\}\_\{a\},\(7\)whereκa\\kappa\_\{a\}is the horizontal diffusion coefficient, andℛa\\mathcal\{R\}\_\{a\}collects variable\-specific surface forcing or closure terms\.

For*wind dynamics*,ℛa\\mathcal\{R\}\_\{a\}includes the Coriolis force and surface friction:

∂u∂t\\displaystyle\\frac\{\\partial u\}\{\\partial t\}=−𝐕⋅∇u\+fv−rmu\+κu∇2u,\\displaystyle=\-\\mathbf\{V\}\\cdot\\nabla u\+fv\-r\_\{m\}u\+\\kappa\_\{u\}\\nabla^\{2\}u,\(8\)∂v∂t\\displaystyle\\frac\{\\partial v\}\{\\partial t\}=−𝐕⋅∇v−fu−rmv\+κv∇2v\.\\displaystyle=\-\\mathbf\{V\}\\cdot\\nabla v\-fu\-r\_\{m\}v\+\\kappa\_\{v\}\\nabla^\{2\}v\.Here,f=2​Ω​sin⁡\(ϕ\)f=2\\Omega\\sin\(\\phi\)is the Coriolis parameter, whereΩ\\Omegais the Earth angular velocity andϕ\\phiis the grid latitude\. The surface friction coefficientrmr\_\{m\}and wind diffusion coefficientsκu,κv\\kappa\_\{u\},\\kappa\_\{v\}are learnable positive parameters\.

For the remaining state variables, the same transport\-diffusion form gives*pressure transport*,*thermodynamic state*, and*humidity evolution*:

∂p∂t\\displaystyle\\frac\{\\partial p\}\{\\partial t\}=−𝐕⋅∇p\+κp∇2p\+Sp,\\displaystyle=\-\\mathbf\{V\}\\cdot\\nabla p\+\\kappa\_\{p\}\\nabla^\{2\}p\+S\_\{p\},\(9\)∂θ∂t\\displaystyle\\frac\{\\partial\\theta\}\{\\partial t\}=−𝐕⋅∇θ\+κθ∇2θ\+Sθ,\\displaystyle=\-\\mathbf\{V\}\\cdot\\nabla\\theta\+\\kappa\_\{\\theta\}\\nabla^\{2\}\\theta\+S\_\{\\theta\},∂q∂t\\displaystyle\\frac\{\\partial q\}\{\\partial t\}=−𝐕⋅∇q\+κq∇2q\+Sq−λcmax\(q−qs,0\)\.\\displaystyle=\-\\mathbf\{V\}\\cdot\\nabla q\+\\kappa\_\{q\}\\nabla^\{2\}q\+S\_\{q\}\-\\lambda\_\{c\}\\max\(q\-q\_\{s\},0\)\.Here,SpS\_\{p\},SθS\_\{\\theta\}, andSqS\_\{q\}are closure terms generated from the current state, terrain, and time features to represent unresolved pressure, heat, and moisture effects\. In implementation, they are produced by a two\-layer convolutional closure network with terrain\-conditioned hidden features and channel\-wise output scales\. The saturation specific humidityqsq\_\{s\}is computed from temperature and pressure, andλc​max⁡\(q−qs,0\)\\lambda\_\{c\}\\max\(q\-q\_\{s\},0\)removes supersaturated moisture\. The diffusion coefficientsκp,κθ,κq\\kappa\_\{p\},\\kappa\_\{\\theta\},\\kappa\_\{q\}and condensation coefficientλc\\lambda\_\{c\}are learnable positive parameters\.

The above evolution equations are used to advance the current surface state through repeated PDE substeps\. As shown in Fig\.[2](https://arxiv.org/html/2609.22123#S2.F2)\(b\), given the current field𝐗t\\mathbf\{X\}\_\{t\}, we set𝐗\(0\)=𝐗t\\mathbf\{X\}^\{\(0\)\}=\\mathbf\{X\}\_\{t\}and chooseKKsuch thatK​Δ​tK\\Delta tmatches one forecast interval\. At therr\-th substep, the PDE tendencies are converted into a range\-scaled Euler update:

𝐗\(r\+1\)\\displaystyle\\mathbf\{X\}^\{\(r\+1\)\}=𝐗\(r\)\+𝜸⊙Scale⁡\(Δ​t​ℱPDE​\(𝐗\(r\)\),𝐗\(r\)\),\\displaystyle=\\mathbf\{X\}^\{\(r\)\}\+\\boldsymbol\{\\gamma\}\\odot\\mathrm\{Scale\}\\big\(\\Delta t\\,\\mathcal\{F\}\_\{\\mathrm\{PDE\}\}\(\\mathbf\{X\}^\{\(r\)\}\);\\mathbf\{X\}^\{\(r\)\}\\big\),\(10\)r=0,…,K−1\.\\displaystyle r=0,\\ldots,K\-1\.Here,Δ​t=1800​s\\Delta t=1800\\ \\mathrm\{s\}, andℱPDE​\(⋅\)\\mathcal\{F\}\_\{\\mathrm\{PDE\}\}\(\\cdot\)concatenates the tendencies of all state variables\. Following the range\-scaling strategy in WeatherGFT\([16](https://arxiv.org/html/2609.22123#bib.bib13)\),Scale⁡\(⋅\)\\mathrm\{Scale\}\(\\cdot\)rescales the Euler update according to the spatial range of each state channel\. The factor𝜸\\boldsymbol\{\\gamma\}is a learnable channel\-wise update factor\. AfterKKsubsteps, the station\-oriented PDE branch outputs𝐗t\+1PDE=𝐗\(K\)\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{PDE\}\}=\\mathbf\{X\}^\{\(K\)\}\.

### 4\.3Data\-Driven Diffusion Branch

Although the PDE branch provides explicit surface evolution, the simplified surface PDEs cannot fully cover data\-driven motion and residual patterns\. As shown in Fig\.[2](https://arxiv.org/html/2609.22123#S2.F2)\(c\), StationPDE therefore uses a data\-driven diffusion branch to provide adaptive correction on the continuous field\. Given the current field𝐗t\\mathbf\{X\}\_\{t\}and the step\-wise condition field𝐁t\\mathbf\{B\}\_\{t\}containing location, time, and lead\-time features, the branch encodes a condition\-aware hidden field and then evolves it through motion\-guided warping\. This branch starts by generating a candidate state:

𝐗~t=𝐗t\+τc​𝒜ω​\(\[𝐗t,𝐁t\]\),\\widetilde\{\\mathbf\{X\}\}\_\{t\}=\\mathbf\{X\}\_\{t\}\+\\tau\_\{c\}\\mathcal\{A\}\_\{\\omega\}\(\[\\mathbf\{X\}\_\{t\},\\mathbf\{B\}\_\{t\}\]\),\(11\)where𝒜ω\\mathcal\{A\}\_\{\\omega\}is a convolutional candidate network andτc\\tau\_\{c\}is a learnable scale\. Then,a state gate controls the correction from𝐗t\\mathbf\{X\}\_\{t\}to𝐗~t\\widetilde\{\\mathbf\{X\}\}\_\{t\}, producing the encoded hidden field𝐇t\\mathbf\{H\}\_\{t\}:

𝐇t=𝐗t\+𝐌ts⊙\(𝐗~t−𝐗t\),\\mathbf\{H\}\_\{t\}=\\mathbf\{X\}\_\{t\}\+\\mathbf\{M\}\_\{t\}^\{s\}\\odot\(\\widetilde\{\\mathbf\{X\}\}\_\{t\}\-\\mathbf\{X\}\_\{t\}\),\(12\)where𝐌ts\\mathbf\{M\}\_\{t\}^\{s\}is a sigmoid state gate generated from𝐗t\\mathbf\{X\}\_\{t\},𝐗~t\\widetilde\{\\mathbf\{X\}\}\_\{t\}, and𝐁t\\mathbf\{B\}\_\{t\}\.

From𝐇t\\mathbf\{H\}\_\{t\}, the branch estimates a two\-channel motion field:

𝐃t=ℳω​\(\[𝐇t,𝐁t\]\),\\mathbf\{D\}\_\{t\}=\\mathcal\{M\}\_\{\\omega\}\(\[\\mathbf\{H\}\_\{t\},\\mathbf\{B\}\_\{t\}\]\),\(13\)whereℳω\\mathcal\{M\}\_\{\\omega\}is a motion network and𝐃t\\mathbf\{D\}\_\{t\}denotes horizontal displacement\.Implementation details of these networks are provided in the appendix\.

The hidden field𝐇t\\mathbf\{H\}\_\{t\}is then transported by differentiable bilinear warping:

𝐖t=Warp⁡\(𝐇t,𝐃t\),\\mathbf\{W\}\_\{t\}=\\mathrm\{Warp\}\(\\mathbf\{H\}\_\{t\},\\mathbf\{D\}\_\{t\}\),\(14\)where𝐖t\\mathbf\{W\}\_\{t\}is the warped field\.

Finally, the warped change is combined with a residual correction through a gate:

𝐗t\+1AI=𝐇t\+𝐌td⊙\(𝐖t−𝐇t\)\+𝐑t,\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{AI\}\}=\\mathbf\{H\}\_\{t\}\+\\mathbf\{M\}\_\{t\}^\{d\}\\odot\(\\mathbf\{W\}\_\{t\}\-\\mathbf\{H\}\_\{t\}\)\+\\mathbf\{R\}\_\{t\},\(15\)where𝐌td\\mathbf\{M\}\_\{t\}^\{d\}is the warp gate generated from𝐇t\\mathbf\{H\}\_\{t\},𝐗~t\\widetilde\{\\mathbf\{X\}\}\_\{t\}, and𝐁t\\mathbf\{B\}\_\{t\}, and𝐑t\\mathbf\{R\}\_\{t\}is a learned residual correction generated from𝐇t\\mathbf\{H\}\_\{t\}and𝐁t\\mathbf\{B\}\_\{t\}\. The output of this branch is the data\-driven field forecast𝐗t\+1AI\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{AI\}\}\.

### 4\.4Adaptive Router

As shown in Fig\.[2](https://arxiv.org/html/2609.22123#S2.F2)\(d\), the adaptive router integrates the PDE and data\-driven branch outputs, and then maps the fused field back to station\-level forecasts\. Given𝐗t\+1PDE\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{PDE\}\}and𝐗t\+1AI\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{AI\}\}, the router first estimates a fusion weight:

𝜶t=σ⁡\(ℛρ​\(\[𝐗t\+1AI,𝐗t\+1PDE,𝐁t\]\)\),\\boldsymbol\{\\alpha\}\_\{t\}=\\sigma\\\!\\left\(\\mathcal\{R\}\_\{\\rho\}\(\[\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{AI\}\},\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{PDE\}\},\\mathbf\{B\}\_\{t\}\]\)\\right\),\(16\)whereℛρ\\mathcal\{R\}\_\{\\rho\}is the router network,σ⁡\(⋅\)\\sigma\(\\cdot\)is the sigmoid function, and𝜶t\\boldsymbol\{\\alpha\}\_\{t\}controls the contribution of the data\-driven branch\. The fused field is then obtained by

𝐗t\+1=𝜶t⊙𝐗t\+1AI\+\(1−𝜶t\)⊙𝐗t\+1PDE\.\\mathbf\{X\}\_\{t\+1\}=\\boldsymbol\{\\alpha\}\_\{t\}\\odot\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{AI\}\}\+\(1\-\\boldsymbol\{\\alpha\}\_\{t\}\)\\odot\\mathbf\{X\}\_\{t\+1\}^\{\\mathrm\{PDE\}\}\.\(17\)
The fused field is then sampled at station locations and refined by a unified station residual head:

𝐒^t\+1=Ψ−1\(Readout\(𝐗t\+1\)\+ℋξ\(Ψ\(𝐒t−Tin:t−1\),𝐁t\)\)\.\\hat\{\\mathbf\{S\}\}\_\{t\+1\}=\\Psi^\{\-1\}\\left\(\\mathrm\{Readout\}\(\\mathbf\{X\}\_\{t\+1\}\)\+\\mathcal\{H\}\_\{\\xi\}\(\\Psi\(\\mathbf\{S\}\_\{t\-T\_\{\\mathrm\{in\}\}:t\-1\}\),\\mathbf\{B\}\_\{t\}\)\\right\)\.\(18\)Here,Readout⁡\(⋅\)\\mathrm\{Readout\}\(\\cdot\)denotes sampling the continuous field at station locations, andℋξ\\mathcal\{H\}\_\{\\xi\}is a unified station residual head shared across variables\. It is implemented as a lightweight MLP with two SiLU hidden layers, which takes the historical station state sequence, station geographic features, time features, and lead step as input and outputs channel\-wise scaled station residuals for all state channels\. The inverse transformΨ−1​\(⋅\)\\Psi^\{\-1\}\(\\cdot\)converts the corrected physical state back to observable weather variables\. Repeating this process autoregressively gives the final multi\-station forecast𝐒^t:t\+Tout−1\\hat\{\\mathbf\{S\}\}\_\{t:t\+T\_\{\\mathrm\{out\}\}\-1\}\.

ModelsStationPDE\(Ours\)CDPNetTimeFilterMultiPatchFormerTimeXerTimeMixeriTransformerCorrformerDLinearweather2kpressure4\.413±0\.0909\.840±0\.644 8\.625±0\.137 6\.750±0\.122 7\.738±0\.080 7\.684±0\.119 7\.700±0\.077 5\.732±0\.047 11\.317±0\.129 temp4\.607±0\.0555\.856±0\.101 7\.058±0\.022 6\.694±0\.038 7\.121±0\.020 7\.653±0\.020 7\.118±0\.015 5\.897±0\.016 8\.356±0\.006 humidity127\.240±0\.807150\.312±2\.075 171\.486±0\.824 168\.889±0\.440 165\.972±0\.204 174\.633±0\.756 171\.770±0\.373 148\.102±0\.213 187\.482±0\.078 u\-wind2\.155±0\.0072\.234±0\.011 2\.597±0\.002 2\.573±0\.004 2\.516±0\.005 2\.581±0\.006 2\.606±0\.003 2\.374±0\.003 2\.747±0\.001 v\-wind2\.194±0\.0082\.322±0\.013 2\.820±0\.010 2\.789±0\.009 2\.768±0\.005 2\.864±0\.013 2\.863±0\.006 2\.534±0\.003 3\.083±0\.001 metonettemp4\.298±0\.0954\.756±0\.077 5\.348±0\.069 5\.310±0\.030 5\.389±0\.067 6\.298±0\.037 5\.314±0\.039 7\.444±0\.164 6\.762±0\.046 humidity87\.115±0\.94188\.562±1\.293 96\.163±0\.213 92\.915±0\.170 94\.075±1\.054 102\.617±0\.316 95\.460±0\.290 113\.500±0\.760 103\.519±0\.182 u\-wind4\.241±0\.0874\.429±0\.055 4\.697±0\.031 4\.875±0\.012 4\.703±0\.028 5\.391±0\.013 4\.832±0\.009 5\.255±0\.106 5\.359±0\.012 v\-wind4\.316±0\.0664\.462±0\.033 5\.104±0\.028 5\.314±0\.030 5\.365±0\.023 5\.706±0\.018 5\.311±0\.029 5\.964±0\.107 5\.708±0\.014

Table 1:Experiment results \(MSE\) on the Weather2K and MeteoNet datasets\. All models use 48 hours of historical observations to forecast the next 24 hours\. Best results are marked inbold, and second\-best results areunderlined\. Complete MAE results are reported in the appendix\.
### 4\.5Model Training

We train StationPDE end\-to\-end with station\-level supervision\. The main loss is computed in the normalized physical state space:

ℒstate=1∑c=1Cwc​∑c=1Cwc​‖\[Ψ⁡\(𝐒^\)\]norm\(c\)−\[Ψ⁡\(𝐒\)\]norm\(c\)‖1,\\mathcal\{L\}\_\{\\mathrm\{state\}\}=\\frac\{1\}\{\\sum\_\{c=1\}^\{C\}w\_\{c\}\}\\sum\_\{c=1\}^\{C\}w\_\{c\}\\,\\left\\\|\[\\Psi\(\\hat\{\\mathbf\{S\}\}\)\]\_\{\\mathrm\{norm\}\}^\{\(c\)\}\-\[\\Psi\(\\mathbf\{S\}\)\]\_\{\\mathrm\{norm\}\}^\{\(c\)\}\\right\\\|\_\{1\},\(19\)wherewcw\_\{c\}is the state\-channel weight\.

To align training with observable weather variables, we additionally use an auxiliary loss in the observable variable space:

ℒobs=1∑j=1Cvj​∑j=1Cvj​‖\[𝐒^\]\(j\)−\[𝐒\]\(j\)sj‖1,\\mathcal\{L\}\_\{\\mathrm\{obs\}\}=\\frac\{1\}\{\\sum\_\{j=1\}^\{C\}v\_\{j\}\}\\sum\_\{j=1\}^\{C\}v\_\{j\}\\,\\left\\\|\\frac\{\[\\hat\{\\mathbf\{S\}\}\]^\{\(j\)\}\-\[\\mathbf\{S\}\]^\{\(j\)\}\}\{s\_\{j\}\}\\right\\\|\_\{1\},\(20\)wherevjv\_\{j\}is the observable\-variable weight, andsjs\_\{j\}is the batch\-wise target standard deviation used to balance variables with different units\. The total objective is

ℒ=ℒstate\+λobs​ℒobs,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{state\}\}\+\\lambda\_\{\\mathrm\{obs\}\}\\mathcal\{L\}\_\{\\mathrm\{obs\}\},\(21\)whereλobs\\lambda\_\{\\mathrm\{obs\}\}controls the strength of the auxiliary observable\-space loss\. All neural modules, adaptive router parameters, station residual head, and learnable PDE coefficients are optimized jointly\.

## 5Experiments

### 5\.1Experimental Settings

Dataset\.We evaluate StationPDE on two real\-world multi\-station multivariate weather datasets\. The Weather2K dataset\([20](https://arxiv.org/html/2609.22123#bib.bib9)\)contains1,8661\{,\}866ground weather stations across China from January 1, 2017 to August 31, 2021, with 3\-hour temporal resolution\. Each station provides observable surface weather variables\[u,v,p,T,R​H\]\[u,v,p,T,RH\], including air pressure, air temperature, relative humidity, and horizontal wind components converted from wind observations\. The MeteoNet dataset\([4](https://arxiv.org/html/2609.22123#bib.bib2)\)provides hourly surface station observations in northwestern France from January 1, 2016 to December 31, 2018\. After filtering stations with high missing rates, we retain133133stations and use the same types of surface weather variables as Weather2K, except that pressure is excluded from the final evaluation because MeteoNet records pressure reduced to sea level and its missing rate reaches64%64\\%\. For both datasets, we chronologically split the sequences into training, validation, and test sets with a6:1:36\{:\}1\{:\}3ratio\.

ModelsStationPDE\(Ours\)CDPNetTimeFilterMultiPatchFormerTimeXerTimeMixeriTransformerCorrformerDLineartemp486\.978±0\.1797\.288±0\.057 8\.556±0\.183 8\.397±0\.120 8\.483±0\.055 9\.274±0\.091 8\.323±0\.075 10\.449±0\.374 9\.969±0\.136 969\.849±0\.27810\.685±0\.144 12\.891±0\.154 12\.845±0\.148 12\.726±0\.118 13\.352±0\.065 12\.761±0\.051 14\.452±0\.188 13\.940±0\.063 humidity48103\.910±3\.200107\.508±0\.805 122\.763±1\.142 117\.479±0\.315 118\.568±0\.778 126\.363±0\.822 120\.362±0\.260 135\.347±1\.055 122\.470±0\.283 96133\.674±2\.090138\.263±0\.455 146\.440±2\.454 141\.550±0\.817 142\.138±1\.698 148\.828±0\.826 144\.545±0\.504 155\.502±1\.214 139\.269±0\.108 u\-wind485\.794±0\.0815\.865±0\.039 6\.669±0\.118 6\.686±0\.020 6\.630±0\.014 7\.137±0\.018 6\.724±0\.025 7\.138±0\.076 6\.669±0\.025 967\.174±0\.2257\.219±0\.037 8\.433±0\.172 8\.453±0\.033 8\.513±0\.032 8\.970±0\.024 8\.480±0\.028 9\.028±0\.307 7\.881±0\.007 v\-wind485\.966±0\.0535\.991±0\.012 7\.110±0\.135 7\.100±0\.031 7\.235±0\.009 7\.388±0\.042 7\.109±0\.056 7\.693±0\.083 6\.980±0\.027 967\.178±0\.1507\.227±0\.014 8\.649±0\.161 8\.788±0\.038 9\.001±0\.036 9\.063±0\.031 8\.674±0\.042 9\.210±0\.195 8\.087±0\.007

Table 2:Experiment results \(MSE\) on the MeteoNet dataset with different forecast horizons \(48h and 96h\)\. All models use 48 hours of historical observations\. Best results are marked inbold, and second\-best results areunderlined\. Complete MAE results are reported in the appendix\.Methodpressuretemphumidityu\-windv\-windFull / StationPDE4\.41±0\.094\.61±0\.06127\.24±0\.812\.16±0\.012\.19±0\.01w/o PDE4\.56±0\.03 5\.42±0\.06 150\.62±1\.55 2\.39±0\.01 2\.41±0\.00 w/o Data\-Driven4\.84±0\.06 5\.87±0\.11 161\.17±0\.70 2\.41±0\.06 2\.42±0\.01 w/o Station Residual4\.72±0\.06 5\.36±0\.11 148\.95±0\.74 2\.38±0\.01 2\.40±0\.00 w/oℒobs\\mathcal\{L\}\_\{\\mathrm\{obs\}\}4\.45±0\.07 4\.76±0\.01 134\.88±0\.61 2\.16±0\.01 2\.20±0\.01

Table 3:Ablation study results \(MSE\) on the Weather2K dataset\. All models use 48 hours of historical observations to forecast the next 24 hours\. Best results are marked inboldbased on higher\-precision values\.Baselines\.We compare StationPDE with representative baselines from two groups: multi\-station weather forecasting models, including Corrformer\([14](https://arxiv.org/html/2609.22123#bib.bib7)\)and CDPNet\([15](https://arxiv.org/html/2609.22123#bib.bib17)\); and recent general multivariate time series forecasting models, including DLinear\([18](https://arxiv.org/html/2609.22123#bib.bib6)\), iTransformer\([5](https://arxiv.org/html/2609.22123#bib.bib14)\), TimeMixer\([10](https://arxiv.org/html/2609.22123#bib.bib10)\), TimeXer\([12](https://arxiv.org/html/2609.22123#bib.bib15)\), MultiPatchFormer\([7](https://arxiv.org/html/2609.22123#bib.bib19)\), and TimeFilter\([3](https://arxiv.org/html/2609.22123#bib.bib20)\)\.

Implementation details\.Except for Corrformer and CDPNet, all baselines are implemented based on Time\-Series\-Library\([11](https://arxiv.org/html/2609.22123#bib.bib16)\)\. Corrformer and CDPNet are implemented following the hyperparameter settings recommended in their original papers\. For the MeteoNet dataset, pressure is recorded as pressure reduced to sea level and has a missing rate of64%64\\%\. Therefore, we exclude pressure from the final evaluation, but estimate a surface pressure reference from station altitude using the standard atmosphere approximation to drive PDE evolution\. For StationPDE, the grid heightHHis set to3232, and the grid width is automatically adjusted according to the regional aspect ratio\. The PDE substep is fixed asΔ​t=1800\\Delta t=1800seconds, withK=6K=6for Weather2K andK=2K=2for MeteoNet to match their forecast intervals\. The observable\-variable auxiliary loss weight is set toλobs=0\.2\\lambda\_\{\\mathrm\{obs\}\}=0\.2\. We train all trainable models for1010epochs using AdamW with a learning rate of1×10−41\\times 10^\{\-4\}\. Each experiment is repeated with five random seeds, and we report the mean and standard deviation\. Other hyperparameters, including channel\-wise loss weights, are provided in the appendix\. All experiments were conducted on a single NVIDIA RTX 4090 GPU \(24GB\)\.

### 5\.2Result Analysis

Main results\.Table[1](https://arxiv.org/html/2609.22123#S4.T1)reports the main results on Weather2K and MeteoNet\. Overall, StationPDE achieves the best performance across all evaluated variables and both datasets in terms of MSE, demonstrating the effectiveness of introducing station\-oriented surface PDE evolution into multi\-station multivariate forecasting\. Compared with the strongest baseline for each variable, StationPDE reduces MSE by about9\.6%9\.6\\%on average, showing consistent gains across different datasets and weather variables\. Compared with CDPNet, the closest station\-to\-field baseline, StationPDE brings particularly clear improvements on Weather2K pressure and humidity, reducing MSE from9\.8409\.840to4\.4134\.413and from150\.312150\.312to127\.240127\.240, respectively\. This indicates that terrain\-aware field construction and surface PDE evolution are especially beneficial for variables that are sensitive to terrain effects and cross\-variable physical coupling\. For wind variables, StationPDE also consistently outperforms CDPNet, showing that the proposed PDE branch and adaptive fusion provide gains beyond data\-driven continuous field modeling\. Complete MAE results show similar trends and are reported in the appendix\.

Performance across forecast horizons\.Table[2](https://arxiv.org/html/2609.22123#S5.T2)evaluates longer forecast horizons on MeteoNet\. When the forecast horizon extends to4848and9696hours, StationPDE achieves the best MSE on all variables\. This suggests that StationPDE is robust under longer forecast horizons, where accumulated errors make station\-level forecasting more challenging\.

### 5\.3Ablation Study

Table[3](https://arxiv.org/html/2609.22123#S5.T3)reports the ablation results on Weather2K\. All ablation variants perform worse than the full model, confirming the contribution of each component\.w/o PDEremoves the station\-oriented surface PDE branch and causes clear degradation, especially on humidity and wind, showing the importance of explicit surface evolution\.w/o Data\-Drivenleads to the largest overall degradation, indicating that the data\-driven branch is necessary for capturing residual dynamics beyond the PDE process\.w/o Station Residualworsens performance on most variables, suggesting that station\-level residual correction remains important after sampling from the fused field\.w/oℒobs\\mathcal\{L\}\_\{\\mathrm\{obs\}\}yields smaller but consistent degradation, showing that observable\-variable supervision helps align the learned physical state with weather observations\.

Figure 3:\(a\) Average fusion weights of the PDE branch and data\-driven branch for different weather variables\. \(b\) Lead\-time\-wise fusion weights for humidity from 3h to 24h\. \(c\) Relative MSE increase after removing key PDE components\. \(d\) Sensitivity to the altitude weightwaltw\_\{\\mathrm\{alt\}\}in terrain\-aware surface field initialization\.
### 5\.4Visualization Study

Fig\.[3](https://arxiv.org/html/2609.22123#S5.F3)visualizes the adaptive router weights of the two branches\. As shown in Fig\.[3](https://arxiv.org/html/2609.22123#S5.F3)\(a\), the router assigns variable\-specific fusion weights: pressure relies more on the PDE branch, temperature and humidity use both branches, while wind variables rely more on the data\-driven branch\. This suggests that the router adaptively balances explicit surface evolution and learned correction according to variable characteristics\. Fig\.[3](https://arxiv.org/html/2609.22123#S5.F3)\(b\) shows the lead\-time\-wise weights for humidity, where the data\-driven weight gradually increases with the forecast horizon\. This indicates that longer forecasts require stronger adaptive correction to handle accumulated uncertainty\. Fig\.[3](https://arxiv.org/html/2609.22123#S5.F3)\(c\) further analyzes the PDE components\. Removing horizontal diffusion or closure terms increases errors on most variables, while removing the whole PDE branch causes the largest degradation, especially for temperature and wind\. This supports the contribution of both upper\-air inference and residual surface forcing in the surface PDE process\. Fig\.[3](https://arxiv.org/html/2609.22123#S5.F3)\(d\) shows the sensitivity to the altitude weightwaltw\_\{\\mathrm\{alt\}\}\. The performance is stable aroundwalt=5w\_\{\\mathrm\{alt\}\}=5, which is used in our experiments, indicating that a moderate altitude contribution helps initialize the continuous surface field\.

## 6Conclusion

We proposed StationPDE, a station\-oriented surface PDE learning model for multi\-station multivariate weather forecasting\. StationPDE lifts station observations into a terrain\-aware continuous field, evolves it with surface PDE processes, and integrates data\-driven correction through an adaptive router\. Experiments on real\-world datasets demonstrate its effectiveness over representative baselines\.

## References

- Caoet al\.\(2020\)D\. Cao, Y\. Wang, J\. Duan, C\. Zhang, X\. Zhu, C\. Huang, Y\. Tong, B\. Xu, J\. Bai, J\. Tong, and Q\. ZhangSpectral temporal graph neural network for multivariate time\-series forecasting\.InProceedings of the Advances in Neural Information Processing Systems,Vol\.33,Red Hook, NY, USA,pp\. 17766–17778\.Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p1.1)\.
- Hanet al\.\(2023\)J\. Han, H\. Liu, H\. Zhu, and H\. XiongKill two birds with one stone: a multi\-view multi\-adversarial learning approach for joint air quality and weather prediction\.IEEE Transactions on Knowledge and Data Engineering35\(11\),pp\. 11515–11528\.Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.22123#S2.SS1.p1.1)\.
- Huet al\.\(2025\)Y\. Hu, G\. Zhang, P\. Liu, D\. Lan, N\. Li, D\. Cheng, T\. Dai, S\. Xia, and S\. PanTimeFilter: patch\-specific spatial\-temporal graph filtration for time series forecasting\.InProceedings of the 42nd International Conference on Machine Learning,Vol\.267,pp\. 24893–24911\.Cited by:[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Larvoret al\.\(2020\)G\. Larvor, L\. Berthomier, V\. Chabot, B\. L\. Pape, B\. Pradel, and L\. PerezMeteoNet, an open reference weather dataset by meteo france\.Note:[https://github\.com/meteofrance/meteonet](https://github.com/meteofrance/meteonet)Accessed: 2025\-01\-12Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p2.1),[§1](https://arxiv.org/html/2609.22123#S1.p3.1),[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p1.1)\.
- Liuet al\.\(2024\)Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. LongITransformer: inverted transformers are effective for time series forecasting\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 11116–11140\.Cited by:[§2\.1](https://arxiv.org/html/2609.22123#S2.SS1.p2.1),[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Luoet al\.\(2025\)Y\. Luo, S\. Fang, B\. Wu, Q\. Wen, and L\. SunPhysics\-guided learning of meteorological dynamics for weather downscaling and forecasting\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2,New York, NY, USA,pp\. 2010–2020\.Cited by:[§2\.2](https://arxiv.org/html/2609.22123#S2.SS2.p1.1)\.
- Naghashiet al\.\(2025\)V\. Naghashi, M\. Boukadoum, and A\. B\. DialloA multiscale model for multivariate time series forecasting\.Scientific Reports15\(1\),pp\. 1565\.Cited by:[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Skamarocket al\.\(2019\)W\. C\. Skamarock, J\. B\. Klemp, J\. Dudhia, D\. O\. Gill, Z\. Liu, J\. Berner, W\. Wang, J\. G\. Powers, M\. G\. Duda, D\. M\. Barker,et al\.A description of the advanced research wrf version 4\.NCAR tech\. note ncar/tn\-556\+ str145\(10\.5065\)\.Cited by:[§4\.2](https://arxiv.org/html/2609.22123#S4.SS2.p3.1)\.
- Sotoet al\.\(2024\)Á\. M\. Soto, A\. Cervantes, and M\. SolerPhysics\-informed neural networks for high\-resolution weather reconstruction from sparse weather stations\.Open Research Europe4,pp\. 99\.Cited by:[§2\.2](https://arxiv.org/html/2609.22123#S2.SS2.p1.1)\.
- Wanget al\.\(2024a\)S\. Wang, H\. Wu, X\. Shi, T\. Hu, H\. Luo, L\. Ma, J\. Zhang, and J\. ZHOUTimeMixer: decomposable multiscale mixing for time series forecasting\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 38626–38652\.Cited by:[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Wanget al\.\(2026\)Y\. Wang, H\. Wu, J\. Dong, Y\. Liu, C\. Wang, M\. Long, and J\. WangDeep time series models: a comprehensive survey and benchmark\.IEEE Transactions on Pattern Analysis and Machine Intelligence\(\),pp\. 1–20\.Cited by:[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p3.1)\.
- Wanget al\.\(2024b\)Y\. Wang, H\. Wu, J\. Dong, G\. Qin, H\. Zhang, Y\. Liu, Y\. Qiu, J\. Wang, and M\. LongTimeXer: empowering transformers for time series forecasting with exogenous variables\.InProceedings of the Advances in Neural Information Processing Systems,Vol\.37,pp\. 469–498\.Cited by:[§2\.1](https://arxiv.org/html/2609.22123#S2.SS1.p2.1),[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Wuet al\.\(2021\)H\. Wu, J\. Xu, J\. Wang, and M\. LongAutoformer: decomposition transformers with auto\-correlation for long\-term series forecasting\.InProceedings of the Advances in Neural Information Processing Systems,Vol\.34,Red Hook, NY, USA,pp\. 22419–22430\.Cited by:[§2\.1](https://arxiv.org/html/2609.22123#S2.SS1.p1.1)\.
- Wuet al\.\(2023\)H\. Wu, H\. Zhou, M\. Long, and J\. WangInterpretable weather forecasting for worldwide stations with a unified deep model\.Nature Machine Intelligence5\(6\),pp\. 602–611\.Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.22123#S2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Xuet al\.\(2025\)C\. Xu, Y\. Ma, H\. Deng, Y\. Gao, Y\. Wang, K\. Lv, and X\. LiuContinuous diffusive prediction network for multi\-station weather prediction\.InProceedings of the International Joint Conference on Artificial Intelligence,Montreal, Canada,pp\. 6714–6722\.Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.22123#S2.SS1.p3.1),[§4\.1](https://arxiv.org/html/2609.22123#S4.SS1.p2.1),[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Xuet al\.\(2024a\)W\. Xu, F\. Ling, W\. Zhang, T\. Han, H\. Chen, W\. Ouyang, and L\. BaiGeneralizing weather forecast to fine\-grained temporal scales via physics\-ai hybrid modeling\.InProceedings of the Advances in Neural Information Processing Systems,Red Hook, NY, USA\.Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.22123#S2.SS2.p2.1),[§4\.2](https://arxiv.org/html/2609.22123#S4.SS2.p7.2)\.
- Xuet al\.\(2024b\)Z\. Xu, X\. Wei, J\. Hao, J\. Han, H\. Li, C\. Liu, Z\. Li, D\. Tian, and N\. ZhangDGFormer: a physics\-guided station level weather forecasting model with dynamic spatial\-temporal graph neural network\.GeoInformatica28\(3\),pp\. 499–533\.Cited by:[§2\.2](https://arxiv.org/html/2609.22123#S2.SS2.p1.1)\.
- Zenget al\.\(2023\)A\. Zeng, M\. Chen, L\. Zhang, and Q\. XuAre transformers effective for time series forecasting?\.InProceedings of the AAAI conference on artificial intelligence,Vol\.37,pp\. 11121–11128\.Cited by:[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p2.1)\.
- Zhanget al\.\(2023\)Y\. Zhang, M\. Long, K\. Chen, L\. Xing, R\. Jin, M\. I\. Jordan, and J\. WangSkilful nowcasting of extreme precipitation with nowcastnet\.Nature619\(7970\),pp\. 526–532\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06184-4)Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p1.1)\.
- Zhuet al\.\(2023\)X\. Zhu, Y\. Xiong, M\. Wu, G\. Nie, B\. Zhang, and Z\. YangWeather2K: a multivariate spatio\-temporal benchmark dataset for meteorological forecasting based on real\-time observation data from ground weather stations\.InProceedings of The 26th International Conference on Artificial Intelligence and Statistics,pp\. 2704–2722\.Cited by:[§1](https://arxiv.org/html/2609.22123#S1.p3.1),[§5\.1](https://arxiv.org/html/2609.22123#S5.SS1.p1.1)\.

相似文章