UniGIO:基于时空不完整观测的统一生成式全球原位天气建模
摘要
UniGIO是一种生成式框架,基于不完整观测对全球原位天气动态进行建模,统一了预测、插补和生成任务。该框架在Weather-5K数据集上达到了最先进的性能,并在准确性和极端事件捕获方面有显著提升。
arXiv:2609.22217v1 Announce Type: new
Abstract: Global In-situ Observation (GIO) provides fine-scale, direct records of the global weather system from sparse point stations, making it an indispensable source for capturing localized and transient dynamics beyond the reach of satellite gridded data, and playing a critical role in key fields such as numerical weather prediction, disaster prevention, and agriculture. However, GIO exhibits strong spatiotemporal incompleteness, severely impairing accurate and real-time in-situ weather modeling. Unlike existing methods waiting for completed AI-ready data with extra introduced errors, in this work, we explore UniGIO, a novel generative framework for directly modeling global in-situ weather dynamics from native incomplete GIO. By generating missing data from observed ones annotated by masks, it unifies the coexisting forecasting, imputation, and generation under arbitrary missing ratios. Between the missing and observed, UniGIO captures station and region level complementarity through the Observation Mixer and Event Aligner, which diffuse discrete observations into continuous spaces where weather processes naturally span multiple stations. We further establish temporal dependencies with pattern shifts using the Adaptive Temporal Mixer, and track extreme events in chaotic local weather systems through a Mixture-ofExperts structure. Steady and extreme events are adapted in decoder by a Local Refiner. Extensive experiments on the up-todate largest global station weather dataset Weather-5K validate its SOTA performance with 11%, 12%, and 5% advantages on accuracy, fidelity, and extreme event capture, delivering a novel holistic solution for weather modeling in GIO networks.
查看缓存全文
缓存时间: 2026/09/22 09:19
# UniGIO: Unified Generative Global In-situ Weather Modeling from Spatiotemporal Incomplete Observations
Source: [https://arxiv.org/html/2609.22217](https://arxiv.org/html/2609.22217)
Zili Liu⋆Tao HanBen FeiLei BaiChang LiuZhengxia Zou⋆Xiangyang JiWanli OuyangZhenwei Shi††thanks:This work was supported in part by the National Natural Science Foundation of China under Grants U24B20177, 62125102, U25A20401, and 62471014\.††thanks:Songru Yang, Zhenwei Shi and Zhengxia Zou are with the Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University, and with the Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technologies, Ministry of Education, Beihang University, Beijing 100191, China\. Zili Liu, Tao Han, Lei Bai, Wanli Ouyang are with Shanghai Artificial Intelligence Laboratory, Shanghai 200232, China\. Ben Fei is with Department of Information Engineering, The Chinese University of Hong Kong, Hong Kong 999077, China Chang Liu, Xiangyang Ji are with Department of Automation, Tsinghua University, Beijing 100084, China††thanks:Corresponding authors: Zili Liu \(liuzili@pjlab\.org\.cn\), Zhengxia Zou \(zhengxiazou@buaa\.edu\.cn\)\.
###### Abstract
Global In\-situ Observation \(GIO\) provides fine\-scale, direct records of the global weather system from sparse point stations, making it an indispensable source for capturing localized and transient dynamics beyond the reach of satellite gridded data, and playing a critical role in key fields such as numerical weather prediction, disaster prevention, and agriculture\. However, GIO exhibits strong spatiotemporal incompleteness, severely impairing accurate and real\-time in\-situ weather modeling\. Unlike existing methods waiting for completed AI\-ready data with extra introduced errors, in this work, we explore UniGIO, a novel generative framework for directly modeling global in\-situ weather dynamics from native incomplete GIO\. By generating missing data from observed ones annotated by masks, it unifies the coexisting forecasting, imputation, and generation under arbitrary missing ratios\. Between the missing and observed, UniGIO captures station and region level complementarity through the Observation Mixer and Event Aligner, which diffuse discrete observations into continuous spaces where weather processes naturally span multiple stations\. We further establish temporal dependencies with pattern shifts using the Adaptive Temporal Mixer, and track extreme events in chaotic local weather systems through a Mixture\-of\-Experts structure\. Steady and extreme events are adapted in decoder by a Local Refiner\. Extensive experiments on the up\-to\-date largest global station weather dataset Weather\-5K validate its SOTA performance with 11%, 12%, and 5% advantages on accuracy, fidelity, and extreme event capture, delivering a novel holistic solution for weather modeling in GIO networks\.
###### Index Terms:
In\-situ weather observation, Data\-driven weather system modeling, Time series forecasting, imputation, generation
## IIntroduction
Global In\-situ Observation \(GIO\) provides timely and accurate direct observations of the meteorological variables at the finest scale\.\[[1](https://arxiv.org/html/2609.22217#bib.bib32)\]\. Depicting local weather systems’ evolution, volatility, and extreme events, in\-situ weather modeling underpins the assimilation and validation in numerical weather prediction, and serves as a critical tool for various social sectors requiring efficient regional weather services, such as aviation, agriculture, and disaster monitoring\[[2](https://arxiv.org/html/2609.22217#bib.bib34)\]\. Existing methods have already supported weather modeling for key scenarios such as the Winter Olympics\[[3](https://arxiv.org/html/2609.22217#bib.bib6)\]\.
Fig\. 1:Spatiotemporal incompleteness in GIO makes it difficult to model local weather systems and detect extreme events at the finest scale\. We analyzed the station density, annual missing rate, and the AI\-ready rate to demonstrate its severity\. AI\-ready rate: The ratio of non\-missing to all samples when station data is spliced in windows\.Fig\. 2:\(a\) Mask\-driven observation\-to\-missing generation, \(b\) UniGIO achieves leading performance while maintaining lightweight in data and model scaleGIO naturally provides lightweight sub\-grid, point\-level observations of instantaneous atmospheric dynamics\. Although mainstream heavy grid\-based reanalysis and satellite data\[[4](https://arxiv.org/html/2609.22217#bib.bib7),[5](https://arxiv.org/html/2609.22217#bib.bib8),[6](https://arxiv.org/html/2609.22217#bib.bib9),[7](https://arxiv.org/html/2609.22217#bib.bib56),[8](https://arxiv.org/html/2609.22217#bib.bib57)\]offer global dense coverage, they rely on interpolation to sub\-grid scales, which leads to extreme event and transient signal loss with data inefficiency\. GIO serves as a primary source for local weather services and plays a critical role in bridging global forecasts toward local dynamics, providing a foundation for fine\-scale weather modeling\[[9](https://arxiv.org/html/2609.22217#bib.bib61),[10](https://arxiv.org/html/2609.22217#bib.bib60)\]\.
Nevertheless, GIO exhibits strong spatiotemporal incompleteness, a common and challenging issue in real\-world scenarios\[[11](https://arxiv.org/html/2609.22217#bib.bib28),[12](https://arxiv.org/html/2609.22217#bib.bib29)\]\. In Fig\.[1](https://arxiv.org/html/2609.22217#S1.F1), for Weather\-5K dataset\[[13](https://arxiv.org/html/2609.22217#bib.bib37)\]across 5k\+ global in\-situ stations in 10 years, the data missing rate is 27\.5%, yet AI\-ready samples with complete 24h\-120h window are less than 25%\-10%, manifested in three patterns: \(1\) data missing due to sensor failures or transmission loss; \(2\) uneven spatial data distribution due to geographical constraints; \(3\) Temporal resolution differences among various stations\[[1](https://arxiv.org/html/2609.22217#bib.bib32),[14](https://arxiv.org/html/2609.22217#bib.bib33)\]\. This inherent, inevitable data gap leaves the GIO system ill\-suited to cutting\-edge AI frameworks\.
Existing methods in GIO systems remain limited\. Forecasting methods\[[3](https://arxiv.org/html/2609.22217#bib.bib6),[15](https://arxiv.org/html/2609.22217#bib.bib27)\]often bypass observation gaps by relying on non\-learnable pre\-processing or cascaded models to construct AI\-ready complete series, which introduces data distortion, time delays, cascading errors, and sensitivity to task variations such as different prediction windows\. time series generation offers a possible way to handle different missing patterns\[[16](https://arxiv.org/html/2609.22217#bib.bib35),[17](https://arxiv.org/html/2609.22217#bib.bib36)\], but often yields over\-smoothed or noisy results\. Recent observation foundation models\[[18](https://arxiv.org/html/2609.22217#bib.bib55),[7](https://arxiv.org/html/2609.22217#bib.bib56),[8](https://arxiv.org/html/2609.22217#bib.bib57)\]theoretically support arbitrary\-location modeling, yet they operate at the global scale and still treat GIO modeling as latent interpolation, leading to averaged results, redundant multi\-source inputs, and unaffordable computation for in\-situ weather services\.
Beyond accuracy degradation, incomplete GIO can obscure local extreme weather events, which account for up to 90% of global extremes and have caused 70% of 2 million deaths and 3\.64 trillion in losses over the past 50 years\[[19](https://arxiv.org/html/2609.22217#bib.bib5),[20](https://arxiv.org/html/2609.22217#bib.bib4)\]\. The risk is amplified in Africa, where weather\-related mortality is the highest, station density is only7e−6/km27e^\{\-6\}/km^\{2\}\(∼1/84\\sim 1/84of Germany\), and only 22% of stations meet the GBON standard\[[21](https://arxiv.org/html/2609.22217#bib.bib1),[22](https://arxiv.org/html/2609.22217#bib.bib2),[23](https://arxiv.org/html/2609.22217#bib.bib3),[14](https://arxiv.org/html/2609.22217#bib.bib33)\]\. As stations are more likely to fail under extremes, it is urgent to capture abrupt local events and hidden risks from incomplete observations\[[24](https://arxiv.org/html/2609.22217#bib.bib31),[19](https://arxiv.org/html/2609.22217#bib.bib5)\]\.
These unresolved challenges motivate us to propose a novel unified GIO\-centric modeling framework named UniGIO, operating directly on interpolation\-free distributed weather stations, emphasizing 3 insights: \(1\) GIO complementarity; \(2\) Extreme patterns; \(3\) Lightweight\. UniGIO regards observation gaps as coexisting forecasting, imputation, and generation tasks, and unifies them within a mask\-driven generation from observation to missing parts as Fig\.[2](https://arxiv.org/html/2609.22217#S1.F2)\(a\)\. Specifically, for complementarity, an Observation Mixer is designed to complement missingness in series observations and continuous regions where weather events evolve; an Event Aligner is added to alleviate the delays of the same event reaching different stations\. To handle extreme mutations, an Adaptive Temporal Mixer is designed, leveraging an uncertainty\-weighted state space model to establish robust temporal dependencies, and a Pattern Decoupled MoE to group global affinity and separate local heterogeneity\. A Local Refiner is added to maintain local continuity\. For lightweight design, UniGIO is efficient in both data and model scale\. It is trained directly on 40G sparse GIO data instead of PB\-level dense global observations, and can be compressed to fewer than 1M parameters while still achieving leading performance, as Fig\.[2](https://arxiv.org/html/2609.22217#S1.F2)\(b\)\. All tasks are unified within a conditional VAE \(CVAE\)\[[25](https://arxiv.org/html/2609.22217#bib.bib40)\], which quantifies extreme event risks through variance, supporting its assessment and controllable generation\.
We conducted quantitative and qualitative experiments under both native and curated spatiotemporal incompleteness to benchmark the in\-situ weather modeling task on the up\-to\-date largest GIO dataset Weather\-5K\[[13](https://arxiv.org/html/2609.22217#bib.bib37)\], recording hourly temperature, dew point, wind direction, wind rate, and sea\-level pressure from 5k\+ surface weather stations worldwide, covering 6 continents and 10 years\. UniGIO uniformly handles 7 tasks, including fix/variable\-window forecasting, random&period/resolution imputation, variable/in\-region/cross\-region station generation, and outperforms baselines by an average 9% across 7 metrics, respectively 11%, 12%, and 5% improvement on accuracy, fidelity, and extreme event capture\. Our contributions are as follows:
- •We propose UniGIO, a novel unified framework for versatile and lightweight global in\-situ weather modeling from incomplete GIO through an observation\-to\-missing generation to advance all\-time in\-situ service, even when real\-world stations are absent or inoperable\.
- •We design Observation Mixer, Event Aligner to capture station and region complement; Adaptive Temporal Mixer, Local Refiner are designed for extreme pattern adaptation; CVAE is specially adopted for observation constraints and evaluate extreme risks\.
- •We benchmark across 5k\+ surface weather stations worldwide, covering a 10\-year period to establish strong baselines for in\-situ weather modeling\. Extensive experiments validate our consistent and outstanding performance across tasks, regions, and mask formats\.
## IIRelated Work
### II\-AData\-driven Weather Prediction
Since 2022, the AI and atmospheric science communities have seen a rapidly growing interest in data\-driven Numerical Weather Prediction \(NWP\) models like Pangu\-Weather\[[4](https://arxiv.org/html/2609.22217#bib.bib7)\], GraphCast\[[5](https://arxiv.org/html/2609.22217#bib.bib8)\], GenCast\[[6](https://arxiv.org/html/2609.22217#bib.bib9)\], and FuXi\[[26](https://arxiv.org/html/2609.22217#bib.bib24)\]et al running on gridded reanalysis data \(e,g\.,0\.25∘0\.25^\{\\circ\}and0\.09∘0\.09^\{\\circ\}resolution\)\. In the case that these large\-scale, region\-averaged methods may not match the station level weather systems and can ignore highly destructive local extreme weather, some initial attempts, such as Corrformer\[[3](https://arxiv.org/html/2609.22217#bib.bib6)\]and WSSM\[[15](https://arxiv.org/html/2609.22217#bib.bib27)\], have treated station weather forecasting as an independent task and achieved promising results\. However, these methods operate on spatiotemporal complete data, remaining a gap to the incomplete GIO, which greatly limits their real\-world applications\.
Recently, spatiotemporal weather foundation models\[[27](https://arxiv.org/html/2609.22217#bib.bib25),[28](https://arxiv.org/html/2609.22217#bib.bib26)\]and observation foundation models\[[18](https://arxiv.org/html/2609.22217#bib.bib55),[7](https://arxiv.org/html/2609.22217#bib.bib56),[8](https://arxiv.org/html/2609.22217#bib.bib57)\]have shown the potential to unify station and grid weather modeling at arbitrary locations\. Nevertheless, this paradigm remains poorly suited to GIO\. It relies on massive multi\-source data and intensive computation for global\-scale autoregressive forecasting, which is incompatible with in\-situ weather applications\. Also, although they implicitly incorporate data assimilation into grid\-like structured latent spaces to support observation inputs, their GIO forecasts are still derived from interpolation over smoothed latent fields\. As a result, they remain fundamentally close to grid\-based forecasting and inherit its key limitations\. These challenges indicate that a unified weather modeling framework tailored to GIO is not only indispensable for practical in\-situ services, but also essential for bridging the gap in the grid\-station and global\-local weather modeling\.
### II\-BMultivariate Time Series Forecasting and Imputation
Forecasting and imputation are two fundamental tasks in multivariate time series modeling\. Forecasting aims to infer future dynamics from historical observations, while imputation focuses on recovering missing values from partially observed sequences\. In early studies, CNN\-based methods\[[29](https://arxiv.org/html/2609.22217#bib.bib30)\]were adopted to extract local temporal patterns, and RNN\-based methods\[[30](https://arxiv.org/html/2609.22217#bib.bib44)\]were widely used to model sequential dependencies\. Recently, Transformer\-based architectures\[[31](https://arxiv.org/html/2609.22217#bib.bib51)\]have enabled more flexible long\-range correlation modeling; this trend is consistent with non\-local attention\[[32](https://arxiv.org/html/2609.22217#bib.bib11)\], hierarchical vision transformers\[[33](https://arxiv.org/html/2609.22217#bib.bib12),[34](https://arxiv.org/html/2609.22217#bib.bib14)\], and deformable attention\[[35](https://arxiv.org/html/2609.22217#bib.bib13)\]\. State space models\[[36](https://arxiv.org/html/2609.22217#bib.bib52)\]further introduce temporal inductive biases and recurrent\-like computation to construct efficient and robust long\-term dependencies\. Imputation is often regarded as a preprocessing step for slightly incomplete data\. Classical non\-learnable methods, such as interpolation\[[37](https://arxiv.org/html/2609.22217#bib.bib47)\], regression\[[38](https://arxiv.org/html/2609.22217#bib.bib48)\], and EM algorithms\[[39](https://arxiv.org/html/2609.22217#bib.bib49)\], are simple and efficient, but they usually rely on smoothness or distributional assumptions and struggle to recover complex non\-stationary sequences\. To improve imputation quality, recent learning\-based frameworks\[[40](https://arxiv.org/html/2609.22217#bib.bib50),[41](https://arxiv.org/html/2609.22217#bib.bib45)\]introduce neural structures such as recurrent networks, attention mechanisms, and graph neural networks to exploit temporal continuity and cross\-variable dependencies\.
More recent efforts attempt to perform forecasting under missing observations\[[42](https://arxiv.org/html/2609.22217#bib.bib53),[43](https://arxiv.org/html/2609.22217#bib.bib54)\], indicating the growing need for models that are robust to incomplete inputs\. However, most existing methods still treat forecasting and imputation as separate tasks, where imputation is either performed as an independent preprocessing stage or forecasting is designed for a specific missing pattern\. This separation limits their applicability to Global In\-situ Observation \(GIO\), where missingness is not slight or fixed but arbitrarily coexists across variables, stations, regions, and forecast horizons\. From a broader learning perspective, current data synthesis and reconstruction methods have shown that imperfect observations and distributions can substantially impair generalization and reliability\[[44](https://arxiv.org/html/2609.22217#bib.bib15),[45](https://arxiv.org/html/2609.22217#bib.bib16),[46](https://arxiv.org/html/2609.22217#bib.bib17)\]\. To solve these problems, our work formulates these incomplete patterns with unified observation mask and models forecasting, imputation, and generation within a single framework\.
### II\-CGenerative Time Series Modeling
Generative time series modeling aims to synthesize realistic time series by learning the underlying data distribution, often generating sequences from random noise\. Early studies mainly explored GAN\-based\[[47](https://arxiv.org/html/2609.22217#bib.bib38)\]and VAE\-based\[[48](https://arxiv.org/html/2609.22217#bib.bib41)\]frameworks, which respectively model temporal distributions through adversarial learning and latent\-variable reconstruction\. More recently, diffusion models\[[49](https://arxiv.org/html/2609.22217#bib.bib21),[50](https://arxiv.org/html/2609.22217#bib.bib22),[51](https://arxiv.org/html/2609.22217#bib.bib23)\]have become a dominant paradigm due to their stable training and strong distribution modeling capability in related applications\. These generative paradigms\[[52](https://arxiv.org/html/2609.22217#bib.bib18),[53](https://arxiv.org/html/2609.22217#bib.bib19),[54](https://arxiv.org/html/2609.22217#bib.bib20)\]focusing on data distribution learning, laid the foundation for task unification\. By controlling generation conditions, generative methods can naturally support forecasting, imputation, and generation simultaneously under arbitrary masks, providing a flexible formulation for incomplete time series modeling\.
Despite this natural unification capability, most existing generative time series methods are still primarily designed for unconditional or weakly conditional generation\. They focus on fitting the global data manifold, which becomes problematic for Global In\-situ Observation \(GIO\), where the global weather manifold is often dominated by smooth evolution, whereas local dynamics are frequently affected by strong extreme noise\. Simply landing on this manifold may lead to over\-smoothed predictions that miss local extremes; worse, under sparse and noisy conditions, models often fit spurious noise instead of real dynamics, resulting in uncontrolled fluctuations\.
We empirically apply advanced generative time series methods, such as diffusion models, to GIO and find that their conditional information is often weakly exploited or even largely discarded, resulting in degraded performance\. This indicates that available observations serve as critical anchors for preserving realistic local states and constraining physically plausible generation\. Motivated by this, UniGIO adopts a CVAE that tightly anchors valid observations to the latent prior mean and explicitly leverages the variance to model the uncertainty of extreme events, thereby balancing smooth temporal evolution and localized abrupt mutations in a decoupled manner\.
## IIIMethod
Fig\. 3:Overall pipeline of UniGIO\. The framework is organized as a lightweight CVAE, where key designs for GIO complementation are marked inblueand key designs for extreme pattern capture are marked ingreen\.### III\-ATask Formulation
To unify multiple in\-situ weather modeling tasks within a single framework, we use an observation mask to represent different incomplete patterns\. Within the time windowTT, weather series withNNadjacent stations andCCvariables can be denoted as𝐗∈ℝN×T×C\\mathbf\{X\}\\in\\mathbb\{R\}^\{N\\times T\\times C\}\. The observation mask𝐌n∈ℝT×C\\mathbf\{M\}\_\{n\}\\in\\mathbb\{R\}^\{T\\times C\}fornn\-th station can be obtained as:
𝐌n,t,c∈\{0,1\},\\displaystyle\\mathbf\{M\}\_\{n,t,c\}\\in\\\{0,1\\\},s\.t\.𝐌n,t,c=𝟏\(Xn,t,cis observed\),\\displaystyle\\text\{s\.t\.\}\\ \\mathbf\{M\}\_\{n,t,c\}=\\mathbf\{1\}\(X\_\{n,t,c\}\\text\{ is observed\}\),\(1\)∀t∈\{1,…,T\},c∈\{1,…,C\}\\displaystyle\\forall t\\in\\\{1,\\dots,T\\\},c\\in\\\{1,\\dots,C\\\}
Given weather series𝐗\\mathbf\{X\}, observation mask𝐌\\mathbf\{M\}, and metadata𝐈∈ℝT×dm\\mathbf\{I\}\\in\\mathbb\{R\}^\{T\\times d\_\{m\}\}, our task is to train a modelfθ\(⋅\)f\_\{\\theta\}\(\\cdot\)to generate the missing data from the incomplete input𝐗M\\mathbf\{X\}\_\{M\}, which can be formulated as a self\-supervised learning task:
minθℒ\(𝐗,fθ\(𝐗M,𝐌,𝐈\)\)\\min\_\{\\theta\}\\mathcal\{L\}\(\\mathbf\{X\},f\_\{\\theta\}\(\\mathbf\{X\}\_\{M\},\\mathbf\{M\},\\mathbf\{I\}\)\)\(2\)
Through the various designs of observation masks, we can cover a series of tasks from broad to strict\. Taking the forecast mask with look\-back window sizeiias an example:
𝐌n,t,c=1\(t≤i\),∀t∈\{1,…,T\},c∈\{1,…,C\}\\mathbf\{M\}\_\{n,t,c\}=1\(t\\leq i\),\\forall t\\in\\\{1,\\dots,T\\\},c\\in\\\{1,\\dots,C\\\}\(3\)
Ifiiis randomly sampled, variable\-windows can be achieved; fixed\-windows are implemented by fixedii\.
In this work, we have identified 7 coexisting forecasting, imputation, and generation tasks: 1\) fixed\-window forecasting, 2\) variable\-window forecasting, 3\) random&period imputation, 4\) resolution imputation, 5\) variable generation, 6\) in\-region station generation, and 7\) cross\-region station generation\. Mask formulations are as follows:
Fixed\-window/Variable\-window Forecasting:
𝐌n,t,c=1\(t≤i\)\\mathbf\{M\}\_\{n,t,c\}=1\(t\\leq i\)\(4\)Ifiiis randomly sampled, variable\-window forecasting can be achieved; fixed\-window forecasting is implemented by fixedii\. Notably, the missing future in the forecast task should not exhibit complementarity\. To achieve this,NNstations are aligned to Station 0 to determine the minimum mask length\.
Random&period Imputation:
𝐌n,t,c=\{0,\(t,c\)∈𝒯point\(ℙ=ppoint\)∨t∈⋃l=1L\[al,bl\]1,otherwise\\mathbf\{M\}\_\{n,t,c\}=\\begin\{cases\}0,&\(t,c\)\\in\\mathcal\{T\}\_\{\\text\{point\}\}\\ \(\\mathbb\{P\}=p\_\{\\text\{point\}\}\)\\\\ &\\lor t\\in\\bigcup\_\{l=1\}^\{L\}\[a\_\{l\},b\_\{l\}\]\\\\ 1,&\\text\{otherwise\}\\end\{cases\}\(5\)where𝒯point⊆\{\(1,1\),…,\(T,C\)\}\\mathcal\{T\}\_\{\\text\{point\}\}\\subseteq\\\{\(1,1\),\\dots,\(T,C\)\\\}is the set of isolated points randomly selected from time windowTTwithCCvariables, each point is selected with probabilityppointp\_\{\\text\{point\}\}for random imputation\.⋃l=1L\[al,bl\]\\bigcup\_\{l=1\}^\{L\}\[a\_\{l\},b\_\{l\}\]is the union of randomly sampledLLsegments, and the length of each segment\[al,bl\]\[a\_\{l\},b\_\{l\}\]is also randomly sampled for period imputation\. Notably, considering the independence between variable\-specific sensors, we randomly trigger variable unmasking to remove mask for randomly select variables within random sampled data blocks\.
This masking strategy accounts for scenarios where stations or sensors malfunction over a period of time or data is randomly unavailable\.
Resolution Imputation:
𝐌n,t=\{0,t≢0\(modS\)1,t≡0\(modS\)\\mathbf\{M\}\_\{n,t\}=\\begin\{cases\}0,&t\\not\\equiv 0\\pmod\{S\}\\\\ 1,&t\\equiv 0\\pmod\{S\}\\end\{cases\}\(6\)whereSSis randomly sampled resolution step size, time stepsttnot an integer multiple ofSSare masked to simulate variations in the temporal resolution\.
Variable Generation
𝐌n,c=\{0,c=c∗1,c≠c∗\\mathbf\{M\}\_\{n,c\}=\\begin\{cases\}0,&c=c^\{\*\}\\\\ 1,&c\\neq c^\{\*\}\\end\{cases\}\(7\)wherec∗c^\{\*\}is a single variable to mask randomly chosen fromCC\. This reflects the variation in sensor types for different variables among stations\.
In\-region/Cross\-region station generation
𝐌n=\{0,n=n∗,n∗∈𝒮train\(in\-region\),0,n=n∗,n∗∈𝒮∖𝒮train\(cross\-region\),1,n≠n∗\.\\mathbf\{M\}\_\{n\}=\\begin\{cases\}0,&n=n^\{\*\},\\ n^\{\*\}\\in\\mathcal\{S\}\_\{\\mathrm\{train\}\}\\quad\\text\{\(in\-region\)\},\\\\ 0,&n=n^\{\*\},\\ n^\{\*\}\\in\\mathcal\{S\}\\setminus\\mathcal\{S\}\_\{\\mathrm\{train\}\}\\quad\\text\{\(cross\-region\)\},\\\\ 1,&n\\neq n^\{\*\}\.\\end\{cases\}\(8\)wheren∗n^\{\*\}is randomly selected station index,𝒮\\mathcal\{S\}is all stations,𝒮train\\mathcal\{S\}\_\{\\text\{train\}\}is training set \(in\-region scope\), and∖\\setminusis set minus\. This masking strategy is primarily used for station interpolation and extrapolation to overcome observational gaps in spatial scales, which aligns with the core objectives of in\-region and cross\-region station generation tasks\. It can also represent scenarios where all data within a certain data block is missing\.
### III\-BUniGIO
The overall framework of UniGIO is illustrated in Fig\.[3](https://arxiv.org/html/2609.22217#S3.F3)\. It consists of two main stages: observation\-driven complementary initialization and temporal dependency modeling\.
In the first stage, the Observation Mixer \(OM\) integrates available GIO records from station, variable, and regional perspectives, while the Event Aligner \(EA\) further compensates for temporal phase shifts among neighboring stations\. This stage aims to propagate sparse observed evidence to missing entries and provide a more complete and physically meaningful initialization\. In the second stage, the Adaptive Temporal Mixer \(ATM\) models temporal dependencies in the initialized sequence with uncertainty\-aware state transitions and pattern\-decoupled MoE routing, enabling the model to adaptively handle drifting stable\-extreme patterns under different observation reliabilities\. Meanwhile, the Local Refiner \(LR\) reconstructs local temporal features from sampled latent representations through multi\-receptive\-field convolutions and selective gating, preserving local continuity weakened by MoE routing and CVAE sampling\. These modules are organized within a simple CVAE framework, enabling unified and extreme event\-controllable forecasting, imputation, and generation under arbitrary observation masks\.
Algorithm 1UniGIO Pipeline0:Incomplete sequence
𝐗M\\mathbf\{X\}\_\{M\}, mask
𝐌\\mathbf\{M\}, metadata
𝐈\\mathbf\{I\}, ground truth
𝐗\\mathbf\{X\}
0:Completed sequence
𝐗^\\hat\{\\mathbf\{X\}\}
1:Prior branch:
2:
𝐟C←OM\(𝐗M,𝐌,𝐈\)\\mathbf\{f\}\_\{\\text\{C\}\}\\leftarrow\\operatorname\{OM\}\(\\mathbf\{X\}\_\{M\},\\mathbf\{M\},\\mathbf\{I\}\)
3:
𝐟T←ATM\(𝐟C\)\\mathbf\{f\}\_\{\\text\{T\}\}\\leftarrow\\operatorname\{ATM\}\(\\mathbf\{f\}\_\{\\text\{C\}\}\)
4:
\(𝝁p,𝝈p\)←PriorMLP\(𝐟T\)\(\\boldsymbol\{\\mu\}\_\{p\},\\boldsymbol\{\\sigma\}\_\{p\}\)\\leftarrow\\operatorname\{PriorMLP\}\(\\mathbf\{f\}\_\{\\text\{T\}\}\)
5:
pψ\(𝐳\|𝐗M,𝐌,𝐈\)=𝒩\(𝝁p,𝝈p2\)p\_\{\\psi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\_\{M\},\\mathbf\{M\},\\mathbf\{I\}\)=\\mathcal\{N\}\(\\boldsymbol\{\\mu\}\_\{p\},\\boldsymbol\{\\sigma\}\_\{p\}^\{2\}\)
6:Posterior branch:
7:
𝐟S←EA\(𝐗,𝐈\)\\mathbf\{f\}\_\{\\text\{S\}\}\\leftarrow\\operatorname\{EA\}\(\\mathbf\{X\},\\mathbf\{I\}\)
8:
𝐟V←Variable−wise−attention\(𝐟S\)\\mathbf\{f\}\_\{\\text\{V\}\}\\leftarrow\\operatorname\{Variable\-wise\-attention\}\(\\mathbf\{f\}\_\{\\text\{S\}\}\)
9:
𝐟T′←Bi−SSM\(𝐟V\)\\mathbf\{f\}\_\{\\text\{T\}\}^\{\\prime\}\\leftarrow\\operatorname\{Bi\-SSM\}\(\\mathbf\{f\}\_\{\\text\{V\}\}\)
10:
\(𝝁q,𝝈q\)←PostMLP\(𝐟T′\)\(\\boldsymbol\{\\mu\}\_\{q\},\\boldsymbol\{\\sigma\}\_\{q\}\)\\leftarrow\\operatorname\{PostMLP\}\(\\mathbf\{f\}\_\{\\text\{T\}\}^\{\\prime\}\)
11:
qϕ\(𝐳\|𝐗\)=𝒩\(𝝁q,𝝈q2\)q\_\{\\phi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\)=\\mathcal\{N\}\(\\boldsymbol\{\\mu\}\_\{q\},\\boldsymbol\{\\sigma\}\_\{q\}^\{2\}\)
12:Generation:
13:Sample
ϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
14:Set pointwise sampling strength
𝜶∈ℝN×T×C\\boldsymbol\{\\alpha\}\\in\\mathbb\{R\}^\{N\\times T\\times C\}
15:
𝐳←𝝁p\+𝜶⊙𝝈p⊙ϵ\\mathbf\{z\}\\leftarrow\\boldsymbol\{\\mu\}\_\{p\}\+\\boldsymbol\{\\alpha\}\\odot\\boldsymbol\{\\sigma\}\_\{p\}\\odot\\boldsymbol\{\\epsilon\}
16:Adjust
𝜶\\boldsymbol\{\\alpha\}on target positions to control the intensity of extreme event generation
17:Decode
𝐗^←pθ\(LR\(𝐳\),𝐟T\)\\hat\{\\mathbf\{X\}\}\\leftarrow p\_\{\\theta\}\(\\operatorname\{LR\}\(\\mathbf\{z\}\),\\mathbf\{f\}\_\{\\text\{T\}\}\)
Unlike existing time series methods that directly model incomplete sequences, UniGIO first recovers station and region level complementary representations before temporal modeling\. This decoupled design reduces the burden of simultaneously inferring missing states and learning temporal evolution, allowing ATM to focus on a more complete and observation\-constrained sequence\. As a result, UniGIO preserves observation evidence while enabling efficient recovery of complete weather sequences within a lightweight unified framework\. The overall training and inference pipeline is summarized in Algorithm[1](https://arxiv.org/html/2609.22217#alg1)\.
Fig\. 4:Detailed designs in Observation Mixer: \(a\) Station\-Centric Field Complement to diffuse discrete station level observations into continuous local space with field and memory tokens; \(b\) Event Aligner corrects inter\-station temporal delays via frequency\-correlation alignment\.
### III\-CObservation Mixer and Event Aligner
The Observation Mixer \(OM\) in Fig\.[3](https://arxiv.org/html/2609.22217#S3.F3)\(b\) is designed to initialize unobserved positions by exploiting the complementarity hidden in available GIO records\. Since GIO is sparse and irregular, a missing value cannot be reliably recovered from broken temporal information alone\. Instead, useful evidence may come from three levels: variables observed at the same station, neighboring stations at the same time, and the regional weather field formed by multiple stations\. Therefore, OM performs self\-complementarity, adjacent\-complementarity, and field\-complementarity in a progressive manner\. The first two levels recover local station\-wise information, while the field\-level latent expands discrete station observations into a local continuous space where weather processes naturally evolve across multiple stations\.
For self\- and adjacent\-complementarity, the OM applies self\-attentionMHSAvar\(⋅\)\\operatorname\{MHSA\}\_\{\\text\{var\}\}\(\\cdot\)over variables𝐟varst\\mathbf\{f\}\_\{\\text\{vars\}\}^\{t\}in each station feature𝐟\\mathbf\{f\}, and aggregates neighboring stations via a distance\-based graph attentionGATsta\(⋅\)\\operatorname\{GAT\}\_\{\\text\{sta\}\}\(\\cdot\)with adjacency𝐀\\mathbf\{A\}calculated bydist\(⋅\)\\operatorname\{dist\}\(\\cdot\)with threshτ\\tauat time stept=1,…,Tt=1,\\dots,T\. For these station\-wise complementarity, the calculation is given by:
𝐟=MLP\(𝐗M\)\+STE\(𝐗M\)∈ℝN×T×C×d,\\displaystyle\\mathbf\{f\}=\\operatorname\{MLP\}\(\\mathbf\{X\}\_\{M\}\)\+\\operatorname\{STE\}\(\\mathbf\{X\}\_\{M\}\)\\in\\mathbb\{R\}^\{N\\times T\\times C\\times d\},\(9\)𝐟varst=𝐟:,t,:,:∈ℝN×C×d,t∈\{1,…,T\},\\displaystyle\\mathbf\{f\}\_\{\\text\{vars\}\}^\{t\}=\\mathbf\{f\}\_\{:,t,:,:\}\\in\\mathbb\{R\}^\{N\\times C\\times d\},\\quad t\\in\\\{1,\\dots,T\\\},dij=dist\(𝐬i,𝐬j\),\\displaystyle d\_\{ij\}=\\operatorname\{dist\}\(\\mathbf\{s\}\_\{i\},\\mathbf\{s\}\_\{j\}\),𝐀ij=\{exp\(−dij/τ\),dij≤τ,0,dij\>τ,\\displaystyle\\mathbf\{A\}\_\{ij\}=\\begin\{cases\}\\exp\(\-d\_\{ij\}/\\tau\),&d\_\{ij\}\\leq\\tau,\\\\ 0,&d\_\{ij\}\>\\tau,\\end\{cases\}𝐟selft=MHSAvar\(𝐟varst\),\\displaystyle\\mathbf\{f\}\_\{\\text\{self\}\}^\{t\}=\\operatorname\{MHSA\}\_\{\\text\{var\}\}\\big\(\\mathbf\{f\}\_\{\\text\{vars\}\}^\{t\}\\big\),𝐟adjt=GATsta\(𝐟selft,𝐀\),t=1,…,T\.\\displaystyle\\mathbf\{f\}\_\{\\text\{adj\}\}^\{t\}=\\operatorname\{GAT\}\_\{\\text\{sta\}\}\\big\(\\mathbf\{f\}\_\{\\text\{self\}\}^\{t\},\\mathbf\{A\}\\big\),\\quad t=1,\\dots,T\.where𝐟self\\mathbf\{f\}\_\{\\text\{self\}\}are self variables complementary features,𝐟adjt\\mathbf\{f\}\_\{\\text\{adj\}\}^\{t\}are adjacent station complement features,ddis latent size\.
Then, beyond station\-wise complementarity, our field\-complementarity is detailed in Fig\.[4](https://arxiv.org/html/2609.22217#S3.F4)\(a\)\. Instead of projecting observations onto dense global grid at low resolution, OM constructs a small number of local scale field tokens𝐳field\\mathbf\{z\}\_\{\\text\{field\}\}over gridded \(Grid\(⋅\)\\operatorname\{Grid\}\(\\cdot\)\) input station regions𝐆s\\mathbf\{G\}\_\{s\}with boundaryϕ\\phiandλ\\lambda,δ\\deltais padding\. Unlike heavy grids for global data, These tokens serve as lightweight regional queries that preserve station\-centered dynamics in a continuous space with fewH×WH\\times Wtokens\. The field tokens are initialized with spatiotemporal encodings and interact with the adjacent\-complemented station features through the cross\-attention Field\-formerFf\(⋅\)\\operatorname\{Ff\}\(\\cdot\), which expands sparse GIO records into a continuous latent field\. Meanwhile,KKlearnable memory tokens𝐳mem\\mathbf\{z\}\_\{\\text\{mem\}\}are introduced to store global and event\-level context shared across regional fields, enabling the model to maintain broader weather consistency beyond local station neighborhoods\. Notably, both station tokens and field tokens share the same spatiotemporal encoding functionSTE\(⋅\)\\operatorname\{STE\}\(\\cdot\), which ensures that information propagation between discrete observations and continuous field latents remains spatially and temporally consistent throughout the process\. For field\-complementarity, calculation is given by:
ϕmin=minilati−δ,ϕmax=maxilati\+δ,\\displaystyle\\phi\_\{\\min\}=\\min\_\{i\}\\text\{lat\}\_\{i\}\-\\delta,\\quad\\phi\_\{\\max\}=\\max\_\{i\}\\text\{lat\}\_\{i\}\+\\delta,\(10\)λmin=miniloni−δ,λmax=maxiloni\+δ,\\displaystyle\\lambda\_\{\\min\}=\\min\_\{i\}\\text\{lon\}\_\{i\}\-\\delta,\\quad\\lambda\_\{\\max\}=\\max\_\{i\}\\text\{lon\}\_\{i\}\+\\delta,𝐆s=Grid\(\[ϕmin,ϕmax\],\[λmin,λmax\],H,W\),\\displaystyle\\mathbf\{G\}\_\{s\}=\\operatorname\{Grid\}\\big\(\[\\phi\_\{\\min\},\\phi\_\{\\max\}\],\[\\lambda\_\{\\min\},\\lambda\_\{\\max\}\],H,W\\big\),𝐳field=STE\(𝐆s\)∈ℝH×W×d,\\displaystyle\\mathbf\{z\}\_\{\\text\{field\}\}=\\operatorname\{STE\}\(\\mathbf\{G\}\_\{s\}\)\\in\\mathbb\{R\}^\{H\\times W\\times d\},𝐳mem∈ℝK×d,𝐳memis learnable,\\displaystyle\\mathbf\{z\}\_\{\\text\{mem\}\}\\in\\mathbb\{R\}^\{K\\times d\},\\quad\\mathbf\{z\}\_\{\\text\{mem\}\}\\ \\text\{is learnable\},\[𝐳field′,𝐳mem′\]=Ff\(q=\[𝐳field,𝐳mem\],k/v=\{𝐟adjt\}t=1T\),\\displaystyle\[\\mathbf\{z\}\_\{\\text\{field\}\}^\{\\prime\},\\mathbf\{z\}\_\{\\text\{mem\}\}^\{\\prime\}\]=\\operatorname\{Ff\}\\big\(q=\[\\mathbf\{z\}\_\{\\text\{field\}\},\\mathbf\{z\}\_\{\\text\{mem\}\}\],k/v=\\\{\\mathbf\{f\}\_\{\\text\{adj\}\}^\{t\}\\\}\_\{t=1\}^\{T\}\\big\),
After the Field\-former aggregates station evidence into field and memory tokens, CNNs are applied to enhance spatial continuity within the local field latent\. Finally, OM uses masked tokens as queries in inverse cross\-attentionICAtt\(⋅\)\\operatorname\{ICAtt\}\(\\cdot\), projecting the continuous field representation back to the missing GIO positions and obtaining the complemented feature𝐟C\\mathbf\{f\}\_\{\\text\{C\}\}\. The calculation is given by:
𝐳ref=CNNs\(\[𝐳field′,𝐳mem′\]\),\\displaystyle\\mathbf\{z\}\_\{\\text\{ref\}\}=\\operatorname\{CNNs\}\\big\(\[\\mathbf\{z\}\_\{\\text\{field\}\}^\{\\prime\},\\mathbf\{z\}\_\{\\text\{mem\}\}^\{\\prime\}\]\\big\),\(11\)𝐟C=ICAtt\(q=𝐟mask,k/v=𝐳ref\)\.\\displaystyle\\mathbf\{f\}\_\{\\text\{C\}\}=\\operatorname\{ICAtt\}\\big\(q=\\mathbf\{f\}\_\{\\text\{mask\}\},k/v=\\mathbf\{z\}\_\{\\text\{ref\}\}\\big\)\.
Although OM captures spatial and variable complementarity at the same time step, weather events usually reach different stations with temporal delays\. Directly aggregating neighboring stations without considering such phase shifts may mix misaligned event states, especially for rapidly evolving local weather\. To address this issue, we design the Event Aligner \(EA\) in Fig\.[4](https://arxiv.org/html/2609.22217#S3.F4)\(b\)\. EA estimates the pairwise delay matrix𝐃∈ℝN×N\\mathbf\{D\}\\in\\mathbb\{R\}^\{N\\times N\}between stations using Fast Fourier Transformℱ\(⋅\)\\mathcal\{F\}\(\\cdot\)cross\-correlation with complex conjugationℱ\(⋅\)¯\\overline\{\\mathcal\{F\}\(\\cdot\)\}\. Then, neighboring station features are temporally shifted according to𝐃\\mathbf\{D\}before attention aggregationMHA\(⋅\)\\operatorname\{MHA\}\(\\cdot\)\. This allows each station to attend to event\-aligned neighboring features rather than raw features at the same timestamp\.
𝐃i,j=argmax\(ℱ−1\(ℱ\(𝐗i\)⊙ℱ\(𝐗j\)¯\)\),\\displaystyle\\mathbf\{D\}\_\{i,j\}=\\operatorname\{argmax\}\\Big\(\\mathcal\{F\}^\{\-1\}\\big\(\\mathcal\{F\}\(\\mathbf\{X\}\_\{i\}\)\\odot\\overline\{\\mathcal\{F\}\(\\mathbf\{X\}\_\{j\}\)\}\\big\)\\Big\),\(12\)𝐟a=Shift\(𝐟C,𝐃\)∈ℝN×T×N×C×d,\\displaystyle\\mathbf\{f\}\_\{a\}=\\operatorname\{Shift\}\(\\mathbf\{f\}\_\{\\text\{C\}\},\\mathbf\{D\}\)\\in\\mathbb\{R\}^\{N\\times T\\times N\\times C\\times d\},𝐟ego=Diag\(𝐟a\)∈ℝN×T×C×d,𝐟adj=OffDiag\(𝐟a\),\\displaystyle\\mathbf\{f\}\_\{\\text\{ego\}\}=\\operatorname\{Diag\}\(\\mathbf\{f\}\_\{a\}\)\\in\\mathbb\{R\}^\{N\\times T\\times C\\times d\},\\quad\\mathbf\{f\}\_\{\\text\{adj\}\}=\\operatorname\{OffDiag\}\(\\mathbf\{f\}\_\{a\}\),𝐟S=MHA\(Q:𝐟ego;K,V:𝐟adj;Mask:𝐌s\)\.\\displaystyle\\mathbf\{f\}\_\{\\text\{S\}\}=\\operatorname\{MHA\}\(\\text\{Q\}:\\mathbf\{f\}\_\{\\text\{ego\}\};\\text\{K,V\}:\\mathbf\{f\}\_\{\\text\{adj\}\};\\text\{Mask\}:\\mathbf\{M\}\_\{s\}\)\.
Here,⊙\\odotis dot product,Diag\(⋅\)\\operatorname\{Diag\}\(\\cdot\)extracts each station’s self\-features from the aligned tensor, andOffDiag\(⋅\)\\operatorname\{OffDiag\}\(\\cdot\)removes self\-features to retain only neighboring features\. Since reliable phase\-delay estimation requires complete temporal sequences, EA is applied only in the posterior branch of the CVAE\. In this way, EA uses complete target sequences during training to provide an aligned posterior correction, while the prior branch remains applicable to incomplete inputs during inference\.
### III\-DAdaptive Temporal Mixer and Local Refiner
Fig\. 5:Details for extreme pattern capture, assessment, and generation\. \(a\) Uncertainty Score and Pattern Decoupled MoE suppress missing\-induced state noise and decouple global similarity from local heterogeneity\. \(b\) Extreme event Assessment and Generation use CVAE variance and sampling\-noise control to estimate and generate extremes\.After complementary initialization, temporal modeling focuses on learning sequential dependencies between observed and unobserved parts\. GIO sequences contain both smooth low\-frequency evolution and abrupt local extremes, while missing observations introduce unreliable inputs\. A single temporal mechanism may either overfit noisy missing regions or over\-smooth short\-term extremes\. Therefore, we design the Adaptive Temporal Mixer \(ATM\) in Fig\.[3](https://arxiv.org/html/2609.22217#S3.F3)with a ‘robust SSM and adaptive MoE’ strategy detailed in Fig\.[5](https://arxiv.org/html/2609.22217#S3.F5)\(a\)\. Specifically, ATM first uses a Bi\-Mamba block to construct a denoised and robust temporal backbone from reliability\-weighted inputs, and then employs a Pattern Decoupled MoE \(PD\-MoE\) to adapt this backbone to heterogeneous stable\-extreme weather patterns\. This design separates trend extraction from pattern adaptation, allowing the model to preserve stable temporal dependencies while remaining sensitive to abrupt local changes\.
For the state space model stage, ATM introduces an Uncertainty Score𝐔s\\mathbf\{U\}\_\{s\}to control how missing or low\-reliability observations enter the state space\. Different from treating all initialized tokens equally,𝐔s\\mathbf\{U\}\_\{s\}explicitly measures the reliability of each temporal position according to its distance to the nearest observed data\. Positions farther from available observations are assigned lower reliability and are suppressed before state transition, preventing uncertain missing regions from dominating the temporal state\. Given the observed temporal index set𝒯obs\\mathcal\{T\}\_\{obs\}, the uncertainty\-aware input is formulated as:
dt=mint′∈𝒯obs\|t−t′\|,t∈\{1,…,T\},\\displaystyle d\_\{t\}=\\min\_\{t^\{\\prime\}\\in\\mathcal\{T\}\_\{obs\}\}\|t\-t^\{\\prime\}\|,\\quad t\\in\\\{1,\\dots,T\\\},\(13\)𝐔st=exp\(−ReLU\(MLP\(dt\)\)\),\\displaystyle\\mathbf\{U\}\_\{s\}^\{t\}=\\exp\\Big\(\-\\operatorname\{ReLU\}\\big\(\\operatorname\{MLP\}\(d\_\{t\}\)\\big\)\\Big\),𝐟~Ct=𝐔st⊙𝐟Ct\.\\displaystyle\\tilde\{\\mathbf\{f\}\}\_\{\\text\{C\}\}^\{t\}=\\mathbf\{U\}\_\{s\}^\{t\}\\odot\\mathbf\{f\}\_\{\\text\{C\}\}^\{t\}\.
Based on these reliability\-weighted inputs𝐟~C\\tilde\{\\mathbf\{f\}\}\_\{\\text\{C\}\}, the Bi\-Mamba block performs temporal modeling through compressed transitions on state space𝐡\\mathbf\{h\}\. Such compression acts as a denoising bottleneck for chaotic weather systems, filtering out severe perturbations that are weakly related to the dominant trend\. In this way, the state space stage first obtains a stable low\-frequency temporal representation𝐟T\\mathbf\{f\}\_\{\\text\{T\}\}, which serves as a robust dependency backbone for subsequent pattern adaptation:
𝐡→t=SSMf\(𝐟~C1,…,𝐟~Ct\),\\displaystyle\\overrightarrow\{\\mathbf\{h\}\}^\{t\}=\\operatorname\{SSM\}\_\{f\}\\big\(\\tilde\{\\mathbf\{f\}\}\_\{\\text\{C\}\}^\{1\},\\dots,\\tilde\{\\mathbf\{f\}\}\_\{\\text\{C\}\}^\{t\}\\big\),\(14\)𝐡←t=SSMb\(𝐟~CT,…,𝐟~Ct\),\\displaystyle\\overleftarrow\{\\mathbf\{h\}\}^\{t\}=\\operatorname\{SSM\}\_\{b\}\\big\(\\tilde\{\\mathbf\{f\}\}\_\{\\text\{C\}\}^\{T\},\\dots,\\tilde\{\\mathbf\{f\}\}\_\{\\text\{C\}\}^\{t\}\\big\),𝐟Tt=MLP\(\[𝐡→t;𝐡←t\]\),∈ℝT×d\.\\displaystyle\\mathbf\{f\}\_\{\\text\{T\}\}^\{t\}=\\operatorname\{MLP\}\\big\(\[\\overrightarrow\{\\mathbf\{h\}\}^\{t\};\\overleftarrow\{\\mathbf\{h\}\}^\{t\}\]\\big\),\\in\\mathbb\{R\}^\{T\\times d\}\.
For the MoE stage, PD\-MoE adapts the denoised temporal dependency to heterogeneous local weather patterns withNeN\_\{e\}independent experts\. Although Bi\-Mamba captures a stable low\-frequency trend, local weather may still switch among smooth evolution, periodic fluctuations, and abrupt extreme mutations\. Instead of using a single shared projection or an unconstrained MLP router, PD\-MoE introduces a structured router composed of pattern clustering and region clustering\. The pattern clustering functionP\(𝐟Tt\)\\operatorname\{P\}\(\\mathbf\{f\}\_\{\\text\{T\}\}^\{t\}\)maps temporal features intoNeN\_\{e\}separated pattern subspaces withpdimp\_\{dim\}size, grouping globally similar temporal states while separating consecutive mutations\. The region clustering functionG\(𝐈\)\\operatorname\{G\}\(\\mathbf\{I\}\)maps station metadata and location features intoSeS\_\{e\}geo\-related subspaces withpdimp\_\{dim\}size, allowing the routing process to distinguish regional responses under similar temporal patterns:
𝐩t=P\(𝐟Tt\)∈ℝNe×pdim,\\displaystyle\\mathbf\{p\}\_\{t\}=\\operatorname\{P\}\(\\mathbf\{f\}\_\{\\text\{T\}\}^\{t\}\)\\in\\mathbb\{R\}^\{N\_\{e\}\\times p\_\{dim\}\},\(15\)𝐠=G\(𝐈\)∈ℝSe×pdim,\\displaystyle\\mathbf\{g\}=\\operatorname\{G\}\(\\mathbf\{I\}\)\\in\\mathbb\{R\}^\{S\_\{e\}\\times p\_\{dim\}\},𝐫t=Softmax\(∑s∈Se\(𝐩t𝐠⊤\)s\)∈ℝNe,\\displaystyle\\mathbf\{r\}\_\{t\}=\\operatorname\{Softmax\}\\big\(\\sum\_\{s\\in S\_\{e\}\}\\big\(\\mathbf\{p\}\_\{t\}\\mathbf\{g\}^\{\\top\}\\big\)\_\{s\}\\big\)\\in\\mathbb\{R\}^\{N\_\{e\}\},where𝐩t\\mathbf\{p\}\_\{t\}denotes the temporal pattern affinity,𝐠\\mathbf\{g\}denotes the regional affinity, and𝐫t\\mathbf\{r\}\_\{t\}is the final routing score at time steptt\.
Moreover, MoE complements the sequential Mamba with a global pattern\-level view\. Unlike direct global attention, which may overemphasize large\-magnitude jumps and lose stable temporal dependencies, the SSM\-MoE structure separates the representation spaces of different patterns through Mamba and specialized experts, thereby balancing stable temporal dependency modeling and abrupt extreme event sensitivity\. The number of activated expertsEtE\_\{t\}is dynamically determined by relative𝐔st\\mathbf\{U\}\_\{s\}^\{t\}: reliable tokens use fewer experts for sharper prediction, while high\-uncertainty tokens aggregate more experts for flexible reconstruction\. The expertℰt\\mathcal\{E\}\_\{t\}selection and feature𝐟T\\mathbf\{f\}\_\{\\text\{T\}\}aggregation are formulated as:
Et=Relative\(𝐔st\),\\displaystyle E\_\{t\}=\\operatorname\{Relative\}\(\\mathbf\{U\}\_\{s\}^\{t\}\),\(16\)ℰt=TopK\(𝐫t,K=Et\),\\displaystyle\\mathcal\{E\}\_\{t\}=\\operatorname\{TopK\}\\big\(\\mathbf\{r\}\_\{t\},K=E\_\{t\}\\big\),𝐰t,e=exp\(𝐫t,e\)∑j∈ℰtexp\(𝐫t,j\),e∈ℰt,\\displaystyle\\mathbf\{w\}\_\{t,e\}=\\frac\{\\exp\(\\mathbf\{r\}\_\{t,e\}\)\}\{\\sum\_\{j\\in\\mathcal\{E\}\_\{t\}\}\\exp\(\\mathbf\{r\}\_\{t,j\}\)\},\\quad e\\in\\mathcal\{E\}\_\{t\},𝐟Tt=∑e∈ℰt𝐰t,eMLPe\(𝐟Tt\)\\displaystyle\\mathbf\{f\}\_\{\\text\{T\}\}^\{t\}=\\sum\_\{e\\in\\mathcal\{E\}\_\{t\}\}\\mathbf\{w\}\_\{t,e\}\\operatorname\{MLP\}\_\{e\}\\big\(\\mathbf\{f\}\_\{\\text\{T\}\}^\{t\}\\big\)𝐟T=\[𝐟T1,𝐟T2,…,𝐟TT\]\\displaystyle\\mathbf\{f\}\_\{\\text\{T\}\}=\[\\mathbf\{f\}\_\{\\text\{T\}\}^\{1\},\\mathbf\{f\}\_\{\\text\{T\}\}^\{2\},\\dots,\\mathbf\{f\}\_\{\\text\{T\}\}^\{T\}\]
After ATM captures robust dependency and adaptive pattern transitions, the time\-independent operations in MoE routing and CVAE sampling may weaken local temporal continuity\. To address this issue, we design the Local Refiner \(LR\)\. LR applies convolutions with multiple receptive fieldsℛ=\{r1,r2,…,rNr\}\\mathcal\{R\}=\\\{r\_\{1\},r\_\{2\},\\dots,r\_\{N\_\{r\}\}\\\}to the sampled latent representation, reconstructing local temporal features at different scales\. A selective MLP gateSG\(⋅\)\\operatorname\{SG\}\(\\cdot\)then adaptively selects among these components, allowing the model to preserve smooth local evolution while retaining necessary extreme variations\.
For latent features𝐟T∈ℝN×T×C×d\\mathbf\{f\}\_\{\\text\{T\}\}\\in\\mathbb\{R\}^\{N\\times T\\times C\\times d\}, LR proceeds as follows:
𝐳=Sample\(PriorMLP\(𝐟T\)\)∈ℝN×T×C×d,\\displaystyle\\mathbf\{z\}=\\operatorname\{Sample\}\(\\operatorname\{PriorMLP\}\(\\mathbf\{f\}\_\{\\text\{T\}\}\)\)\\in\\mathbb\{R\}^\{N\\times T\\times C\\times d\},\(17\)𝐳c=Concat\[Conv\(𝐳,rnr\)\],nr∈\{1,…,Nr\},\\displaystyle\\mathbf\{z\}\_\{c\}=\\operatorname\{Concat\}\\left\[\\operatorname\{Conv\}\(\\mathbf\{z\},r\_\{n\_\{r\}\}\)\\right\],\\quad\{n\_\{r\}\}\\in\\\{1,\\dots,N\_\{r\}\\\},𝐒=Softmax\(SG\(𝐟T\)\)∈ℝN×T×C×\(Nr×d/Nr\),\\displaystyle\\mathbf\{S\}=\\operatorname\{Softmax\}\(\\operatorname\{SG\}\(\\mathbf\{f\}\_\{\\text\{T\}\}\)\)\\in\\mathbb\{R\}^\{N\\times T\\times C\\times\(\{N\_\{r\}\}\\times d/\{N\_\{r\}\}\)\},𝐗^=MLP\(Concat\[𝐟T;𝐒⊙𝐳c\]\)∈ℝN×T×C\.\\displaystyle\\mathbf\{\\hat\{X\}\}=\\operatorname\{MLP\}\\left\(\\operatorname\{Concat\}\\left\[\\mathbf\{f\}\_\{\\text\{T\}\};\\mathbf\{S\}\\odot\\mathbf\{z\}\_\{c\}\\right\]\\right\)\\in\\mathbb\{R\}^\{N\\times T\\times C\}\.
In this way, ATM first obtains a denoised and robust low\-frequency temporal backbone through compressed state space modeling, then uses uncertainty\-aware expert routing to adapt to stable\-extreme pattern shifts\. LR further restores local temporal coherence after stochastic sampling and expert mixing, enabling UniGIO to model both smooth weather evolution and abrupt local changes\.
### III\-ELoss Function
Our total loss function is composed of the CVAE reconstruction loss, KL divergence loss, and orthonormal constraint regulations for pattern decoupling and spatial clustering:
Specifically, given the posterior distributionqϕ\(𝐳\|𝐗\)q\_\{\\phi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\), the prior distributionpψ\(𝐳\|𝐗M,𝐌,𝐈\)p\_\{\\psi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\_\{M\},\\mathbf\{M\},\\mathbf\{I\}\)and decoderpθ\(𝐗\|𝐳\)p\_\{\\theta\}\(\\mathbf\{X\}\|\\mathbf\{z\}\), the total objective is formulated as:
ℒ\\displaystyle\\mathcal\{L\}=−𝔼𝐳∼qϕ\(𝐳\|𝐗\)\[logpθ\(𝐗\|𝐳\)\]⏟ℒrec\\displaystyle=\\underbrace\{\-\\mathbb\{E\}\_\{\\mathbf\{z\}\\sim q\_\{\\phi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\)\}\\left\[\\log p\_\{\\theta\}\(\\mathbf\{X\}\|\\mathbf\{z\}\)\\right\]\}\_\{\\mathcal\{L\}\_\{rec\}\}\(18\)\+βDKL\(qϕ\(𝐳\|𝐗\)∥pψ\(𝐳\|𝐗M,𝐌,𝐈\)\)⏟ℒKL\\displaystyle\+\\beta\\underbrace\{D\_\{KL\}\\left\(q\_\{\\phi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\)\\\|p\_\{\\psi\}\(\\mathbf\{z\}\|\\mathbf\{X\}\_\{M\},\\mathbf\{M\},\\mathbf\{I\}\)\\right\)\}\_\{\\mathcal\{L\}\_\{KL\}\}\+γ‖𝐏⊤𝐏−𝐈‖F2⏟ℒorth\.\\displaystyle\+\\gamma\\underbrace\{\\left\\\|\\mathbf\{P\}^\{\\top\}\\mathbf\{P\}\-\\mathbf\{I\}\\right\\\|\_\{F\}^\{2\}\}\_\{\\mathcal\{L\}\_\{orth\}\}\.
where𝐏\\mathbf\{P\}denotes the learnable weights associated with the pattern decoupling functionP\(⋅\)\\operatorname\{P\}\(\\cdot\), the geo\-related routing functionG\(⋅\)\\operatorname\{G\}\(\\cdot\), and the learnable memory tokens𝐳mem\\mathbf\{z\}\_\{\\text\{mem\}\}\. These basis functions are constrained to be orthonormal\.
## IVExperiments
### IV\-ABenchmark and Setup
Dataset:We curate and benchmark Global In\-situ Weather Modeling tasks on theGIO\-Udataset built from the Weather\-5K\[[13](https://arxiv.org/html/2609.22217#bib.bib37)\]\(quality controlled on HadISD, ICOADS, et al\.\), which covers 5672 weather stations worldwide, recording hourly in\-situ temperature, dew point, wind direction, wind rate, and sea\-level pressure at each station from 2014 to 2024\. We consider time windows of 72, 120, and 240 hours and use the 120\-hour window as our main testbed\.
Mask Settings:We highlight bothnativeandcuratedmasks\. Lacking supervision, training on native missing data is inappropriate\. The model was therefore trained on curated masks andzero\-shot testedon native masks to assess under real\-world incompleteness\. Curated masks also support overall and task\-specific evaluation\.
Baselines:We compare UniGIO with representative generative time series models and station\-interaction\-aware trajectory models\. For general time series generation, we include TimeGAN\[[47](https://arxiv.org/html/2609.22217#bib.bib38)\], TimeVAE\[[48](https://arxiv.org/html/2609.22217#bib.bib41)\], DiffWave\[[55](https://arxiv.org/html/2609.22217#bib.bib43)\], SSSD\[[56](https://arxiv.org/html/2609.22217#bib.bib42)\], and Diffusion\-TS\[[16](https://arxiv.org/html/2609.22217#bib.bib35)\], covering GAN\-, VAE\-, and diffusion\-based paradigms\. These methods are widely used for time series generation, imputation, or distribution modeling, but they mainly focus on unconditional or weakly conditional sequence modeling and rarely consider the coupled station\-variable dependencies inherent in GIO\. Since few existing frameworks are directly designed for unified GIO forecasting, imputation, and generation, we further reimplement two recent trajectory generation methods, UniMTD\[[17](https://arxiv.org/html/2609.22217#bib.bib36)\]and SportsTraj\[[57](https://arxiv.org/html/2609.22217#bib.bib46)\], which naturally support interactions among multiple agents\. We extend their interaction modeling from spatial trajectories to station and variable dependencies to align them with our task setting\. In addition, we introduce a deterministic variant of UniGIO with the CVAE generative structure removed to verify the effectiveness of the generative paradigm\.
Evaluation Metrics:We evaluate model performance from three perspectives: accuracy, fidelity, and extreme event capture\. For accuracy, we report MAE and MSE on masked regions to measure pointwise reconstruction or forecasting errors\. For fidelity, we adopt FID\[[16](https://arxiv.org/html/2609.22217#bib.bib35)\], statistical error \(STE\), and Correlation Score \(CroS\)\[[16](https://arxiv.org/html/2609.22217#bib.bib35)\]to assess distribution\-level consistency between generated sequences and ground truth, including feature distribution, statistical variation, and cross\-variable correlations\. For extreme event capture, we improve SEDI\[[13](https://arxiv.org/html/2609.22217#bib.bib37)\]into F1\-SEDI at the 99\.5th and 90\.0th percentiles by combining SEDI\-based recall with precision in an F1\-score manner, since practical weather services require not only detecting extremes but also reducing false alarms\. In task\-specific comparisons, including forecasting and cross\-region generation, we mainly report accuracy and F1\-SEDI because practical GIO applications prioritize pointwise reliability and extreme event detection over distributional similarity\. Since wind direction is a circular variable ranging from 0 to 360 degrees, F1\-SEDI is not computed for wind direction\.
ForAccuracy, we adopt Mean Absolute Error \(MAE\) and Mean Squared Error \(MSE\) on mask area:
MAE\\displaystyle\\text\{MAE\}=1∑i\(1−𝐌i\)∑i\(1−𝐌i\)\|𝐗^i−𝐗i\|\\displaystyle=\\frac\{1\}\{\\sum\_\{i\}\(1\-\\mathbf\{M\}\_\{i\}\)\}\\sum\_\{i\}\(1\-\\mathbf\{M\}\_\{i\}\)\\,\\left\|\\hat\{\\mathbf\{X\}\}\_\{i\}\-\\mathbf\{X\}\_\{i\}\\right\|\(19\)MSE\\displaystyle\\text\{MSE\}=1∑i\(1−𝐌i\)∑i\(1−𝐌i\)\(𝐗^i−𝐗i\)2\\displaystyle=\\frac\{1\}\{\\sum\_\{i\}\(1\-\\mathbf\{M\}\_\{i\}\)\}\\sum\_\{i\}\(1\-\\mathbf\{M\}\_\{i\}\)\\,\\left\(\\hat\{\\mathbf\{X\}\}\_\{i\}\-\\mathbf\{X\}\_\{i\}\\right\)^\{2\}where∑i\(1−Mi\)\\sum\_\{i\}\(1\-M\_\{i\}\)is the total number of masked values\.iiis element\-wise index
ForExtreme Event Capture, previous methods use Symmetric Extremal Dependence Index \(SEDI\) to represent the recall rate of extreme events\. However, in practical applications, a large number of false alarms is also unacceptable\. Therefore, we propose F1\-SEDI, which evaluates both precision and recall in an F1\-score manner:
Precision=1∑i𝕀\(𝐗^i<Qlowerp\)\+∑i𝕀\(𝐗^i\>Qupperp\)\\displaystyle\\text\{Precision\}=\\frac\{1\}\{\\sum\_\{i\}\\mathbb\{I\}\(\\hat\{\\mathbf\{X\}\}\_\{i\}<Q^\{p\}\_\{\\text\{lower\}\}\)\+\\sum\_\{i\}\\mathbb\{I\}\(\\hat\{\\mathbf\{X\}\}\_\{i\}\>Q^\{p\}\_\{\\text\{upper\}\}\)\}\(20\)\(∑i𝕀\(𝐗^i<Qlowerp∩𝐗i<Qlowerp\)\+CLOSE\\displaystyle\\Bigg\(\\sum\_\{i\}\\mathbb\{I\}\(\\hat\{\\mathbf\{X\}\}\_\{i\}<Q^\{p\}\_\{\\text\{lower\}\}\\cap\\mathbf\{X\}\_\{i\}<Q^\{p\}\_\{\\text\{lower\}\}\)\+\{\}OPEN∑i𝕀\(𝐗^i\>Qupperp∩𝐗i\>Qupperp\)\)\\displaystyle\\sum\_\{i\}\\mathbb\{I\}\(\\hat\{\\mathbf\{X\}\}\_\{i\}\>Q^\{p\}\_\{\\text\{upper\}\}\\cap\\mathbf\{X\}\_\{i\}\>Q^\{p\}\_\{\\text\{upper\}\}\)\\Bigg\)SEDI=1∑i𝕀\(𝐗i<Qlowerp\)\+∑i𝕀\(𝐗i\>Qupperp\)\\displaystyle\\text\{SEDI\}=\\frac\{1\}\{\\sum\_\{i\}\\mathbb\{I\}\(\\mathbf\{X\}\_\{i\}<Q^\{p\}\_\{\\text\{lower\}\}\)\+\\sum\_\{i\}\\mathbb\{I\}\(\\mathbf\{X\}\_\{i\}\>Q^\{p\}\_\{\\text\{upper\}\}\)\}\(∑i𝕀\(𝐗^i<Qlowerp∩𝐗i<Qlowerp\)\+CLOSE\\displaystyle\\Bigg\(\\sum\_\{i\}\\mathbb\{I\}\(\\hat\{\\mathbf\{X\}\}\_\{i\}<Q^\{p\}\_\{\\text\{lower\}\}\\cap\\mathbf\{X\}\_\{i\}<Q^\{p\}\_\{\\text\{lower\}\}\)\+\{\}OPEN∑i𝕀\(𝐗^i\>Qupperp∩𝐗i\>Qupperp\)\)\\displaystyle\\sum\_\{i\}\\mathbb\{I\}\(\\hat\{\\mathbf\{X\}\}\_\{i\}\>Q^\{p\}\_\{\\text\{upper\}\}\\cap\\mathbf\{X\}\_\{i\}\>Q^\{p\}\_\{\\text\{upper\}\}\)\\Bigg\)F1\-SEDI=2×SEDI×PrecisionSEDI\+Precision\\displaystyle\\text\{F1\-SEDI\}=\\frac\{2\\times\\text\{SEDI\}\\times\\text\{Precision\}\}\{\\text\{SEDI\}\+\\text\{Precision\}\}whereQlowerpQ^\{p\}\_\{\\text\{lower\}\}andQupperpQ^\{p\}\_\{\\text\{upper\}\}are theppth lower and upper percentiles,𝕀\\mathbb\{I\}is indicator function\.
ForFidelity, we adopt Fréchet Inception Distance \(FID\)\[[58](https://arxiv.org/html/2609.22217#bib.bib39)\], statistical errors \(STE\), and Correlation Score \(CroS\)\[[47](https://arxiv.org/html/2609.22217#bib.bib38),[16](https://arxiv.org/html/2609.22217#bib.bib35)\]between generated data distribution and ground truth data distribution:
FID\\displaystyle\\text\{FID\}=‖𝝁r−𝝁f‖22\+Tr\(𝚺r\+𝚺f−2\(𝚺r𝚺f\)1/2\)\\displaystyle=\\left\\\|\\boldsymbol\{\\mu\}\_\{r\}\-\\boldsymbol\{\\mu\}\_\{f\}\\right\\\|\_\{2\}^\{2\}\+\\operatorname\{Tr\}\\left\(\\boldsymbol\{\\Sigma\}\_\{r\}\+\\boldsymbol\{\\Sigma\}\_\{f\}\-2\\left\(\\boldsymbol\{\\Sigma\}\_\{r\}\\boldsymbol\{\\Sigma\}\_\{f\}\\right\)^\{1/2\}\\right\)\(21\)STE\\displaystyle\\text\{STE\}=∑i\(\(1−𝐌i\)\|σrw\(𝐗i\)−σfw\(𝐗i^\)\|\)∑i\(1−𝐌i\)\\displaystyle=\\frac\{\\sum\_\{i\}\\Bigg\(\(1\-\\mathbf\{M\}\_\{i\}\)\\Bigg\|\\sigma\_\{r\}^\{w\}\(\\mathbf\{X\}\_\{i\}\)\-\\sigma\_\{f\}^\{w\}\(\\hat\{\\mathbf\{X\}\_\{i\}\}\)\\Bigg\|\\Bigg\)\}\{\\sum\_\{i\}\(1\-\\mathbf\{M\}\_\{i\}\)\}CroS\\displaystyle\\text\{CroS\}=∑i\(\(1−𝐌i\)\|𝔼N,T\[𝐗i^~⊙𝐗i^~′\]−𝔼N,T\[𝐗i~⊙𝐗i~′\]\|\)∑i\(1−𝐌i\)\\displaystyle=\\frac\{\\sum\_\{i\}\\Bigg\(\(1\-\\mathbf\{M\}\_\{i\}\)\\left\|\\begin\{aligned\} &\\mathbb\{E\}\_\{N,T\}\\left\[\\tilde\{\\hat\{\\mathbf\{X\}\_\{i\}\}\}\\odot\\tilde\{\\hat\{\\mathbf\{X\}\_\{i\}\}\}^\{\\prime\}\\right\]\\\\ &\\quad\-\\mathbb\{E\}\_\{N,T\}\\left\[\\tilde\{\\mathbf\{X\}\_\{i\}\}\\odot\\tilde\{\\mathbf\{X\}\_\{i\}\}^\{\\prime\}\\right\]\\end\{aligned\}\\right\|\\Bigg\)\}\{\\sum\_\{i\}\(1\-\\mathbf\{M\}\_\{i\}\)\}
For FID,𝝁r\\boldsymbol\{\\mu\}\_\{r\},𝝁f\\boldsymbol\{\\mu\}\_\{f\},𝚺r\\boldsymbol\{\\Sigma\}\_\{r\}, and𝚺f\\boldsymbol\{\\Sigma\}\_\{f\}are the mean and covariance of ground truth and model output features\. For STE,σrw\(𝐗\)\\sigma\_\{r\}^\{w\}\(\\mathbf\{X\}\)andσfw\(𝐗^\)\\sigma\_\{f\}^\{w\}\(\\hat\{\\mathbf\{X\}\}\)are the standard deviation of ground truth and model output in windowww\. For CroS,𝐗^~\\tilde\{\\hat\{\\mathbf\{X\}\}\}and𝐗~\\tilde\{\\mathbf\{X\}\}are normalized𝐗^\\hat\{\\mathbf\{X\}\}and𝐗\\mathbf\{X\};𝐗^~′\\tilde\{\\hat\{\\mathbf\{X\}\}\}^\{\\prime\}and𝐗~′\\tilde\{\\mathbf\{X\}\}^\{\\prime\}are lower\-triangular variable pairs of normalized data;𝔼N,T\[⋅\]\\mathbb\{E\}\_\{N,T\}\[\\cdot\]is the expectation over station and time dimensions\.
### IV\-BOverall performance
#### IV\-B1Quantitative Results
Quantitative results on native masks and unified tasks with curated masks are visualized in Fig\.[6](https://arxiv.org/html/2609.22217#S4.F6)and reported in Tab\.[I](https://arxiv.org/html/2609.22217#S4.T1), Tab\.[II](https://arxiv.org/html/2609.22217#S4.T2), and Tab\.[III](https://arxiv.org/html/2609.22217#S4.T3)in terms of accuracy, extreme event capture, and fidelity, respectively\. The best results are highlighted inred, and the second\-best results are highlighted inblue\. Our UniGIO achieves state\-of\-the\-art performance over all baselines, leading by a large margin on the vast majority of 7 metrics and 5 variables on both curated and native masks\.
For different mask settings, UniGIO outperforms the second\-best baseline by an average of 12\.1% on accuracy, 3\.5% on F1\-SEDI, and 11\.1% on fidelity over curated masks\. And for the zero\-shot transfer on native masks, UniGIO has achieved 10\.8%, 6\.5%, and 12\.6% advantage\. The numerical performance is comparable with that on curated masks, which validates that our model can directly learn the intrinsic GIO patterns with real\-world obscures from our synthetic mask designs\.
For different methods, regrettably, both the deterministic method \(UniGIO w/o CVAE\) and time series generators fail on GIO data\. The deterministic method is pulled to the mean value due to the MSE loss\. Time series generators fall into two extremes: either mining only low\-frequency and generating over\-smoothed results, or focusing solely on high\-frequency and producing spurious fluctuations on erroneous trends\. Surprisingly, two trajectory methods survived and generated second\-best results; the CVAE even achieved superiority on F1\-SEDI compared with diffusion\. The finding highlights the significance of both condition emphasized generative frameworks and GIO complementary\.
We further compare representative frameworks under window lengths of 72h and 240h, with results reported in Tab\.[IV](https://arxiv.org/html/2609.22217#S4.T4), Tab\.[V](https://arxiv.org/html/2609.22217#S4.T5), and Tab\.[VI](https://arxiv.org/html/2609.22217#S4.T6)\. UniGIO achieves the best performance across different temporal ranges\. Under the 72h Curated/Native settings, it improves accuracy by 8\.8%/13\.5%, F1\-SEDI by 3\.6%/3\.6%, and fidelity by 7\.1%/5\.3%\. Under the 240h longer\-window setting, the gains further increase to 11\.1%/10\.0% in accuracy, 5\.1%/6\.0% in F1\-SEDI, and 21\.4%/16\.4% in fidelity, demonstrating its effectiveness under both short\- and long\-window scenarios, and UniGIO’s advantage becomes more pronounced as the temporal window length increases\. Moreover, as the window size increases, most metrics deteriorate due to longer temporal dependencies and higher uncertainty\. CroS is an exception, as it measures inter\-variable consistency and steadily improves with longer windows\. This may be because longer windows provide richer temporal context for modeling variable relationships and reduce the sensitivity of this steady\-state metric to short\-window random fluctuations\.
Fig\. 6:UniGIO achieves SOTA performance on accuracy, extreme event capture, and fidelity across all time windows\.As shown in Table[VII](https://arxiv.org/html/2609.22217#S4.T7), the parameter comparison between UniGIO and the baseline methods demonstrates the lightness of our model design\. Even among time series models, UniGIO remains compact, while UniGIO\-M can be further compressed to less than 1M parameters\. This further indicates that the improvement of UniGIO mainly stems from effective modeling designs rather than parameter scaling\.
TABLE I:Quantitative Accuracy Comparison on Curated / Native Masks \(120h\)†\{\\dagger\}means re\-implementedTABLE II:Quantitative F1\-SEDI Comparison on Curated / Native Masks \(120h\)†\{\\dagger\}means re\-implementedTABLE III:Quantitative Fidelity Comparison on Curated / Native Masks \(120h\)†\{\\dagger\}means re\-implementedTABLE IV:Quantitative Accuracy Comparison on Curated / Native Masks \(72h/240h\)†\{\\dagger\}means re\-implementedTABLE V:Quantitative F1\-SEDI Comparison on Curated / Native Masks \(72h/240h\)†\{\\dagger\}means re\-implementedTABLE VI:Quantitative Fidelity Comparison on Curated / Native Masks \(72h/240h\)†\{\\dagger\}means re\-implementedTABLE VII:Parameter Comparison of Unified and Task\-specific GIO models\.MethodTimeGANTimeVAEDiffwaveSSSDDiffusion\-TSUniMTDSportsTrajUniGIO\-MUniGIOESFM\-SChronos2WSSMiTransformerCorrformerParams\.3\.4M2\.3M4\.1M4\.3M4\.9M10\.1M7\.3M0\.9M3\.4M115M120M1\.8M0\.9M55\.6M
#### IV\-B2Qualitative Results
In Fig\.[7](https://arxiv.org/html/2609.22217#S4.F7), we qualitatively compared the forecasting, imputation, and generation results between UniGIO, SportsTraj, UniMTD, and TimeGAN\. The results in row 2 illustrate that our UniGIO generates weather series with superior accuracy and fidelity, especially for temperature \(red\) and sea\-level pressure \(violet\)\. In row 3, SportsTraj generates a large number of spurious fluctuations, which are mainly attributed to the noise introduced by its sampling\. Although this may help it simulate the random fluctuations of wind direction, it takes a toll on variables and tasks need inherent stability and continuity \(e\.g\., temperature and imputation\)\. Rows 4\-5 show over\-smoothing and over\-fluctuation results, UniMTD can only generate low\-frequency results, and this deficiency is not limited to high\-frequency variables; it also leads to the loss of small\-scale patterns in sea\-level pressure\. In contrast, TimeGAN merely fits high\-frequency noise, with completely erroneous trends\.
The qualitative evaluation results confirm the significant superiority of our method in quantitative results and also demonstrate heterogeneity among variables\.
Fig\. 7:Due to space constraints, only wind rate and sea\-level pressure are displayed for imputation and generation tasksFig\. 8:UniGIO achieves SOTA performance on Forecasting, Imputation, and Generation tasks, even beating Foundation Models\.
### IV\-CTask\-specific Analysis
TABLE VIII:Quantitative comparison under variable\-length, fixed\-length, and noisy forecasting settings\.†\{\\dagger\}denotes re\-implemented methods, and⋆\\stardenotes forecasting with an imputed look\-back window\.TABLE IX:Quantitative Comparison with Finetuned Observation and Time Series Foundation models on 24h Forecast TaskTABLE X:Quantitative Comparison on Imputation Task†\{\\dagger\}means re\-implementedTABLE XI:Quantitative Comparison on Generation Task†\{\\dagger\}means re\-implementedTo evaluate UniGIO as a versatile framework across different tasks on 120h time windows, in Fig\.[8](https://arxiv.org/html/2609.22217#S4.F8), the accuracy and F1\-SEDI performance of 7 tasks grouped into fixed/variable\-windowForecasting, random&period/resolutionImputation, and variable/in\-region/cross\-regionGenerationare presented from left to right\. UniGIO achieved the optimal performance on 92% of all 108 metrics\.
For forecasting, we obtained an advantage of 18\.5% and 8\.13% on accuracy and F1\-SEDI\. Notably, we significantly outperformed SOTA forecast methods, including WSSM\[[15](https://arxiv.org/html/2609.22217#bib.bib27)\], iTransformer\[[59](https://arxiv.org/html/2609.22217#bib.bib10)\], Corrformer\[[3](https://arxiv.org/html/2609.22217#bib.bib6)\], and Chronos2\[[60](https://arxiv.org/html/2609.22217#bib.bib58)\], in the 48h\-72h fixed\-window setting\. More importantly, for forecasting with incomplete observations, our unified paradigm significantly outperforms existing approaches that first impute a complete look\-back window before forecasting, yielding improvements of 24\.5% and 23\.3%, respectively\. These results demonstrate that UniGIO is better suited for real\-world forecasting under incomplete\-observation GIO scenarios\. For imputation and generation, our advantages reached 9\.4%, 6\.0% and 6\.8%, 4\.4%, respectively, also outperforming specialized imputation\[[40](https://arxiv.org/html/2609.22217#bib.bib50),[41](https://arxiv.org/html/2609.22217#bib.bib45)\]and generation models\[[16](https://arxiv.org/html/2609.22217#bib.bib35),[17](https://arxiv.org/html/2609.22217#bib.bib36),[57](https://arxiv.org/html/2609.22217#bib.bib46)\]\.
Further, UniGIO outperforms the fine\-tuned global multi\-source observation foundation model ESFM\[[8](https://arxiv.org/html/2609.22217#bib.bib57)\], time series foundation model Chronos2\[[60](https://arxiv.org/html/2609.22217#bib.bib58)\], and operational NWP product HRES\[[61](https://arxiv.org/html/2609.22217#bib.bib59)\]under the 24h forecast setting in ESFM, achieving an improvement of 14\.7%\. These results verify that UniGIO, as a lightweight unified framework, is broadly effective across different sub\-tasks\. It can even surpass task\-specific models and, despite model and data volume, outperform foundation models and operational NWP that rely on global multi\-source observations or generic time series\. Full tables are in Tab\.[VIII](https://arxiv.org/html/2609.22217#S4.T8),[IX](https://arxiv.org/html/2609.22217#S4.T9),[X](https://arxiv.org/html/2609.22217#S4.T10),[XI](https://arxiv.org/html/2609.22217#S4.T11)\.
### IV\-DSpatial Generalization Analysis
TABLE XII:Quantitative Evaluation on cross\-region \(South America / Africa\) GenerationTo evaluate spatial generalization, we conducted cross\-domain generation on 120h time windows summarized in Tab\.[XII](https://arxiv.org/html/2609.22217#S4.T12), generating African and South American stations using UniGIO trained on other stations\. Focusing on temperature, wind direction, and wind rate—core variables in tropical coastal regions prone to storms and high temperatures\.
Using in\-region generation with fewer spatial gaps as a baseline ‘World wide’, we randomly partitioned global stations into training and generation sets\. For cross\-region generation, first, the model was trained outside the target regions and tested on all stations in Africa and South America\. Then, with 90%, 60%, or 30% of stations randomly selected for generation and the remainder for fine\-tuning\.
Results in Tab\.[XII](https://arxiv.org/html/2609.22217#S4.T12)show South America is more challenging than Africa\. Despite reduced performance under higher spatial incompleteness, the model maintains reasonable zero\-shot accuracy, improving markedly when fine\-tuned on only 10% of stations\. Fine\-tuning on\>\>50% stations leads cross\-region generation to outperform in\-region generation on 35% of 40 metrics, though extreme event capture still lags by 10%\. Overall, these results demonstrate the model’s substantial spatial generalization capability\.
Fig\. 9:Distribution coverage ofresultsagainstGTacross baselines and different mask settings\. Less red area is better\. TSNE plots look unclustered because GIO is largely continuous with no ‘category’ between randomly selected samples for visualization\.
### IV\-EFidelity and Robustness Analysis
In Fig\.[9](https://arxiv.org/html/2609.22217#S4.F9), we present the distribution coverage of results on 120h time windows against ground truth across baselines and different mask settings under t\-SNE reduction, comparing different baselines and mask settings\. In this visualization, the exposed red regions indicate distributional discrepancies between the generated samples and the ground truth; therefore, a smaller visible red area suggests stronger distributional consistency and better coverage of the target data manifold\.
As shown in the upper part, our method achieves more comprehensive coverage of the ground\-truth distribution than the competing baselines\. The generated samples produced by our method are more closely aligned with the ground\-truth regions, while the baselines exhibit larger uncovered or mismatched areas\. This observation indicates that our approach is better able to capture the underlying data distribution\. Meanwhile, combined with the lower part, we observe that the native masks and the curated masks exhibit consistent alignment with the ground\-truth distribution, indicating that our method is not sensitive to specific mask formats\. These results confirm the performance advantages in the quantitative experiments\.
For mask setting, despite mask type change or mask rates varying from 30% \(imputation\-dominated\) to 90% \(generation\-dominated\) with increasing red regions, our method maintains distribution consistency\. Together with Exp[IV\-C](https://arxiv.org/html/2609.22217#S4.SS3), proving our robustness across varying conditions\.
Fig\. 10:Wind ratewith multiple extreme events and stablesea\-level pressureunder different sampling noise intensities
### IV\-FExtreme Event Risk Analysis
To verify whether the prior distribution variance of UniGIO serves as an indicator of the risk of extreme events, samples are generated under different noise intensities for two representative variables in Fig\.[10](https://arxiv.org/html/2609.22217#S4.F10): wind rate, which contains multiple extreme events, and sea\-level pressure, which remains relatively stable\.
It can be observed that when the intensity is00, both variables converge to a smooth mean trend\. As the intensity increases, the former rapidly exhibits strong fluctuations, while the latter remains unchanged\. When the intensity becomes excessively high, both variables fluctuate\.
Notably, stable patterns exist even in high\-risk variables, as highlighted by blue boxes, and vice versa\. This observation suggests that the relationship between latent\-space variance and extreme event risk is not uniformly distributed across all time steps, but instead reflects fine\-grained, pointwise differences in uncertainty and event sensitivity\.
These results indicate that we can, to a certain extent, quantify the risk of extreme events using the variance in the latent space\. UniGIO also allows for pointwise control of extreme events as Fig\.[5](https://arxiv.org/html/2609.22217#S3.F5)\(b\), demonstrating not only interpretability in uncertainty modeling but also flexibility in controllable generation\.
Fig\. 11:Ablation study results \(right\) and the representative expert specializations across variables, locations, uncertainties, and temporal patterns \(left\), different experts are denoted by colors
### IV\-GModel Structure Analysis
TABLE XIII:Ablation Study on Accuracy and F1\-SEDITABLE XIV:Ablation Study on FidelityThe right part of Fig\.[11](https://arxiv.org/html/2609.22217#S4.F11)shows our ablation results on 120h time windows, corresponding to Tab\.[XIII](https://arxiv.org/html/2609.22217#S4.T13),[XIV](https://arxiv.org/html/2609.22217#S4.T14)\.
Regarding model capacity, performance increases with model size, increasing the number of experts from 16 \(default\) to 32 provides an effective way to improve performance\. For varying sampling intensities, low intensity improves accuracy but degrades F1\-SEDI, while high intensity has the opposite effect, both compromising fidelity\. Despite using 73\.5% fewer parameters, UniGIO\-M incurs only a 5% performance compromise on average, while retaining largely comparable fidelity, demonstrating the strong parameter efficiency and compressibility of our framework\.
Progressive structure ablation confirms the effectiveness of our design: The OM brings substantial improvements across all metrics, especially in accuracy and fidelity\. This is consistent with the nature of weather events as field patterns rather than isolated pointwise series\. MoEs consistently improve the performance across all variables, particularly in F1\-SEDI\. This validates that the ATM design is effective in capturing extreme events, rather than merely improving accuracy\. LR enhances fidelity at the cost of smoothing extremes, and EA aids variables with distinct periodicity\. The above results verify that the core design intentions of UniGIO are well fulfilled
The left and bottom parts of Fig\.[11](https://arxiv.org/html/2609.22217#S4.F11)illustrate ATM expert specialization\. Variable\-wise,Expert0favors temperature and dew point, whileExpert7handles wind direction\. Location\-wise, stations processed primarily byExpert0cluster in Asia, whereas those handled byExpert2are concentrated in North America, Europe\. Token\-wise uncertainty showsExpert0manages long\-range high\-uncertainty tokens,Expert4short\-range low\-uncertainty tokens, andExpert3focuses on observed values\. Expert shift in samples show reasonable pattern\-expert allocation, despite token\-wise routing, experts achieve segment\-wise consistency due to robust long\-term dependencies\.
## VConclusion
In this work, we aim to address global in\-situ weather modeling under incomplete observations and establish a strong benchmark comprehensively evaluating accuracy, fidelity, and extreme event capture performance\. Further, we propose UniGIO, a novel unified framework that allows in\-situ weather modeling from incomplete GIO input through observation to missing generation\. Essentially, we highlight \(1\) GIO complementarity; \(2\) Extreme patterns; \(3\) Lightweight to address the incompleteness and chaotic nature of GIO data and operate directly on interpolation\-free distributed weather stations\. We design an Observation Mixer and Event Aligner to capture the GIO complementarity at station and region level, and propose an Adaptive Temporal Mixer to divide\-and\-conquer stable evolution with chaotic drifting patterns\. A Local Refiner is added to bridge the separated local patterns\. Through extensive experiments and comprehensive evaluations, UniGIO consistently outperforms baselines across curated masks, native masks, and 7 forecasting, imputation, and generation tasks with an average 11%, 12%, and 5% advantage on accuracy, fidelity, and extreme event capture\. Establishing itself as a SOTA solution with significant potential for exploration in the climate foundation model\.
## References
- \[1\]WMO\(2018\)Manual on the global observing system\.World Meteorological Organization,Geneva, Switzerland\.External Links:[Link](https://library.wmo.int/doc_num.php?explnum_id=1000000000067)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p1.1),[§I](https://arxiv.org/html/2609.22217#S1.p3.1)\.
- \[2\]X\. X\. Zhu, Z\. Xiong, Y\. Wang, A\. J\. Stewart, K\. Heidler, Y\. Wang, Z\. Yuan, T\. Dujardin, Q\. Xu, and Y\. Shi\(2026\)On the foundations of earth foundation models\.Communications Earth & Environment7\(1\),pp\. 103\.External Links:[Document](https://dx.doi.org/10.1038/s43247-025-03127-x)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p1.1)\.
- \[3\]H\. Wu, H\. Zhou, M\. Long, and J\. Wang\(2023\)Interpretable weather forecasting for worldwide stations with a unified deep model\.Nature Machine Intelligence5\(6\),pp\. 602–611\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p1.1),[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p1.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[4\]K\. Bi, L\. Xie, H\. Zhang, X\. Chen, X\. Gu, and Q\. Tian\(2023\)Accurate medium\-range global weather forecasting with 3d neural networks\.Nature619,pp\. 533–538\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06185-3)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p1.1)\.
- \[5\]R\. Lam, A\. Sanchez\-Gonzalez, M\. Willson, P\. Wirnsberger, M\. Fortunato, F\. Alet, S\. Ravuri, T\. Ewalds, Z\. Eaton\-Rosen, W\. Hu,et al\.\(2023\)Learning skillful medium\-range global weather forecasting\.Science382\(6677\),pp\. 1416–1421\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p1.1)\.
- \[6\]I\. Price, A\. Sanchez\-Gonzalez, F\. Alet, T\. R\. Andersson, A\. El\-Kadi, D\. Masters, T\. Ewalds, J\. Stott, S\. Mohamed, P\. Battaglia,et al\.\(2025\)Probabilistic weather forecasting with machine learning\.Nature637\(8044\),pp\. 84–90\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p1.1)\.
- \[7\]A\. Vaughan, S\. Markou, W\. Tebbutt, J\. Requeima, W\. P\. Bruinsma, T\. R\. Andersson, M\. Herzog, N\. D\. Lane, M\. Chantry, J\. S\. Hosking,et al\.\(2025\)End\-to\-end data\-driven weather prediction\.Nature641,pp\. 1172–1179\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-08897-0)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1),[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p2.1)\.
- \[8\]F\. Ozdemir, Y\. Cheng, S\. Mohebi, F\. Lehmann, S\. Adamov, Z\. Zhang, L\. Trentini, D\. Grund, O\. Fuhrer, T\. Hoefler,et al\.\(2026\)Earth system foundation model \(esfm\): a unified framework for heterogeneous data integration and forecasting\.arXiv preprint arXiv:2605\.00850\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1),[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p2.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p3.1)\.
- \[9\]E\. Kalnay\(2003\)Atmospheric modeling, data assimilation and predictability\.Cambridge University Press\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1)\.
- \[10\]P\. Bauer, A\. Thorpe, and G\. Brunet\(2015\)The quiet revolution of numerical weather prediction\.Nature525,pp\. 47–55\.External Links:[Document](https://dx.doi.org/10.1038/nature14956)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p2.1)\.
- \[11\]W\. Cao, D\. Wang, J\. Li, H\. Zhou, L\. Li, and Y\. Li\(2018\)BRITS: bidirectional recurrent imputation for time series\.InAdvances in Neural Information Processing Systems,Vol\.31,pp\. 6776–6786\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p3.1)\.
- \[12\]Y\. Tashiro, J\. Song, Y\. Song, and S\. Ermon\(2021\)CSDI: conditional score\-based diffusion models for probabilistic time series imputation\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 24804–24816\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p3.1)\.
- \[13\]T\. Han, S\. Guo, Z\. Chen, W\. Xu, and L\. Bai\(2024\)WEATHER\-5k: a large\-scale global station weather dataset towards comprehensive time\-series forecasting benchmark\.CoRR\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p3.1),[§I](https://arxiv.org/html/2609.22217#S1.p7.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p1.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p4.1)\.
- \[14\]WMO\(2021\)Technical regulations for the global basic observing network\.World Meteorological Organization,Geneva, Switzerland\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p3.1),[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[15\]S\. Yang, Z\. Liu, Z\. Shi, and Z\. Zou\(2025\)WSSM: geographic\-enhanced hierarchical state\-space model for global station weather forecast\.arXiv preprint arXiv:2501\.11238\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p1.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[16\]X\. Yuan and Y\. Qiao\(2024\)Diffusion\-ts: interpretable diffusion for general time series generation\.InThe Twelfth International Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p4.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p8.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[17\]S\. Yang, Z\. Shi, and Z\. Zou\(2025\)Unified multi\-agent trajectory modeling with masked trajectory diffusion\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 27563–27574\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[18\]J\. Gong, K\. Xu, W\. Wei, S\. Tu, J\. Xu, Z\. Liu, H\. Fan, Z\. Zhou, T\. Han, Y\. Xiao,et al\.\(2026\)Earth\-o1: a grid\-free observation\-native atmospheric world model\.arXiv preprint arXiv:2605\.06337\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p2.1)\.
- \[19\]WMO\(2021\)WMO atlas of mortality and economic losses from weather, climate and water extremes \(1970–2019\)\.Technical ReportTechnical ReportWMO\-No\. 1267,World Meteorological Organization,Geneva, Switzerland\.External Links:ISBN 978\-92\-63\-11267\-5Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[20\]UNDRR\(2022\)Global assessment report on disaster risk reduction 2022: our world at risk: transforming governance for a resilient future\.UNDRR,Geneva, Switzerland\.Note:Referenced for extensive risk and high\-frequency, low\-severity event statisticsExternal Links:[Link](https://www.undrr.org/gar/gar2022-our-world-risk)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[21\]World Meteorological Organization\(2019\)WMO Strategic Plan 2020–2023\.Technical reportTechnical ReportWMO\-No\. 1225,World Meteorological Organization,Geneva, Switzerland\.External Links:ISBN 978\-92\-63\-11225\-5,[Link](https://etrp.wmo.int/pluginfile.php/25842/mod_resource/content/1/WMO%20Strategic%20Plan%202020-2023%20%28WMO-No.1225%29.pdf)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[22\]World Meteorological OrganizationSystematic observations financing facility \(SOFF\)\(Website\)World Meteorological Organization\.External Links:[Link](https://wmo.int/activities/systematic-observations-financing-facility-soff)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[23\]D\. P\. Rogers, V\. V\. Tsirkunov, H\. Kootval, A\. Soares, A\. Khamidov, M\. Staudinger, and M\. Kadi\(2019\)Weathering the change: how to improve hydromet services in developing countries\.World Bank,Washington, DC\.External Links:[Document](https://dx.doi.org/10.1596/31507),[Link](https://www.gfdrr.org/en/publication/weathering-change-how-improve-hydromet-services-developing-countries)Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[24\]World Meteorological Organization\(2024\)Guide to instruments and methods of observation: volume i – measurement of meteorological variables\.World Meteorological Organization,Geneva, Switzerland\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p5.1)\.
- \[25\]K\. Sohn, H\. Lee, and X\. Yan\(2015\)Learning structured output representation using deep conditional generative models\.InAdvances in Neural Information Processing Systems,Vol\.28\.Cited by:[§I](https://arxiv.org/html/2609.22217#S1.p6.1)\.
- \[26\]L\. Chen, X\. Zhong, F\. Zhang, Y\. Cheng, Y\. Xu, Y\. Qi, H\. Li, M\. Tang, R\. Gao, M\. Wang,et al\.\(2023\)FuXi: a cascade machine learning forecasting system for 15\-day global weather forecast\.npj Climate and Atmospheric Science6\(1\),pp\. 190\.Cited by:[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p1.1)\.
- \[27\]Z\. Gao, X\. Shi, H\. Wang, Y\. Zhu, Y\. Wang, M\. Li, and D\. Yeung\(2022\)Earthformer: exploring space\-time transformers for earth system forecasting\.InAdvances in Neural Information Processing Systems,Cited by:[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p2.1)\.
- \[28\]T\. Nguyen, R\. Shah, H\. Bansal, T\. Arcomano, S\. Madireddy, R\. Maulik, K\. Kashinath,et al\.\(2023\)ClimaX: a foundation model for weather and climate\.InInternational Conference on Machine Learning,Cited by:[§II\-A](https://arxiv.org/html/2609.22217#S2.SS1.p2.1)\.
- \[29\]B\. Lim and S\. Zohren\(2021\)Time\-series forecasting with deep learning: a survey\.Philosophical transactions of the royal society a: mathematical, physical and engineering sciences379\(2194\)\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[30\]H\. Hewamalage, C\. Bergmeir, and K\. Bandara\(2021\)Recurrent neural networks for time series forecasting: current status and future directions\.International Journal of Forecasting37\(1\),pp\. 388–427\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[31\]Y\. Wang, Y\. Qiu, P\. Chen, Y\. Shu, Z\. Rao, L\. Pan, B\. Yang, and C\. Guo\(2025\)LightGTS: a lightweight general time series forecasting model\.arXiv preprint arXiv:2506\.06005\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[32\]X\. Wang, R\. Girshick, A\. Gupta, and K\. He\(2018\)Non\-local neural networks\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 7794–7803\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2018.00813)Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[33\]Z\. Liu, Y\. Lin, Y\. Cao, H\. Hu, Y\. Wei, Z\. Zhang, S\. Lin, and B\. Guo\(2021\)Swin transformer: hierarchical vision transformer using shifted windows\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 10012–10022\.External Links:[Document](https://dx.doi.org/10.1109/ICCV48922.2021.00986)Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[34\]K\. Han, Y\. Wang, H\. Chen, X\. Chen, J\. Guo, Z\. Liu, Y\. Tang, A\. Xiao, C\. Xu, Y\. Xu,et al\.\(2023\)A survey on vision transformer\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(1\),pp\. 87–110\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[35\]X\. Zhu, W\. Su, L\. Lu, B\. Li, X\. Wang, and J\. Dai\(2021\)Deformable detr: deformable transformers for end\-to\-end object detection\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gZ9hCDWe6ke)Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[36\]X\. Ma, Z\. Ni, S\. Xiao, and X\. Chen\(2025\)TimePro: efficient multivariate long\-term time series forecasting with variable\-and time\-aware hyper\-state\.arXiv preprint arXiv:2505\.20774\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[37\]N\. M\. Noor, M\. M\. Al Bakri Abdullah, A\. S\. Yahaya, and N\. A\. Ramli\(2015\)Comparison of linear interpolation method and mean method to replace the missing values in environmental data set\.InMaterials science forum,Vol\.803,pp\. 278–281\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[38\]L\. F\. Burgette and J\. P\. Reiter\(2010\)Multiple imputation for missing data via sequential regression trees\.American journal of epidemiology172\(9\),pp\. 1070–1076\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[39\]M\. G\. Rahman and M\. Z\. Islam\(2016\)Missing value imputation using a fuzzy clustering\-based em approach\.Knowledge and Information Systems46\(2\),pp\. 389–422\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1)\.
- \[40\]Y\. Liu, R\. Yu, S\. Zheng, E\. Zhan, and Y\. Yue\(2019\)NAOMI: non\-autoregressive multiresolution sequence imputation\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[41\]Y\. Xu, A\. Bazarjani, H\. Chi, C\. Choi, and Y\. Fu\(2023\)Uncovering the missing pattern: unified framework towards trajectory imputation and prediction\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 9632–9643\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p1.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[42\]C\. Yu, F\. Wang, C\. Yang, Z\. Shao, T\. Sun, T\. Qian, W\. Wei, Z\. An, and Y\. Xu\(2025\)Merlin: multi\-view representation learning for robust multivariate time series forecasting with unfixed missing rates\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 3633–3644\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p2.1)\.
- \[43\]Y\. Huang, X\. Mao, S\. Guo, Y\. Chen, J\. Shen, T\. Li, Y\. Lin, and H\. Wan\(2025\)Std\-plm: understanding both spatial and temporal properties of spatial\-temporal data with plm\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 11817–11825\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p2.1)\.
- \[44\]S\. J\. Pan and Q\. Yang\(2010\)A survey on transfer learning\.IEEE Transactions on Knowledge and Data Engineering22\(10\),pp\. 1345–1359\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2009.191)Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p2.1)\.
- \[45\]S\. I\. Nikolenko\(2021\)Synthetic data for deep learning\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(9\),pp\. 3106–3124\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p2.1)\.
- \[46\]Z\. Wang, J\. Chen, and S\. C\. H\. Hoi\(2021\)Deep learning for image super\-resolution: a survey\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(10\),pp\. 3365–3387\.Cited by:[§II\-B](https://arxiv.org/html/2609.22217#S2.SS2.p2.1)\.
- \[47\]J\. Yoon, D\. Jarrett, and M\. Van der Schaar\(2019\)Time\-series generative adversarial networks\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p8.1)\.
- \[48\]A\. Desai, C\. Freeman, Z\. Wang, and I\. Beaver\(2021\)Timevae: a variational auto\-encoder for multivariate time series generation\.arXiv preprint arXiv:2111\.08095\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1),[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1)\.
- \[49\]L\. Yang, Z\. Zhang, Y\. Song, S\. Hong, R\. Xu, Y\. Zhao, Y\. Shao, W\. Zhang, B\. Cui, and M\. Yang\(2023\)Diffusion models: a comprehensive survey of methods and applications\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(12\),pp\. 15245–15268\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1)\.
- \[50\]F\. Croitoru, V\. Hondru, R\. T\. Ionescu, and M\. Shah\(2023\)Diffusion models in vision: a survey\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(9\),pp\. 10850–10869\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1)\.
- \[51\]H\. Cao, C\. Tan, Z\. Gao, Y\. Xu, G\. Chen, P\. Heng, and S\. Z\. Li\(2024\)A survey on generative diffusion models\.IEEE Transactions on Pattern Analysis and Machine Intelligence\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1)\.
- \[52\]Z\. Pan, W\. Yu, X\. Yi, A\. Khan, F\. Yuan, and Y\. Zheng\(2021\)A survey on generative adversarial networks: variants, applications, and training\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(10\),pp\. 3213–3232\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1)\.
- \[53\]J\. Gui, Z\. Sun, Y\. Wen, D\. Tao, and J\. Ye\(2023\)A review on generative adversarial networks: algorithms, theory, and applications\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(4\),pp\. 4189–4210\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1)\.
- \[54\]S\. Bond\-Taylor, A\. Leach, Y\. Long, and C\. G\. Willcocks\(2022\)Deep generative modelling: a comparative review of vaes, gans, normalizing flows, energy\-based and autoregressive models\.IEEE Transactions on Pattern Analysis and Machine Intelligence44\(11\),pp\. 7327–7347\.Cited by:[§II\-C](https://arxiv.org/html/2609.22217#S2.SS3.p1.1)\.
- \[55\]Z\. Kong, W\. Ping, J\. Huang, K\. Zhao, and B\. Catanzaro\(2021\)DiffWave: a versatile diffusion model for audio synthesis\.InInternational Conference on Learning Representations,Cited by:[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1)\.
- \[56\]J\. L\. Alcaraz and N\. Strodthoff\(2022\)Diffusion\-based time series imputation and forecasting with structured state space models\.Transactions on Machine Learning Research\.External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=hHiIbk7ApW)Cited by:[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1)\.
- \[57\]Y\. Xu and Y\. Fu\(2025\)Sports\-traj: a unified trajectory generation model for multi\-agent movement in sports\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p3.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[58\]M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. Hochreiter\(2017\)GANs trained by a two time\-scale update rule converge to a local nash equilibrium\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 6629–6640\.Cited by:[§IV\-A](https://arxiv.org/html/2609.22217#S4.SS1.p8.1)\.
- \[59\]Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. Long\(2024\)ITransformer: inverted transformers are effective for time series forecasting\.InThe Twelfth International Conference on Learning Representations,Cited by:[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1)\.
- \[60\]A\. F\. Ansari, L\. Stella, C\. Turkmen, X\. Zhang, P\. Mercado, H\. Shen, O\. Shchur, S\. S\. Rangapuram, S\. Arango, S\. Kapoor, D\. C\. Maddix,et al\.\(2024\)Chronos: learning the language of time series\.Transactions on Machine Learning Research\.Cited by:[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p2.1),[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p3.1)\.
- \[61\]European Centre for Medium\-Range Weather Forecasts\(2024\)IFS Documentation CY49R1 – Part V: Ensemble Prediction System\.Note:[https://www\.ecmwf\.int/en/elibrary/81373\-ifs\-documentation\-cy49r1\-part\-v\-ensemble\-prediction\-system](https://www.ecmwf.int/en/elibrary/81373-ifs-documentation-cy49r1-part-v-ensemble-prediction-system)Cited by:[§IV\-C](https://arxiv.org/html/2609.22217#S4.SS3.p3.1)\.
Songru Yangreceived the B\.S\. degree from Beihang University, Beijing, China in 2023\. He is pursuing the Ph\.D\. degree in the Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University\. His research interests include machine learning, parttern recognition and AI4S\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/ZiliLiu.jpg)Zili Liureceived the Ph\.D\. degree from the Image Processing Center, School of Astronautics, Beihang University, in 2025\.He is currently a Researcher at Shanghai Artificial Intelligence \(AI\) Laboratory, Shanghai, China\. His research interests include deep learning, AI for meteorology, remote sensing image processing, and efficient deep learning\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/Taohan.jpg)Tao Hanreceived a B\.E\. degree in transportation equipment and control engineering and an M\.S\. degree in computer science and technology from Northwestern Polytechnical University, Xi’an, China, in 2019 and 2022\. He is currently pursuing a Ph\.D\. degree in computer science and engineering at the Hong Kong University of Science and Technology\. His research interests include computer vision, ai4science, and AIGC\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/BenFei.jpg)Ben Feireceived the M\.S\. degree in the Department of Materials Science from Fudan University, Shanghai, China, in 2021\. He received a Ph\.D\. degree in the School of Computer Science at Fudan University\. He is currently a postdoctoral fellow at Multimedia Lab, Department of Information Engineering, The Chinese University of Hong Kong, Hong Kong SAR\. His research interests include generative models, 3D computer vision and AI for Science\. He is an IEEE Young Professional\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/LeiBai.png)Lei Baireceived the Ph\.D\. degree from the University of New South Wales, Sydney, NSW, Australia, in 2021\.He was a Post\-Doctoral Researcher with the University of Sydney, Camperdown, NSW, Australia\. He is currently a Research Scientist with Shanghai Artificial Intelligence \(AI\) Laboratory, Shanghai, China\. He has authored or co\-authored a set of peer\-reviewed papers in top AI conferences and journals, such as Neural Information Processing Systems \(NeurIPS\), Conference on Computer Vision and Pattern Recognition \(CVPR\), International Joint Conference on Artificial Intelligence \(IJCAI\), Knowledge Discovery and Data Mining \(KDD\), International Conference on Computer Vision \(ICCV\), International Conference on Ubiquitous Computing \(Ubicomp\), IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE \(TPAMI\), and IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS \(TITS\)\. His research interests include machine learning, spatial\-temporal learning, and their applications \(e\.g\., Earth System Science and Smart City\)\.Dr\. Bai is or was a Program Committee Member or Reviewer for IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, NeurIPS, International Conference on Machine Learning \(ICML\), International Conference on Learning Representations \(ICLR\), CVPR, ICCV, Association for the Advancement of Artificial Intelligence \(AAAI\), IJCAI, KDD, European Conference on Computer Vision \(ECCV\), IEEE TRANSACTIONS ON IMAGE PROCESSING, IEEE TRANSACTIONS ON MULTIMEDIA, and ACM Transactions on Sensor Networks\. He was a recipient of the 2020 Google Ph\.D\. Fellowship, the 2020 UNSW Engineering Excellence Award, and the 2021 Dean’s Award for Outstanding Ph\.D\. Theses\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/ChangLiu.png)Chang Liureceived the B\.S\. degree from Jilin University, Changchun, Jilin, China, in 2012, and the Ph\.D\. degree from the University of Chinese Academy of Sciences, Beijing, China, in 2022\. He is currently a Post\-Doctoral Researcher with the Department of Automation, School of Information Science and Technology, Tsinghua University, Beijing\. He has published more than 50 papers in refereed conferences and journals\. His research interests include computer vision and machine learning\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/ZhengxiaZou.jpg)Zhengxia Zou\(Senior Member, IEEE\) received his BS degree and his Ph\.D\. degree from Beihang University in 2013 and 2018\. He is currently a Professor at the Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University\. During 2018\-2021, he was a postdoc research fellow at the University of Michigan, Ann Arbor\. His research interests include computer vision and related problems in remote sensing\. He has published over 30 peer\-reviewed papers in top\-tier journals and conferences, including Proceedings of the IEEE, Nature Communications, IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Transactions on Geoscience and Remote Sensing, and IEEE / CVF Computer Vision and Pattern Recognition\. Dr\. Zou serves as the Associate Editor for IEEE Transactions on Image Processing\. His personal website is[https://zhengxiazou\.github\.io/](https://zhengxiazou.github.io/)\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/XiangyangJi.jpg)Xiangyang Ji\(Member, IEEE\) received the BE degree in materials science and the MS degree in computer science from the Harbin Institute of Technology, Harbin, China, in 1999 and 2001, respectively, and the PhD degree in computer science from the Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China\. He joined Tsinghua University, Beijing, in 2008, where he is currently a professor with the Department of Automation, School of Information Science and Technology\. He has authored more than 200 refereed conference and journal papers\. His current research interests include signal processing, computer vision, and computational photography\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/WanliOuyang.png)Wanli Ouyangreceived the Ph\.D\. degree from the Department of Electronic Engineering, Chinese University of Hong Kong, Hong Kong, in 2010\.He was an Associate Professor with The University of Sydney, Camperdown, NSW, Australia\. He is now a Professor with Shanghai Artificial Intelligence \(AI\) Laboratory, Shanghai, China\. His research interests include pattern recognition, machine learning, and AI for Science\.Dr\. Ouyang served as an Associate Editor for International Journal of Computer Vision \(IJCV\) and Pattern Recognition \(PR\), the Senior Area Chair for Conference on Computer Vision and Pattern Recognition \(CVPR\), and the Guest Editor for IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE \(TPAMI\)\.![[Uncaptioned image]](https://arxiv.org/html/2609.22217v1/Bios/ZhenweiShi.jpg)Zhenwei Shi\(Senior Member, IEEE\) is currently a Professor and Dean of the Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University\. He has authored or co\-authored over 200 scientific articles in refereed journals and proceedings, including the IEEE Transactions on Pattern Analysis and Machine Intelligence, the IEEE Transactions on Image Processing, the IEEE Transactions on Geoscience and Remote Sensing, the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\) and the IEEE International Conference on Computer Vision \(ICCV\)\. His current research interests include remote sensing image processing and analysis, computer vision, pattern recognition, and machine learning\.Prof\. Shi serves as an Editor for IEEE Transactions on Geoscience and Remote Sensing, Pattern Recognition, ISPRS Journal of Photogrammetry and Remote Sensing, Infrared Physics and Technology, etc\. His personal website is http://levir\.buaa\.edu\.cn/\.相似文章
Tianmu-TC: 用于全球热带气旋预报的物理约束生成式人工智能
本文介绍了Tianmu-TC,一个用于全球热带气旋预报的物理约束生成式人工智能框架,它在可靠性和计算效率方面优于传统系统。
GenONet:一种用于高分辨率降水临近预报的生成算子网络
GenONet是一种新颖的深度学习架构,它在生成对抗网络框架内使用深度算子网络进行高分辨率降水临近预报,能够生成清晰且物理一致的预测。
EO-WM:一种物理信息驱动的概率地球观测预测世界模型
EO-WM提出了一种视频扩散变换器,用于概率性地球观测预测,该模型融入了物理信息条件,以捕捉天气驱动的不确定性,从而在极端天气下实现了对植被指数的更好预测。
基于图信息流匹配的时空插补
GiFlow 是一个用于时空插补的图信息流匹配框架,它用图信息先验取代高斯先验,并使用结合了空间注意力、时间注意力和时空传播的混合向量场模型。在合成数据集和真实世界数据集上均优于现有最先进方法。
UniGD:一种用于工业检索的统一生成-判别框架
快手研究者提出UniGD,一个用于工业检索的统一生成-判别框架,将检索与相关性评分整合到单一模型中,并采用CAGE与CAM等技术来提升效果并降低延迟。在线A/B测试显示,广告收入提升5.78%,推理延迟降低33.1%。