SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

arXiv cs.LG Papers

Summary

SimCast-S2S is a generative latent-diffusion framework for probabilistic subseasonal precipitation forecasting that leverages transfer learning from climate simulations to outperform deep learning baselines and compete with operational systems.

arXiv:2608.26594v1 Announce Type: new Abstract: Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity. We introduce SimCast-S2S, a generative latent-diffusion framework for probabilistic S2S precipitation forecasting that addresses three major bottlenecks in data-driven prediction. First, because S2S prediction requires uncertainty quantification rather than only deterministic point forecasts, SimCast-S2S is the first data-driven system that uses a diffusion-based generative pipeline for S2S prediction, enabling effective sampling from the underlying conditional distribution. Second, since generating large probabilistic ensembles is computationally costly in physical space, SimCast-S2S instead operates in a compact latent space learned by variational autoencoders, enabling efficient large-ensemble generation. Third, diffusion models typically require large training datasets; SimCast-S2S overcomes this via transfer learning with low-rank adaptation (LoRA), pretraining on large ensembles of climate simulations before fine-tuning on limited reanalysis data. On reanalysis data, SimCast-S2S outperforms deep learning baselines, including convolutional neural networks and U-Net architectures. Notably, despite using only a subset of atmospheric input variables and no post-processing, bias correction, or calibration, SimCast-S2S remains competitive with, and in many cases outperforms, state-of-the-art operational systems such as the ECMWF-S2S baseline. These results indicate that latent generative modeling combined with simulation-to-reanalysis transfer learning offers an efficient and scalable path toward data-driven probabilistic S2S precipitation forecasting.
Original Article
View Cached Full Text

Cached at: 08/28/26, 09:42 AM

# SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations
Source: [https://arxiv.org/html/2608.26594](https://arxiv.org/html/2608.26594)
Hiep V\. DangAffiliation:Department of Environmental Sciences,Affiliation:University of Virginia, Charlottesville, VA 22903, United StatesEmail:[zgp2ps@virginia\.edu](mailto:)Antonios MamalakisAffiliation:School of Data Science, Department of Environmental Sciences,Affiliation:University of Virginia, Charlottesville, VA 22903, United StatesEmail:[npa4tg@virginia\.edu](mailto:)

###### Abstract

Subseasonal\-to\-seasonal \(S2S\) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity\. We introduce SimCast\-S2S, a generative latent\-diffusion framework for probabilistic S2S precipitation forecasting that addresses three major bottlenecks in data\-driven prediction\. First, because S2S prediction requires uncertainty quantification rather than only deterministic point forecasts, SimCast\-S2S is the first data\-driven system that uses a diffusion\-based generative pipeline for S2S prediction, enabling effective sampling from the underlying conditional distribution\. Second, since generating large probabilistic ensembles is computationally costly in physical space, SimCast\-S2S instead operates in a compact latent space learned by variational autoencoders, enabling efficient large\-ensemble generation\. Third, diffusion models typically require large training datasets; SimCast\-S2S overcomes this via transfer learning with low\-rank adaptation \(LoRA\), pretraining on large ensembles of climate simulations before fine\-tuning on limited reanalysis data\. On reanalysis data, SimCast\-S2S outperforms deep learning baselines, including convolutional neural networks and U\-Net architectures\. Notably, despite using only a subset of atmospheric input variables and no post\-processing, bias correction, or calibration, SimCast\-S2S remains competitive with, and in many cases outperforms, state\-of\-the\-art operational systems such as the ECMWF\-S2S baseline\. These results indicate that latent generative modeling combined with simulation\-to\-reanalysis transfer learning offers an efficient and scalable path toward data\-driven probabilistic S2S precipitation forecasting\.

††footnotetext:Source code for this study is available at:[https://github\.com/hiepdang\-ml/SimCast\-S2S](https://github.com/hiepdang-ml/SimCast-S2S)\.## 1Introduction

Subseasonal to seasonal \(S2S\) forecasting remains among the most important challenges in climate science, as it refers to the transition zone between weather and climate timescales\([National Academies of Sciences, Engineering, and Medicine, 2016](https://arxiv.org/html/2608.26594#bib.bib79);[Vitart et al\., 2017](https://arxiv.org/html/2608.26594#bib.bib1);[Pegion et al\., 2019](https://arxiv.org/html/2608.26594#bib.bib80);[Robertson et al\., 2020](https://arxiv.org/html/2608.26594#bib.bib3);[White and Coauthors, 2022](https://arxiv.org/html/2608.26594#bib.bib84)\)\. At short \(weather\) lead times, forecasts depend mainly on the initial conditions of the atmosphere, and weather prediction models can track the evolution of storms, fronts, and circulation patterns with relatively high accuracy before errors in initial conditions accumulate significantly\([Lorenz, 1969](https://arxiv.org/html/2608.26594#bib.bib15);[Bauer et al\., 2015](https://arxiv.org/html/2608.26594#bib.bib76);[Leutbecher and Palmer, 2008](https://arxiv.org/html/2608.26594#bib.bib77)\)\. At seasonal and longer lead times, the focus shifts away from individual weather events and toward slowly varying boundary conditions enforced by climate variability in sea surface temperature, soil moisture, snow cover, and other land–ocean interactions that provide sources of predictability\([Palmer and Anderson, 1994](https://arxiv.org/html/2608.26594#bib.bib85);[Koster et al\., 2010](https://arxiv.org/html/2608.26594#bib.bib81);[Mariotti and Coauthors, 2020](https://arxiv.org/html/2608.26594#bib.bib86);[Merryfield and Coauthors, 2020](https://arxiv.org/html/2608.26594#bib.bib87)\)\. S2S forecasting, usually about 3 to 6 weeks ahead, lies between these two regimes\([National Academies of Sciences, Engineering, and Medicine, 2016](https://arxiv.org/html/2608.26594#bib.bib79);[Vitart et al\., 2017](https://arxiv.org/html/2608.26594#bib.bib1);[Pegion et al\., 2019](https://arxiv.org/html/2608.26594#bib.bib80);[Robertson et al\., 2020](https://arxiv.org/html/2608.26594#bib.bib3);[White and Coauthors, 2022](https://arxiv.org/html/2608.26594#bib.bib84)\), where small errors in the initial atmospheric conditions have grown substantially, while at the same time, slower sources of variability do not yet yield stable predictive signals\. Skill therefore depends on whether models can capture a mixture of fine\-scale variability in atmospheric dynamics and slower Earth\-system processes\([Zhang, 2005](https://arxiv.org/html/2608.26594#bib.bib83);[Baldwin and Dunkerton, 2001](https://arxiv.org/html/2608.26594#bib.bib82);[Koster et al\., 2010](https://arxiv.org/html/2608.26594#bib.bib81);[Mariotti and Coauthors, 2020](https://arxiv.org/html/2608.26594#bib.bib86);[Merryfield and Coauthors, 2020](https://arxiv.org/html/2608.26594#bib.bib87);[Richter et al\., 2024](https://arxiv.org/html/2608.26594#bib.bib88)\)\.

Recent progress in physical modeling and computation has made it possible to produce S2S forecasts operationally on a global scale[Vitart et al\. \(2017\)](https://arxiv.org/html/2608.26594#bib.bib1);[Vitart and Robertson \(2018\)](https://arxiv.org/html/2608.26594#bib.bib2)\. A central development in modern S2S forecasting is the use of ensemble prediction systems \(EPSs\)\([Toth and Kalnay, 1993](https://arxiv.org/html/2608.26594#bib.bib89);[Molteni et al\., 1996](https://arxiv.org/html/2608.26594#bib.bib90);[Buizza et al\., 1999](https://arxiv.org/html/2608.26594#bib.bib29);[Gneiting and Raftery, 2005](https://arxiv.org/html/2608.26594#bib.bib91);[Leutbecher and Palmer, 2008](https://arxiv.org/html/2608.26594#bib.bib77);[Bougeault et al\., 2010](https://arxiv.org/html/2608.26594#bib.bib92);[Swinbank et al\., 2016](https://arxiv.org/html/2608.26594#bib.bib93);[Palmer, 2019](https://arxiv.org/html/2608.26594#bib.bib94);[Vitart et al\., 2017](https://arxiv.org/html/2608.26594#bib.bib1)\)\. In chaotic atmospheric systems, a single deterministic forecast is usually insufficient[Lorenz \(1969\)](https://arxiv.org/html/2608.26594#bib.bib15);[Palmer \(2000\)](https://arxiv.org/html/2608.26594#bib.bib16);[Leutbecher and Palmer \(2008\)](https://arxiv.org/html/2608.26594#bib.bib77)\. Thus, an EPS generates multiple forecast members by perturbing initial conditions and model physics\([Toth and Kalnay, 1993](https://arxiv.org/html/2608.26594#bib.bib89);[Molteni et al\., 1996](https://arxiv.org/html/2608.26594#bib.bib90);[Buizza et al\., 1999](https://arxiv.org/html/2608.26594#bib.bib29);[Gneiting and Raftery, 2005](https://arxiv.org/html/2608.26594#bib.bib91);[Leutbecher and Palmer, 2008](https://arxiv.org/html/2608.26594#bib.bib77);[Palmer, 2019](https://arxiv.org/html/2608.26594#bib.bib94)\)\. This approach is especially important at S2S lead times because forecast errors often grow exponentially as the atmosphere evolves before nonlinear saturation occurs\([Lorenz, 1963](https://arxiv.org/html/2608.26594#bib.bib95);[Lorenz, 1969](https://arxiv.org/html/2608.26594#bib.bib15);[Leith, 1974](https://arxiv.org/html/2608.26594#bib.bib96);[Leutbecher and Palmer, 2008](https://arxiv.org/html/2608.26594#bib.bib77)\), and an EPS provides a probabilistic description of possible future states\. Operational centers, such as the European Centre for Medium\-Range Weather Forecasts \(ECMWF\), now issue S2S ensemble forecasts out to 46 days, with products often summarized as weekly anomalies relative to model climatology\. The value of EPS forecasting is particularly clear for precipitation, which is one of the most difficult variables to model and predict\. Precipitation is strongly affected by nonlinear processes, including convection, moisture transport, land–atmosphere feedbacks, and tropical variability[Robertson et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib3)\. It is widely observed that small errors in circulation, moisture transport, and humidity fields usually lead to large errors in the location and intensity of precipitation[Fritsch and Carbone \(2004\)](https://arxiv.org/html/2608.26594#bib.bib13);[Ebert and McBride \(2000\)](https://arxiv.org/html/2608.26594#bib.bib14);[Arakawa \(2004\)](https://arxiv.org/html/2608.26594#bib.bib5)\. Forecasting centers partially address this problem with post\-processing, calibration, multi\-model combination, and statistical correction[Vannitsem et al\. \(2021\)](https://arxiv.org/html/2608.26594#bib.bib4), but they still face major limitations regarding systematic model biases and high computational costs\([Bauer et al\., 2015](https://arxiv.org/html/2608.26594#bib.bib76);[Watt\-Meyer et al\., 2021](https://arxiv.org/html/2608.26594#bib.bib97);[Ben\-Bouallègue et al\., 2024](https://arxiv.org/html/2608.26594#bib.bib98)\)\.

Statistical and machine\-learning \(ML\) models have become an increasingly important alternative, and in many cases a complementary approach, to conventional physics\-based forecasting systems[Rasp et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib7);[Hwang et al\. \(2019\)](https://arxiv.org/html/2608.26594#bib.bib8);[Dang and Nguyen \(2026\)](https://arxiv.org/html/2608.26594#bib.bib17);[Chen et al\. \(2024\)](https://arxiv.org/html/2608.26594#bib.bib9)\. These data\-driven models can exploit large archives of observations, reanalysis products, and dynamical model output to learn nonlinear relationships between atmospheric predictors and future target states\. Instead of numerically solving the governing equations at every forecast step, ML models are trained to approximate the forecast mapping directly from data\. This allows fast and highly parallelizable predictions during inference\. Recent studies have shown that such approaches can improve S2S forecasting when they are trained on large reanalysis datasets and multi\-model forecast archives[Hwang et al\. \(2019\)](https://arxiv.org/html/2608.26594#bib.bib8);[Chen et al\. \(2024\)](https://arxiv.org/html/2608.26594#bib.bib9)\. Furthermore, as will be shown later, ML models pretrained on large amounts of simulated climate data can be transferred to real\-world forecasting tasks through fine\-tuning on observations or reanalysis products; see also[Weyn et al\. \(2021\)](https://arxiv.org/html/2608.26594#bib.bib10);[Nguyen et al\. \(2023b\)](https://arxiv.org/html/2608.26594#bib.bib11);[Bodnar et al\. \(2025\)](https://arxiv.org/html/2608.26594#bib.bib12)\. This pretraining strategy can help the model learn general atmospheric structures before adapting to the observed climate system, and it may improve forecast skill compared to models trained from scratch on limited observational data\.

Traditional EPS outputs are very computationally expensive because they require repeated model integrations with perturbed initial conditions and model physics[Leutbecher and Palmer \(2008\)](https://arxiv.org/html/2608.26594#bib.bib77);[Buizza et al\. \(1999\)](https://arxiv.org/html/2608.26594#bib.bib29)\. A state\-of\-the\-art subcategory of ML methods, Generative AI, provides a framework for probabilistic S2S forecasting\. Generative models aim to approximate the entire conditional distribution of future atmospheric states conditioned on relevant large\-scale current climate states, allowing multiple plausible forecast samples to be drawn efficiently after training\. This makes them especially useful for variables such as precipitation, where forecast uncertainty is strongly shaped by nonlinear dynamics\. Recent diffusion\-based forecasting systems have shown that generative models can produce realistic probabilistic forecasts and emulate large ensemble distributions at substantially lower computational cost than traditional EPSs[Li et al\. \(2024\)](https://arxiv.org/html/2608.26594#bib.bib6);[Price et al\. \(2025a\)](https://arxiv.org/html/2608.26594#bib.bib20)\.

Despite their promise, generative models also have important limitations\. As noted earlier, purely data\-driven models are data\-thirsty, requiring large training datasets to infer atmospheric dynamics[Rasp et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib7);[Dueben and Bauer \(2018\)](https://arxiv.org/html/2608.26594#bib.bib21), and such datasets are often unavailable or limited, particularly beyond weather timescales\. In addition, not all generative frameworks are equally stable\. Generative adversarial networks \(GANs\), for example, have been widely used for realistic sample generation[Goodfellow et al\. \(2014\)](https://arxiv.org/html/2608.26594#bib.bib22), but their adversarial training procedure is known to suffer from instability, mode collapse, and sensitivity to optimization settings[Salimans et al\. \(2016\)](https://arxiv.org/html/2608.26594#bib.bib23);[Mescheder et al\. \(2018\)](https://arxiv.org/html/2608.26594#bib.bib24)\. Encoder–decoder probabilistic frameworks, including models related to variational autoencoders[Kingma and Welling \(2014a\)](https://arxiv.org/html/2608.26594#bib.bib27), provide a more stable alternative and have been used in recent machine\-learning S2S forecasting systems\. For example, FuXi\-S2S combines an encoder–decoder architecture with a perturbation module in the learned latent space, which has been shown to outperform ECMWF\-S2S for selected variables, including precipitation and outgoing longwave radiation[Chen et al\. \(2024\)](https://arxiv.org/html/2608.26594#bib.bib9), although the authors did not assess how well the model captured uncertainty\. However, such models still require substantial input variables, which makes training expensive and limits their accessibility\.

To address the above challenges, namely, i\) the need for large forecast ensembles to characterize S2S uncertainty, ii\) the large training datasets required by deep generative models, particularly under weak predictive signals and training stability issues, and iii\) the high computational cost of ensemble generation, we propose SimCast\-S2S, an efficient generative model based on the diffusion framework[Ho et al\. \(2020a\)](https://arxiv.org/html/2608.26594#bib.bib25);[Song et al\. \(2021b\)](https://arxiv.org/html/2608.26594#bib.bib26)\. SimCast\-S2S uses multiple domain\-specific variational autoencoders \(VAEs\) to compress high\-dimensional physical input fields into low\-dimensional Gaussian latent representations[Kingma and Welling \(2014a\)](https://arxiv.org/html/2608.26594#bib.bib27)\. A compact diffusion network is then trained in this latent space to map standard Gaussian noise, conditioned on past and current atmospheric states, to the future target latent representation\. In the final step, a separate decoder maps the generated latent representation back to the physical space to produce the final forecast \(see Methods\)\. This design follows the general idea of latent generative modeling, where operating in a compressed representation can substantially reduce computational cost compared to generating directly in the original high\-dimensional space[Rombach et al\. \(2022a\)](https://arxiv.org/html/2608.26594#bib.bib28)\.

This computational advantage is substantial in practice: on a standard A100 GPU, generating one ensemble member takes approximately 12 seconds, and different ensemble members can be produced in parallel across multiple GPUs\. As a result, forecasting speed scales almost linearly with the number of GPUs and generating a e\.g\., 100\-member forecast ensemble to better characterize forecast uncertainty becomes inexpensive compared to repeatedly running a full physics\-based numerical model\. Most importantly, this high degree of efficiency also makes it feasible to pre\-train SimCast\-S2S on a large ensemble of climate simulation output, which, we will show, addresses the need for large training datasets and substantially improves accuracy\. Specifically, we train the model on 28 ensemble members of the Community Earth System Model version 2 Large Ensemble \(CESM2\-LE\), each initialized from perturbed initial conditions[Rodgers et al\. \(2021\)](https://arxiv.org/html/2608.26594#bib.bib18)\. By pre\-training SimCast\-S2S on multiple simulation\-based realizations, the model can learn a broader mapping between large\-scale climate fields and future S2S precipitation variability rather than constraining training on a single and limited reanalysis dataset\. The learned representation is then used for real\-world forecasting by fine\-tuning the model on ECMWF Reanalysis v5 \(ERA5\), a global atmospheric reanalysis widely used as a reference dataset for weather and climate applications[Hersbach et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib30)\. Our results indicate that SimCast\-S2S outperforms other machine\-learning baselines and challenges the state\-of\-the\-art ECMWF\-S2S physics\-based forecast system on the ERA5 reanalysis dataset\. The improvements of SimCast\-S2S are shown in terms of both deterministic forecast accuracy and probabilistic uncertainty quantification\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig1.png)Figure 1:An ensemble of precipitation\-anomaly forecasts generated by SimCast\-S2S for 2\-week period from February 12, 2021 to February 25, 2021 \(ERA5 dataset\), conditioned on the atmospheric inputs from January 01, 2021 to January 28, 2021\. The ensemble mean and the corresponding ERA5 reanalysis are also provided\. All fields are expressed in units of standard deviation\. This example shows stronger ensemble agreement in the tropics, especially along the equatorial Pacific, and larger member\-to\-member variability over the extratropics\. However, coherent extratropical signals remain visible over the North Pacific, the North Atlantic, and much of the Southern Hemisphere storm\-track region\. These areas of high consensus are also consistent with the ERA5 target\.
## 2Results

SimCast\-S2S is first pretrained on 28 ensemble members from CESM2\-LE using five atmospheric variables across three pressure levels 200 hPa, 500 hPa, and 850 hPa, along with five single\-level variables \(all considered during the last 28 days\) to forecast precipitation during the 3rd and 4th week into the future\. The model is trained from 1950 to 2000, validated from 2001 to 2010, and tested in years from 2011 to 2014\. After pretraining, SimCast\-S2S is transferred to ERA5 dataset where it is fine\-tuned in the years from 1940 to 2010, validated from 2011 to 2020, and tested from 2021 to 2025 \(see Section[4](https://arxiv.org/html/2608.26594#S4)for more details\)\.

SimCast\-S2S is computationally efficient, making it feasible to generate large forecast ensembles\. A 100\-member SimCast\-S2S ensemble reproduces many of the large\-scale precipitation anomaly structures while still retaining realistic member\-to\-member variability \(see Fig[1](https://arxiv.org/html/2608.26594#S1.F1)\)\. This section provides a comprehensive quantitative evaluation of the quality of the forecasts using all the test samples\. Subsection[2\.1](https://arxiv.org/html/2608.26594#S2.SS1)compares the deterministic skill of SimCast\-S2S with other ML baselines \(e\.g\., convolutional neural networks, UNets etc\.\) and with the operational ECMWF\-S2S system\. Subsection[2\.2](https://arxiv.org/html/2608.26594#S2.SS2)evaluates probabilistic forecast skill across the 2021–2025 ERA5 test samples\. Subsection[2\.4](https://arxiv.org/html/2608.26594#S2.SS4)assesses the spatial realism of the generated precipitation fields\. Subsection[2\.3](https://arxiv.org/html/2608.26594#S2.SS3)assesses the uncertainty reproduced by SimCast\-S2S and ECMWF\-S2S, and evaluates how well the forecast ranges cover the true events\. Finally, subsection[2\.5](https://arxiv.org/html/2608.26594#S2.SS5)discusses the computational efficiency and scalability of SimCast\-S2S under different hardware configurations\.

### 2\.1Deterministic Skill Evaluation

Table 1:Deterministic forecast skill measured by the mean absolute error \(MAE\), expressed in units of standard deviation and scaled by10−210^\{\-2\}\. We report the average MAE across all test samples and the inter\-sample standard deviation\. The first result column evaluates models on held\-out CESM2 test data, and the remaining three result columns evaluate models on ERA5 test data under different training settings: CESM2\-only training, ERA5\-only training, and CESM2 pretraining followed by ERA5 fine\-tuning\. SimCast\-S2S and ECMWF\-S2S are evaluated using the ensemble mean of 100 generated members\. Bold values indicate the best performance on ERA5 among different models and settings\.Unit: 10\-2stdTrained onCESM2Trained onERA5Pretrained on CESM2& Finetuned on ERA5Model / Tested onCESM2ERA5ERA5ERA5Proposed generative modelSimCast\-S2S \(η=0\.0\\eta=0\.0\)1\.25±0\.061\.25\\pm 0\.061\.32±0\.061\.32\\pm 0\.061\.37±0\.061\.37\\pm 0\.061\.30±0\.06\\mathbf\{1\.30\\pm 0\.06\}SimCast\-S2S \(η=0\.5\\eta=0\.5\)1\.25±0\.061\.25\\pm 0\.061\.32±0\.061\.32\\pm 0\.061\.37±0\.061\.37\\pm 0\.061\.30±0\.06\\mathbf\{1\.30\\pm 0\.06\}SimCast\-S2S \(η=1\.0\\eta=1\.0\)1\.25±0\.061\.25\\pm 0\.061\.32±0\.071\.32\\pm 0\.071\.38±0\.071\.38\\pm 0\.071\.30±0\.06\\mathbf\{1\.30\\pm 0\.06\}Neural\-network baselinesCNN\-Small1\.43±0\.081\.43\\pm 0\.081\.42±0\.061\.42\\pm 0\.061\.34±0\.071\.34\\pm 0\.071\.38±0\.061\.38\\pm 0\.06CNN\-Medium1\.31±0\.071\.31\\pm 0\.071\.32±0\.061\.32\\pm 0\.061\.46±0\.061\.46\\pm 0\.061\.33±0\.061\.33\\pm 0\.06CNN\-Large1\.51±0\.091\.51\\pm 0\.091\.47±0\.081\.47\\pm 0\.081\.47±0\.071\.47\\pm 0\.071\.50±0\.081\.50\\pm 0\.08UNet\-Small1\.35±0\.081\.35\\pm 0\.081\.35±0\.061\.35\\pm 0\.061\.36±0\.071\.36\\pm 0\.071\.35±0\.071\.35\\pm 0\.07UNet\-Medium1\.32±0\.081\.32\\pm 0\.081\.32±0\.061\.32\\pm 0\.061\.36±0\.081\.36\\pm 0\.081\.33±0\.061\.33\\pm 0\.06UNet\-Large1\.31±0\.071\.31\\pm 0\.071\.32±0\.061\.32\\pm 0\.061\.35±0\.071\.35\\pm 0\.071\.33±0\.061\.33\\pm 0\.06Operational baselineECMWF\-S2S1\.32±0\.061\.32\\pm 0\.06Although the atmosphere is highly chaotic and a deterministic forecast cannot represent the full range of possible future states, deterministic skill remains an essential first\-order measure of forecast quality, as it assesses the conditional mean of the underlying distribution of future states\. In an ensemble forecasting system, the ensemble mean summarizes the predictable component shared across members by averaging out member\-specific variability\. If the individual ensemble members are centered around physically meaningful future states, the ensemble mean should provide an accurate estimate of the expected precipitation anomalous field\. Evaluating deterministic skill tests whether SimCast\-S2S captures the dominant predictable signal and is also necessary for fair comparison with popular Deep Learning \(DL\) baselines, which often produce deterministic predictions\. We have evaluated the deterministic skill of SimCast\-S2S on the ERA5 test samples using its ensemble\-mean prediction, and compared it against the ensemble\-mean prediction of ECMWF\-S2S and deterministic predictions from different variants of Convolutional Neural Networks \(CNNs\)[LeCun et al\. \(1998a\)](https://arxiv.org/html/2608.26594#bib.bib31)and UNets[Ronneberger et al\. \(2015\)](https://arxiv.org/html/2608.26594#bib.bib32)\. Each CNN and UNet variant corresponds to a specific model size—small, medium, or large—determined by the number of layers and hidden features\. Architectural details are provided in Section[4](https://arxiv.org/html/2608.26594#S4)\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig2.png)Figure 2:Validation anomaly correlation coefficient \(ACC\) for SimCast\-S2S under different sampling parametersη\\etaand for ECMWF\-S2S\. ACC is calculated using the ensemble mean forecast\. Dark bars show unweighted ACC, and light bars show latitude\-weighted ACC\. Red outlines indicate cases where SimCast\-S2S has significantly higher ACC than ECMWF\-S2S under a one\-sided bootstrap test at the2\.5%2\.5\\%significance level\.On the CESM2 test set, SimCast\-S2S achieves the lowest mean absolute error \(MAE\) among all evaluated models and training settings \(see the first column in Table[1](https://arxiv.org/html/2608.26594#S2.T1)\)\. All three sampling settings obtain\(1\.25±0\.06\)×10−2\(1\.25\\pm 0\.06\)\\times 10^\{\-2\}std, while the best CNN and UNet baselines reach only\(1\.31±0\.07\)×10−2\(1\.31\\pm 0\.07\)\\times 10^\{\-2\}std\. The baseline results also show that increasing model size does not guarantee better performance\. CNN\-Large performs worse than CNN\-Medium, indicating possible overfitting\. Similarly, UNet\-Large has substantially more parameters than UNet\-Medium, but it produces nearly the same error, which suggests that the UNet architecture gains little from further scaling\. SimCast\-S2S also exhibits competitive simulation\-to\-reanalysis transfer without ERA5 fine\-tuning\. When trained only on CESM2 and evaluated directly on ERA5, SimCast\-S2S achieves an MAE of1\.32×10−21\.32\\times 10^\{\-2\}std, lower than most CNN and UNet baselines and comparable to the best\-performing UNet variants \(see the second column in Table[1](https://arxiv.org/html/2608.26594#S2.T1)\)\. This result indicates that CESM2 training alone learns precipitation features that partially transfer to ERA5\. However, the remaining error also shows a clear simulation\-to\-reanalysis gap\. This gap \(although marginal\) may be partially attributable to imperfect/biased representations of the S2S relevant processes in CESM2 simulations that have led to a slightly compromised learning that is revealed when models are tested on real\-world data\. Regarding different sizes of the baseline DL models, results mirror the behavior of when tested on the CESM2 test set\. SimCast\-S2S performs worst when trained only on ERA5 \(see the third column in[1](https://arxiv.org/html/2608.26594#S2.T1)\), with MAE increasing to1\.371\.37–1\.38×10−21\.38\\times 10^\{\-2\}std, a large degradation especially when compared to1\.32×10−21\.32\\times 10^\{\-2\}std in CESM2 training\. This result indicates that the diffusion\-based SimCast\-S2S requires more training data, and training on ERA5 alone was not sufficient\. Large simulated ensembles are therefore important for training the generative mapping before adapting the model to reanalysis data[Rasp and Thuerey \(2021\)](https://arxiv.org/html/2608.26594#bib.bib104);[Ham et al\. \(2019\)](https://arxiv.org/html/2608.26594#bib.bib105);[González\-Abad et al\. \(2026\)](https://arxiv.org/html/2608.26594#bib.bib106);[Unal et al\. \(2023\)](https://arxiv.org/html/2608.26594#bib.bib107);[Martin et al\. \(2025\)](https://arxiv.org/html/2608.26594#bib.bib108)\. Pretraining on CESM2 followed by ERA5 fine\-tuning yields the best performance on the ERA5 test set \(see the last column in Table[1](https://arxiv.org/html/2608.26594#S2.T1)\)\. Under this setting, all three SimCast\-S2S variants achieve\(1\.30±0\.06\)×10−2\(1\.30\\pm 0\.06\)\\times 10^\{\-2\}std, which is the lowest MAE on ERA5 in the table\. SimCast\-S2S outperforms all CNN and UNet baselines and also improves over the operational ECMWF\-S2S benchmark, which obtains\(1\.32±0\.06\)×10−2\(1\.32\\pm 0\.06\)\\times 10^\{\-2\}std on the same ERA5 test samples\. It is clear that CESM2 pretraining exposes the model to a larger set of physically plausible atmospheric evolutions and ERA5 fine\-tuning adapts this learned representation to the observation\-constrained atmosphere\. Transfer learning is the key strategy that allows SimCast\-S2S to overcome a central limitation of generative models, namely their dependence on large training datasets\. The simulation\-to\-reanalysis transfer strategy, rather than the architecture alone, enables SimCast\-S2S to outperform the operational baseline\. More details on the transfer learning method are discussed in Section[4\.4](https://arxiv.org/html/2608.26594#S4.SS4)\.

The deterministic results also show that the choice of sampling parameterη\\eta, which controls the trade\-off between diversity and stability \(see Subsection[4\.3](https://arxiv.org/html/2608.26594#S4.SS3)in Methods\), has little effect on the ensemble\-mean MAE\. Across all training settings, the three SimCast\-S2S variants produce nearly identical deterministic errors\. This is somewhat expected because MAE is evaluated using the ensemble mean, which averages out member\-specific sampling variability\. However, differentη\\etavalues result in different anomaly correlation coefficients \(ACC\)\. SimCast\-S2S achieves substantially higher mean ACC than ECMWF\-S2S forη=0\\eta=0andη=0\.5\\eta=0\.5, whereas its performance forη=1\.0\\eta=1\.0is closer to that of ECMWF\-S2S over the 2021–2025 validation period \(see Fig\.[2](https://arxiv.org/html/2608.26594#S2.F2)\)\. Given the similar MAE across sampling settings, the higher ACC forη=0\\eta=0andη=0\.5\\eta=0\.5indicates better spatial correspondence between the predicted and the true precipitation patterns rather than simply smaller point\-wise error magnitudes\. As described in Section[4](https://arxiv.org/html/2608.26594#S4),η\\etacontrols the sampling behavior of SimCast\-S2S and is not relevant to training\. We therefore treatη\\etaas an untrainable hyperparameter and select it using validation ACC reported in Fig[2](https://arxiv.org/html/2608.26594#S2.F2)\. Among the three experimented values,η=0\\eta=0achieves the highest ACC\. We useη=0\\eta=0as the default sampling setting for SimCast\-S2S, and all remaining analyses in this section report SimCast\-S2S results under this choice\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig3.png)Figure 3:Probabilistic forecast skill of ECMWF\-S2S and SimCast\-S2S for precipitation anomalies\.Top row:Spatial distribution of tercile\-based Ranked Probability Skill Score \(RPSS\)\.Bottom row:Continuous Ranked Probability Skill Score \(CRPSS\)\. For each metric, the left and middle columns show the skill scores of ECMWF\-S2S and SimCast\-S2S, respectively, while the right column shows the difference between SimCast\-S2S and ECMWF\-S2S\. Positive values in the left and middle columns indicate skill above the climatological baseline, whereas negative values indicate performance below it\. Positive values in the right column indicate higher skill for SimCast\-S2S relative to ECMWF\-S2S\. Dots indicate regions where the corresponding positive skill score, or positive difference in skill score, is statistically significant at the2\.5%2\.5\\%level based on 1,000 bootstrap resamples\.![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig4.png)Figure 4:Seasonal CRPSS for SimCast\-S2S and ECMWF\-S2S precipitation forecasts during the 2021–2025 test period over the global, tropical, and extratropical domains\. Seasons are defined as winter \(December–February\), spring \(March–May\), summer \(June–August\), and fall \(September–November\)\. Positive CRPSS indicates better probabilistic skill than the climatological baseline\. Higher values indicate better performance\.
### 2\.2Probabilistic Skill Evaluation

Deterministic skill evaluation is limiting for subseasonal forecasts because small perturbations in the initial state can lead to substantially different precipitation outcomes; thus, a point forecast can at best approximate the expected value of the conditional distribution and may differ substantially from any individual realization because of the system’s inherent internal variability\. Since SimCast\-S2S is inherently stochastic, it produces a distribution of possible forecasts rather than a point estimate\. We next evaluate how well the forecast distribution produced by SimCast\-S2S covers the observed ERA5 target\.

Probabilistic forecast skill is strongest in the tropics for both ECMWF\-S2S and SimCast\-S2S, particularly along the equatorial Pacific, as measured by the ranked probability skill score \(RPSS\) for tercile forecasts \(see Fig\.[3](https://arxiv.org/html/2608.26594#S2.F3)\)\. SimCast\-S2S not only exhibits stronger tropical skill pattern but also shows more widespread increased skill relative to ECMWF\-S2S across many extratropical regions, especially in the North Pacific, North Asia, North America, the North Atlantic, and the Southern Ocean\. This advantage is more clearly visible in the difference map, where positive values indicate higher skill for SimCast\-S2S \(see red shading\)\. The continuous ranked probability skill score \(CRPSS\) in Fig\.[3](https://arxiv.org/html/2608.26594#S2.F3), which evaluates the full predictive distribution rather than discrete tercile categories, shows a broadly consistent spatial pattern\. SimCast\-S2S exhibits higher CRPSS than ECMWF\-S2S across most of the globe, indicating that its advantage extends beyond tercile\-category prediction to the continuous forecast distribution\. It should also be noted that some persistently dry regions, such as the Sahara, the southeastern Pacific off South America, and Antarctica, have precipitation values that are zero or near zero for much of the time, with low temporal variability\. In these regions, the verifying precipitation anomalies are small, and both models’ forecasts remain close to the reference climatology\.

The spatial patterns in Fig\.[3](https://arxiv.org/html/2608.26594#S2.F3)show where probabilistic skill is concentrated\. Fig\.[4](https://arxiv.org/html/2608.26594#S2.F4)examines which season this skill occurs\. Spring and summer appear to be the most difficult seasons to predict subseasonal precipitation variability\. ECMWF\-S2S shows negative CRPSS over the extratropics in these seasons; however, SimCast\-S2S still manages to maintain positive CRPSS\. In fact, SimCast\-S2S shows higher CRPSS than ECMWF\-S2S across all seasons and regions, with the largest differences occurring in the tropics\. This agrees with the spatial maps in Fig\.[3](https://arxiv.org/html/2608.26594#S2.F3), where the strongest probabilistic skill is concentrated along the tropical regions\. In the extratropics, both models exhibit weaker skill\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig5.png)Figure 5:Probability integral transform \(PIT\) histograms for SimCast\-S2S and ECMWF\-S2S precipitation forecasts\. The red dashed line indicates the uniform distribution expected under perfect calibration\. Both models show U\-shaped histograms, indicating underdispersive ensemble forecasts\. SimCast\-S2S shows nearly constantχ2\\chi^\{2\}distance across regions, whereas ECMWF\-S2S shows better extratropical calibration but larger tropical deviation from uniform behavior\.
### 2\.3Uncertainty Quantification

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig6.png)Figure 6:Spread–error and interval\-score diagnostics for SimCast\-S2S and ECMWF\-S2S precipitation forecasts over global, tropical, and extratropical domains\. Thetop rowcompares mean ensemble spread with the RMSE of the ensemble mean\. Both models are above the ideal one\-to\-one dashed line, indicating forecast errors are larger than ensemble spread\. Spread–error correlations are high\. This suggests that larger spread is still associated with larger error and the ensemble spread truly reflects useful uncertainty information\. Thebottom rowshows mean interval scores for different central interval coverages, where lower values indicate better probabilistic forecasts\. SimCast\-S2S consistently exhibits better score, particularly in tropical regions\.Beyond aggregate probabilistic skill, it is important to examine how well each model represents forecast uncertainty\. A useful ensemble forecast should not only produce accurate probabilistic predictions but also assign uncertainty in a way that is statistically consistent with the reanalysis data\. In this section, we evaluate uncertainty quantification from two complementary perspectives\. First, we examine the percentile histogram to assess the calibration of the ensemble distribution across global, tropical, and extratropical regions\. Second, we analyze the relationship between ensemble spread and forecast error, together with interval scores for different nominal coverage levels\.

A well\-calibrated probabilistic forecast should produce an approximately uniform probability integral transform \(PIT\) distribution, indicating that the verifying observed/reanalysis estimates are evenly distributed across the forecast distribution\. Both SimCast\-S2S and ECMWF\-S2S exhibit U\-shaped PIT histograms, with elevated frequencies near the lowest and highest percentiles across the global, tropical, and extratropical regions \(see Fig\.[5](https://arxiv.org/html/2608.26594#S2.F5)\)\. This pattern indicates underdispersion, meaning that the forecast distributions are narrow and assign insufficient probability to outcomes in the tails\. In order words, both SimCast\-S2S and ECMWF\-S2S tend to underestimate forecast uncertainty\.

The degree of underdispersion varies by region and model\. Globally, ECMWF\-S2S shows a smallerχ2\\chi^\{2\}distance from uniform behavior than SimCast\-S2S\. This global difference is mainly driven by the extratropics, where ECMWF\-S2S has a lowerχ2\\chi^\{2\}distance\. In the tropics, however, ECMWF\-S2S has a higherχ2\\chi^\{2\}distance than SimCast\-S2S\. That being said, this comparison should be interpreted carefully\. SimCast\-S2S is evaluated directly from its raw ensemble output, without additional statistical calibration or post\-processing\. ECMWF\-S2S, in contrast, is accompanied by post\-processed forecasts that are commonly used to adjust systematic forecast biases and construct calibrated anomaly or probabilistic products\. Therefore, the ECMWF\-S2S ensemble evaluated here may not represent an equally raw forecast product in the same sense as in SimCast\-S2S\. The relatively stableχ2\\chi^\{2\}distance across regions for SimCast\-S2S suggests more spatially uniform ensemble dispersion, whereas the larger tropical–extratropical contrast for ECMWF\-S2S indicates stronger regional dependence signaling regional overfitting in the post\-processing calibration\.

Another key diagnostic of ensemble reliability is the relationship between ensemble spread and forecast error\. In a well\-scaled ensemble, these two quantities should be close to the one\-to\-one line, so that larger forecast uncertainty corresponds to larger realized error\. Both SimCast\-S2S and ECMWF\-S2S lie slightly above this line in all three regions of focus, indicating that the realized errors are larger than the ensemble spread \(see the top row of Fig\.[6](https://arxiv.org/html/2608.26594#S2.F6)\)\. Although high spread–error correlations are observed in both models, the main difference is that SimCast\-S2S lies closer to the one\-to\-one line in most cases over the tropics, and hence provides a better\-scaled estimate of forecast uncertainty than ECMWF\-S2S in these regions\. Over the extratropics, SimCast\-S2S slightly underestimates the uncertainty compared to ECMWF\-S2S\.

The ensemble prediction intervals are further evaluated using the mean interval score \(see the bottom row of Fig\.[6](https://arxiv.org/html/2608.26594#S2.F6)\)\. The x\-axis shows the nominal central interval coverage\. For example, a90%90\\%central interval spans the 5th to 95th percentiles of the ensemble distribution, corresponding to the central90%90\\%of the forecast distribution around the median\. The y\-axis shows the mean interval score\. This score penalizes two types of behavior: intervals that are unnecessarily wide and intervals that fail to contain the ground truth\. Therefore, a lower interval score indicates both better sharpness and coverage\. A formal description of mean interval score can be found in Section \. As expected, the interval score increases as the nominal coverage becomes larger because higher coverage requires wider prediction intervals\. Across the global, tropical, and extratropical regions, SimCast\-S2S shows slightly lower error than ECMWF\-S2S, particularly at the highest coverage levels over the tropical regions\. This is consistent with the PIT histogram in Fig\.[5](https://arxiv.org/html/2608.26594#S2.F5)\.

### 2\.4Spatial Structure Realism

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig7.png)Figure 7:Spatial autocorrelation of ERA5, SimCast\-S2S, and ECMWF\-S2S precipitation fields along zonal \(west–east\), meridional \(south–north\), main\-diagonal \(northwest–southeast\), and anti\-diagonal \(southwest–northeast\) directions\. The x\-axis shows spatial lag, and the y\-axis shows autocorrelation\. Smaller differences from the ERA5 distributions indicate better agreement with the ground truth regarding spatial structure of precipitation\. SimCast\-S2S better follows the true meridional and diagonal autocorrelation, especially at smaller spatial lags where fine\-scale precipitation structure remains meaningful and most impactful\.We next examine whether the forecasted precipitation fields have realistic spatial structure\. This is important because a model can achieve good aggregate scores while still producing fields that are too smooth, too noisy, or spatially inconsistent with true precipitation\. Spatial autocorrelation is evaluated in four directions: zonal, meridional, main diagonal, and anti\-diagonal \(see Fig\.[7](https://arxiv.org/html/2608.26594#S2.F7)\)\. The meridional direction is especially important because it provides a strong test of whether the models capture realistic transitions across latitude\-dependent precipitation regimes\.

Across all four directions, the ERA5 fields show decreasing autocorrelation with increasing spatial lag, and both SimCast\-S2S and ECMWF\-S2S are able to reproduce this basic decay\. In the zonal direction, SimCast\-S2S and ECMWF\-S2S show nearly equal performance\. SimCast\-S2S and ECMWF\-S2S both track the observed decay reasonably well although both models show slightly stronger west\-east persistence than the reanalysis data\. The meridional direction provides a more demanding test because precipitation regimes vary strongly with latitude, spanning e\.g\., deep convective precipitation tied to the intertropical convergence zone in the tropics, suppressed precipitation beneath the subsiding branch of the Hadley circulation in the subtropics, frontal and baroclinic precipitation along midlatitude storm tracks, and weaker orographically\-driven precipitation in the high\-latitudes\. In this direction, SimCast\-S2S follows the observed decay more closely than ECMWF\-S2S across most lags whereas ECMWF\-S2S unfavorably persists more spatial structures across latitude bands\.

The main diagonal and anti\-diagonal directions follow the northwest–southeast and southwest–northeast directions, respectively\. In both directions, SimCast\-S2S is closer to the true autocorrelation than ECMWF\-S2S at smaller spatial lags, where local precipitation structure is still meaningfully present\. As the lag increases, the autocorrelation quickly collapses to zero\. This is the distance range where precipitation at one location has little statistical connection to precipitation farther away\. Consequently, the advantage of SimCast\-S2S diminishes \(although still present\), not because the models become equally realistic, but because there is little remaining spatial dependence for either model to reproduce\.

### 2\.5Computational Efficiency and Scalability

A practical advantage of SimCast\-S2S is that probabilistic forecasts can be generated rapidly once the model is deployed\. Conventional subseasonal systems such as ECMWF\-S2S generate probabilistic forecasts by repeatedly integrating a full numerical model forward in time for many ensemble members\. SimCast\-S2S instead produces ensemble members through stochastic sampling during the denoising process of the diffusion model \(see Subsection[4\.3](https://arxiv.org/html/2608.26594#S4.SS3)\), so each member can be generated independently and parallelized efficiently across GPUs\.

SimCast\-S2S can generate large precipitation ensembles efficiently at192×288192\\times 288resolution\. On a single A100 GPU, it takes approximately 12\.3 seconds per ensemble member, or 123 seconds for 10 members, 246 seconds for 20 members, 615 seconds for 50 members, and 1230 seconds, or 20\.5 minutes, for 100 members \(see Fig\.[8](https://arxiv.org/html/2608.26594#S2.F8)\)\. With three A100 GPUs, the same 100\-member ensemble requires 438\.21 seconds, or about 7\.3 minutes\. With four A100 GPUs, it decreases further to 342\.19 seconds, or about 5\.7 minutes\. Thus, a 100\-member ensemble can be generated in less than ten minutes using a small number of industry\-grade A100 GPUs\.

It should be noted that the scaling behavior is close to linear because ensemble generation is highly parallelizable\. Moving from one to two A100 GPUs reduces the 100\-member runtime from 1230\.0 seconds to 634\.19 seconds, corresponding to a parallel efficiency of12302×634\.19=0\.9697\\frac\{1230\}\{2\\times 634\.19\}=0\.9697\. With three GPUs, the efficiency is 0\.9356, and with four GPUs it is 0\.8986\. The efficiency decreases gradually as more GPUs are added, reaching 0\.7392 at eight GPUs \(see panel \(b\) of Fig\.[8](https://arxiv.org/html/2608.26594#S2.F8)\)\. This departure from ideal linear scaling is small and expected[Li et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib102);[Amdahl \(1967\)](https://arxiv.org/html/2608.26594#bib.bib101)\. It reflects fixed overheads from model loading, data movement, GPU synchronization, and inter\-device communication\. These overheads are more visible when the ensemble size is small, but they are increasingly amortized as the number of ensemble members grows[Gustafson \(1988\)](https://arxiv.org/html/2608.26594#bib.bib103)\. SimCast\-S2S also benefits directly from newer accelerator hardware\. For a 100\-member ensemble on one GPU, the wall\-clock time is 2213 seconds on V100, 1477 seconds on A6000, 1231 seconds on A100, and 683 seconds on H200 \(see panel \(a\) of Fig\.[8](https://arxiv.org/html/2608.26594#S2.F8)\)\. With eight H200 GPUs, the 100\-member runtime decreases to only 115\.55 seconds, or less than two minutes\. This is a strong indication that SimCast\-S2S can exploit both per\-GPU improvement and multi\-GPU parallelism\.

To our knowledge, ECMWF does not publicly disclose the wall\-clock integration time\. Products corresponding to the 00:00 UTC forecast base time are scheduled to become available at 20:00 UTC[European Centre for Medium\-Range Weather Forecasts \(\)](https://arxiv.org/html/2608.26594#bib.bib99)\. However, this 20\-hour interval likely encompasses an end\-to\-end operational workflow, including initialization, numerical integration, post\-processing, and dissemination, not just model runtime\. An earlier ECMWF benchmark nevertheless provides some indication of the computational expense\. Its previous 51\-member, 15\-day ensemble required an average wall\-clock time of 82 minutes on 1,530 compute nodes[Bauer et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib100)\. Although this historical benchmark is not directly comparable with the runtime of SimCast\-S2S, it illustrates the large\-scale computational resources required by conventional dynamical ensemble forecasting\. In contrast, generating a 100\-member probabilistic ensemble with SimCast\-S2S is just a matter of minutes on a single GPU node\. This efficiency makes large\-ensemble S2S prediction widely accessible to typical research groups and also allows uncertainty estimates and extreme\-event probabilities to be produced with much lower latency\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig8.png)Figure 8:Computational efficiency and scalability of SimCast\-S2S ensemble generation\.Panel \(a\)shows wall\-clock inference time as a function of number of GPUs, GPU types and ensemble sizes\.Panel \(b\)compares measured A100 runtimes with ideal linear scaling; closer agreement with the dashed lines indicates better multi\-GPU scalability\.

## 3Discussion

We tackle three key bottlenecks that have limited data\-driven models for S2S forecasting\. First, generative models are notoriously data\-hungry; we address this through transfer learning via LoRA, pretraining on large ensembles of climate simulations before fine\-tuning on limited reanalysis data\. Second, generating large probabilistic ensembles at high resolution is computationally prohibitive; we address this by operating in a compact latent space rather than physical space, which makes large\-ensemble generation efficient\. Third, S2S forecasting fundamentally requires representing forecast uncertainty, not just a single deterministic outcome; we address this by using a generative \(diffusion\-based\) framework that directly models the forecast distribution rather than a point estimate\. We show that each of these design choices works: our model outperforms all deterministic AI baselines, and using only a subset of input variables and without any post\-processing or bias correction, it challenges a leading operational forecasting system\.

A central result of this study is that the higher probabilistic skill of SimCast\-S2S also translates into stronger deterministic performance\. SimCast\-S2S exhibits higher probabilistic skill across much of the globe as shown in Figs\.[3](https://arxiv.org/html/2608.26594#S2.F3)and[4](https://arxiv.org/html/2608.26594#S2.F4), consistent with the lower deterministic errors reported in Table[1](https://arxiv.org/html/2608.26594#S2.T1)\. This result suggests that the generated forecast members are not merely adding stochastic variability around a weak deterministic forecast; rather, they contain a coherent predictive signal that survives ensemble averaging and is closer to the true conditional distribution\.

Our uncertainty diagnostics in Subsection[2\.3](https://arxiv.org/html/2608.26594#S2.SS3)indicate that both SimCast\-S2S and ECMWF\-S2S remain underdispersive, i\.e\., their ensemble spread is narrow relative to realized true variability\. However, the degree and regional structure of this underdispersion differ\. ECMWF\-S2S shows stronger uncertainty behavior in the extratropics, while SimCast\-S2S is more competitive in the tropics and it overall lies closer to the spread–RMSE1:11\{:\}1reference\. This comparison is demanding for SimCast\-S2S because ECMWF\-S2S uses reforecasts to estimate model\-specific systematic errors; bias correction then reduces persistent offsets, and probabilistic calibration adjusts ensemble probabilities toward observed event frequencies\. In contrast, SimCast\-S2S is evaluated directly from its raw ensemble output without comparable post\-processing or calibration\. Therefore, its competitive tropical performance and improved spread–error scaling suggest that the generative ensemble already contains meaningful uncertainty information before any statistical correction is applied\.

Importantly, we show that SimCast\-S2S better captures the spatial decay of precipitation dependence, especially in the meridional and diagonal directions\. This is notable because precipitation regimes vary strongly with latitude, from tropical convection to subtropical suppression, midlatitude storm tracks and high\-latitude stratiform precipitation\. It is well known that deep convolutional forecast models trained with point\-wise loss functions favor overly smooth fields and suppress fine\-scale precipitation structure\([Mathieu et al\., 2016](https://arxiv.org/html/2608.26594#bib.bib33);[Ravuri et al\., 2021](https://arxiv.org/html/2608.26594#bib.bib34);[Harris et al\., 2022](https://arxiv.org/html/2608.26594#bib.bib35)\)\. In contrast, SimCast\-S2S learns the spatial organization of precipitation across dynamically distinct regimes\. This is one of the strongest arguments for using a diffusion\-based generative model in this setting\. The model is trained to sample from a learned distribution of atmospheric states, and therefore can preserve spatial dependence in ways that cannot be done by point\-wise loss functions alone\.

SimCast\-S2S also changes the practical cost structure of subseasonal ensemble prediction\. Conventional dynamical systems generate ensemble forecasts by repeatedly integrating a full numerical model forward in time\. SimCast\-S2S directly models the one\-to\-one mapping, from the current atmospheric state to the target weeks\. This is done through neural\-network stochastic sampling, where individual members can be generated independently and parallelized efficiently across accelerators\. As shown in Subsection[2\.5](https://arxiv.org/html/2608.26594#S2.SS5), a 100\-member ensemble can be generated in minutes on a small GPU node\. The same ensemble size would require thousands of CPU nodes and several hours of wall\-clock time in a conventional dynamical forecasting system[Bauer et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib100)\. This efficiency opens opportunities for research settings where operational\-scale computing is unavailable, and for applications that require many ensemble members, such as uncertainty quantification, tail\-risk assessment, scenario generation, and impact modeling\([Gneiting and Katzfuss, 2014](https://arxiv.org/html/2608.26594#bib.bib36);[Kurth et al\., 2023](https://arxiv.org/html/2608.26594#bib.bib38);[Price et al\., 2025b](https://arxiv.org/html/2608.26594#bib.bib39)\)\.

A limitation is that the realism of SimCast\-S2S is currently statistical rather than explicitly physical\. The model can learn realistic distributions, spatial structures, and predictor–target relationships, but it is not guaranteed to satisfy conservation laws or dynamical balances\. This limitation is closely tied to the same design choice that makes the model efficient: latent\-space sampling\. Working in a low\-dimensional latent representation largely reduces computational cost, but physical laws such as mass conservation, moisture conservation, and momentum balance only make sense in physical space\. There is no obvious way to impose these laws directly on abstract latent variables\. As a result, SimCast\-S2S may generate fields that are statistically plausible but not strictly dynamically consistent\. A natural next step is therefore to combine latent\-space generation with \(occasional\) physical\-space correction\. One possible direction is a hybrid sampling scheme in which the model generates most of the trajectory in latent space, periodically decodes the state into physical variables, evaluates physical residuals, applies corrections, and then maps the corrected state back into latent space\. For precipitation forecasting, this could involve constraints derived from the moisture budget, physical bounds on humidity and precipitation, or consistency between circulation and moisture convergence\. Such a scheme is feasible because neural networks are inherently differentiable, and physical residuals can be computed directly from their computation graph\([Baydin et al\., 2018](https://arxiv.org/html/2608.26594#bib.bib40);[Raissi et al\., 2019](https://arxiv.org/html/2608.26594#bib.bib41);[Karniadakis et al\., 2021](https://arxiv.org/html/2608.26594#bib.bib42);[Nguyen et al\., 2023a](https://arxiv.org/html/2608.26594#bib.bib43);[Nguyen et al\., 2024](https://arxiv.org/html/2608.26594#bib.bib44)\)\. Incorporating physical constraints into the generative model may also help improve representation of extremes, which remain a key weakness of both SimCast\-S2S and ECMWF\-S2S\.

Another important direction is interpretability\. SimCast\-S2S uses many atmospheric predictors, but the present analysis does not fully explain which fields, regions, or dynamical patterns the model relies on when generating forecasts\. Explainable AI methods could help determine whether the model uses physically meaningful sources of predictability, such as tropical moisture anomalies, upper\-level circulation, midlatitude wave activity, atmospheric rivers, or teleconnection patterns\([Mamalakis et al\., 2022b](https://arxiv.org/html/2608.26594#bib.bib49);[Mamalakis et al\., 2022a](https://arxiv.org/html/2608.26594#bib.bib50);[Higgins et al\., 2023](https://arxiv.org/html/2608.26594#bib.bib54);[Higgins et al\., 2024](https://arxiv.org/html/2608.26594#bib.bib55);[Singh et al\., 2023](https://arxiv.org/html/2608.26594#bib.bib51);[Singh and Goyal, 2023](https://arxiv.org/html/2608.26594#bib.bib56);[Goyal and Singh, 2024](https://arxiv.org/html/2608.26594#bib.bib52);[Corner et al\., 2026](https://arxiv.org/html/2608.26594#bib.bib53)\)\. Methods such as input perturbation, integrated gradients, latent\-space sensitivity analysis, and counterfactual sampling could be adapted to our generative setting\([Ribeiro et al\., 2016](https://arxiv.org/html/2608.26594#bib.bib46);[Sundararajan et al\., 2017](https://arxiv.org/html/2608.26594#bib.bib47);[Wachter et al\., 2017](https://arxiv.org/html/2608.26594#bib.bib48)\)\. This would be especially valuable because generative models can be right for the wrong reasons, and such failures are difficult to diagnose from forecast scores alone\. Therefore, connecting SimCast\-S2S skill to interpretable physical mechanisms becomes crucial for making the model more scientifically useful and for identifying regimes in which it is likely to fail\.

Our results demonstrate the potential of latent, diffusion\-based, generative modeling combined with simulation\-to\-reanalysis transfer learning as an efficient and scalable framework for a new generation of probabilistic S2S precipitation forecasts\. Further advances in physical consistency and interpretability could strengthen their reliability and scientific utility\.

## 4Methods

### 4\.1Data

SimCast\-S2S is designed to maximize computational efficiency so that it can be trained on large ensembles of climate\-model simulations and then transferred to observation\-constrained reanalysis data\. We first pretrain the model on 28 ensembles from the Community Earth System Model version 2 \(CESM2\), a fully coupled Earth system model that represents interactions among the atmosphere, ocean, land surface, sea ice, and other climate components[Danabasoglu et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib19);[Rodgers et al\. \(2021\)](https://arxiv.org/html/2608.26594#bib.bib18)\. After pretraining on CESM2 simulations, SimCast\-S2S is fine\-tuned and evaluated on the fifth\-generation European Centre for Medium\-Range Weather Forecasts atmospheric reanalysis, ERA5[Hersbach et al\. \(2020\)](https://arxiv.org/html/2608.26594#bib.bib30)\. ERA5 combines a global numerical weather prediction model with a large collection of real\-world observations through data assimilation, including information from satellites, radiosondes, aircraft, and surface stations\. Therefore, ERA5 is not a strictly observational dataset, but rather an observation\-constrained reconstruction of the atmosphere\. As a result, it provides one of the closest available approximations to the observed atmospheric state and can be used to transfer the model from idealized climate\-model simulations to realistic atmospheric conditions\.

To make transfer learning between CESM2 and ERA5 physically meaningful, the input variables must be defined consistently across the two datasets\. We therefore select variables that are available, or have close physical counterparts, in both CESM2 and ERA5\. The selected variables are also chosen to describe the atmospheric column across three representative pressure levels: 850 hPa for the lower troposphere, 500 hPa for the middle troposphere, and 200 hPa for the upper troposphere\. In addition, single\-level variables are included to describe surface and column\-integrated conditions that are directly linked to precipitation, including surface temperature, sea\-level pressure, outgoing longwave radiation, total precipitable water, and precipitation itself\.

The list of all CESM2 variables used for pretraining and their ERA5 counterparts used for fine\-tuning are provided in Table[2](https://arxiv.org/html/2608.26594#S4.T2)\. At each pressure level, we include wind, geopotential, temperature, and humidity\. These variables jointly describe the dynamical and thermodynamical state of the atmosphere: winds capture horizontal and vertical transport, geopotential represents large\-scale circulation structure, temperature describes thermal stratification, and humidity provides the moisture supply needed for precipitation\. At the surface or column\-integrated level, the selected variables provide additional constraints on boundary\-layer conditions, radiative state, pressure patterns, and atmospheric moisture content\.

Table 2:CESM2 input variables used for SimCast\-S2S pretraining and their corresponding ERA5 variables used for fine\-tuning\.Pressure levelVariable typeCESM2 variableERA5 variable200 hPaZonal windU200u\_200Meridional windV200v\_200TemperatureT200t\_200Specific humidityQ200q\_200Geopotential heightZ200z\_200500 hPaZonal windU500u\_500Meridional windV500v\_500Vertical velocityOMEGA500w\_500TemperatureT500t\_500Specific humidityQ500q\_500Geopotential heightZ500z\_500850 hPaZonal windU850u\_850Meridional windV850v\_850TemperatureT850t\_850Specific humidityQ850q\_850Geopotential heightZ850z\_850Single levelSurface temperatureTSt2mSea\-level pressurePSLmslPrecipitationPRECTtpTop longwave fluxFLUTavg\_tnlwrfTotal precipitable waterTMQtcwvSome variables are physically related but not identical across the two datasets, so simple conversions are required\. For example, ERA5 reports average top net longwave radiation flux \(avg\_tnlwrf\), following the ECMWF sign convention in which downward flux is positive\. This has the opposite sign convention from the CESM2 upwelling longwave flux at the top of the model \(FLUT\), so the ERA5 field is multiplied by−1\-1\. Similarly, CESM2 reports geopotential height \(Z\), whereas ERA5 reports geopotential \(z\)\. To make the variables equivalent, ERA5 geopotential is divided by gravitational acceleration,g≈9\.80665​m​s−2g\\approx 9\.80665~\\mathrm\{m~s^\{\-2\}\}, to obtain geopotential height\. Precipitation also requires unit adjustment\. ERA5 reports accumulated precipitation over an hourly interval \(tp\), whereas CESM2 reports total precipitation as a per\-second rate \(PRECT\)\. Therefore, ERA5 precipitation is divided by3,6003,600\. These conversions are made to ensure that SimCast\-S2S is pretrained and fine\-tuned on variables with comparable units\.

All input and target variables were transformed into standardized anomalies before model training\. From a meteorological standpoint, this step separates subseasonal variability from the much stronger background structure of the climate system, including the seasonal cycle, spatially varying climatology, and slow long\-term trends\. Raw atmospheric fields are dominated by predictable differences between seasons and regions; however, subseasonal forecasting is primarily concerned with departures from those expected states\. Removing the local trend and day\-of\-year climatology therefore encourages the model to learn physically meaningful anomaly relationships rather than simply reproducing the mean annual cycle or the background climate state\. The same transformation was applied to both input variables and the precipitation target to preserve a consistent meteorological reference frame\. In this form, the model learns how anomalous large\-scale atmospheric conditions, such as circulation, temperature, and moisture anomalies, are associated with anomalous precipitation\. Specifically, for each variable, a grid\-point\-wise linear trend was first estimated from the daily fields and removed to reduce slow nonstationary changes unrelated to subseasonal variability\. A smoothed day\-of\-year climatology was then obtained\. After that, each field was standardized by subtracting the corresponding climatological mean and dividing by the climatological standard deviation\. For validation and testing, the detrending coefficients and climatological statistics estimated from the training period were reused to avoid information leakage\.

### 4\.2Physical and Latent Space

SimCast\-S2S represents the atmospheric state in two coupled spaces: the gridded physical space and a learned probabilistic latent space\. In physical space, each sample is a collection of standardized anomaly fields defined on a latitude–longitude grid, with separate variable groups describing circulation, mass, thermodynamic, moisture, and precipitation processes\. Although this representation preserves direct meteorological interpretability, it is high\-dimensional and highly redundant because large\-scale atmospheric fields exhibit strong spatial covariance and coherent dynamical structures\. Directly modeling the conditional distribution of future precipitation in this space would therefore require learning stochastic relationships across a large number of mutually dependent grid points\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig9.png)Figure 9:Schematic illustration of the VAE mapping between the high\-dimensional physical space and the low\-dimensional latent space for variable groupgg\. The encoderℰg\\mathcal\{E\}\_\{g\}maps a anomaly field𝐱g\\mathbf\{x\}\_\{g\}to a Gaussian latent distribution𝐳g∼𝒩⁡\(𝝁g,𝝈g2\)\\mathbf\{z\}\_\{g\}\\sim\\mathcal\{N\}\(\\boldsymbol\{\\mu\}\_\{g\},\\boldsymbol\{\\sigma\}\_\{g\}^\{2\}\)\. Sampling from this distribution introduces stochasticity in latent space, and the decoder𝒟g\\mathcal\{D\}\_\{g\}maps the sampled latent representation back to a reconstructed physical field𝐱^g\\hat\{\\mathbf\{x\}\}\_\{g\}\. The latent space has a much lower dimension than the physical space, and stochastic sampling introduces a small perturbation between𝐱^g\\hat\{\\mathbf\{x\}\}\_\{g\}and𝐱g\\mathbf\{x\}\_\{g\}\.To reduce dimensionality, SimCast\-S2S uses variational autoencoders \(VAEs\)[Kingma and Welling \(2014b\)](https://arxiv.org/html/2608.26594#bib.bib45)to map each physical variable group from its gridded physical space to a compact probabilistic latent space\. For variable groupgg, let𝐱g∈𝒳g⊂ℝCg×H×W\\mathbf\{x\}\_\{g\}\\in\\mathcal\{X\}\_\{g\}\\subset\\mathbb\{R\}^\{C\_\{g\}\\times H\\times W\}denote a physical\-space field group, whereCgC\_\{g\}is the number of variables in the group andH×WH\\times Wis the latitude–longitude grid\. The corresponding latent representation is defined in a lower\-dimensional space𝐳g∈𝒵g⊂ℝdg\\mathbf\{z\}\_\{g\}\\in\\mathcal\{Z\}\_\{g\}\\subset\\mathbb\{R\}^\{d\_\{g\}\}, wheredg≪Cg​H​Wd\_\{g\}\\ll C\_\{g\}HW\. Under a parameterizationϕ\\phi, the encoder therefore learns a probabilistic mapping between the high\-dimensional physical space𝒳g\\mathcal\{X\}\_\{g\}and the compact latent space𝒵g\\mathcal\{Z\}\_\{g\}, i\.e\.ℰϕg:𝒳g→𝒵g\\mathcal\{E\}\_\{\\phi\_\{g\}\}:\\mathcal\{X\}\_\{g\}\\to\\mathcal\{Z\}\_\{g\}\.

Although the encoderℰϕg\\mathcal\{E\}\_\{\\phi\_\{g\}\}models a probabilistic latent representation, it is still a deterministic neural network\. Given an input field group𝐱g\\mathbf\{x\}\_\{g\}, the encoder first predicts the mean and standard deviation of the approximate posterior:

\(𝝁g,𝝈g2\)=ℰϕg​\[𝐱g\]\.\\left\(\\boldsymbol\{\\mu\}\_\{g\},\\boldsymbol\{\\sigma\}\_\{g\}^\{2\}\\right\)=\\mathcal\{E\}\_\{\\phi\_\{g\}\}\\left\[\\mathbf\{x\}\_\{g\}\\right\]\.\(1\)
These parameters define a Gaussian approximate posteriorqϕ​\(𝐳g∣𝐱g\)=𝒩⁡\(𝝁g,diag⁡\(𝝈g2\)\)\.q\_\{\\phi\}\\left\(\\mathbf\{z\}\_\{g\}\\mid\\mathbf\{x\}\_\{g\}\\right\)=\\mathcal\{N\}\\left\(\\boldsymbol\{\\mu\}\_\{g\},\\,\\operatorname\{diag\}\\left\(\\boldsymbol\{\\sigma\}\_\{g\}^\{2\}\\right\)\\right\)\.The stochasticity is introduced by sampling𝐳g\\mathbf\{z\}\_\{g\}from this approximate posterior\. To allow gradients to pass through, the sampling is done using a continuous reparameterization trick:

𝐳g=𝝁g\+𝝈g⊙ϵ,\\mathbf\{z\}\_\{g\}=\\boldsymbol\{\\mu\}\_\{g\}\+\\boldsymbol\{\\sigma\}\_\{g\}\\odot\\boldsymbol\{\\epsilon\},\(2\)
where⊙\\odotis the element\-wise multiplication, andϵ∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.

The decoder𝒟θg:𝒵g→𝒳g\\mathcal\{D\}\_\{\\theta\_\{g\}\}:\\mathcal\{Z\}\_\{g\}\\to\\mathcal\{X\}\_\{g\}, parameterized byθ\\theta, maps the sampled latent vector back to physical space:

𝐱^g=𝒟θg​\(𝐳g\)\.\\hat\{\\mathbf\{x\}\}\_\{g\}=\\mathcal\{D\}\_\{\\theta\_\{g\}\}\\left\(\\mathbf\{z\}\_\{g\}\\right\)\.\(3\)
where𝐱^g∈𝒳g\\hat\{\\mathbf\{x\}\}\_\{g\}\\in\\mathcal\{X\}\_\{g\}is the reconstructed standardized anomaly field\.

In probabilistic form, the decoder parameterizes the conditional reconstruction distributionpθ​\(𝐱g∣𝐳g\)p\_\{\\theta\}\\left\(\\mathbf\{x\}\_\{g\}\\mid\\mathbf\{z\}\_\{g\}\\right\), which describes the distribution of physical\-space fields that can be reconstructed from a given latent state𝐳g\\mathbf\{z\}\_\{g\}\.

There is always some discrepancy between𝐱^g\\hat\{\\mathbf\{x\}\}\_\{g\}and𝐱g\\mathbf\{x\}\_\{g\}for two reasons\. First, the latent vector𝐳g∼qϕ​\(𝐳g∣𝐱g\)\\mathbf\{z\}\_\{g\}\\sim q\_\{\\phi\}\\left\(\\mathbf\{z\}\_\{g\}\\mid\\mathbf\{x\}\_\{g\}\\right\)is a random variable, different samples from the same approximate posterior can lead to slightly different decoded fields\. Second, some information is inevitably lost when a high\-dimensional physical field is compressed into a much lower\-dimensional latent space\. As a result, the same anomaly field𝐱g\\mathbf\{x\}\_\{g\}, when passed through the encoding and decoding processes multiple times, can produce multiple possible reconstructions𝐱^g\\hat\{\\mathbf\{x\}\}\_\{g\}\. This introduces a natural source of perturbation in the reconstructed physical fields, which supports the stochastic nature of SimCast\-S2S\. Fig\.[9](https://arxiv.org/html/2608.26594#S4.F9)illustrates this encoding–decoding process\.

The learning objective of a VAE is to maximize the marginal likelihood of the observed datapϕ,θ​\(𝐱\)p\_\{\\phi,\\theta\}\(\\mathbf\{x\}\)\. Although this marginal likelihood is intractable, it can be showed that this objective can be obtained equivalently by minimizing its negative evidence lower bound \(ELBO\), which leads to the following loss function:

ℒVAE\(𝜽,ϕ,𝐱\):=−𝔼qϕ​\(𝐳∣𝐱\)\[logp𝜽\(𝐱∣𝐳\)\]\+DKL\[qϕ\(𝐳∣𝐱\)∥p\(𝐳\)\]\.\\mathcal\{L\}\_\{\\text\{VAE\}\}\(\\boldsymbol\{\\theta\},\\boldsymbol\{\\phi\},\\mathbf\{x\}\):=\-\\mathbb\{E\}\_\{q\_\{\\boldsymbol\{\\phi\}\}\(\\mathbf\{z\}\\mid\\mathbf\{x\}\)\}\[\\log p\_\{\\boldsymbol\{\\theta\}\}\(\\mathbf\{x\}\\mid\\mathbf\{z\}\)\]\+D\_\{\\mathrm\{KL\}\}\[q\_\{\\boldsymbol\{\\phi\}\}\(\\mathbf\{z\}\\mid\\mathbf\{x\}\)\\,\\\|\\,p\(\\mathbf\{z\}\)\]\.\(4\)
The first term in Eq\.[4](https://arxiv.org/html/2608.26594#S4.E4)is the negative log\-likelihood of the observed dataset, which encourages the model to reconstruct the data accurately\. The second term is the KL divergence between two distributions, which acts as a regularizer that encourages the approximate posteriorqϕ​\(𝐳∣𝐱\)q\_\{\\boldsymbol\{\\phi\}\}\(\\mathbf\{z\}\\mid\\mathbf\{x\}\)to remain close to the priorp⁡\(𝐳\)p\(\\mathbf\{z\}\)\.

In practice, it is common to further assume that the reconstruction error follows a Gaussian likelihood, i\.e\.,𝐱−𝐱^∼𝒩⁡\(𝟎,σ2​𝐈\)\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\sigma^\{2\}\\mathbf\{I\}\), where andσ∈ℝ\\sigma\\in\\mathbb\{R\}is a fixed constant\. This is equivalent to modeling the likelihood aspθ​\(𝐱∣𝐳\)=𝒩⁡\(𝐱,𝐱^,σ2​𝐈\)p\_\{\\theta\}\(\\mathbf\{x\}\\mid\\mathbf\{z\}\)=\\mathcal\{N\}\(\\mathbf\{x\};\\,\\hat\{\\mathbf\{x\}\},\\sigma^\{2\}\\mathbf\{I\}\)\. Under this assumption, the negative log\-likelihood simplifies to a scaled MSE loss, and minimizing the negative log\-likelihood is equivalent to minimizing the MSE between the reconstruction and the input, up to a constant scaling factor\.

−log⁡pθ​\(𝐱∣𝐳\)=12​σ2​‖𝐱−𝐱^‖2\+const\.\-\\log p\_\{\\theta\}\(\\mathbf\{x\}\\mid\\mathbf\{z\}\)=\\frac\{1\}\{2\\sigma^\{2\}\}\\\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\\\|^\{2\}\+\\text\{const\}\.\(5\)
Also, with the priorp⁡\(𝐳\)p\(\\mathbf\{z\}\)assumed to be a standard Gaussian, the KL divergence in Eq\.[4](https://arxiv.org/html/2608.26594#S4.E4)becomes a special case and can be written in a compact form:

DKL\[qϕ\(𝐳∣𝐱\)∥p𝜽\(𝐳\)\]=12∑j=1dg\(σj2\+μj2−1−logσj2\),D\_\{\\mathrm\{KL\}\}\\left\[q\_\{\\boldsymbol\{\\phi\}\}\(\\mathbf\{z\}\\mid\\mathbf\{x\}\)\\,\\\|\\,p\_\{\\boldsymbol\{\\theta\}\}\(\\mathbf\{z\}\)\\right\]=\\frac\{1\}\{2\}\\sum\_\{j=1\}^\{d\_\{g\}\}\\left\(\\sigma\_\{j\}^\{2\}\+\\mu\_\{j\}^\{2\}\-1\-\\log\\sigma\_\{j\}^\{2\}\\right\),\(6\)
whereμi\\mu\_\{i\}andσi\\sigma\_\{i\}are the components of𝝁g\\boldsymbol\{\\mu\}\_\{g\}and𝝈g\\boldsymbol\{\\sigma\}\_\{g\}, respectively\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig10.png)Figure 10:VAE architecture used to map physical\-space anomaly fields into a compact latent representation and reconstruct them back to physical space\. The encoderℰ\\mathcal\{E\}progressively compresses the input field𝐱\\mathbf\{x\}through ConvStack and DownSampling blocks to obtain the latent vector𝐳\\mathbf\{z\}\. The decoder𝒟\\mathcal\{D\}mirrors this process using UpSampling and ConvStack blocks to reconstruct the field𝐱^\\hat\{\\mathbf\{x\}\}\.Combining Eq\.[4](https://arxiv.org/html/2608.26594#S4.E4),[6](https://arxiv.org/html/2608.26594#S4.E6),[5](https://arxiv.org/html/2608.26594#S4.E5), the training procedure of VAE is summarized in Algo\.[1](https://arxiv.org/html/2608.26594#alg1)\. The encoder parametersϕ\\boldsymbol\{\\phi\}and decoder parameters𝜽\\boldsymbol\{\\theta\}are optimized jointly using the Adam optimizer\([Kingma and Ba, 2015](https://arxiv.org/html/2608.26594#bib.bib63)\)\. Fig\.[10](https://arxiv.org/html/2608.26594#S4.F10)describe the structure ofℰϕg\\mathcal\{E\}\_\{\\phi\_\{g\}\}and𝒟θg\\mathcal\{D\}\_\{\\theta\_\{g\}\}, which together form the complete VAE architecture\. Both the encoder and decoder use a standard convolutional architecture\([LeCun et al\., 1998b](https://arxiv.org/html/2608.26594#bib.bib64);[Krizhevsky et al\., 2012](https://arxiv.org/html/2608.26594#bib.bib65)\)\. The encoder is built from stacked convolutional layers, each followed by a downsampling block to progressively reduce the spatial resolution by a factor for four\. The decoder mirrors this structure but uses upsampling blocks\. Each upsampling block contains a transposed convolutional layer to expand the spatial resolution by a factor for four\.

The number of ConvStack–DownScaling blocks directly controls the latent dimension\. LetCgC\_\{g\}denote the number of physical variables in groupgg, letEgE\_\{g\}denote the number of latent channels per variable, and letNgN\_\{g\}denote the number of ConvStack–DownScaling blocks, the latent dimension isdg=Cg​Eg​H​W4Ngd\_\{g\}=\\frac\{C\_\{g\}E\_\{g\}HW\}\{4^\{N\_\{g\}\}\}, whereH​WHWis the original physical\-space grid\. The corresponding compression ratio is4NgEg\\frac\{4^\{N\_\{g\}\}\}\{E\_\{g\}\}\. For example, for a group ofCg=7C\_\{g\}=7variables, settingNg=3N\_\{g\}=3andEg=28E\_\{g\}=28gives a16×16\\timescompression\. For a group ofCg=1C\_\{g\}=1variable, settingNg=3N\_\{g\}=3andEg=16E\_\{g\}=16gives a4×4\\timescompression\.

Algorithm 1VAE Training Procedure for variable groupggDataset

𝔻\\mathbb\{D\}, encoder

ℰϕg\\mathcal\{E\}\_\{\\phi\_\{g\}\}, decoder

𝒟θg\\mathcal\{D\}\_\{\\theta\_\{g\}\}, loss balancing term

λ∈ℝ\+\\lambda\\in\\mathbb\{R\}^\{\+\}
repeat

Sample

𝐱∈𝔻\\mathbf\{x\}\\in\\mathbb\{D\}
Encoder: compute

\(𝝁g,𝝈g2\)=ℰϕg​\[𝐱g\]\\left\(\\boldsymbol\{\\mu\}\_\{g\},\\boldsymbol\{\\sigma\}\_\{g\}^\{2\}\\right\)=\\mathcal\{E\}\_\{\\phi\_\{g\}\}\\left\[\\mathbf\{x\}\_\{g\}\\right\]
Sample

ϵ∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
Compute

𝐳g=𝝁g\+𝝈g⊙ϵ\\mathbf\{z\}\_\{g\}=\\boldsymbol\{\\mu\}\_\{g\}\+\\boldsymbol\{\\sigma\}\_\{g\}\\odot\\boldsymbol\{\\epsilon\}
Decoder: compute

𝐱^g=𝒟θg​\(𝐳g\)\\hat\{\\mathbf\{x\}\}\_\{g\}=\\mathcal\{D\}\_\{\\theta\_\{g\}\}\\left\(\\mathbf\{z\}\_\{g\}\\right\)
Take a gradient step on:

∇θ,ϕ\[λ​‖𝐱−𝐱^‖2\+12​∑j=1dg\(σj2\+μj2−1−log⁡σj2\)\]\\nabla\_\{\\theta,\\phi\}\\left\[\\lambda\\\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\\\|^\{2\}\+\\frac\{1\}\{2\}\\sum\_\{j=1\}^\{d\_\{g\}\}\\left\(\\sigma\_\{j\}^\{2\}\+\\mu\_\{j\}^\{2\}\-1\-\\log\\sigma\_\{j\}^\{2\}\\right\)\\right\]

untilconverged

Table 3:Variable groups used in SimCast\-S2S and their corresponding CESM2 and ERA5 variables\. The 21 input variables are divided into five physically motivated groups: wind, mass, thermal, hydro, and precipitation\. Each group has physical\-space dimensionCg​H​WC\_\{g\}HW, latent\-space dimensiondgd\_\{g\}, and compression ratioCg​H​W/dgC\_\{g\}HW/d\_\{g\}\. Here,H×W=192×288H\\times W=192\\times 288\.GroupCESM2 variablesERA5 variables𝑪𝒈​𝑯​𝑾\\boldsymbol\{C\_\{g\}HW\}𝒅𝒈\\boldsymbol\{d\_\{g\}\}CompressionwindU200, V200, U500, V500, OMEGA500, U850, V850u200, v200, u500, v500, w500, u850, v850387,072387\{,\}07224,19224\{,\}1921616timesmassZ200, Z500, Z850, PSLz200, z500, z850, msl221,184221\{,\}18413,82413\{,\}8241616timesthermalTS, T200, T500, T850, FLUTskt, t200, t500, t850, avgtnlwrf276,480276\{,\}48017,28017\{,\}2801616timeshydroTMQ, Q200, Q500, Q850tcwv, q200, q500, q850221,184221\{,\}18413,82413\{,\}8241616timesprecipPRECTtp55,29655\{,\}29613,82413\{,\}82444timesIn this study, 21 variables listed in Table[2](https://arxiv.org/html/2608.26594#S4.T2)are divided into five physically motivated groups, namelywind,mass,thermal,hydro, andprecip\. Table[3](https://arxiv.org/html/2608.26594#S4.T3)details the variables included in each group for both CESM2 and ERA5\. Each group has a physical\-space dimension ofCg​H​WC\_\{g\}HWwhereH​W=192×288=55,296HW=192\\times 288=55\{,\}296\. Choosing the compression level involves a trade\-off between efficiency and reconstruction accuracy: a smaller latent dimension gives a more compact representation, while a larger latent dimension reduces reconstruction error\. Empirically, we find reconstruction error decreases approximately exponentially as the latent dimension increases\. Based on our analysis, we choose a16×16\\timescompression for the wind, mass, thermal, and hydro groups, and a4×4\\timescompression for precipitation\. The lower compression ratio for precipitation is necessary because its distribution is highly intermittent, strongly skewed, and dominated by localized extremes\. More aggressive compression would therefore oversmooth the latent distribution and degrade reconstruction fidelity\.

### 4\.3SimCast\-S2S model

SimCast\-S2S is a latent diffusion model\([Sohl\-Dickstein et al\., 2015](https://arxiv.org/html/2608.26594#bib.bib57);[Song and Ermon, 2019](https://arxiv.org/html/2608.26594#bib.bib58);[Ho et al\., 2020b](https://arxiv.org/html/2608.26594#bib.bib59);[Song et al\., 2021c](https://arxiv.org/html/2608.26594#bib.bib60);[Rombach et al\., 2022b](https://arxiv.org/html/2608.26594#bib.bib61);[Yang et al\., 2023](https://arxiv.org/html/2608.26594#bib.bib62)\)\. The central idea is to model the stochastic evolution in the low\-dimensional latent space\. During the forward diffusion process, Gaussian noise is gradually added to the target latent vector until it is fully corrupted into a standard Gaussian\. A neural network is then trained to reverse this process: starting from a noisy target latent state, it progressively removes noise by conditioning on the latent representations of the past atmospheric states\. The objective is to recover the original target latent representation\.

Consider a variance schedule0<β1,β2,⋯,βK<10<\\beta\_\{1\},\\beta\_\{2\},\\cdots,\\beta\_\{K\}<1\. For a precipitation latent𝐳0∼p⁡\(𝐳\)\\mathbf\{z\}\_\{0\}\\sim p\(\\mathbf\{z\}\), a discrete Markov chain\{𝐳0,𝐳1,⋯,𝐳K\}\\\{\\mathbf\{z\}\_\{0\},\\mathbf\{z\}\_\{1\},\\cdots,\\mathbf\{z\}\_\{K\}\\\}is defined by:

𝐳k:=1−βt​𝐳k−1\+βk​𝜺k,𝜺k∼𝒩⁡\(𝟎,𝐈\),k:1→K,\\mathbf\{z\}\_\{k\}:=\\sqrt\{1\-\\beta\_\{t\}\}\\,\\mathbf\{z\}\_\{k\-1\}\+\\sqrt\{\\beta\_\{k\}\}\\,\\mathbf\{\\boldsymbol\{\\varepsilon\}\}\_\{k\},\\quad\\mathbf\{\\boldsymbol\{\\varepsilon\}\}\_\{k\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),\\quad k:1\\to K,\(7\)
or equivalently:

q⁡\(𝐳k∣𝐳k−1\)=𝒩⁡\(𝐳k,1−βk​𝐳k−1,βk​𝐈\)\.q\(\\mathbf\{z\}\_\{k\}\\mid\\mathbf\{z\}\_\{k\-1\}\)=\\mathcal\{N\}\\left\(\\mathbf\{z\}\_\{k\};\\sqrt\{1\-\\beta\_\{k\}\}\\,\\mathbf\{z\}\_\{k\-1\},\\beta\_\{k\}\\mathbf\{I\}\\right\)\.\(8\)
Eq\.[7](https://arxiv.org/html/2608.26594#S4.E7)defines a first\-order Markov chain where the signal is gradually degraded by1−βk\\sqrt\{1\-\\beta\_\{k\}\}, and Gaussian noise with varianceβk\\beta\_\{k\}is injected at each step\. Ask→Kk\\to K,𝐳K∼𝒩⁡\(𝟎,𝐈\)\\mathbf\{z\}\_\{K\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.

Because each transition is Gaussian and linear, we can computeq⁡\(𝐳k∣𝐳0\)q\(\\mathbf\{z\}\_\{k\}\\mid\\mathbf\{z\}\_\{0\}\)in closed form\. Definingα¯k:=∏j=1k\(1−βj\)\\bar\{\\alpha\}\_\{k\}:=\\prod\_\{j=1\}^\{k\}\(1\-\\beta\_\{j\}\)whereα¯0=1\\bar\{\\alpha\}\_\{0\}=1, which quantifies how much signal remains from𝐳0\\mathbf\{z\}\_\{0\}at stepkk, Eq\.[7](https://arxiv.org/html/2608.26594#S4.E7)is now simplified as:

𝐳k=α¯k​𝐳0\+1−α¯k​ϵ,\\mathbf\{z\}\_\{k\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\mathbf\{z\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\},\(9\)
in whichϵ∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)is a fresh standard Gaussian variable\. This implies thatq⁡\(𝐳k∣𝐳0\)q\(\\mathbf\{z\}\_\{k\}\\mid\\mathbf\{z\}\_\{0\}\)is Gaussian with a closed form:

q⁡\(𝐳k∣𝐳0\)=𝒩⁡\(𝐳k,α¯k​𝐳0,\(1−α¯k\)​𝐈\),q\(\\mathbf\{z\}\_\{k\}\\mid\\mathbf\{z\}\_\{0\}\)=\\mathcal\{N\}\(\\mathbf\{z\}\_\{k\};\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\},\(1\-\\bar\{\\alpha\}\_\{k\}\)\\mathbf\{I\}\),\(10\)
Eq\.[10](https://arxiv.org/html/2608.26594#S4.E10)is important because it enables direct sampling of𝐱k\\mathbf\{x\}\_\{k\}from the clean data point𝐱0\\mathbf\{x\}\_\{0\}without iteratively going through all the diffusion steps\.

We can train a neural networkfθ​\(𝐙cond,t,𝐳k,k\)f\_\{\\theta\}\(\\mathbf\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\)to approximateϵ\\boldsymbol\{\\epsilon\}, where𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}is the concatenated latent representation of the initial physical state andttis the day of year as detailed in section[4\.2](https://arxiv.org/html/2608.26594#S4.SS2)\. The training objective reduces to a simpleℓ2\\ell\_\{2\}loss function:

ℒSimCast\-S2S​\(θ\)=𝔼𝐳0,t​\[‖ϵ−fθ​\(𝒁cond,t,𝐳k,k\)‖22\],\\mathcal\{L\}\_\{\\text\{SimCast\-S2S\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\mathbf\{z\}\_\{0\},t\}\\left\[\\left\\\|\\boldsymbol\{\\epsilon\}\-f\_\{\\theta\}\(\\boldsymbol\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\)\\right\\\|\_\{2\}^\{2\}\\right\],\(11\)
Because the forward processq⁡\(𝐳k∣𝐳k−1\)q\(\\mathbf\{z\}\_\{k\}\\mid\\mathbf\{z\}\_\{k\-1\}\)is Markov and Gaussian, it can be shown that the reverse processp⁡\(𝐳k−1∣𝐳k\)p\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\}\)is also Markov and Gaussian:

pθ​\(𝐳k−1∣𝐳k\)=𝒩⁡\(𝐳k−1,𝝁θ​\(𝒁cond,t,𝐳k,k\),Σk​𝐈\)\.p\_\{\\theta\}\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\}\)=\\mathcal\{N\}\\left\(\\mathbf\{z\}\_\{k\-1\};\\boldsymbol\{\\mu\}\_\{\\theta\}\(\\boldsymbol\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\),\\,\\Sigma\_\{k\}\\mathbf\{I\}\\right\)\.\(12\)
The mean𝝁θ​\(𝒁cond,t,𝐳k,k\)\\boldsymbol\{\\mu\}\_\{\\theta\}\(\\boldsymbol\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\)is not a separate neural network by itself, but is analytically computed from the noise predictionϵ^k−1=fθ​\(𝒁cond,t,𝐳k,k\)\\boldsymbol\{\\hat\{\\epsilon\}\}\_\{k\-1\}=f\_\{\\theta\}\(\\boldsymbol\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\)\. In practice, the varianceΣk=Var⁡\[𝐱k−1∣𝐱k\]\\Sigma\_\{k\}=\\mathrm\{Var\}\[\\mathbf\{x\}\_\{k\-1\}\\mid\\mathbf\{x\}\_\{k\}\]is either fixed or learned\. In this study, we make a common choice of settingΣk=β~k=1−α¯k−11−α¯k​βk\\Sigma\_\{k\}=\\tilde\{\\beta\}\_\{k\}=\\frac\{1\-\\bar\{\\alpha\}\_\{k\-1\}\}\{1\-\\bar\{\\alpha\}\_\{k\}\}\\beta\_\{k\}\. This choice ensures the reverse variance matches the true posterior varianceβ~k=Var\[𝐳k−1∣𝐳k,𝐳0\]\\tilde\{\\beta\}\_\{k\}=\\mathrm\{Var\}\[\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\]\.

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig11.png)Figure 11:Schematic illustration of the SimCast\-S2S latent diffusion framework\. Historical atmospheric fields are first separated into five physical groups: wind, thermal, mass, hydro, and precipitation, as described in Table[3](https://arxiv.org/html/2608.26594#S4.T3)\. Each group is encoded by its pretrained VAE encoder into a low\-dimensional latent space\. The resulting conditioning latents are concatenated to form𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}\. Starting from a random Gaussian latent𝐳^K∼𝒩⁡\(𝟎,𝐈\)\\hat\{\\mathbf\{z\}\}\_\{K\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\), the neural network iteratively denoises the latent state, conditioned on𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}, the denoising stepkk, and seasonal informationtt, until it produces the target precipitation latent𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}\. The pretrained precipitation decoder𝒟precip\\mathcal\{D\}\_\{\\mathrm\{precip\}\}then maps𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}back to physical space to obtain the subseasonal precipitation\-anomaly forecast field𝐲^\\hat\{\\mathbf\{y\}\}\.Recall that the generative processpθ​\(𝐳k−1∣𝐳k\)p\_\{\\theta\}\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\}\)is meant to approximate the true reverse process of the forward Markov chain\. The ideal reverse transition is the posterior distributionq⁡\(𝐳k−1∣𝐳k\)q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\}\), but this is intractable\. However, we can exactly compute the posteriorq⁡\(𝐳k−1∣𝐳k,𝐳0\)q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\), which is the distribution over𝐳k−1\\mathbf\{z\}\_\{k\-1\}conditioned on both the noisy input𝐳k\\mathbf\{z\}\_\{k\}and the original clean sample𝐳0\\mathbf\{z\}\_\{0\}\. Here, we useq⁡\(𝐳k−1∣𝐳k,𝐳0\)q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\)as a tractable surrogate for the intractable marginal posteriorq⁡\(𝐳k−1∣𝐳k\)q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\}\)\. Because affine transformations and marginalizations of Gaussians preserve Gaussianity, the triplet\(𝐳0,𝐳k−1,𝐳k\)\(\\mathbf\{z\}\_\{0\},\\mathbf\{z\}\_\{k\-1\},\\mathbf\{z\}\_\{k\}\)jointly forms a multivariate Gaussian distribution, and conditioning on any subset—including both𝐳0\\mathbf\{z\}\_\{0\}and𝐳k\\mathbf\{z\}\_\{k\}—also yields another Gaussian\. This allows us to analytically compute the posteriorq⁡\(𝐳k−1∣𝐳k,𝐳0\)q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\)in closed form using standard results from Gaussian conditioning:

q⁡\(𝐳k−1∣𝐳k,𝐳0\)=𝒩⁡\(𝐳k−1,𝝁~​\(𝐳k,𝐳0\),Σk​𝐈\),q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\)=\\mathcal\{N\}\\left\(\\mathbf\{z\}\_\{k\-1\};\\tilde\{\\boldsymbol\{\\mu\}\}\(\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\),\\Sigma\_\{k\}\\mathbf\{I\}\\right\),\(13\)
where

𝝁~\(𝐳k,𝐳0\)=𝔼\[𝐳k−1∣𝐳k,𝐳0\]=α¯k−1​βk1−α¯k𝐳0\+αk​\(1−α¯k−1\)1−α¯k𝐳k,\\tilde\{\\boldsymbol\{\\mu\}\}\(\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\)=\\mathbb\{E\}\[\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\]=\\frac\{\\sqrt\{\\bar\{\\alpha\}\_\{k\-1\}\}\\beta\_\{k\}\}\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\}\+\\frac\{\\sqrt\{\\alpha\_\{k\}\}\(1\-\\bar\{\\alpha\}\_\{k\-1\}\)\}\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{k\},\(14\)
During the generative process, we do not have access to the true clean sample𝐳0\\mathbf\{z\}\_\{0\}\. To approximate the reverse transition distributionq⁡\(𝐳k−1∣𝐳k,𝐳0\)q\(\\mathbf\{z\}\_\{k\-1\}\\mid\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\), we leverage the reparameterized form of the forward process, which expresses𝐳k\\mathbf\{z\}\_\{k\}as a linear combination of𝐳0\\mathbf\{z\}\_\{0\}and Gaussian noise \(see Eq\.[9](https://arxiv.org/html/2608.26594#S4.E9)\)\. Using the predicted noiseϵ^=fθ​\(𝐙cond,t,𝐳k,k\)\\boldsymbol\{\\hat\{\\epsilon\}\}=f\_\{\\theta\}\(\\mathbf\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\), we can estimate𝐳0\\mathbf\{z\}\_\{0\}as:

𝐳^0:=1α¯k​\[𝐳k−1−α¯k​ϵ^\]\.\\hat\{\\mathbf\{z\}\}\_\{0\}:=\\frac\{1\}\{\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\}\\left\[\\mathbf\{z\}\_\{k\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\hat\{\\epsilon\}\}\\right\]\.\(15\)
This estimator effectively inverts the forward process at timestepkkunder the assumption thatϵ^\\boldsymbol\{\\hat\{\\epsilon\}\}accurately predicts the noise component added to𝐳0\\mathbf\{z\}\_\{0\}to obtain𝐳k\\mathbf\{z\}\_\{k\}\. We then substitute this estimate𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}into the analytic expression for the true posterior mean𝝁~​\(𝐳k,𝐳0\)\\tilde\{\\boldsymbol\{\\mu\}\}\(\\mathbf\{z\}\_\{k\},\\mathbf\{z\}\_\{0\}\)\. With algebraic simplification, this expression becomes a compact and efficient form that is commonly used in practice:

𝝁~​\(𝐳k,𝐳^0\)=11−βk​\(𝐳k−βk1−α¯k​ϵ^\)\.\\tilde\{\\boldsymbol\{\\mu\}\}\(\\mathbf\{z\}\_\{k\},\\hat\{\\mathbf\{z\}\}\_\{0\}\)=\\frac\{1\}\{\\sqrt\{1\-\\beta\_\{k\}\}\}\\left\(\\mathbf\{z\}\_\{k\}\-\\frac\{\\beta\_\{k\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\}\\,\\boldsymbol\{\\hat\{\\epsilon\}\}\\right\)\.\(16\)
With the mean in Eq\.[13](https://arxiv.org/html/2608.26594#S4.E13)now fully defined, the full stochastic reverse process at timestepkkis obtained by sampling from a Gaussian centered at this predicted mean and with variance given byΣk=β~k=1−α¯k−11−α¯k​βk\\Sigma\_\{k\}=\\tilde\{\\beta\}\_\{k\}=\\frac\{1\-\\bar\{\\alpha\}\_\{k\-1\}\}\{1\-\\bar\{\\alpha\}\_\{k\}\}\\beta\_\{k\}\. This yields the update rule used during sampling:

𝐳k−1=11−βk​\(𝐳k−βk1−α¯k​ϵ^\)\+Σk​ϵ,where​ϵ∼𝒩⁡\(𝟎,𝐈\),\\mathbf\{z\}\_\{k\-1\}=\\frac\{1\}\{\\sqrt\{1\-\\beta\_\{k\}\}\}\\left\(\\mathbf\{z\}\_\{k\}\-\\frac\{\\beta\_\{k\}\}\{\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\}\\,\\boldsymbol\{\\hat\{\\epsilon\}\}\\right\)\+\\sqrt\{\\Sigma\_\{k\}\}\\,\\boldsymbol\{\\epsilon\},\\quad\\text\{where \}\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),\(17\)
which defines one step of the generative chain, gradually transforming pure noise into a clean latent\.

Eq\.[17](https://arxiv.org/html/2608.26594#S4.E17)is known as the Denoising Diffusion Probabilistic Models \(DDPM\)[Ho et al\. \(2020b\)](https://arxiv.org/html/2608.26594#bib.bib59)\. In this study, we also employ a method to further generalize this update rule by decomposing the stochasticity in Eq\.[9](https://arxiv.org/html/2608.26594#S4.E9)into a component along the data\-aligned directionϵθ\\boldsymbol\{\\epsilon\}\_\{\\theta\}, and a fresh isotropic noise𝜻∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\zeta\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\):

𝐳k−1=α¯k−1​𝐳^0\+ck​ϵ^\+σk​𝜻,\\mathbf\{z\}\_\{k\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\-1\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{0\}\+c\_\{k\}\\,\\boldsymbol\{\\hat\{\\epsilon\}\}\+\\sigma\_\{k\}\\,\\boldsymbol\{\\zeta\},\(18\)
with the scalar coefficientsck≥0c\_\{k\}\\geq 0andσk≥0\\sigma\_\{k\}\\geq 0to be determined\. To ensure variance matching, the variance of𝐳k−1\\mathbf\{z\}\_\{k\-1\}conditioned on𝐳0\\mathbf\{z\}\_\{0\}under Eq\.[18](https://arxiv.org/html/2608.26594#S4.E18)must satisfy the constraint:

Var⁡\[𝐱k−1∣𝐱0\]=ck2​𝐈\+σk2​𝐈\.\\mathrm\{Var\}\[\\mathbf\{x\}\_\{k\-1\}\\mid\\mathbf\{x\}\_\{0\}\]\\;=\\;c\_\{k\}^\{2\}\\,\\mathbf\{I\}\\;\+\\;\\sigma\_\{k\}^\{2\}\\,\\mathbf\{I\}\.\(19\)
or equivalently,

ck2\+σk2=1−α¯k−1c\_\{k\}^\{2\}\\;\+\\;\\sigma\_\{k\}^\{2\}=1\-\\bar\{\\alpha\}\_\{k\-1\}\(20\)
Among all pairs\(ck,σk\)\(c\_\{k\},\\sigma\_\{k\}\)that satisfy the constraint in Eq\.[20](https://arxiv.org/html/2608.26594#S4.E20), we introduce a parameterη∈\[0,1\]\\eta\\in\[0,1\]to control the split betweenckc\_\{k\}andσk\\sigma\_\{k\}\. This is done by defining

σk2:=η2​β~k,\\sigma\_\{k\}^\{2\}:=\\eta^\{2\}\\,\\tilde\{\\beta\}\_\{k\},\(21\)
The remaining variance is allocated tockc\_\{k\}:

ct2=1−α¯t−1−σt2\.c\_\{t\}^\{2\}=1\-\\bar\{\\alpha\}\_\{t\-1\}\-\\sigma\_\{t\}^\{2\}\.\(22\)
There are two important special cases\. Whenη=1\\eta=1, we haveσk2=β~k\\sigma\_\{k\}^\{2\}=\\tilde\{\\beta\}\_\{k\}\. It can be shown that this recovers the DDPM sampling rule described in Eq\.[17](https://arxiv.org/html/2608.26594#S4.E17)\. Whenη=0\\eta=0, we haveσk=0\\sigma\_\{k\}=0andck=1−α¯k−1c\_\{k\}=\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\-1\}\}\. In this case, the reverse process becomes deterministic and corresponds to Denoising Diffusion Implicit Model \(DDIM\) sampling\([Song et al\., 2021a](https://arxiv.org/html/2608.26594#bib.bib66)\)\. Eq\.[18](https://arxiv.org/html/2608.26594#S4.E18)then reduces to

𝐳k−1=α¯k−1​𝐳^0\+1−α¯k−1​ϵ^\.\\mathbf\{z\}\_\{k\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\-1\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\-1\}\}\\,\\boldsymbol\{\\hat\{\\epsilon\}\}\.\(23\)
When0<η<10<\\eta<1, we obtain an interpolation between the deterministic path \(η=0\\eta=0\) and the fully stochastic path \(η=1\\eta=1\):

𝐳k−1=α¯k−1​𝐳^0\+1−α¯k−1−σk2​𝐳k−α¯k​𝐳^01−α¯k⏟ϵ^\+σk​𝜻,𝜻∼𝒩⁡\(𝟎,𝐈\),\\mathbf\{z\}\_\{k\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\-1\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{0\}\+\\sqrt\{\\,1\-\\bar\{\\alpha\}\_\{k\-1\}\-\\sigma\_\{k\}^\{2\}\\,\}\\;\\underbrace\{\\frac\{\\mathbf\{z\}\_\{k\}\-\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{0\}\}\{\\sqrt\{\\,1\-\\bar\{\\alpha\}\_\{k\}\\,\}\}\}\_\{\\boldsymbol\{\\hat\{\\epsilon\}\}\}\+\\sigma\_\{k\}\\,\\boldsymbol\{\\zeta\},\\quad\\boldsymbol\{\\zeta\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\),\(24\)
in which,σt2\\sigma^\{2\}\_\{t\}takes the general form:σt2=η2​β~t=η2⋅1−α¯t−11−α¯t​\(1−α¯tα¯t−1\)\\sigma\_\{t\}^\{2\}=\\eta^\{2\}\\,\\tilde\{\\beta\}\_\{t\}=\\eta^\{2\}\\cdot\\frac\{1\-\\bar\{\\alpha\}\_\{t\-1\}\}\{\\,1\-\\bar\{\\alpha\}\_\{t\}\\,\}\\left\(1\-\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{t\-1\}\}\\right\)\.

The interpolated form from Eq\.[24](https://arxiv.org/html/2608.26594#S4.E24)has its mean still depending on𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}andϵ^\\boldsymbol\{\\hat\{\\epsilon\}\}, but the variance injected is scaled down byη2\\eta^\{2\}\. Thus,η\\etainterpolates between two anchors: whenη=0\\eta=0, the process is purely deterministic, whenη=1\\eta=1, we recover the standard stochastic sampling; and for0<η<10<\\eta<1, the process offers a trade\-off between diversity and stability\.

So far, this formulation assumes that the neural network predicts the noise termϵ\\boldsymbol\{\\epsilon\}directly\. However, this is not the only possible parameterization of the denoising objective\. Instead of predictingϵ\\boldsymbol\{\\epsilon\}, SimCast\-S2S is trained to predict a velocity variable𝐯\\mathbf\{v\}, which represents a rotated coordinate in the two\-dimensional plane spanned by the clean latent𝐳0\\mathbf\{z\}\_\{0\}and the noiseϵ\\boldsymbol\{\\epsilon\}\. This𝐯\\mathbf\{v\}\-prediction formulation provides an alternative way to describe the same diffusion trajectory\. However, as will be shown later in this subsection,𝐯\\mathbf\{v\}\-prediction yields a more stable training scheme thanϵ\\boldsymbol\{\\epsilon\}\-prediction\.

Recall that the forward process is defined as𝐳k=α¯k​𝐳0\+1−α¯k​ϵ\\mathbf\{z\}\_\{k\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}\(Eq\.[9](https://arxiv.org/html/2608.26594#S4.E9)\), whereϵ∼𝒩⁡\(0,I\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(0,I\)is pure Gaussian and𝐳0\\mathbf\{z\}\_\{0\}denotes the clean data sample\. Because𝐳0\\mathbf\{z\}\_\{0\}andϵ\\boldsymbol\{\\epsilon\}are independent and orthogonal in expectation, i\.e\.𝔼⁡\[𝐳0⊤​ϵ\]=0\\mathbb\{E\}\[\\mathbf\{z\}\_\{0\}^\{\\top\}\\boldsymbol\{\\epsilon\}\]=0, their covariance vanishes:

Cov⁡\(𝐳0,ϵ\)=𝔼⁡\[\(𝐳0−𝔼​𝐳0\)⊤​\(ϵ−𝔼​ϵ\)\]=𝔼​\(𝐳0−𝔼​𝐳0\)⊤​𝔼​\(ϵ−𝔼​ϵ\)=𝟎,\\mathrm\{Cov\}\(\\mathbf\{z\}\_\{0\},\\boldsymbol\{\\epsilon\}\)=\\mathbb\{E\}\\\!\\big\[\(\\mathbf\{z\}\_\{0\}\-\\mathbb\{E\}\\mathbf\{z\}\_\{0\}\)^\{\\top\}\(\\boldsymbol\{\\epsilon\}\-\\mathbb\{E\}\\boldsymbol\{\\epsilon\}\)\\big\]=\\mathbb\{E\}\(\\mathbf\{z\}\_\{0\}\-\\mathbb\{E\}\\mathbf\{z\}\_\{0\}\)^\{\\top\}\\mathbb\{E\}\(\\boldsymbol\{\\epsilon\}\-\\mathbb\{E\}\\boldsymbol\{\\epsilon\}\)=\\mathbf\{0\},\(25\)
α¯k\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}1−α¯k\\sqrt\{\\,1\-\\bar\{\\alpha\}\_\{k\}\\,\}k=Kk=Kk=0k=0𝐳k=α¯k​𝐳0\+1−α¯k​ϵ\\mathbf\{z\}\_\{k\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\mathbf\{z\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}ωk′​\(k\)​𝐯k\\omega\_\{k\}^\{\\prime\}\(k\)\\,\\mathbf\{v\}\_\{k\}𝐯k=−1−α¯k​𝐳0\+α¯k​ϵ\\mathbf\{v\}\_\{k\}=\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\mathbf\{z\}\_\{0\}\+\\,\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\boldsymbol\{\\epsilon\}𝐳0\\mathbf\{z\}\_\{0\}ϵ\\boldsymbol\{\\epsilon\}ωk\\omega\_\{k\}Figure 12:Geometric illustration of thevv\-prediction formulation on the unit circle\.Because𝐳0\\mathbf\{z\}\_\{0\}is a member in a Gaussian latent space, it follows that each𝐳k\\mathbf\{z\}\_\{k\}is itself Gaussian for allkkwith mean𝟎\\mathbf\{0\}and constant unit variance𝐈\\mathbf\{I\}:

𝔼⁡\[𝐳k\]=α¯k​𝔼​\[𝐳0\]\+1−α¯k​𝔼​\[ϵ\]=α¯k​0\+1−α¯k​0=𝟎,\\mathbb\{E\}\[\\mathbf\{z\}\_\{k\}\]=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbb\{E\}\[\\mathbf\{z\}\_\{0\}\]\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbb\{E\}\[\\boldsymbol\{\\epsilon\}\]=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{0\}=\\mathbf\{0\},\(26\)
Var⁡\(𝐳k\)=α¯k​Var​\(𝐳0\)\+\(1−α¯k\)​Var​\(ϵ\)=α¯k​𝐈\+\(1−α¯k\)​𝐈=𝐈\.\\mathrm\{Var\}\(\\mathbf\{z\}\_\{k\}\)=\\bar\{\\alpha\}\_\{k\}\\,\\mathrm\{Var\}\(\\mathbf\{z\}\_\{0\}\)\+\(1\-\\bar\{\\alpha\}\_\{k\}\)\\,\\mathrm\{Var\}\(\\boldsymbol\{\\epsilon\}\)=\\bar\{\\alpha\}\_\{k\}\\,\\mathbf\{I\}\+\(1\-\\bar\{\\alpha\}\_\{k\}\)\\,\\mathbf\{I\}=\\mathbf\{I\}\.\(27\)
Hence, the diffusion process can be viewed as a rotation on the unit\(𝐳0,ϵ\)\(\\mathbf\{z\}\_\{0\},\\boldsymbol\{\\epsilon\}\)plane\. By introducing an angular parameterωk\\omega\_\{k\}such thatα¯k=cos⁡ωk\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}=\\cos\\omega\_\{k\}and1−α¯k=sin⁡ωk\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}=\\sin\\omega\_\{k\}, the sample𝐳k\\mathbf\{z\}\_\{k\}can be expressed as

𝐳k=cosωk𝐳0\+sinωkϵ,\\mathbf\{z\}\_\{k\}=\\cos\\omega\_\{k\}\\,\\mathbf\{z\}\_\{0\}\+\\sin\\omega\_\{k\}\\,\\boldsymbol\{\\epsilon\},\(28\)
which parameterizes a circular trajectory from pure data \(ω0=0\\omega\_\{0\}=0\) to pure noise \(ωK=π2\\omega\_\{K\}=\\tfrac\{\\pi\}\{2\}\)\. Differentiating this expression with respect toωk\\omega\_\{k\}yields

∂𝐳k∂θk=−sinωk𝐳0\+cosωkϵ=−1−α¯k𝐳0\+α¯kϵ≡𝐯k\.\\frac\{\\partial\\mathbf\{z\}\_\{k\}\}\{\\partial\\theta\_\{k\}\}=\-\\sin\\omega\_\{k\}\\,\\mathbf\{z\}\_\{0\}\+\\cos\\omega\_\{k\}\\,\\boldsymbol\{\\epsilon\}=\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\}\+\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}\\equiv\\mathbf\{v\}\_\{k\}\.\(29\)
Because𝐳0\\mathbf\{z\}\_\{0\}andϵ\\boldsymbol\{\\epsilon\}are orthogonal and have unit variance,𝐳k\\mathbf\{z\}\_\{k\}and𝐯k\\mathbf\{v\}\_\{k\}also form an orthonormal pair\. Geometrically,𝐳k\\mathbf\{z\}\_\{k\}represents the position on the unit circle \(or hypersphere\) and the vector𝐯k\\mathbf\{v\}\_\{k\}defines the tangent \(velocity\) direction of the probability\-flow trajectory\. The derivative with respect tokkshows us that the magnitude of the diffusion motion is modulated byωk′​\(k\)\\omega\_\{k\}^\{\\prime\}\(k\)and the direction is fully determined by𝐯k\\mathbf\{v\}\_\{k\}:

d​𝐳kd​k=d​ωkd​k​∂𝐱k∂ωk=ωt′​\(k\)​𝐯k,\\frac\{d\\mathbf\{z\}\_\{k\}\}\{dk\}=\\frac\{d\\omega\_\{k\}\}\{dk\}\\,\\frac\{\\partial\\mathbf\{x\}\_\{k\}\}\{\\partial\\omega\_\{k\}\}=\\omega\_\{t\}^\{\\prime\}\(k\)\\,\\mathbf\{v\}\_\{k\},\(30\)\.

Algorithm 2SimCast\-S2S TrainingDataset

𝔻\\mathbb\{D\}
Pretrained VAE encoders

ℰg\\mathcal\{E\}\_\{g\}for

g∈\{wind,mass,thermal,hydro,precip\}g\\in\\\{\\mathrm\{wind\},\\mathrm\{mass\},\\mathrm\{thermal\},\\mathrm\{hydro\},\\mathrm\{precip\}\\\}
Neural denoiser

f𝜽f\_\{\\boldsymbol\{\\theta\}\}
Diffusion schedule

\{α¯k\}k=1K\\\{\\bar\{\\alpha\}\_\{k\}\\\}\_\{k=1\}^\{K\}, with

α¯0=1\\bar\{\\alpha\}\_\{0\}=1
repeat

Sample

\(𝐱wind,𝐱mass,𝐱thermal,𝐱hydro,𝐱precip,t,𝐲precip\)∈𝔻\\left\(\\mathbf\{x\}\_\{\\mathrm\{wind\}\},\\mathbf\{x\}\_\{\\mathrm\{mass\}\},\\mathbf\{x\}\_\{\\mathrm\{thermal\}\},\\mathbf\{x\}\_\{\\mathrm\{hydro\}\},\\mathbf\{x\}\_\{\\mathrm\{precip\}\},t,\\mathbf\{y\}\_\{\\mathrm\{precip\}\}\\right\)\\in\\mathbb\{D\}
for

g∈\{wind,mass,thermal,hydro,precip\}g\\in\\\{\\mathrm\{wind\},\\mathrm\{mass\},\\mathrm\{thermal\},\\mathrm\{hydro\},\\mathrm\{precip\}\\\}do

Compute:

\(𝝁g,𝝈g2\)=ℰg​\(𝐱g\)\\left\(\\boldsymbol\{\\mu\}\_\{g\},\\boldsymbol\{\\sigma\}\_\{g\}^\{2\}\\right\)=\\mathcal\{E\}\_\{g\}\\left\(\\mathbf\{x\}\_\{g\}\\right\)
Sample:

ϵg∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\_\{g\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
Compute conditioning latent:

𝐳g=𝝁g\+𝝈g⊙ϵg\\mathbf\{z\}\_\{g\}=\\boldsymbol\{\\mu\}\_\{g\}\+\\boldsymbol\{\\sigma\}\_\{g\}\\odot\\boldsymbol\{\\epsilon\}\_\{g\}
endfor

Concatenate:

𝐙cond=𝐳wind⊕𝐳mass⊕𝐳thermal⊕𝐳hydro⊕𝐳precip\.\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}=\\mathbf\{z\}\_\{\\mathrm\{wind\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{mass\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{thermal\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{hydro\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{precip\}\}\.
Compute:

\(𝝁y,𝝈y2\)=ℰprecip​\(𝐲precip\)\\left\(\\boldsymbol\{\\mu\}\_\{y\},\\boldsymbol\{\\sigma\}\_\{y\}^\{2\}\\right\)=\\mathcal\{E\}\_\{\\mathrm\{precip\}\}\\left\(\\mathbf\{y\}\_\{\\mathrm\{precip\}\}\\right\)
Sample:

ϵy∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\_\{y\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
Compute clean latent:

𝐳0=𝝁y\+𝝈y⊙ϵy\\mathbf\{z\}\_\{0\}=\\boldsymbol\{\\mu\}\_\{y\}\+\\boldsymbol\{\\sigma\}\_\{y\}\\odot\\boldsymbol\{\\epsilon\}\_\{y\}
Sample:

k∼Uniform⁡\(\{1,…,K\}\)k\\sim\\mathrm\{Uniform\}\\left\(\\\{1,\\ldots,K\\\}\\right\)
Sample:

ϵ∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
Compute noisy latent:

𝐳k=α¯k​𝐳0\+1−α¯k​ϵ\.\\mathbf\{z\}\_\{k\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}\.
Compute true velocity:

𝐯k=−1−α¯k​𝐳0\+α¯k​ϵ\\mathbf\{v\}\_\{k\}=\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\}\+\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}\.

Predict velocity:

𝐯^k=f𝜽​\(𝐙cond,t,𝐳k,k\)\\hat\{\\mathbf\{v\}\}\_\{k\}=f\_\{\\boldsymbol\{\\theta\}\}\\left\(\\mathbf\{Z\}\_\{\\mathrm\{cond\}\},t,\\mathbf\{z\}\_\{k\},k\\right\)
Take a gradient step on:

∇𝜽‖𝐯k−𝐯^k‖22\.\\nabla\_\{\\boldsymbol\{\\theta\}\}\\left\\\|\\mathbf\{v\}\_\{k\}\-\\hat\{\\mathbf\{v\}\}\_\{k\}\\right\\\|\_\{2\}^\{2\}\.
untilconverged

Since𝐯k\\mathbf\{v\}\_\{k\}is still a unit vector, i\.e\.,Var⁡\(𝐯k\)=𝐈\\mathrm\{Var\}\(\\mathbf\{v\}\_\{k\}\)=\\mathbf\{I\}, the transition from theϵ\\boldsymbol\{\\epsilon\}\-prediction to thevv\-prediction formulation can be interpreted as a change of coordinate basis from the original\(𝐳0,ϵ\)\(\\mathbf\{z\}\_\{0\},\\boldsymbol\{\\epsilon\}\)system to the rotated\(𝐳k,𝐯k\)\(\\mathbf\{z\}\_\{k\},\\mathbf\{v\}\_\{k\}\)system\. As shown in Fig\.[12](https://arxiv.org/html/2608.26594#S4.F12), this transformation corresponds to a90∘90^\{\\circ\}counterclockwise rotation on the unit circle in the\(𝐳0,ϵ\)\(\\mathbf\{z\}\_\{0\},\\boldsymbol\{\\epsilon\}\)plane, preserving orthogonality and unit norm\. Consequently, conversion between two basis becomes immediately straightforward:

The signal\-to\-noise ratio \(SNR\) at timestepkkis defined as the ratio between the signal variance and the noise variance:

SNRk=Var⁡\(α¯k​𝐳0\)Var⁡\(1−α¯k​ϵ\)=α¯k1−α¯k\.\\mathrm\{SNR\}\_\{k\}=\\frac\{\\mathrm\{Var\}\(\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\mathbf\{z\}\_\{0\}\)\}\{\\mathrm\{Var\}\(\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}\)\}=\\frac\{\\bar\{\\alpha\}\_\{k\}\}\{1\-\\bar\{\\alpha\}\_\{k\}\}\.\(32\)
The standardϵ\\boldsymbol\{\\epsilon\}\-prediction objective minimizes the MSE between the true and predicted noise:ℒϵ=𝔼k​‖ϵ−ϵ^‖2\\mathcal\{L\}\_\{\\boldsymbol\{\\epsilon\}\}=\\mathbb\{E\}\_\{k\}\\,\\\|\\boldsymbol\{\\epsilon\}\-\\boldsymbol\{\\hat\{\\epsilon\}\}\\\|^\{2\}\. To examine how this loss contributes to the construction lossℒ𝐳0=𝔼k​‖𝐳0−𝐳^0‖2\\mathcal\{L\}\_\{\\mathbf\{z\}\_\{0\}\}=\\mathbb\{E\}\_\{k\}\\,\\\|\\mathbf\{z\}\_\{0\}\-\\hat\{\\mathbf\{z\}\}\_\{0\}\\\|^\{2\}at different denoising stepkk, we rewriteℒ𝐳0\\mathcal\{L\}\_\{\\mathbf\{z\}\_\{0\}\}in terms of the true clean sample𝐳0=1α¯k​\[𝐳k−1−α¯k​ϵ\]\\mathbf\{z\}\_\{0\}=\\frac\{1\}\{\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\}\[\\mathbf\{z\}\_\{k\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\epsilon\}\]and the predicted clean sample𝐳^0=1α¯k​\[𝐳k−1−α¯k​ϵ^\]\\hat\{\\mathbf\{z\}\}\_\{0\}=\\frac\{1\}\{\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\}\[\\mathbf\{z\}\_\{k\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\boldsymbol\{\\hat\{\\epsilon\}\}\]\. The reconstruction loss is then:

ℒ𝐳0=𝔼k​‖𝐳0−𝐳^0‖2=𝔼k​\[1−α¯kα¯k​‖ϵ−ϵ^‖2\]=𝔼k​\[1SNRk​‖ϵ−ϵ^‖2\]\.\\mathcal\{L\}\_\{\\mathbf\{z\}\_\{0\}\}=\\mathbb\{E\}\_\{k\}\\big\\\|\\mathbf\{z\}\_\{0\}\-\\hat\{\\mathbf\{z\}\}\_\{0\}\\big\\\|^\{2\}=\\mathbb\{E\}\_\{k\}\\left\[\\frac\{1\-\\bar\{\\alpha\}\_\{k\}\}\{\\bar\{\\alpha\}\_\{k\}\}\\big\\\|\\boldsymbol\{\\epsilon\}\-\\boldsymbol\{\\hat\{\\epsilon\}\}\\big\\\|^\{2\}\\right\]=\\mathbb\{E\}\_\{k\}\\left\[\\frac\{1\}\{\\text\{SNR\}\_\{k\}\}\\,\\big\\\|\\boldsymbol\{\\epsilon\}\-\\boldsymbol\{\\hat\{\\epsilon\}\}\\big\\\|^\{2\}\\right\]\.\(33\)
We see that each term inℒϵ\\mathcal\{L\}\_\{\\boldsymbol\{\\epsilon\}\}contributes differently to the target reconstruction lossℒ𝐳0\\mathcal\{L\}\_\{\\mathbf\{z\}\_\{0\}\}by1SNRk\\frac\{1\}\{\\text\{SNR\}\_\{k\}\}\. This weight blows up asα¯k→0\\bar\{\\alpha\}\_\{k\}\\to 0\(i\.e\.k→Kk\\to K\)\. Althoughϵ\\boldsymbol\{\\epsilon\}\-prediction still provides an unbiased estimator of the score function, its inverse\-SNR imbalance implies poor numerical conditioning and motivates thevv\-prediction formulation, whose contribution remains stable over time\.

Indeed, writing the reconstruction lossℒ𝐳0\\mathcal\{L\}\_\{\\mathbf\{z\}\_\{0\}\}with𝐳0=α¯k​𝐳k−1−α¯k​𝐯k\\mathbf\{z\}\_\{0\}=\\sqrt\{\\mathbf\{\\bar\{\\alpha\}\}\_\{k\}\}\\,\\mathbf\{z\}\_\{k\}\-\\sqrt\{1\-\\mathbf\{\\bar\{\\alpha\}\}\_\{k\}\}\\,\\mathbf\{v\}\_\{k\}and𝐳^0=α¯k​𝐳k−1−α¯k​𝐯^k\\hat\{\\mathbf\{z\}\}\_\{0\}=\\sqrt\{\\mathbf\{\\bar\{\\alpha\}\}\_\{k\}\}\\,\\mathbf\{z\}\_\{k\}\-\\sqrt\{1\-\\mathbf\{\\bar\{\\alpha\}\}\_\{k\}\}\\,\\hat\{\\mathbf\{v\}\}\_\{k\}, it now becomes:

ℒ𝐳0=𝔼k​‖𝐳0−𝐳^0‖2=𝔼k​\[\(1−α¯k\)​‖𝒗k−𝒗^k‖2\]\.\\mathcal\{L\}\_\{\\mathbf\{z\}\_\{0\}\}=\\mathbb\{E\}\_\{k\}\\big\\\|\\mathbf\{z\}\_\{0\}\-\\hat\{\\mathbf\{z\}\}\_\{0\}\\big\\\|^\{2\}=\\mathbb\{E\}\_\{k\}\\left\[\(1\-\\bar\{\\alpha\}\_\{k\}\)\\,\\big\\\|\\boldsymbol\{v\}\_\{k\}\-\\hat\{\\boldsymbol\{v\}\}\_\{k\}\\big\\\|^\{2\}\\right\]\.\(34\)
Therefore, minimizingℒ𝐯=𝔼k​‖𝒗k−𝒗^θ​\(𝐳k,k\)‖2\\mathcal\{L\}\_\{\\mathbf\{v\}\}=\\mathbb\{E\}\_\{k\}\\big\\\|\\boldsymbol\{v\}\_\{k\}\-\\hat\{\\boldsymbol\{v\}\}\_\{\\theta\}\(\\mathbf\{z\}\_\{k\},k\)\\big\\\|^\{2\}induces𝐳0\\mathbf\{z\}\_\{0\}\-errors that scale only by the bounded factors1−α¯k∈\[0,1\]1\-\\bar\{\\alpha\}\_\{k\}\\in\\left\[0,1\\right\]\. As a result, althoughvv\-prediction optimizes a step\-stationary target \(unit variance for allkk\) just likeϵ\\epsilon\-prediction, it has no inverse\-SNR amplification\. The per\-step gradient magnitudes remain commensurate acrosskk, which implies much more stable optimization\.

From the five pretrained VAEs corresponding to the five groups of physical fields listed in Table[3](https://arxiv.org/html/2608.26594#S4.T3), together with the loss function in Eq\.[11](https://arxiv.org/html/2608.26594#S4.E11), the general sampling rule in Eq\.[24](https://arxiv.org/html/2608.26594#S4.E24), and the basis conversions in Eq\.[31a](https://arxiv.org/html/2608.26594#S4.E31.x1)–[31b](https://arxiv.org/html/2608.26594#S4.E31.x2), we obtain the complete training and generating algorithms for SimCast\-S2S, summarized in Algorithms[2](https://arxiv.org/html/2608.26594#alg2)–[3](https://arxiv.org/html/2608.26594#alg3), respectively\.

Algorithm 3SimCast\-S2S GeneratingInput fields:

𝐱wind,𝐱mass,𝐱thermal,𝐱hydro,𝐱precip\\mathbf\{x\}\_\{\\mathrm\{wind\}\},\\mathbf\{x\}\_\{\\mathrm\{mass\}\},\\mathbf\{x\}\_\{\\mathrm\{thermal\}\},\\mathbf\{x\}\_\{\\mathrm\{hydro\}\},\\mathbf\{x\}\_\{\\mathrm\{precip\}\}
Day\-of\-year information

tt
Pretrained VAE encoders

ℰg\\mathcal\{E\}\_\{g\}for

g∈\{wind,mass,thermal,hydro,precip\}g\\in\\\{\\mathrm\{wind\},\\mathrm\{mass\},\\mathrm\{thermal\},\\mathrm\{hydro\},\\mathrm\{precip\}\\\}
Pretrained precipitation decoder

𝒟precip\\mathcal\{D\}\_\{\\mathrm\{precip\}\}
Trained neural denoiser

f𝜽f\_\{\\boldsymbol\{\\theta\}\}
Diffusion schedule

\{α¯k\}k=0K\\\{\\bar\{\\alpha\}\_\{k\}\\\}\_\{k=0\}^\{K\}, with

α¯0=1\\bar\{\\alpha\}\_\{0\}=1
Stochasticity parameter

η∈\[0,1\]\\eta\\in\[0,1\]
for

g∈\{wind,mass,thermal,hydro,precip\}g\\in\\\{\\mathrm\{wind\},\\mathrm\{mass\},\\mathrm\{thermal\},\\mathrm\{hydro\},\\mathrm\{precip\}\\\}do

Compute:

\(𝝁g,𝝈g2\)=ℰg​\(𝐱g\)\\left\(\\boldsymbol\{\\mu\}\_\{g\},\\boldsymbol\{\\sigma\}\_\{g\}^\{2\}\\right\)=\\mathcal\{E\}\_\{g\}\\left\(\\mathbf\{x\}\_\{g\}\\right\)
Sample:

ϵg∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\_\{g\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
Compute conditioning latent:

𝐳g=𝝁g\+𝝈g⊙ϵg\\mathbf\{z\}\_\{g\}=\\boldsymbol\{\\mu\}\_\{g\}\+\\boldsymbol\{\\sigma\}\_\{g\}\\odot\\boldsymbol\{\\epsilon\}\_\{g\}
endfor

Concatenate:

𝐙cond=𝐳wind⊕𝐳mass⊕𝐳thermal⊕𝐳hydro⊕𝐳precip\.\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}=\\mathbf\{z\}\_\{\\mathrm\{wind\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{mass\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{thermal\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{hydro\}\}\\oplus\\mathbf\{z\}\_\{\\mathrm\{precip\}\}\.
Initialize:

𝐳^K∼𝒩⁡\(𝟎,𝐈\)\\hat\{\\mathbf\{z\}\}\_\{K\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.

for

k=K,K−1,…,1k=K,K\-1,\\ldots,1do

Predict velocity:

𝐯^k=f𝜽​\(𝐙cond,t,𝐳^k,k\)\\hat\{\\mathbf\{v\}\}\_\{k\}=f\_\{\\boldsymbol\{\\theta\}\}\\left\(\\mathbf\{Z\}\_\{\\mathrm\{cond\}\},t,\\hat\{\\mathbf\{z\}\}\_\{k\},k\\right\)\.

Compute clean latent estimate:

𝐳^0=α¯k​𝐳^k−1−α¯k​𝐯^k\\hat\{\\mathbf\{z\}\}\_\{0\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{k\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\hat\{\\mathbf\{v\}\}\_\{k\}\.

Compute noise estimate:

ϵ^k=1−α¯k​𝐳^k\+α¯k​𝐯^k\\hat\{\\boldsymbol\{\\epsilon\}\}\_\{k\}=\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{k\}\+\\sqrt\{\\bar\{\\alpha\}\_\{k\}\}\\,\\hat\{\\mathbf\{v\}\}\_\{k\}\.

Compute:

σk=η​1−α¯k−11−α¯k​1−α¯kα¯k−1\\sigma\_\{k\}=\\eta\\sqrt\{\\frac\{1\-\\bar\{\\alpha\}\_\{k\-1\}\}\{1\-\\bar\{\\alpha\}\_\{k\}\}\}\\sqrt\{1\-\\frac\{\\bar\{\\alpha\}\_\{k\}\}\{\\bar\{\\alpha\}\_\{k\-1\}\}\}\.

Compute:

ck=1−α¯k−1−σk2c\_\{k\}=\\sqrt\{1\-\\bar\{\\alpha\}\_\{k\-1\}\-\\sigma\_\{k\}^\{2\}\}\.

Sample:

𝜻k∼𝒩⁡\(𝟎,𝐈\)\\boldsymbol\{\\zeta\}\_\{k\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)
Compute:

𝐳^k−1=α¯k−1​𝐳^0\+ck​ϵ^k\+σk​𝜻k\\hat\{\\mathbf\{z\}\}\_\{k\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{k\-1\}\}\\,\\hat\{\\mathbf\{z\}\}\_\{0\}\+c\_\{k\}\\hat\{\\boldsymbol\{\\epsilon\}\}\_\{k\}\+\\sigma\_\{k\}\\boldsymbol\{\\zeta\}\_\{k\}\.

endfor

returnprecipitation\-anomaly forecast:

𝐲^=𝒟precip​\(𝐳^0\)\\hat\{\\mathbf\{y\}\}=\\mathcal\{D\}\_\{\\mathrm\{precip\}\}\\left\(\\hat\{\\mathbf\{z\}\}\_\{0\}\\right\)\.

In this study,𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}andttencode the physical atmospheric states over the previous2828days and their associated seasonal context, respectively\. Both𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}andttare step\-independent and remain fixed throughout the denoising process\. The neural network uses these inputs to guide the reverse process so that the generated latent𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}represents the precipitation anomaly over lead days 15–28, corresponding to the 14\-day subseasonal forecast window\. A conceptual illustration of the latent diffusion process is provided in Fig\.[11](https://arxiv.org/html/2608.26594#S4.F11)\. During generation, the neural network denoises the latent state sequentially from stepKKto step00\. It should be noted that the generation process is initialized randomly and𝐳^K\\hat\{\\mathbf\{z\}\}\_\{K\}can be far apart from𝐳K\\mathbf\{z\}\_\{K\}\. However, the neural network uses the conditioning information to guide the denoising trajectory so that𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}is expected to be close to𝐳0\\mathbf\{z\}\_\{0\}\. Additionally, ifη=0\\eta=0, the denoising trajectory \(yellow\) is deterministic for a fixed initial latent𝐳^K\\hat\{\\mathbf\{z\}\}\_\{K\}\. If0<η≤10<\\eta\\leq 1, additional Gaussian noise is injected at each reverse step, which makes the denoising trajectory stochastic\.

The neural networkfθf\_\{\\theta\}is designed as a compact denoising engine for latent\-space forecasting\. It follows an encoder–decoder structure with multiple downsampling blocks, each combining convolutional and transformer layers, followed by a symmetric set of upsampling blocks\. The convolutional layers refine local spatial structure within the latent fields and the transformer layers handle the temporal exchange of information between the conditioning history and the forecast target\.

This temporal structure is explicit\. The conditioning sequence contains 28 tokens, one for each input day, and the target sequence contains 14 tokens, one for each forecast day\. Through cross\-attention\([Vaswani et al\., 2017](https://arxiv.org/html/2608.26594#bib.bib67)\), the 14 noisy target tokens query the 28 conditioning tokens and extract the parts of the past atmospheric evolution most relevant for denoising the future precipitation latent\. In this formulation,𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}provides the key–value pairs, and the intermediate representation of𝐳^k\\hat\{\\mathbf\{z\}\}\_\{k\}provides the queries\. The denoiser therefore selectively learns which parts of the 28\-day atmospheric history should inform each forecast day\.

The diffusion stepkkand day\-of\-yearttare passed through separate sinusoidal embedding blocks[Vaswani et al\. \(2017\)](https://arxiv.org/html/2608.26594#bib.bib67), which convert scalar inputs into vector representations while preserving their relative proximity in Euclidean space\. Given the noisy target latent𝐳^​k\\hat\{\\mathbf\{z\}\}k, the conditioning latent sequence𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}, the diffusion\-step embedding, and the day\-of\-year embedding, the denoiser predicts the velocity𝐯^k\\hat\{\\mathbf\{v\}\}\_\{k\}\. This predicted velocity defines the direction used to move the latent state from𝐳^k\\hat\{\\mathbf\{z\}\}\_\{k\}toward𝐳^k−1\\hat\{\\mathbf\{z\}\}\_\{k\-1\}at each reverse step\. The detailed architecture offθf\_\{\\theta\}is provided in Fig\.[13](https://arxiv.org/html/2608.26594#S4.F13)

![Refer to caption](https://arxiv.org/html/2608.26594v1/assets/fig13.png)Figure 13:Architecture of the SimCast\-S2S denoiser\. At stepkk, the noisy target precipitation latent𝐳^k\\hat\{\\mathbf\{z\}\}\_\{k\}is passed through a sequence of ConvStack and Transformer blocks to predict the velocity𝐯^k\\hat\{\\mathbf\{v\}\}\_\{k\}\. Although the denoiser is trained to predict𝐯^k\\hat\{\\mathbf\{v\}\}\_\{k\}, the next denoised latent state𝐳^k−1\\hat\{\\mathbf\{z\}\}\_\{k\-1\}is obtained using Eqs\.[31a](https://arxiv.org/html/2608.26594#S4.E31.x1)–[31b](https://arxiv.org/html/2608.26594#S4.E31.x2)and Eq\.[24](https://arxiv.org/html/2608.26594#S4.E24)\. The conditioning information is formed by concatenating the latent representations of the five physical groups,𝐳wind\\mathbf\{z\}\_\{\\mathrm\{wind\}\},𝐳mass\\mathbf\{z\}\_\{\\mathrm\{mass\}\},𝐳thermal\\mathbf\{z\}\_\{\\mathrm\{thermal\}\},𝐳hydro\\mathbf\{z\}\_\{\\mathrm\{hydro\}\}, and𝐳precip\\mathbf\{z\}\_\{\\mathrm\{precip\}\}, into𝐙cond\\mathbf\{Z\}\_\{\\mathrm\{cond\}\}, which is injected into the Transformer blocks through cross\-attention\. The diffusion stepkkand seasonal informationttare encoded using separate sinusoidal embedding layers and injected into the ConvStack blocks\. This denoising operation is repeated forKKreverse steps, transforming an initial Gaussian latent sample𝐳^K∼𝒩⁡\(𝟎,𝐈\)\\hat\{\\mathbf\{z\}\}\_\{K\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)into the generated target precipitation latent𝐳^0\\hat\{\\mathbf\{z\}\}\_\{0\}\.
### 4\.4Transfer Learning: LoRA

Perhaps one of the most important features of SimCast\-S2S is its computationally efficient latent\-space framework\. This efficiency allows SimCast\-S2S to be pretrained on 28 CESM2 atmospheric simulations and then transferred to real\-world atmospheric estimates from ERA5\. Traditional transfer learning typically follows one of common strategies: fine\-tuning the full pretrained network with a small learning rate[Pan and Yang \(2010\)](https://arxiv.org/html/2608.26594#bib.bib72);[Yosinski et al\. \(2014\)](https://arxiv.org/html/2608.26594#bib.bib68);[Zhuang et al\. \(2021\)](https://arxiv.org/html/2608.26594#bib.bib73), freezing early layers and retraining only task\-specific layers[Donahue et al\. \(2014\)](https://arxiv.org/html/2608.26594#bib.bib69);[Oquab et al\. \(2014\)](https://arxiv.org/html/2608.26594#bib.bib74), or keeping the pretrained backbone mostly fixed while adding small trainable adaptation modules[Rebuffi et al\. \(2017\)](https://arxiv.org/html/2608.26594#bib.bib75);[Houlsby et al\. \(2019\)](https://arxiv.org/html/2608.26594#bib.bib70)\. SimCast\-S2S follows the third strategy by using low\-rank adaptation \(LoRA\)[Hu et al\. \(2022\)](https://arxiv.org/html/2608.26594#bib.bib71)\.

LoRA transfer\-learning strategy is applied toℰprecip\\mathcal\{E\}\_\{\\mathrm\{precip\}\},𝒟precip\\mathcal\{D\}\_\{\\mathrm\{precip\}\}, andfθf\_\{\\theta\}\. We do not adapt the VAE encoders and decoders for group wind, mass, thermal, hydro because these VAEs show no substantial degradation when evaluated on ERA5 with no fine\-tuning\. In contrast, precipitation exhibits a stronger distributional shift between CESM2 and ERA5, so the precipitation encoder and decoder are fine\-tuned together with the denoising network\.

Each ConvStack layer in Figs\.[10](https://arxiv.org/html/2608.26594#S4.F10)and[13](https://arxiv.org/html/2608.26594#S4.F13)contains multiple convolutional layers[LeCun et al\. \(1998b\)](https://arxiv.org/html/2608.26594#bib.bib64);[Krizhevsky et al\. \(2012\)](https://arxiv.org/html/2608.26594#bib.bib65), and each Transformer layer in Fig\.[13](https://arxiv.org/html/2608.26594#S4.F13)also contains multiple dense layers as their core components[Vaswani et al\. \(2017\)](https://arxiv.org/html/2608.26594#bib.bib67)\. For any such layer, let𝐖∈ℝdin×dout\\mathbf\{W\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\\times d\_\{\\mathrm\{out\}\}\}denote the pretrained weight matrix learned from CESM2, wheredind\_\{\\mathrm\{in\}\}anddoutd\_\{\\mathrm\{out\}\}are the input and output dimensions, respectively\.

Instead of fine\-tuning the full weight matrix𝐖\\mathbf\{W\}, which can be computationally expensive and may overwrite useful structure learned during CESM2 pretraining, LoRA freezes𝐖\\mathbf\{W\}and learns only a low\-rank updateΔ​𝐖\\Delta\\mathbf\{W\}:

𝐖′=𝐖\+Δ​𝐖\.\\mathbf\{W\}^\{\\prime\}=\\mathbf\{W\}\+\\Delta\\mathbf\{W\}\.\(35\)
The matrixΔ​𝐖\\Delta\\mathbf\{W\}can be written in a low\-rank factorization:

Δ​𝐖=𝐁𝐀,\\Delta\\mathbf\{W\}=\\mathbf\{B\}\\mathbf\{A\},\(36\)
where𝐁∈ℝdin×r\\mathbf\{B\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\\times r\},𝐀∈ℝr×dout\\mathbf\{A\}\\in\\mathbb\{R\}^\{r\\times d\_\{\\mathrm\{out\}\}\}, and most importantly,r≪min⁡\(din,dout\)r\\ll\\min\(d\_\{\\mathrm\{in\}\},d\_\{\\mathrm\{out\}\}\)is the rank of matrixΔ​𝐖\\Delta\\mathbf\{W\}\. For an input feature vector𝐡∈ℝdin\\mathbf\{h\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\}, the adapted output becomes

𝐡𝐖′=𝐡𝐖\+𝐡𝐁𝐀\.\\mathbf\{h\}\\mathbf\{W\}^\{\\prime\}=\\mathbf\{h\}\\mathbf\{W\}\+\\mathbf\{h\}\\mathbf\{B\}\\mathbf\{A\}\.\(37\)The first term on the right\-hand side preserves the pretrained CESM2 knowledge\. The second term introduces a small trainable correction that adapts the model from CESM2 to ERA5\. During fine\-tuning, only𝐀\\mathbf\{A\}and𝐁\\mathbf\{B\}are optimized,𝐖\\mathbf\{W\}remains intact\. As a result, we only need to optimize𝒪⁡\(r⁡\(din\+dout\)\)\\mathcal\{O\}\(r\(d\_\{\\mathrm\{in\}\}\+d\_\{\\mathrm\{out\}\}\)\)number of parameters, instead of𝒪⁡\(din​dout\)\\mathcal\{O\}\(d\_\{\\mathrm\{in\}\}d\_\{\\mathrm\{out\}\}\)\. Becauserris chosen to be much smaller than bothdind\_\{\\mathrm\{in\}\}anddoutd\_\{\\mathrm\{out\}\},𝒪⁡\(r⁡\(din\+dout\)\)\\mathcal\{O\}\(r\(d\_\{\\mathrm\{in\}\}\+d\_\{\\mathrm\{out\}\}\)\)is orders of magnitude smaller than𝒪⁡\(din​dout\)\\mathcal\{O\}\(d\_\{\\mathrm\{in\}\}d\_\{\\mathrm\{out\}\}\)\. This makes LoRA more memory efficient and less prone to overfitting\.

LoRA is particularly useful for SimCast\-S2S because the model can exploit a large ensemble of CESM2 simulations\. As shown in Section[2\.3](https://arxiv.org/html/2608.26594#S2.SS3), CESM2 pretraining followed by ERA5 fine\-tuning with LoRA performs substantially better than training on ERA5 alone, and it is one of the main reasons SimCast\-S2S outperforms the process\-based ECMWF\-S2S\. The only task assigned to SimCast\-S2S during fine\-tuning is to adjust for the domain shift from many simulated worlds to one single real world\.

### 4\.5Deep Learning Baselines: CNN, UNet

SimCast\-S2S is benchmarked against two widely used deep learning architectures for spatial prediction: CNN and UNet\. CNN models are built on top of the convolution operator and provide a natural baseline for gridded prediction tasks and provide a natural baseline for gridded prediction tasks because they learn local spatial patterns through shared kernels\([LeCun et al\., 1998b](https://arxiv.org/html/2608.26594#bib.bib64);[Krizhevsky et al\., 2012](https://arxiv.org/html/2608.26594#bib.bib65)\)\. The UNet architecture extends this idea with an encoder–decoder structure and skip connections that help preserve multiscale spatial information\([Ronneberger et al\., 2015](https://arxiv.org/html/2608.26594#bib.bib32)\)\. In this study, we construct three CNN baselines and three U\-Net baselines with different model sizes\. Tables[4](https://arxiv.org/html/2608.26594#S4.T4)–[5](https://arxiv.org/html/2608.26594#S4.T5)summarize the key design choices\.

Table 4:Architectural configurations of the CNN baselines\. Each CNN uses a stack of convolutional layers \(NN\) with a fixed embedding dimension \(dd\)\. The three model sizes increase both the embedding dimension and the number of layers\.ModelddNNcnn\-small25625644cnn\-medium51251288cnn\-large102410241212Table 5:Architectural configurations of the UNet baselines\. Each U\-Net uses a symmetric encoder–decoder structure with four sampling blocks, with embedding dimensions\{di\}i=14\\\{d\_\{i\}\\\}\_\{i=1\}^\{4\}, respectively\. In the encoder, each block downsamples the spatial resolution and doubles the embedding dimension\. In the decoder, the corresponding upsampling blocks recover the spatial resolution and halve the embedding dimension\.Modeld1d\_\{1\}d2d\_\{2\}d3d\_\{3\}d4d\_\{4\}unet\-small12812825625651251210241024unet\-medium2562565125121024102420482048unet\-large512512102410242048204840964096
### 4\.6Physical Baseline: ECMWF\-S2S

To place SimCast\-S2S against a strong process\-based reference, we use ECMWF\-S2S as the operational physical baseline\. ECMWF is widely regarded as one of the world\-leading numerical weather prediction centers, and its subseasonal forecasts, ECMWF\-S2S, represent a highly developed dynamical forecasting system[Bauer et al\. \(2015\)](https://arxiv.org/html/2608.26594#bib.bib76);[Vitart et al\. \(2017\)](https://arxiv.org/html/2608.26594#bib.bib1);[Vitart \(2022\)](https://arxiv.org/html/2608.26594#bib.bib37)\. Unlike neural\-network baselines, ECMWF\-S2S is produced by integrating a physically based Earth\-system model that explicitly resolves atmospheric dynamics, parameterizes unresolved physical processes, and generates ensemble forecasts through perturbed initial conditions and model uncertainties[Leutbecher and Palmer \(2008\)](https://arxiv.org/html/2608.26594#bib.bib77);[Vitart et al\. \(2017\)](https://arxiv.org/html/2608.26594#bib.bib1)\. Outperforming ECMWF\-S2S would indicate that SimCast\-S2S is competitive with one of the most advanced operational forecasting systems currently available\. At the same time, this comparison should be interpreted with care\. SimCast\-S2S is not intended to replace ECMWF\-S2S\. Instead, it provides an inexpensive alternative for subseasonal precipitation forecasting, accessible to research groups and computing centers with smaller\-scale computational resources\.

### 4\.7Evaluation Metrics

#### 4\.7\.1Deterministic Skill Metrics

We evaluate deterministic skill over all test samples using the ensemble\-mean precipitation\-anomaly forecast\. Let𝐲^\(m\)\\hat\{\\mathbf\{y\}\}^\{\(m\)\}and𝐲\(m\)\\mathbf\{y\}^\{\(m\)\}denote the predicted and verifying precipitation\-anomaly field for test samplemm, respectively\. Here,m=1,…,Mm=1,\\ldots,Mindexes the test samples andi=1,…,Ni=1,\\ldots,Nindexes the spatial grid points\.

The first deterministic metric is mean absolute error \(MAE\), which measures the average magnitude of the grid\-level forecast error across samples and space:

MAE=1M​N​∑m=1M∑i=1N\|y^i\(m\)−yi\(m\)\|∈ℝ\.\\mathrm\{MAE\}=\\frac\{1\}\{MN\}\\sum\_\{m=1\}^\{M\}\\sum\_\{i=1\}^\{N\}\\left\|\\hat\{y\}^\{\(m\)\}\_\{i\}\-y^\{\(m\)\}\_\{i\}\\right\|\\quad\\in\\mathbb\{R\}\.\(38\)
Lower MAE indicates smaller absolute error and therefore better deterministic accuracy\. Because all precipitation fields are converted to standardized anomalies, MAE measures error relative to the local climatology\.

The second deterministic metric is anomaly correlation coefficient \(ACC\), which evaluates whether the forecast captures the spatial pattern of the observed precipitation anomaly\. For each test sample, ACC is first computed over the spatial grid:

ACC\(m\)=∑i=1Ny^i\(m\)​yi\(m\)∑i=1N\(y^i\(m\)\)2​∑i=1N\(yi\(m\)\)2∈ℝ\.\\mathrm\{ACC\}^\{\(m\)\}=\\frac\{\\sum\_\{i=1\}^\{N\}\\hat\{y\}^\{\(m\)\}\_\{i\}y^\{\(m\)\}\_\{i\}\}\{\\sqrt\{\\sum\_\{i=1\}^\{N\}\\left\(\\hat\{y\}^\{\(m\)\}\_\{i\}\\right\)^\{2\}\}\\sqrt\{\\sum\_\{i=1\}^\{N\}\\left\(y^\{\(m\)\}\_\{i\}\\right\)^\{2\}\}\}\\quad\\in\\mathbb\{R\}\.\(39\)
The reported ACC is then averaged across all test samples:

ACC=1M​∑m=1MACC\(m\)∈ℝ\.\\mathrm\{ACC\}=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}\\mathrm\{ACC\}^\{\(m\)\}\\quad\\in\\mathbb\{R\}\.\(40\)
Higher ACC indicates stronger agreement between the predicted and observed anomaly patterns\. A skillful subseasonal precipitation forecast should achieve both low MAE and high ACC, meaning that it is close to the observed anomaly field in magnitude and also reproducing the spatial structure of wet and dry anomalies\.

#### 4\.7\.2Probabilistic Skill Metrics

Probabilistic skill is evaluated using the full ensemble distribution the ensemble mean\. For each test samplemm, SimCast\-S2S produces an ensemble of precipitation\-anomaly forecasts\{𝐲^\(m,e\)\}e=1E\\\{\\hat\{\\mathbf\{y\}\}^\{\(m,e\)\}\\\}\_\{e=1\}^\{E\}, whereEEis the number of ensemble members\. These ensemble members define an empirical predictive distribution at each grid point\. The probabilistic metrics evaluate whether this distribution assigns high probability to the observed outcome and whether it improves over a reference forecast distribution\.

The first probabilistic metric is the ranked probability skill score \(RPSS\), which is used for tercile\-based categorical forecasts\. For each grid point, the precipitation anomaly distribution is divided into three climatological categories: below\-normal, near\-normal, and above\-normal\. Let𝐩\(m\)=\(p1\(m\),p2\(m\),p3\(m\)\)\\mathbf\{p\}^\{\(m\)\}=\(p^\{\(m\)\}\_\{1\},p^\{\(m\)\}\_\{2\},p^\{\(m\)\}\_\{3\}\)denote the forecast probabilities assigned to the three categories for samplemm, and let𝐨\(m\)=\(o1\(m\),o2\(m\),o3\(m\)\)\\mathbf\{o\}^\{\(m\)\}=\(o^\{\(m\)\}\_\{1\},o^\{\(m\)\}\_\{2\},o^\{\(m\)\}\_\{3\}\)denote the observed one\-hot category vector\. The ranked probability score \(RPS\) is

RPS\(m\)=∑c=13\(∑j=1cpj\(m\)−∑j=1coj\(m\)\)2∈ℝN\.\\mathrm\{RPS\}^\{\(m\)\}=\\sum\_\{c=1\}^\{3\}\\left\(\\sum\_\{j=1\}^\{c\}p^\{\(m\)\}\_\{j\}\-\\sum\_\{j=1\}^\{c\}o^\{\(m\)\}\_\{j\}\\right\)^\{2\}\\quad\\in\\mathbb\{R\}^\{N\}\.\(41\)
The RPSS compares the forecast RPS against the climatological reference\. The sample\-level RPSS is defined as

RPSS\(m\)=1−RPS¯forecast\(m\)RPS¯climatology\(m\)∈ℝ\.\\mathrm\{RPSS\}^\{\(m\)\}=1\-\\frac\{\\overline\{\\mathrm\{RPS\}\}^\{\(m\)\}\_\{\\mathrm\{forecast\}\}\}\{\\overline\{\\mathrm\{RPS\}\}^\{\(m\)\}\_\{\\mathrm\{climatology\}\}\}\\quad\\in\\mathbb\{R\}\.\(42\)
Here and throughout the rest of this section, the overbar denotes averaging over allNNgrid points\. The reported RPSS is obtained by averaging the sample\-level skill scores over allMMtest samples\. Positive RPSS indicates that the probabilistic forecast improves upon climatology\. Negative RPSS indicates worse performance than climatology\.

A more general version of RPSS is the continuous ranked probability skill score \(CRPSS\), which evaluates the full continuous predictive distribution rather than discrete tercile categories\. For a predictive cumulative distribution functionF\(m\)F^\{\(m\)\}and verifying observationy\(m\)y^\{\(m\)\}of samplemm, the continuous ranked probability score \(CRPS\) is

CRPS⁡\(F\(m\),y\(m\)\)=∫−∞∞\[F\(m\)​\(x\)−𝟏y\(m\)≤x\]2​𝑑x∈ℝN\.\\mathrm\{CRPS\}\\left\(F^\{\(m\)\},y^\{\(m\)\}\\right\)=\\int\_\{\-\\infty\}^\{\\infty\}\\left\[F^\{\(m\)\}\(x\)\-\\mathbf\{1\}\_\{y^\{\(m\)\}\\leq x\}\\right\]^\{2\}dx\\quad\\in\\mathbb\{R\}^\{N\}\.\(43\)
For ensemble forecasts,F\(m\)F^\{\(m\)\}is represented by the empirical distribution of ensemble members\. The CRPSS of samplemmis then defined as

CRPSS\(m\)=1−CRPS¯forecast\(m\)CRPS¯climatology\(m\)∈ℝ\.\\mathrm\{CRPSS\}^\{\(m\)\}=1\-\\frac\{\\overline\{\\mathrm\{CRPS\}\}^\{\(m\)\}\_\{\\mathrm\{forecast\}\}\}\{\\overline\{\\mathrm\{CRPS\}\}^\{\(m\)\}\_\{\\mathrm\{climatology\}\}\}\\quad\\in\\mathbb\{R\}\.\(44\)
Similar to RPSS, the reported RPSS is the average of all sample\-level skill scores in the test dataset\. Higher CRPSS is better, positive CRPSS indicates improvement from the climatology\.

The third probabilistic metric is the Brier skill score \(BSS\), which is used to evaluate binary event probabilities\. In this study, BSS is applied to extreme precipitation events exceeding the9090th percentile\. Letp\(m\)p^\{\(m\)\}denote the forecast probability of the event estimated from the ensemble forecasts, and let the one\-hot categoryo\(m\)∈\{0,1\}o^\{\(m\)\}\\in\\\{0,1\\\}denote whether the extreme event occurred\. The Brier score of samplemmis

BS\(m\)=\(p\(m\)−o\(m\)\)2∈ℝN\.\\mathrm\{BS\}^\{\(m\)\}=\\left\(p^\{\(m\)\}\-o^\{\(m\)\}\\right\)^\{2\}\\quad\\in\\mathbb\{R\}^\{N\}\.\(45\)
The BSS also compares this score with the corresponding climatological event probability:

BSS\(m\)=1−BS¯forecast\(m\)BS¯climatology\(m\)∈ℝ\.\\mathrm\{BSS\}^\{\(m\)\}=1\-\\frac\{\\overline\{\\mathrm\{BS\}\}^\{\(m\)\}\_\{\\mathrm\{forecast\}\}\}\{\\overline\{\\mathrm\{BS\}\}^\{\(m\)\}\_\{\\mathrm\{climatology\}\}\}\\quad\\in\\mathbb\{R\}\.\(46\)
The reported BSS is also the average score of all test samples\. Positive BSS indicates that the ensemble forecasts capture extreme events better than the climatological reference\.

Although RPSS, CRPSS, and BSS evaluate different aspects of probabilistic forecast quality, they share a common idea\. A skillful subseasonal precipitation forecast should place probability mass near the observed precipitation anomaly\. The three metrics differ mainly in how they break down the forecast probability density function \(PDF\)\. RPSS evaluates probability mass over discretized partitions of the PDF, CRPSS evaluates the continuous PDF as it is, and BSS evaluates the right tail of the PDF for extreme precipitation events\.

#### 4\.7\.3Uncertainty Quantification

A powerful metric for quantifying the uncertainty of ensemble forecasts is the probability integral transform \(PIT\), which directly tests whether the predictive distribution is statistically consistent with the observations\. The idea is simple\. For an unbiased probabilistic forecast, the observation should behave like a random draw from the forecast distribution\. Therefore, the PIT values should be uniformly distributed on\[0,1\]\[0,1\]\.

For test samplemmand grid pointii, letF^m,i\\widehat\{F\}\_\{m,i\}denote the empirical cumulative distribution function \(CDF\) defined by theEE–member ensemble forecast\{y^i\(m,e\)\}e=1E\\\{\\hat\{y\}^\{\(m,e\)\}\_\{i\}\\\}\_\{e=1\}^\{E\}\. The PIT value is computed as

um,i=F^m,i\(yi\(m\)\)=1E∑e=1E𝟏\{y^i\(m,e\)≤yi\(m\)\}\.u\_\{m,i\}=\\widehat\{F\}\_\{m,i\}\\left\(y^\{\(m\)\}\_\{i\}\\right\)=\\frac\{1\}\{E\}\\sum\_\{e=1\}^\{E\}\\mathbf\{1\}\\left\\\{\\hat\{y\}^\{\(m,e\)\}\_\{i\}\\leq y^\{\(m\)\}\_\{i\}\\right\\\}\.\(47\)whereyi\(m\)y^\{\(m\)\}\_\{i\}is the true precipitation anomaly\. An unbiased ensemble should produce PIT values that are approximately uniform\. A∪\\cup–shaped PIT histogram indicates underdispersion, meaning that the ensemble spread is too narrow and observations fall too often in the tails of the forecast distribution\. A∩\\cap–shaped PIT histogram indicates overdispersion, meaning that the ensemble spread is too wide\.

To quantify the deviation of the PIT histogram from uniformity, we compute theχ2\\chi^\{2\}–distance to flatness\. Let the interval\[0,1\]\[0,1\]be divided intoBBequally spaced bins, and let

nb=∑m=1M∑i=1N𝟏\{um,i∈Ib\}n\_\{b\}=\\sum\_\{m=1\}^\{M\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\\left\\\{u\_\{m,i\}\\in I\_\{b\}\\right\\\}\(48\)
denote the number of PIT values falling in binIbI\_\{b\}\. The total number of PIT values isM​NMN, so the expected count under a perfectly uniform PIT distribution isM​NB\\frac\{MN\}\{B\}in each bin\. Theχ2\\chi^\{2\}distance to flatness is then defined as

χ2=∑b=1B\(nb−M​NB\)2M​NB\.\\chi^\{2\}=\\sum\_\{b=1\}^\{B\}\\frac\{\\left\(n\_\{b\}\-\\frac\{MN\}\{B\}\\right\)^\{2\}\}\{\\frac\{MN\}\{B\}\}\.\(49\)
A lower value ofχ2\\chi^\{2\}–distance indicates a PIT histogram closer to uniformity and therefore better capture the true uncertainty\. A largeχ2\\chi^\{2\}–distance value indicates that the ensemble distribution is systematically inconsistent with the observed precipitation anomalies, either because the ensemble is underdispersed, or overdispersed\.

A completely different view on uncertainty quantification is provided by the relationship between ensemble spread and deterministic error\. The spread–RMSE relationship evaluates whether the ensemble spread is consistent with the actual error of the ensemble\-mean forecast\. We want larger forecast uncertainty to correspond to larger ensemble\-mean error\. Consequently, the average ensemble spread should be comparable to the average RMSE\.

For each test samplemmand grid pointii, the ensemble\-mean forecast over allEEmembers is

y¯i\(m\)=1E​∑e=1Ey^i\(m,e\),\\bar\{y\}^\{\(m\)\}\_\{i\}=\\frac\{1\}\{E\}\\sum\_\{e=1\}^\{E\}\\hat\{y\}^\{\(m,e\)\}\_\{i\},\(50\)and the ensemble spread is

si\(m\)=1E−1∑e=1E\(y^\(,e\)i−y¯\(m\)i\)2\.s^\{\(m\)\}\_\{i\}=\\sqrt\{\\frac\{1\}\{E\-1\}\\sum\_\{e=1\}^\{E\}\\left\(\\hat\{y\}^\{\(,e\)\}\_\{i\}\-\\bar\{y\}^\{\(m\)\}\_\{i\}\\right\)^\{2\}\}\.\(51\)
The spatially averaged ensemble spread for samplemmis then defined as

Spread\(m\)=1N​∑i=1N\(si\(m\)\)2,\\mathrm\{Spread\}^\{\(m\)\}=\\sqrt\{\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(s^\{\(m\)\}\_\{i\}\\right\)^\{2\}\},\(52\)
and the corresponding RMSE of the ensemble mean is simply

RMSE\(m\)=1N​∑i=1N\(y¯i\(m\)−yi\(m\)\)2\.\\mathrm\{RMSE\}^\{\(m\)\}=\\sqrt\{\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(\\bar\{y\}^\{\(m\)\}\_\{i\}\-y^\{\(m\)\}\_\{i\}\\right\)^\{2\}\}\.\(53\)
For an ideally statistically consistent ensemble, it can be shown that the spread–RMSE diagram should approximately be the1:11\{:\}1line[Fortin et al\. \(2014\)](https://arxiv.org/html/2608.26594#bib.bib78)\. If the spread is much smaller than the RMSE, the ensemble is underdispersed and does not represent enough forecast uncertainty\. If the spread is much larger than the RMSE, the ensemble is overdispersed and produces an unnecessarily broad forecast distribution\.

We also evaluate uncertainty using the mean interval score \(MIS\)\. For a central\(1−α\)\(1\-\\alpha\)prediction interval, letli\(m\)l^\{\(m\)\}\_\{i\}andui\(m\)u^\{\(m\)\}\_\{i\}denote the lower and upper ensemble quantiles at grid pointiifor samplemm\. The interval score is defined as

ISi\(m\)\(α\)=\(ui\(m\)−li\(m\)\)\+2α\(li\(m\)−yi\(m\)\)𝟏\{yi\(m\)<li\(m\)\}\+2α\(yi\(m\)−ui\(m\)\)𝟏\{yi\(m\)\>ui\(m\)\}\.\\mathrm\{IS\}^\{\(m\)\}\_\{i\}\(\\alpha\)=\\left\(u^\{\(m\)\}\_\{i\}\-l^\{\(m\)\}\_\{i\}\\right\)\+\\frac\{2\}\{\\alpha\}\\left\(l^\{\(m\)\}\_\{i\}\-y^\{\(m\)\}\_\{i\}\\right\)\\mathbf\{1\}\\left\\\{y^\{\(m\)\}\_\{i\}<l^\{\(m\)\}\_\{i\}\\right\\\}\+\\frac\{2\}\{\\alpha\}\\left\(y^\{\(m\)\}\_\{i\}\-u^\{\(m\)\}\_\{i\}\\right\)\\mathbf\{1\}\\left\\\{y^\{\(m\)\}\_\{i\}\>u^\{\(m\)\}\_\{i\}\\right\\\}\.\(54\)
The MIS is obtained by averaging over all test samples and grid points:

MIS⁡\(α\)=1M​N​∑m=1M∑i=1NISi\(m\)​\(α\)\.\\mathrm\{MIS\}\(\\alpha\)=\\frac\{1\}\{MN\}\\sum\_\{m=1\}^\{M\}\\sum\_\{i=1\}^\{N\}\\mathrm\{IS\}^\{\(m\)\}\_\{i\}\(\\alpha\)\.\(55\)
Lower MIS indicates better interval forecasts\. The first term\(ui\(m\)−li\(m\)\)\(u^\{\(m\)\}\_\{i\}\-l^\{\(m\)\}\_\{i\}\)rewards sharp prediction intervals\. The terms2α\(li\(m\)−yi\(m\)\)1\{yi\(m\)<li\(m\)\}\\frac\{2\}\{\\alpha\}\(l^\{\(m\)\}\_\{i\}\-y^\{\(m\)\}\_\{i\}\)\\,\\mathbf\{1\}\\\{y^\{\(m\)\}\_\{i\}<l^\{\(m\)\}\_\{i\}\\\}and2α\(yi\(m\)−ui\(m\)\)𝟏\{yi\(m\)\>ui\(m\)\}\\frac\{2\}\{\\alpha\}\(y^\{\(m\)\}\_\{i\}\-u^\{\(m\)\}\_\{i\}\)\\mathbf\{1\}\\\{y^\{\(m\)\}\_\{i\}\>u^\{\(m\)\}\_\{i\}\\\}penalize observations that fall outside the predicted interval\. Therefore, MIS jointly evaluates sharpness and statistical consistency\. A good ensemble should produce intervals that are narrow but still contain the observed precipitation anomaly\.

#### 4\.7\.4Realism Quantification: Spatial Auto\-correlation

Beyond deterministic accuracy and probabilistic skill, we also care about whether the generated precipitation\-anomaly fields have realistic spatial structure\. This is important because a forecast can obtain reasonable pointwise scores but smooths out the physical structures at the same time\. A good way to quantify realism is comparing the spatial autocorrelation between the forecasts and the true observations\. Spatial autocorrelation quantifies the internal spatial texture of a precipitation field by measuring how strongly the field correlates with a spatially shifted version of itself\.

For a precipitation\-anomaly field𝐲\(m\)\\mathbf\{y\}^\{\(m\)\}, we compute directional autocorrelation by shifting the field by a spatial lagℓ\\elland correlating the original field with the shifted field over their overlapping grid points\. Let𝒮d,ℓ\\mathcal\{S\}\_\{d,\\ell\}denote the shift operator in directionddwith lagℓ\\ell, whered∈\{zonal,meridional,main​\-​diagonal,anti​\-​diagonal\}\.d\\in\\\{\\mathrm\{zonal\},\\mathrm\{meridional\},\\mathrm\{main\\text\{\-\}diagonal\},\\mathrm\{anti\\text\{\-\}diagonal\}\\\}\.

For example, a zonal shift compares each grid point with another grid pointℓ\\ellpixels away in the longitudinal direction, a meridional shift compares grid pointsℓ\\ellpixels apart in the latitudinal direction\. The two diagonal shifts compare grid points displaced simultaneously in both spatial dimensions\.

LetΩd,ℓ\\Omega\_\{d,\\ell\}denote the set of grid points for which both the original field and the shifted field are defined\. The observed directional autocorrelation for test samplemm, in directiondd, and with lagℓ\\ellis defined as

ρobs\(m\)​\(d,ℓ\)=∑i∈Ωd,ℓ\(yi\(m\)−y¯d,ℓ\(m\)\)​\(𝒮d,ℓ​yi\(m\)−𝒮d,ℓ​y¯d,ℓ\(m\)\)∑i∈Ωd,ℓ\(yi\(m\)−y¯d,ℓ\(m\)\)2​∑i∈Ωd,ℓ\(𝒮d,ℓ​yi\(m\)−𝒮d,ℓ​y¯d,ℓ\(m\)\)2,\\rho^\{\(m\)\}\_\{\\mathrm\{obs\}\}\(d,\\ell\)=\\frac\{\\sum\_\{i\\in\\Omega\_\{d,\\ell\}\}\\left\(y^\{\(m\)\}\_\{i\}\-\\bar\{y\}^\{\(m\)\}\_\{d,\\ell\}\\right\)\\left\(\\mathcal\{S\}\_\{d,\\ell\}y^\{\(m\)\}\_\{i\}\-\\overline\{\\mathcal\{S\}\_\{d,\\ell\}y\}^\{\(m\)\}\_\{d,\\ell\}\\right\)\}\{\\sqrt\{\\sum\_\{i\\in\\Omega\_\{d,\\ell\}\}\\left\(y^\{\(m\)\}\_\{i\}\-\\bar\{y\}^\{\(m\)\}\_\{d,\\ell\}\\right\)^\{2\}\}\\sqrt\{\\sum\_\{i\\in\\Omega\_\{d,\\ell\}\}\\left\(\\mathcal\{S\}\_\{d,\\ell\}y^\{\(m\)\}\_\{i\}\-\\overline\{\\mathcal\{S\}\_\{d,\\ell\}y\}^\{\(m\)\}\_\{d,\\ell\}\\right\)^\{2\}\}\},\(56\)
wherey¯d,ℓ\(m\)\\bar\{y\}^\{\(m\)\}\_\{d,\\ell\}and𝒮d,ℓ​y¯d,ℓ\(m\)\\overline\{\\mathcal\{S\}\_\{d,\\ell\}y\}^\{\(m\)\}\_\{d,\\ell\}are the spatial means of the original and shifted fields over the overlapping domainΩd,ℓ\\Omega\_\{d,\\ell\}\. The same calculation is applied to each forecast ensemble member\. For ensemble memberee, the forecast spatial autocorrelation isρforecast\(m,e\)​\(d,ℓ\)\\rho^\{\(m,e\)\}\_\{\\mathrm\{forecast\}\}\(d,\\ell\)

We compute this quantity for multiple lagsℓ\\elland for the four directions shown in Fig\.[7](https://arxiv.org/html/2608.26594#S2.F7)\. The resulting distributions are summarized with boxplots across test samples and ensemble members\.

A realistic precipitation generator should reproduce the observed decay of spatial autocorrelation with increasing lag\. In other words, the forecast autocorrelationρforecast\(m,e\)​\(d,ℓ\)\\rho^\{\(m,e\)\}\_\{\\mathrm\{forecast\}\}\(d,\\ell\)should follow the same lag\-dependent structure as the observed autocorrelationρobs\(m\)​\(d,ℓ\)\\rho^\{\(m\)\}\_\{\\mathrm\{obs\}\}\(d,\\ell\)across directionsddand spatial lagsℓ\\ell\. Ifρforecast\(m,e\)​\(d,ℓ\)\\rho^\{\(m,e\)\}\_\{\\mathrm\{forecast\}\}\(d,\\ell\)remains too high relative toρobs\(m\)​\(d,ℓ\)\\rho^\{\(m\)\}\_\{\\mathrm\{obs\}\}\(d,\\ell\)asℓ\\ellincreases, the generated fields are overly smooth\. Ifρforecast\(m,e\)​\(d,ℓ\)\\rho^\{\(m,e\)\}\_\{\\mathrm\{forecast\}\}\(d,\\ell\)decays too rapidly, the generated fields are too noisy\.

## References

- \[1\]\(2016\)Next generation earth system prediction: strategies for subseasonal to seasonal forecasts\.The National Academies Press,Washington, DC\.External Links:[Document](https://dx.doi.org/10.17226/21873)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[2\]F\. Vitart, C\. Ardilouze, A\. Bonet, A\. Brookshaw, M\. Chen, C\. Codorean, M\. Déqué, L\. Ferranti, E\. Fucile, M\. Fuentes, H\. Hendon, J\. Hodgson, H\.\-S\. Kang, A\. Kumar, H\. Lin, G\. Liu, X\. Liu, P\. Malguzzi, I\. Mallas, M\. Manoussakis, D\. Mastrangelo, C\. MacLachlan, P\. McLean, A\. Minami, R\. Mladek, T\. Nakazawa, S\. Najm, Y\. Nie, M\. Rixen, A\. W\. Robertson, P\. Ruti, C\. Sun, Y\. Takaya, M\. Tlostykh, F\. Venuti, D\. Waliser, S\. Woolnough, T\. Wu, D\.\-J\. Won, H\. Xiao, R\. Zaripov, and L\. Zhang\(2017\)The subseasonal to seasonal \(s2s\) prediction project database\.Bulletin of the American Meteorological Society98\(1\),pp\. 163–173\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-16-0017.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1),[§1](https://arxiv.org/html/2608.26594#S1.p2.1),[§4\.6](https://arxiv.org/html/2608.26594#S4.SS6.p1.1)\.
- \[3\]K\. Pegion, B\. P\. Kirtman, E\. Becker, D\. C\. Collins, E\. LaJoie, R\. Burgman, R\. Bell, T\. DelSole, D\. Min, Y\. Zhu, W\. Li, E\. Sinsky, H\. Guan, J\. Gottschalck, E\. J\. Metzger, N\. P\. Barton, D\. Achuthavarier, J\. Marshak, R\. Koster, H\. Lin, N\. Gagnon, M\. Bell, M\. K\. Tippett, A\. W\. Robertson, S\. Sun, S\. G\. Benjamin, B\. W\. Green, R\. Bleck, H\. Kim, J\. Wang, S\. Henderson, and A\. M\. Tompkins\(2019\)The subseasonal experiment \(SubX\): a multimodel subseasonal prediction experiment\.Bulletin of the American Meteorological Society100\(10\),pp\. 2043–2060\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-18-0270.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[4\]A\. W\. Robertson, F\. Vitart, and S\. J\. Camargo\(2020\)Subseasonal to seasonal prediction of weather to climate with application to tropical meteorology\.Journal of Geophysical Research: Atmospheres125\(7\),pp\. e2018JD029375\.External Links:[Document](https://dx.doi.org/10.1029/2018JD029375)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1),[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[5\]C\. J\. White and Coauthors\(2022\)Advances in the application and utility of subseasonal\-to\-seasonal predictions\.Bulletin of the American Meteorological Society103\(6\),pp\. E1448–E1472\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-20-0224.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[6\]E\. N\. Lorenz\(1969\)The predictability of a flow which possesses many scales of motion\.Tellus21\(3\),pp\. 289–307\.External Links:[Document](https://dx.doi.org/10.3402/tellusa.v21i3.10086)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1),[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[7\]P\. Bauer, A\. Thorpe, and G\. Brunet\(2015\)The quiet revolution of numerical weather prediction\.Nature525,pp\. 47–55\.External Links:[Document](https://dx.doi.org/10.1038/nature14956)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1),[§1](https://arxiv.org/html/2608.26594#S1.p2.1),[§4\.6](https://arxiv.org/html/2608.26594#S4.SS6.p1.1)\.
- \[8\]M\. Leutbecher and T\. N\. Palmer\(2008\)Ensemble forecasting\.Journal of Computational Physics227\(7\),pp\. 3515–3539\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2007.02.014)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1),[§1](https://arxiv.org/html/2608.26594#S1.p2.1),[§1](https://arxiv.org/html/2608.26594#S1.p4.1),[§4\.6](https://arxiv.org/html/2608.26594#S4.SS6.p1.1)\.
- \[9\]T\. N\. Palmer and D\. L\. T\. Anderson\(1994\)The prospects for seasonal forecasting—a review paper\.Quarterly Journal of the Royal Meteorological Society120\(518\),pp\. 755–793\.External Links:[Document](https://dx.doi.org/10.1002/qj.49712051802)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[10\]R\. D\. Koster, S\. P\. P\. Mahanama, T\. J\. Yamada, G\. Balsamo, A\. A\. Berg, M\. Boisserie, P\. A\. Dirmeyer, F\. J\. Doblas\-Reyes, G\. Drewitt, C\. T\. Gordon, Z\. Guo, J\. Jeong, D\. M\. Lawrence, W\. Lee, Z\. Li, L\. Luo, S\. Malyshev, W\. J\. Merryfield, S\. I\. Seneviratne, T\. Stanelle, B\. J\. J\. M\. van den Hurk, F\. Vitart, and E\. F\. Wood\(2010\)Contribution of land surface initialization to subseasonal forecast skill: first results from a multi\-model experiment\.Geophysical Research Letters37\(2\),pp\. L02402\.External Links:[Document](https://dx.doi.org/10.1029/2009GL041677)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[11\]A\. Mariotti and Coauthors\(2020\)Windows of opportunity for skillful forecasts subseasonal to seasonal and beyond\.Bulletin of the American Meteorological Society101\(5\),pp\. E608–E625\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-18-0326.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[12\]W\. J\. Merryfield and Coauthors\(2020\)Current and emerging developments in subseasonal to decadal prediction\.Bulletin of the American Meteorological Society101\(6\),pp\. E869–E896\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-19-0037.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[13\]C\. Zhang\(2005\)Madden\-julian oscillation\.Reviews of Geophysics43\(2\),pp\. RG2003\.External Links:[Document](https://dx.doi.org/10.1029/2004RG000158)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[14\]M\. P\. Baldwin and T\. J\. Dunkerton\(2001\)Stratospheric harbingers of anomalous weather regimes\.Science294\(5542\),pp\. 581–584\.External Links:[Document](https://dx.doi.org/10.1126/science.1063315)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[15\]J\. H\. Richter, A\. A\. Glanville, T\. King, S\. Kumar, S\. G\. Yeager, N\. A\. Davis, Y\. Duan, M\. D\. Fowler, A\. Jaye, J\. Edwards, J\. M\. Caron, P\. A\. Dirmeyer, G\. Danabasoglu, and K\. Oleson\(2024\)Quantifying sources of subseasonal prediction skill in CESM2\.npj Climate and Atmospheric Science7\(59\)\.External Links:[Document](https://dx.doi.org/10.1038/s41612-024-00595-4)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p1.1)\.
- \[16\]F\. Vitart and A\. W\. Robertson\(2018\)The sub\-seasonal to seasonal prediction project \(s2s\) and the prediction of extreme events\.npj Climate and Atmospheric Science1\(1\),pp\. 3\.External Links:[Document](https://dx.doi.org/10.1038/s41612-018-0013-0)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[17\]Z\. Toth and E\. Kalnay\(1993\)Ensemble forecasting at NMC: the generation of perturbations\.Bulletin of the American Meteorological Society74\(12\),pp\. 2317–2330\.External Links:[Document](https://dx.doi.org/10.1175/1520-0477%281993%29074%3C2317%3AEFANTG%3E2.0.CO%3B2)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[18\]F\. Molteni, R\. Buizza, T\. N\. Palmer, and T\. Petroliagis\(1996\)The ECMWF ensemble prediction system: methodology and validation\.Quarterly Journal of the Royal Meteorological Society122\(529\),pp\. 73–119\.External Links:[Document](https://dx.doi.org/10.1002/qj.49712252905)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[19\]R\. Buizza, M\. Miller, and T\. N\. Palmer\(1999\)Stochastic representation of model uncertainties in the ecmwf ensemble prediction system\.Quarterly Journal of the Royal Meteorological Society125\(560\),pp\. 2887–2908\.External Links:[Document](https://dx.doi.org/10.1002/qj.49712556006)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1),[§1](https://arxiv.org/html/2608.26594#S1.p4.1)\.
- \[20\]T\. Gneiting and A\. E\. Raftery\(2005\)Weather forecasting with ensemble methods\.Science310\(5746\),pp\. 248–249\.External Links:[Document](https://dx.doi.org/10.1126/science.1115255)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[21\]P\. Bougeault, Z\. Toth, C\. Bishop, B\. Brown, D\. Burridge, D\. H\. Chen, B\. Ebert, M\. Fuentes, T\. M\. Hamill, K\. Mylne, J\. Nicolau, T\. Paccagnella, Y\. Park, D\. Parsons, B\. Raoult, D\. Schuster, P\. S\. Dias, R\. Swinbank, Y\. Takeuchi, W\. Tennant, L\. Wilson, and S\. Worley\(2010\)The THORPEX interactive grand global ensemble\.Bulletin of the American Meteorological Society91\(8\),pp\. 1059–1072\.External Links:[Document](https://dx.doi.org/10.1175/2010BAMS2853.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[22\]R\. Swinbank, M\. Kyouda, P\. Buchanan, L\. Froude, T\. M\. Hamill, T\. D\. Hewson, J\. H\. Keller, M\. Matsueda, J\. Methven, F\. Pappenberger, M\. Scheuerer, H\. A\. Titley, L\. Wilson, and M\. Yamaguchi\(2016\)The TIGGE project and its achievements\.Bulletin of the American Meteorological Society97\(1\),pp\. 49–67\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-13-00191.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[23\]T\. N\. Palmer\(2019\)The ECMWF ensemble prediction system: looking back \(more than\) 25 years and projecting forward 25 years\.Quarterly Journal of the Royal Meteorological Society145\(S1\),pp\. 12–24\.External Links:[Document](https://dx.doi.org/10.1002/qj.3383)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[24\]T\. N\. Palmer\(2000\)Predicting uncertainty in forecasts of weather and climate\.Reports on Progress in Physics63\(2\),pp\. 71–116\.External Links:[Document](https://dx.doi.org/10.1088/0034-4885/63/2/201)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[25\]E\. N\. Lorenz\(1963\)Deterministic nonperiodic flow\.Journal of the Atmospheric Sciences20\(2\),pp\. 130–141\.External Links:[Document](https://dx.doi.org/10.1175/1520-0469%281963%29020%3C0130%3ADNF%3E2.0.CO%3B2)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[26\]C\. E\. Leith\(1974\)Theoretical skill of monte carlo forecasts\.Monthly Weather Review102\(6\),pp\. 409–418\.External Links:[Document](https://dx.doi.org/10.1175/1520-0493%281974%29102%3C0409%3ATSOMCF%3E2.0.CO%3B2)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[27\]J\. M\. Fritsch and R\. E\. Carbone\(2004\)Improving quantitative precipitation forecasts in the warm season: a uswrp research and development strategy\.Bulletin of the American Meteorological Society85\(7\),pp\. 955–965\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-85-7-955)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[28\]E\. E\. Ebert and J\. L\. McBride\(2000\)Verification of precipitation in weather systems: determination of systematic errors\.Journal of Hydrology239\(1–4\),pp\. 179–202\.External Links:[Document](https://dx.doi.org/10.1016/S0022-1694%2800%2900343-7)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[29\]A\. Arakawa\(2004\)The cumulus parameterization problem: past, present, and future\.Journal of Climate17\(13\),pp\. 2493–2525\.External Links:[Document](https://dx.doi.org/10.1175/1520-0442%282004%29017%3C2493%3ARATCPP%3E2.0.CO%3B2)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[30\]S\. Vannitsem, D\. S\. Wilks, and J\. W\. Messner\(2021\)Statistical postprocessing of ensemble forecasts\.Elsevier\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[31\]O\. Watt\-Meyer, N\. D\. Brenowitz, S\. K\. Clark, B\. Henn, A\. Kwa, J\. McGibbon, W\. A\. Perkins, and C\. S\. Bretherton\(2021\)Correcting weather and climate models by machine learning nudged historical simulations\.Geophysical Research Letters48\(15\),pp\. e2021GL092555\.External Links:[Document](https://dx.doi.org/10.1029/2021GL092555)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[32\]Z\. Ben\-Bouallègue, M\. C\. A\. Clare, L\. Magnusson, E\. Gascon, M\. Maier\-Gerber, M\. Janousek, M\. Rodwell, F\. Pinault, J\. S\. Dramsch, S\. T\. K\. Lang, B\. Raoult, F\. Rabier, M\. Chevallier, I\. Sandu, P\. Dueben, M\. Chantry, and F\. Pappenberger\(2024\)The rise of data\-driven weather forecasting\.Bulletin of the American Meteorological Society105\(6\),pp\. E1057–E1078\.External Links:[Document](https://dx.doi.org/10.1175/BAMS-D-23-0162.1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p2.1)\.
- \[33\]S\. Rasp, P\. D\. Dueben, S\. Scher, J\. A\. Weyn, S\. Mouatadid, and N\. Thuerey\(2020\)WeatherBench: a benchmark dataset for data\-driven weather forecasting\.Journal of Advances in Modeling Earth Systems12\(11\),pp\. e2020MS002203\.External Links:[Document](https://dx.doi.org/10.1029/2020MS002203)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1),[§1](https://arxiv.org/html/2608.26594#S1.p5.1)\.
- \[34\]J\. Hwang, P\. Orenstein, J\. Cohen, K\. Pfeiffer, and L\. Mackey\(2019\)Improving subseasonal forecasting in the western u\.s\. with machine learning\.arXiv preprint arXiv:1809\.07394\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1809.07394)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1)\.
- \[35\]H\. V\. Dang and P\. C\. H\. Nguyen\(2026\)Deep operator learning for high\-fidelity fluid flow field reconstruction from sparse sensor measurements\.Journal of Computing and Information Science in Engineering26\(1\),pp\. 011007\.External Links:[Document](https://dx.doi.org/10.1115/1.4070332)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1)\.
- \[36\]L\. Chen, X\. Zhong, H\. Li, J\. Wu, B\. Lu, D\. Chen, S\.\-P\. Xie, L\. Wu, Q\. Chao, C\. Lin, Z\. Hu, and Y\. Qi\(2024\)A machine learning model that outperforms conventional global subseasonal forecast models\.Nature Communications15\(1\),pp\. 6425\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-50714-1)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1),[§1](https://arxiv.org/html/2608.26594#S1.p5.1)\.
- \[37\]J\. A\. Weyn, D\. R\. Durran, and R\. Caruana\(2021\)Data\-driven medium\-range weather prediction with a resnet pretrained on climate simulations: a new model for weatherbench\.Journal of Advances in Modeling Earth Systems13\(2\),pp\. e2020MS002405\.External Links:[Document](https://dx.doi.org/10.1029/2020MS002405)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1)\.
- \[38\]T\. Nguyen, J\. Brandstetter, A\. Kapoor, J\. K\. Gupta, and A\. Grover\(2023\)ClimaX: a foundation model for weather and climate\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 25904–25938\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1)\.
- \[39\]C\. Bodnar, W\. P\. Bruinsma, A\. Lucic, M\. Stanley, J\. Brandstetter, P\. Garvan, M\. Riechert, J\. Weyn, H\. Dong, A\. Vaughan, J\. K\. Gupta, K\. Tambiratnam, A\. Archibald, E\. Heider, M\. Welling, R\. E\. Turner, and P\. Perdikaris\(2025\)A foundation model for the earth system\.Nature641,pp\. 1180–1187\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09005-y)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p3.1)\.
- \[40\]L\. Li, R\. Carver, I\. Lopez\-Gomez, F\. Sha, and J\. Anderson\(2024\)Generative emulation of weather forecast ensembles with diffusion models\.Science Advances10\(13\),pp\. eadk4489\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adk4489)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p4.1)\.
- \[41\]I\. Price, A\. Sanchez\-Gonzalez, F\. Alet, T\. R\. Andersson, A\. El\-Kadi, D\. Masters, T\. Ewalds, J\. Stott, S\. Mohamed, P\. Battaglia, R\. Lam, and M\. Willson\(2025\)Probabilistic weather forecasting with machine learning\.Nature637,pp\. 84–90\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08252-9)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p4.1)\.
- \[42\]P\. D\. Dueben and P\. Bauer\(2018\)Challenges and design choices for global weather and climate models based on machine learning\.Geoscientific Model Development11\(10\),pp\. 3999–4009\.External Links:[Document](https://dx.doi.org/10.5194/gmd-11-3999-2018)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p5.1)\.
- \[43\]I\. Goodfellow, J\. Pouget\-Abadie, M\. Mirza, B\. Xu, D\. Warde\-Farley, S\. Ozair, A\. Courville, and Y\. Bengio\(2014\)Generative adversarial nets\.InAdvances in Neural Information Processing Systems,Vol\.27\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p5.1)\.
- \[44\]T\. Salimans, I\. Goodfellow, W\. Zaremba, V\. Cheung, A\. Radford, and X\. Chen\(2016\)Improved techniques for training gans\.InAdvances in Neural Information Processing Systems,Vol\.29\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p5.1)\.
- \[45\]L\. Mescheder, A\. Geiger, and S\. Nowozin\(2018\)Which training methods for gans do actually converge?\.InProceedings of the 35th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.80,pp\. 3481–3490\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p5.1)\.
- \[46\]D\. P\. Kingma and M\. Welling\(2014\)Auto\-encoding variational bayes\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p5.1),[§1](https://arxiv.org/html/2608.26594#S1.p6.1)\.
- \[47\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 6840–6851\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p6.1)\.
- \[48\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.International Conference on Learning Representations\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p6.1)\.
- \[49\]R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer\(2022\)High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 10684–10695\.Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p6.1)\.
- \[50\]K\. B\. Rodgers, S\. Lee, N\. Rosenbloom, A\. Timmermann, G\. Danabasoglu, C\. Deser, J\. Edwards, J\. Kim, I\. R\. Simpson, K\. Stein, M\. F\. Stuecker, R\. Yamaguchi, T\. Bódai, E\. Chung, L\. Huang, W\. M\. Kim, J\. Lamarque, D\. L\. Lombardozzi, W\. R\. Wieder, and S\. G\. Yeager\(2021\)Ubiquity of human\-induced changes in climate variability\.Earth System Dynamics12\(4\),pp\. 1393–1411\.External Links:[Document](https://dx.doi.org/10.5194/esd-12-1393-2021)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p7.1),[§4\.1](https://arxiv.org/html/2608.26594#S4.SS1.p1.1)\.
- \[51\]H\. Hersbach, B\. Bell, P\. Berrisford, S\. Hirahara, A\. Horányi, J\. Muñoz\-Sabater, J\. Nicolas, C\. Peubey, R\. Radu, D\. Schepers, A\. Simmons, C\. Soci, S\. Abdalla, X\. Abellan, G\. Balsamo, P\. Bechtold, G\. Biavati, J\. Bidlot, M\. Bonavita, G\. De Chiara, P\. Dahlgren, D\. Dee, M\. Diamantakis, R\. Dragani, J\. Flemming, R\. Forbes, M\. Fuentes, A\. Geer, L\. Haimberger, S\. Healy, R\. J\. Hogan, E\. Hólm, M\. Janisková, S\. Keeley, P\. Laloyaux, P\. Lopez, C\. Lupu, G\. Radnoti, P\. de Rosnay, I\. Rozum, F\. Vamborg, S\. Villaume, and J\.\-N\. Thépaut\(2020\)The era5 global reanalysis\.Quarterly Journal of the Royal Meteorological Society146\(730\),pp\. 1999–2049\.External Links:[Document](https://dx.doi.org/10.1002/qj.3803)Cited by:[§1](https://arxiv.org/html/2608.26594#S1.p7.1),[§4\.1](https://arxiv.org/html/2608.26594#S4.SS1.p1.1)\.
- \[52\]Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner\(1998\)Gradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.External Links:[Document](https://dx.doi.org/10.1109/5.726791)Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p1.1)\.
- \[53\]O\. Ronneberger, P\. Fischer, and T\. Brox\(2015\)U\-net: convolutional networks for biomedical image segmentation\.InMedical Image Computing and Computer\-Assisted Intervention – MICCAI 2015,Lecture Notes in Computer Science, Vol\.9351,pp\. 234–241\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-24574-4%5F28)Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p1.1),[§4\.5](https://arxiv.org/html/2608.26594#S4.SS5.p1.1)\.
- \[54\]S\. Rasp and N\. Thuerey\(2021\)Data\-driven medium\-range weather prediction with a resnet pretrained on climate simulations: a new model for weatherbench\.Journal of Advances in Modeling Earth Systems13\(2\),pp\. e2020MS002405\.External Links:[Document](https://dx.doi.org/10.1029/2020MS002405)Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p2.1)\.
- \[55\]Y\. Ham, J\. Kim, and J\. Luo\(2019\)Deep learning for multi\-year ENSO forecasts\.Nature573\(7775\),pp\. 568–572\.External Links:[Document](https://dx.doi.org/10.1038/s41586-019-1559-7)Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p2.1)\.
- \[56\]J\. González\-Abad, M\. Iturbide, A\. Hernanz, and J\. M\. Gutiérrez\(2026\)Pre\-training for deep statistical climate downscaling: enhancing consistency and robustness across regional datasets\.Geoscientific Model Development19,pp\. 5781–5804\.External Links:[Document](https://dx.doi.org/10.5194/gmd-19-5781-2026)Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p2.1)\.
- \[57\]A\. Unal, B\. Asan, I\. Sezen, B\. Yesilkaynak, Y\. Aydin, M\. Ilicak, and G\. Unal\(2023\)Climate model\-driven seasonal forecasting approach with deep learning\.Environmental Data Science2,pp\. e29\.External Links:[Document](https://dx.doi.org/10.1017/eds.2023.24)Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p2.1)\.
- \[58\]S\. A\. Martin, N\. Brenowitz, D\. Durran, and M\. Pritchard\(2025\)Long\-range distillation: distilling 10,000 years of simulated climate into long timestep AI weather models\.arXiv preprint arXiv:2512\.22814\.External Links:2512\.22814Cited by:[§2\.1](https://arxiv.org/html/2608.26594#S2.SS1.p2.1)\.
- \[59\]S\. Li, Y\. Zhao, R\. Varma, O\. Salpekar, P\. Noordhuis, T\. Li, A\. Paszke, J\. Smith, B\. Vaughan, P\. Damania, and S\. Chintala\(2020\)PyTorch Distributed: experiences on accelerating data parallel training\.Proceedings of the VLDB Endowment13\(12\),pp\. 3005–3018\.External Links:[Document](https://dx.doi.org/10.14778/3415478.3415530)Cited by:[§2\.5](https://arxiv.org/html/2608.26594#S2.SS5.p3.1)\.
- \[60\]G\. M\. Amdahl\(1967\)Validity of the single processor approach to achieving large scale computing capabilities\.InProceedings of the April 18–20, 1967, Spring Joint Computer Conference,AFIPS ’67 \(Spring\),New York, NY, USA,pp\. 483–485\.External Links:[Document](https://dx.doi.org/10.1145/1465482.1465560)Cited by:[§2\.5](https://arxiv.org/html/2608.26594#S2.SS5.p3.1)\.
- \[61\]J\. L\. Gustafson\(1988\)Reevaluating Amdahl’s Law\.Communications of the ACM31\(5\),pp\. 532–533\.External Links:[Document](https://dx.doi.org/10.1145/42411.42415)Cited by:[§2\.5](https://arxiv.org/html/2608.26594#S2.SS5.p3.1)\.
- \[62\]European Centre for Medium\-Range Weather ForecastsAtmospheric model sub\-seasonal forecast \(set VI – sub\-seasonal\)\.Note:[https://www\.ecmwf\.int/en/forecasts/datasets/set\-vi](https://www.ecmwf.int/en/forecasts/datasets/set-vi)Accessed: 28 July 2026Cited by:[§2\.5](https://arxiv.org/html/2608.26594#S2.SS5.p4.1)\.
- \[63\]P\. Bauer, T\. Quintino, N\. Wedi, A\. Bonanni, M\. Chrust, W\. Deconinck, M\. Diamantakis, P\. Düben, S\. English, J\. Flemming, P\. Gillies, I\. Hadade, J\. Hawkes, M\. Hawkins, O\. Iffrig, C\. Kühnlein, M\. Lange, P\. Lean, O\. Marsden, A\. Müller, S\. Saarinen, D\. Sarmany, M\. Sleigh, S\. Smart, P\. Smolarkiewicz, D\. Thiemert, G\. Tumolo, C\. Weihrauch, C\. Zanna, and P\. Maciel\(2020\)The ECMWF scalability programme: progress and plans\.ECMWF Technical MemorandumTechnical Report857,European Centre for Medium\-Range Weather Forecasts\.External Links:[Document](https://dx.doi.org/10.21957/gdit22ulm),[Link](https://www.ecmwf.int/en/elibrary/81155-ecmwf-scalability-programme-progress-and-plans)Cited by:[§2\.5](https://arxiv.org/html/2608.26594#S2.SS5.p4.1),[§3](https://arxiv.org/html/2608.26594#S3.p5.1)\.
- \[64\]M\. Mathieu, C\. Couprie, and Y\. LeCun\(2016\)Deep multi\-scale video prediction beyond mean square error\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/1511.05440)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p4.1)\.
- \[65\]S\. Ravuri, K\. Lenc, M\. Willson, D\. Kangin, R\. Lam, P\. Mirowski, M\. Fitzsimons, M\. Athanassiadou, S\. Kashem, S\. Madge, R\. Prudden, A\. Mandhane, A\. Clark, A\. Brock, K\. Simonyan, R\. Hadsell, N\. Robinson, E\. Clancy, A\. Arribas, and S\. Mohamed\(2021\)Skilful precipitation nowcasting using deep generative models of radar\.Nature597,pp\. 672–677\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-03854-z)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p4.1)\.
- \[66\]L\. Harris, A\. T\. T\. McRae, M\. Chantry, P\. D\. Dueben, and T\. N\. Palmer\(2022\)A generative deep learning approach to stochastic downscaling of precipitation forecasts\.Journal of Advances in Modeling Earth Systems14\(10\),pp\. e2022MS003120\.External Links:[Document](https://dx.doi.org/10.1029/2022MS003120)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p4.1)\.
- \[67\]T\. Gneiting and M\. Katzfuss\(2014\)Probabilistic forecasting\.Annual Review of Statistics and Its Application1,pp\. 125–151\.External Links:[Document](https://dx.doi.org/10.1146/annurev-statistics-062713-085831)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p5.1)\.
- \[68\]T\. Kurth, S\. Subramanian, P\. Harrington, J\. Pathak, M\. Mardani, D\. Hall, A\. Miele, K\. Kashinath, and A\. Anandkumar\(2023\)FourCastNet: accelerating global high\-resolution weather forecasting using adaptive fourier neural operators\.InProceedings of the Platform for Advanced Scientific Computing Conference,External Links:[Document](https://dx.doi.org/10.1145/3592979.3593412)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p5.1)\.
- \[69\]I\. Price, A\. Sanchez\-Gonzalez, F\. Alet, T\. Ewalds, A\. El\-Kadi, J\. Stott, S\. Mohamed, P\. Battaglia, R\. Lam, M\. Willson,et al\.\(2025\)Probabilistic weather forecasting with machine learning\.Nature637,pp\. 84–90\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08252-9)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p5.1)\.
- \[70\]A\. G\. Baydin, B\. A\. Pearlmutter, A\. A\. Radul, and J\. M\. Siskind\(2018\)Automatic differentiation in machine learning: a survey\.Journal of Machine Learning Research18\(153\),pp\. 1–43\.External Links:[Link](https://jmlr.org/papers/v18/17-468.html)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p6.1)\.
- \[71\]M\. Raissi, P\. Perdikaris, and G\. E\. Karniadakis\(2019\)Physics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational Physics378,pp\. 686–707\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2018.10.045)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p6.1)\.
- \[72\]G\. E\. Karniadakis, I\. G\. Kevrekidis, L\. Lu, P\. Perdikaris, S\. Wang, and L\. Yang\(2021\)Physics\-informed machine learning\.Nature Reviews Physics3,pp\. 422–440\.External Links:[Document](https://dx.doi.org/10.1038/s42254-021-00314-5)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p6.1)\.
- \[73\]P\. C\. Nguyen, Y\. Nguyen, J\. B\. Choi, P\. K\. Seshadri, H\. Udaykumar, and S\. S\. Baek\(2023\)PARC: physics\-aware recurrent convolutional neural networks to assimilate meso scale reactive mechanics of energetic materials\.Science advances9\(17\),pp\. eadd6868\.Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p6.1)\.
- \[74\]P\. C\. Nguyen, X\. Cheng, S\. Azarfar, P\. Seshadri, Y\. T\. Nguyen, M\. Kim, S\. Choi, H\. Udaykumar, and S\. Baek\(2024\)PARCv2: physics\-aware recurrent convolutional neural networks for spatiotemporal dynamics modeling\.arXiv preprint arXiv:2402\.12503\.Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p6.1)\.
- \[75\]A\. Mamalakis, I\. Ebert\-Uphoff, and E\. A\. Barnes\(2022\)Explainable artificial intelligence in meteorology and climate science: model fine\-tuning, calibrating trust and learning new science\.InxxAI – Beyond Explainable AI,A\. Holzinger, R\. Goebel, R\. Fong, T\. Moon, K\. Müller, and W\. Samek \(Eds\.\),Lecture Notes in Computer Science, Vol\.13200,pp\. 315–339\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-04083-2%5F16)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[76\]A\. Mamalakis, E\. A\. Barnes, and I\. Ebert\-Uphoff\(2022\)Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience\.Artificial Intelligence for the Earth Systems1\(4\)\.External Links:[Document](https://dx.doi.org/10.1175/AIES-D-22-0012.1)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[77\]T\. B\. Higgins, A\. C\. Subramanian, A\. Graubner, L\. Kapp\-Schwoerer, P\. A\. G\. Watson, S\. Sparrow, K\. Kashinath, S\. Kim, L\. Delle Monache, and W\. Chapman\(2023\)Using deep learning for an analysis of atmospheric rivers in a high\-resolution large ensemble climate data set\.Journal of Advances in Modeling Earth Systems15\(5\),pp\. e2022MS003495\.External Links:[Document](https://dx.doi.org/10.1029/2022MS003495)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[78\]T\. B\. Higgins, A\. C\. Subramanian, W\. E\. Chapman, D\. A\. Lavers, and A\. C\. Winters\(2024\)Subseasonal potential predictability of horizontal water vapor transport and precipitation extremes in the north pacific\.Weather and Forecasting39\(6\),pp\. 833–846\.External Links:[Document](https://dx.doi.org/10.1175/WAF-D-23-0170.1)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[79\]S\. Singh, M\. K\. Goyal, and S\. Jha\(2023\)Role of large\-scale climate oscillations in precipitation extremes associated with atmospheric rivers: nonstationary framework\.Hydrological Sciences Journal68\(3\),pp\. 395–411\.External Links:[Document](https://dx.doi.org/10.1080/02626667.2022.2159412)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[80\]S\. Singh and M\. K\. Goyal\(2023\)An innovative approach to predict atmospheric rivers: exploring convolutional autoencoder\.Atmospheric Research,pp\. 106754\.External Links:[Document](https://dx.doi.org/10.1016/j.atmosres.2023.106754)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[81\]M\. K\. Goyal and S\. Singh\(2024\)Understanding atmospheric rivers using machine learning\.SpringerBriefs in Applied Sciences and Technology,Springer\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-63478-9)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[82\]J\. M\. Corner, A\. Mamalakis, and K\. Schiro\(2026\)Artificial intelligence identifies shortwave troughs as an important local feature of extreme precipitation\.Note:ESS Open Archive preprintExternal Links:[Document](https://dx.doi.org/10.22541/essoar.15001928/v1)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[83\]M\. T\. Ribeiro, S\. Singh, and C\. Guestrin\(2016\)"Why should i trust you?": explaining the predictions of any classifier\.InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,pp\. 1135–1144\.External Links:[Document](https://dx.doi.org/10.1145/2939672.2939778)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[84\]M\. Sundararajan, A\. Taly, and Q\. Yan\(2017\)Axiomatic attribution for deep networks\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 3319–3328\.Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[85\]S\. Wachter, B\. Mittelstadt, and C\. Russell\(2017\)Counterfactual explanations without opening the black box: automated decisions and the gdpr\.Harvard Journal of Law & Technology31\(2\),pp\. 841–887\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.3063289)Cited by:[§3](https://arxiv.org/html/2608.26594#S3.p7.1)\.
- \[86\]G\. Danabasoglu, J\. Lamarque, J\. Bacmeister, D\. A\. Bailey, A\. K\. DuVivier, J\. Edwards, L\. K\. Emmons, J\. Fasullo, R\. Garcia, A\. Gettelman, C\. Hannay, M\. M\. Holland, W\. G\. Large, P\. H\. Lauritzen, D\. M\. Lawrence, J\. T\. M\. Lenaerts, K\. Lindsay, W\. H\. Lipscomb, M\. J\. Mills, R\. Neale, K\. W\. Oleson, B\. Otto\-Bliesner, A\. S\. Phillips, W\. Sacks, S\. Tilmes, L\. van Kampenhout, M\. Vertenstein, A\. Bertini, J\. Dennis, C\. Deser, C\. Fischer, B\. Fox\-Kemper, J\. E\. Kay, D\. Kinnison, P\. J\. Kushner, V\. E\. Larson, M\. C\. Long, S\. Mickelson, J\. K\. Moore, E\. Nienhouse, L\. Polvani, P\. J\. Rasch, and W\. G\. Strand\(2020\)The community earth system model version 2 \(CESM2\)\.Journal of Advances in Modeling Earth Systems12\(2\),pp\. e2019MS001916\.External Links:[Document](https://dx.doi.org/10.1029/2019MS001916)Cited by:[§4\.1](https://arxiv.org/html/2608.26594#S4.SS1.p1.1)\.
- \[87\]D\. P\. Kingma and M\. Welling\(2014\)Auto\-encoding variational bayes\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/1312.6114)Cited by:[§4\.2](https://arxiv.org/html/2608.26594#S4.SS2.p2.1)\.
- \[88\]D\. P\. Kingma and J\. Ba\(2015\)Adam: a method for stochastic optimization\.InInternational Conference on Learning Representations,Cited by:[§4\.2](https://arxiv.org/html/2608.26594#S4.SS2.p21.1)\.
- \[89\]Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner\(1998\)Gradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.External Links:[Document](https://dx.doi.org/10.1109/5.726791)Cited by:[§4\.2](https://arxiv.org/html/2608.26594#S4.SS2.p21.1),[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p3.1),[§4\.5](https://arxiv.org/html/2608.26594#S4.SS5.p1.1)\.
- \[90\]A\. Krizhevsky, I\. Sutskever, and G\. E\. Hinton\(2012\)ImageNet classification with deep convolutional neural networks\.InAdvances in Neural Information Processing Systems,Vol\.25\.Cited by:[§4\.2](https://arxiv.org/html/2608.26594#S4.SS2.p21.1),[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p3.1),[§4\.5](https://arxiv.org/html/2608.26594#S4.SS5.p1.1)\.
- \[91\]J\. Sohl\-Dickstein, E\. A\. Weiss, N\. Maheswaranathan, and S\. Ganguli\(2015\)Deep unsupervised learning using nonequilibrium thermodynamics\.InProceedings of the 32nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.37,pp\. 2256–2265\.Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p1.1)\.
- \[92\]Y\. Song and S\. Ermon\(2019\)Generative modeling by estimating gradients of the data distribution\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p1.1)\.
- \[93\]J\. Ho, A\. Jain, and P\. Abbeel\(2020\)Denoising diffusion probabilistic models\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 6840–6851\.Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p28.1)\.
- \[94\]Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. Poole\(2021\)Score\-based generative modeling through stochastic differential equations\.InInternational Conference on Learning Representations,Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p1.1)\.
- \[95\]R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer\(2022\)High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 10684–10695\.Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p1.1)\.
- \[96\]L\. Yang, Z\. Zhang, Y\. Song, S\. Hong, R\. Xu, Y\. Zhao, Y\. Shao, W\. Zhang, B\. Cui, and M\. Yang\(2023\)Diffusion models: a comprehensive survey of methods and applications\.ACM Computing Surveys56\(4\),pp\. 1–39\.External Links:[Document](https://dx.doi.org/10.1145/3626235)Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p1.1)\.
- \[97\]J\. Song, C\. Meng, and S\. Ermon\(2021\)Denoising diffusion implicit models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=St1giarCHLP)Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p38.1)\.
- \[98\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p68.1),[§4\.3](https://arxiv.org/html/2608.26594#S4.SS3.p69.1),[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p3.1)\.
- \[99\]S\. J\. Pan and Q\. Yang\(2010\)A survey on transfer learning\.IEEE Transactions on Knowledge and Data Engineering22\(10\),pp\. 1345–1359\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2009.191)Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[100\]J\. Yosinski, J\. Clune, Y\. Bengio, and H\. Lipson\(2014\)How transferable are features in deep neural networks?\.InAdvances in Neural Information Processing Systems,Vol\.27\.Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[101\]F\. Zhuang, Z\. Qi, K\. Duan, D\. Xi, Y\. Zhu, H\. Zhu, H\. Xiong, and Q\. He\(2021\)A comprehensive survey on transfer learning\.Proceedings of the IEEE109\(1\),pp\. 43–76\.External Links:[Document](https://dx.doi.org/10.1109/JPROC.2020.3004555)Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[102\]J\. Donahue, Y\. Jia, O\. Vinyals, J\. Hoffman, N\. Zhang, E\. Tzeng, and T\. Darrell\(2014\)DeCAF: a deep convolutional activation feature for generic visual recognition\.InProceedings of the 31st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.32,pp\. 647–655\.Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[103\]M\. Oquab, L\. Bottou, I\. Laptev, and J\. Sivic\(2014\)Learning and transferring mid\-level image representations using convolutional neural networks\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 1717–1724\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2014.222)Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[104\]S\. Rebuffi, H\. Bilen, and A\. Vedaldi\(2017\)Learning multiple visual domains with residual adapters\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[105\]N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. de Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly\(2019\)Parameter\-efficient transfer learning for NLP\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 2790–2799\.Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[106\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,Cited by:[§4\.4](https://arxiv.org/html/2608.26594#S4.SS4.p1.1)\.
- \[107\]F\. Vitart\(2022\)The next extended\-range configuration for IFS cycle 48r1\.ECMWF Newsletter173\.External Links:[Link](https://www.ecmwf.int/en/newsletter/173/earth-system-science/next-extended-range-configuration-ifs-cycle-48r1)Cited by:[§4\.6](https://arxiv.org/html/2608.26594#S4.SS6.p1.1)\.
- \[108\]V\. Fortin, M\. Abaza, F\. Anctil, and R\. Turcotte\(2014\)Why should ensemble spread match the RMSE of the ensemble mean?\.Journal of Hydrometeorology15\(4\),pp\. 1708–1713\.External Links:[Document](https://dx.doi.org/10.1175/JHM-D-14-0008.1)Cited by:[§4\.7\.3](https://arxiv.org/html/2608.26594#S4.SS7.SSS3.p12.1)\.

Similar Articles

Domain-Adaptive Climate Downscaling Under Temporal Distribution Shift

arXiv cs.LG

This paper investigates temporal out-of-distribution shift in deep-learning-based climate downscaling and proposes a domain-adaptive framework that combines supervised reconstruction with domain alignment to improve high-resolution climate projections under non-stationary conditions.