Transferable Dual-Stream Representations for Mesoscale-Preserving Sea Surface Temperature Downscaling

arXiv cs.LG Papers

Summary

This paper introduces EddyFlow, a deep learning framework for kilometer-scale sea surface temperature downscaling that balances predictive accuracy, scale-dependent structure, and regional generalization. It achieves strong zero-shot performance and near-ideal spectral fidelity across multiple ocean regions.

arXiv:2608.04230v1 Announce Type: new Abstract: Deep learning models for scientific spatio-temporal downscaling often minimize reconstruction error while failing to preserve physically meaningful multi-scale structure. For sea surface temperature prediction, this can yield outputs that are numerically plausible yet overly smooth, missing mesoscale variability critical to regional ocean dynamics. Existing methods often focus on pixel-wise objectives or single-context conditioning, which limits their ability to preserve spectral fidelity and generalize across regions. To address this, we propose EddyFlow, a representation learning framework for kilometer-scale sea surface temperature downscaling that balances predictive accuracy, scale-dependent structure, and regional generalization. EddyFlow is trained on the Gulf of St.~Lawrence and evaluated in zero-shot and few-shot settings on the Bay of Fundy and the Gulf of Mexico. EddyFlow demonstrates that physics-informed representation learning reduces zero-shot RMSE by 21%, achieves up to 85.6% skill relative to persistence on unseen domains, and maintains near-ideal spectral fidelity with a PSD ratio of $\approx 1.00$.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:47 AM

# Transferable Dual-Stream Representations for Mesoscale- Preserving Sea Surface Temperature Downscaling
Source: [https://arxiv.org/html/2608.04230](https://arxiv.org/html/2608.04230)
\(26 July 2026\)

###### Abstract\.

Deep learning models for scientific spatio\-temporal downscaling often minimize reconstruction error while failing to preserve physically meaningful multi\-scale structure\. For sea surface temperature prediction, this can yield outputs that are numerically plausible yet overly smooth, missing mesoscale variability critical to regional ocean dynamics\. Existing methods often focus on pixel\-wise objectives or single\-context conditioning, which limits their ability to preserve spectral fidelity and generalize across regions\. To address this, we proposeEddyFlow, a representation learning framework for kilometer\-scale sea surface temperature downscaling that balances predictive accuracy, scale\-dependent structure, and regional generalization\.EddyFlowis trained on the Gulf of St\. Lawrence and evaluated in zero\-shot and few\-shot settings on the Bay of Fundy and the Gulf of Mexico\.EddyFlowdemonstrates that physics\-informed representation learning reduces zero\-shot RMSE by 21%, achieves up to 85\.6% skill relative to persistence on unseen domains, and maintains near\-ideal spectral fidelity with a PSD ratio of≈1\.00\\approx 1\.00\.

SST Downscaling, Diffusion Models, Transfer Learning, Domain Shift, Spectral Fidelity

††copyright:cc††journalyear:2026††doi:XXXXXXX\.XXXXXXX††conference:Proceedings of the 33rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining; August, 2027; San Jose, CA, USA††isbn:978\-1\-4503\-XXXX\-X/2027/08††ccs:Computing methodologies Machine learning††ccs:Computing methodologies Neural networks††ccs:Computing methodologies Learning latent representations††ccs:Applied computing Earth and atmospheric sciences## 1\.Introduction

Sea Surface Temperature \(SST\) is a primary driver of coastal ecosystem health, fisheries productivity, and marine weather hazards\(Stock et al\.,[2015](https://arxiv.org/html/2608.04230#bib.bib2)\)\. However, satellite and reanalysis datasets that provide global SST estimates are far too coarse to capture the frontal structures, eddies, and thermal gradients that govern regional ocean dynamics\(Xing et al\.,[2026](https://arxiv.org/html/2608.04230#bib.bib3); Saxena and Sharma,[2020](https://arxiv.org/html/2608.04230#bib.bib4); Thiria et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib5); Kim et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib6)\)\. In the Gulf of St\. Lawrence \(GSL\), sharp SST fronts associated with the Laurentian Channel and shelf\-break bathymetry shape fish habitats and marine management decisions, yet remain unresolved in standard reanalyses such as ERA5111[https://cds\.climate\.copernicus\.eu/datasets/reanalysis\-era5\-single\-levels](https://cds.climate.copernicus.eu/datasets/reanalysis-era5-single-levels)\(Lavoie et al\.,[2021](https://arxiv.org/html/2608.04230#bib.bib7); Hersbach et al\.,[2020](https://arxiv.org/html/2608.04230#bib.bib8)\)\. Recovering kilometer\-scale SST from coarse inputs is a prerequisite for fisheries management and coastal hazard forecasting, which depend on precise thermal gradients, not approximate magnitudes\(Lavoie et al\.,[2021](https://arxiv.org/html/2608.04230#bib.bib7); Xing et al\.,[2026](https://arxiv.org/html/2608.04230#bib.bib3)\)\.

Deep learning models for structured prediction increasingly infer high\-resolution scientific fields from sparse or coarse observations\(Saxena and Sharma,[2020](https://arxiv.org/html/2608.04230#bib.bib4); Thiria et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib5); Kim et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib6); Lupin\-Jimenez et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib9)\)\. Even at strong predictive accuracy, pixel\-wise losses alone do not guarantee that learned representations capture target\-domain structure\. Representations here mean the latent encoding governing how spatial organization, temporal evolution, and cross\-scale dependencies are preserved\. When these encodings miss physical structures, models can achieve low reconstruction error while producing outputs that are spectrally smooth, temporally uninformative, or unable to generalize across physical regimes, a failure often visible only beyond the training distribution\(Thiria et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib5); Lupin\-Jimenez et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib9)\)\.

SST downscaling exemplifies this limitation as the target application reflects coupled physical processes, so faithful reconstruction requires more than interpolating the coarse input\(Lupin\-Jimenez et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib9)\)\. Models trained to minimize pixel\-wise reconstruction error can match target values while collapsing fine\-scale variance\(Kim et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib6); Thiria et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib5)\)\. This collapse goes unpenalized by pixel\-wise losses, yielding compressed predictions structurally inconsistent with the dynamics and fragile under geographic shift\(Lupin\-Jimenez et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib9)\)\.

The difficulty is sharpest where spatial statistics are not stationary and temporal persistence is strong\. The GSL exemplifies both\. Its semi\-enclosed geometry, freshwater input from major rivers, and strong coupling to shelf circulation produce an SST field whose spatial covariance structure varies across the domain\. Consequently, the mapping from coarse to fine resolution is non\-stationary, forcing region\-dependent relationships rather than a global interpolation rule\. Convolutional or attention\-based architectures\(Mardani et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib10); Wang et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib11)\)often assume spatial homogeneity and struggle to generalize\. Temporal autocorrelation makes persistence a competitive baseline\.

These observations suggest viewing SST downscaling not merely as image super\-resolution, but as representation learning under domain shift\. The learned encoding must remain accurate, temporally informative, and spectrally faithful when the physical regime changes across geographies\. We assess these properties with three metrics\. Root\-mean\-square Error \(RMSE\) measures pixel\-wise reconstruction error, the Skill Score \(SS\) measures improvement over the ERA5 persistence baseline, and the Power Spectral Density \(PSD\) ratio measures preservation of fine\-scale variability\. No single metric captures all three; their joint behavior determines whether learned representations are accurate, informative, and physically consistent\.

We presentEddyFlow, a physics\-informed representation\-learning framework for kilometer\-scale SST downscaling, evaluated under geographic domain shift\. We adopt the GSL as the training domain because its physical complexity, non\-stationary spatial statistics, and strong persistence test whether a model learns generalizable structure rather than dataset\-specific patterns \(Section[3](https://arxiv.org/html/2608.04230#S3)\)\.EddyFlowseparates atmospheric forcing, ocean memory, and diffusion\-based refinement to recover the mesoscale variance that pixel\-wise objectives smooth away \(Section[4](https://arxiv.org/html/2608.04230#S4)\), mesoscale means intermediate size or scale, typically ranging from roughly 2 to 1,000 kilometers\.EddyFlowis evaluated in zero\-shot and few\-shot transfer on the Bay of Fundy \(BOF\) and the Gulf of Mexico \(GOM\) in Section[5](https://arxiv.org/html/2608.04230#S5)\. Physics\-informed representation learning drives generalization across ocean regimes; our contributions:

- •a21% reduction in zero\-shot RMSEon unseen basins relative to a strong downscaling baseline;
- •up to85\.6% skill relative to the ERA5 persistence baselineon domains never observed during training; and
- •near\-ideal spectral fidelity, with a PSD ratio of≈1\.00\\approx 1\.00that preserves fine\-scale variability rather than blurring it\.

The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2608.04230#S2)reviews related work; Section[3](https://arxiv.org/html/2608.04230#S3)formalizes the downscaling task and evaluation criteria; Section[4](https://arxiv.org/html/2608.04230#S4)presentsEddyFlow; Section[5](https://arxiv.org/html/2608.04230#S5)reports in\-domain and transfer results; and Section[6](https://arxiv.org/html/2608.04230#S6)concludes\.

## 2\.Related Work

Learning high\-resolution spatial structure\.Much recent work treats geophysical downscaling as a spatial reconstruction problem, in which diffusion and generative models recover the sharp gradients and high\-frequency variability that pixel\-wise regression smooths away\(Liu et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib12); Hess et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib13)\)\. CorrDiff\(Mardani et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib10)\)pairs a deterministic backbone with a residual diffusion corrector, restricting stochastic generation to unresolved fine scales, improving realism while preserving spectral behavior\. DIFFDS\(Wang et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib11)\)adapts diffusion\-based image restoration to SST, learning a compact high\-resolution prior that guides reconstruction from coarse inputs\. Earlier work established that deep networks can recover ocean\-front structure from satellite imagery\. Lloyd et al\.\(Lloyd et al\.,[2021](https://arxiv.org/html/2608.04230#bib.bib14)\)fuse optical and thermal features for fivefold SST super\-resolution, and Ducournau and Fablet\(Ducournau and Fablet,[2016](https://arxiv.org/html/2608.04230#bib.bib15)\)showed that CNN\-based super\-resolution substantially outperforms classical downscaling methods\. Recent Mediterranean studies with deterministic and generative formulations\(Fanelli et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib16),[2025](https://arxiv.org/html/2608.04230#bib.bib17)\)reinforce that learning a conditional high\-resolution mapping from low\-resolution SST alone, outperforms interpolation, particularly in recovering small\-scale gradients and spectral properties\. These methods excel at spatial fidelity and high\-wavenumber variance recovery, but most refine one or a few concurrent observations with limited temporal conditioning, so improved visual and spectral quality does not guarantee evolution consistent with recent atmospheric forcing\.As they are largely translation\-equivariant architectures, they struggle to capture the fixed, spatially non\-stationary influence of coastlines, bathymetry, and continental shelves\.

Learning temporal ocean evolution\.A second line of work treats ocean prediction as sequence modeling, learning state evolution through autoregressive or multi\-step temporal models that better represent persistence, advection, and slowly varying circulation structure\(Krestenitis et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib18); Hao et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib19)\)\. WenHai\(Cui et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib20)\)shows that explicit temporal modeling is essential for maintaining coherent evolution in autoregressive settings, where small stepwise errors otherwise accumulate\. GraphCast\(Lam et al\.,[2023](https://arxiv.org/html/2608.04230#bib.bib21)\)and GenCast\(Price et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib22)\)demonstrate that graph\- and diffusion\-based temporal emulators achieve strong forecast skill while preserving dynamical coherence across lead times, and FuXi\-Ocean\(Niu et al\.,[2026](https://arxiv.org/html/2608.04230#bib.bib23)\)extends this to sub\-daily, eddy\-resolving global forecasting with improved skill over WenHai\. Forecasting, however, solves a different problem from SST downscaling\. It predicts future ocean states rather than recovering fine\-scale structure missing from coarse observations, capturing large\-scale evolution while missing sharp fronts and kilometer\-scale variability\. Many forecasting models also encode atmospheric forcing and ocean memory in a single latent representation, limiting separate modeling of short\-term dynamics from persistent ocean structure observations\.

Learning regional generalization\.A third line of work emphasizes that ocean prediction models must remain reliable when transferred across water bodies with distinct coastlines, boundary conditions, and circulation regimes, since one coarse input can map to different fine\-scale responses depending on local geometry and context\. Recent regional emulators combine low\-resolution autoregressive forecasting with learned high\-resolution refinement and online bias correction, remaining stable over long horizons while improving spectral behavior and fine\-scale realism\(Lupin\-Jimenez et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib9)\), showing that physical consistency, roll\-out stability, and multiscale correction can coexist in one framework\. Such models, however, are typically trained on a single basin, and their performance degrades substantially when out of the training distribution, making them effective basin\-specific models but unreliable on unseen regions\.

## 3\.Problem Formulation

Global SST products such as ERA5 capture the large\-scale thermal structure of the ocean but are too coarse to resolve fine\-scale features such as fronts, eddies, and coastal gradients\. Statistical downscaling therefore aims to recover a high\-resolution SST field consistent with the observed coarse state\(Michel et al\.,[2022](https://arxiv.org/html/2608.04230#bib.bib24)\)\. This is an ill\-posed problem because multiple plausible high\-resolution fields can correspond to the same coarse input\. As a result, models trained only with pixel\-wise regression losses often predict overly smooth fields that minimize average error but fail to recover realistic fine\-scale variability\. Diffusion models address this by refining an initial estimate towards consistent fine\-scale structure\(Ho et al\.,[2020](https://arxiv.org/html/2608.04230#bib.bib25); Karras et al\.,[2022](https://arxiv.org/html/2608.04230#bib.bib26)\)\.

Letxtx\_\{t\}denote the coarse\-resolution input at timett,oto\_\{t\}denote the optional ocean\-state history, andyty\_\{t\}denote the target high\-resolution SST field\. The objective is a mappingfθf\_\{\\theta\}from coarse input and history to the high\-resolution fieldy^t=fθ​\(xt,ot\)\\hat\{y\}\_\{t\}=f\_\{\\theta\}\(x\_\{t\},o\_\{t\}\)\. Since SST evolves gradually, we estimate the temporal residualrt=yt−yt−1r\_\{t\}=y\_\{t\}\-y\_\{t\-1\}rather than the full field, and recover the target by adding it back to the previous day throughy^t=yt−1\+r^t\\hat\{y\}\_\{t\}=y\_\{t\-1\}\+\\hat\{r\}\_\{t\}\. Predicting the residual focuses the model on day\-to\-day spatial change rather than repeatedly reconstructing the slowly varying background at every timestep\. It also yields a natural persistence reference that carries the previous day’s state forward and thus assumes the SST remains unchanged absent new forcing\.

These observations motivate the three complementary evaluation criteria used throughout this study: RMSE for pixel\-wise accuracy, persistence\-relative skill score for temporal informativeness, and PSD ratio for spectral fidelity\. Beyond metrics, we evaluate in zero\-shot and few\-shot transfer settings\. A model that performs well only within its training basin may simply have learned basin\-specific coastline or circulation patterns rather than the coarse\-to\-fine relationship\. Testing on unseen basins therefore better measures whether the representation generalizes across geographies\.

## 4\.Methodology

### 4\.1\.RMSE and the Skill Score

SST demonstrates strong day\-to\-day persistence\(Bulgin et al\.,[2020](https://arxiv.org/html/2608.04230#bib.bib27)\)\. The simplest predictor,y^tpers=yt−1\\hat\{y\}^\{\\mathrm\{pers\}\}\_\{t\}=y\_\{t\-1\}, achieves low RMSE in most coastal and shelf regions because SST rarely changes significantly within a 24\-hour period\. This situation presents a methodological challenge where a model may report an RMSE that appears excellent in absolute terms while providing no additional predictive information, simply by replicating persistence\. Consequently, RMSE is necessary but insufficient, as it cannot distinguish between a model that learned ocean dynamics and one that reproduces the previous day’s field\.

We therefore define a skill score relative to a persistence reference, following the classical verification framework in which a forecast is judged not against zero, but against a reference forecast requiring no learning component\(Knaff and Landsea,[1997](https://arxiv.org/html/2608.04230#bib.bib28)\)\. LetRMSE​\(y^t,yt\)\\mathrm\{RMSE\}\(\\hat\{y\}\_\{t\},y\_\{t\}\)denote the root\-mean\-square error of the model prediction andRMSE​\(y^tpers,yt\)\\mathrm\{RMSE\}\(\\hat\{y\}^\{\\mathrm\{pers\}\}\_\{t\},y\_\{t\}\)that of the persistence reference\. We define the skill score as:

\(1\)SS=1−RMSE​\(y^t,yt\)RMSE​\(y^tpers,yt\)\.\\mathrm\{SS\}=1\-\\frac\{\\mathrm\{RMSE\}\(\\hat\{y\}\_\{t\},y\_\{t\}\)\}\{\\mathrm\{RMSE\}\(\\hat\{y\}^\{\\mathrm\{pers\}\}\_\{t\},y\_\{t\}\)\}\.
The ratio is directly interpretable\. WhenSS=0\\mathrm\{SS\}=0, the model is exactly as informative as carrying the reference field forward\. WhenSS\>0\\mathrm\{SS\}\>0, the model has extracted genuine predictive signal beyond persistence, and whenSS<0\\mathrm\{SS\}<0, the model is*worse*than the naive baseline despite an RMSE that looks reasonable in isolation\. This last case is the one that a pixel\-wise loss alone cannot detect, which is why persistence\-relative skill, not RMSE alone, is the correct lens for short\-horizon SST prediction\. In our zero\-shot BOF and GOM setting, a positive skill score is the minimum bar for indicating model usefulness\. The skill score, however, is computed on raw pixel values and is blind to*where*error lives across spatial scale\. A model can satisfySS\>0\\mathrm\{SS\}\>0while discarding the fine\-scale structure that makes it useful, a failure visible only in the frequency domain\.

### 4\.2\.Pixel\-Wise Loss and the PSD Ratio

Minimizing squared error drives a network toward the conditional mean of the target at each pixel; when fine\-scale structure is only partially determined by the coarse input, that mean is smoother than any single realization, so a pixel\-wise reconstruction loss \(the loss termℒrec\\mathcal\{L\}\_\{\\mathrm\{rec\}\}in Section[4\.7](https://arxiv.org/html/2608.04230#S4.SS7)\) alone yields pixel\-accurate yet systematically under\-variant predictions at high wavenumbers, a deviation we quantify in the frequency domain, scale by scale\.

Lety​\(𝐮\)y\(\\mathbf\{u\}\)be the target field andy^​\(𝐮\)\\hat\{y\}\(\\mathbf\{u\}\)the prediction over spatial position𝐮=\(u1,u2\)\\mathbf\{u\}=\(u\_\{1\},u\_\{2\}\), each with its spatial mean removed,

\(2\)y′​\(𝐮\)=y​\(𝐮\)−μy\.y^\{\\prime\}\(\\mathbf\{u\}\)=y\(\\mathbf\{u\}\)\-\\mu\_\{y\}\.The 2D Fourier transform of the centered field and its inverse are

\(3\)Y​\(𝐤\)=∬y′​\(𝐮\)​e−i​𝐤⋅𝐮​𝑑𝐮,y′​\(𝐮\)=1\(2​π\)2​∬Y​\(𝐤\)​ei​𝐤⋅𝐮​𝑑𝐤,Y\(\\mathbf\{k\}\)=\\iint y^\{\\prime\}\(\\mathbf\{u\}\)\\,e^\{\-i\\mathbf\{k\}\\cdot\\mathbf\{u\}\}\\,d\\mathbf\{u\},\\quad y^\{\\prime\}\(\\mathbf\{u\}\)=\\tfrac\{1\}\{\(2\\pi\)^\{2\}\}\\\!\\iint Y\(\\mathbf\{k\}\)\\,e^\{i\\mathbf\{k\}\\cdot\\mathbf\{u\}\}\\,d\\mathbf\{k\},with wavenumber𝐤=\(kx,ky\)\\mathbf\{k\}=\(k\_\{x\},k\_\{y\}\), and Parseval’s theorem guarantees

\(4\)∬\|y′​\(𝐮\)\|2​𝑑𝐮=1\(2​π\)2​∬\|Y​\(𝐤\)\|2​𝑑𝐤,\\iint\|y^\{\\prime\}\(\\mathbf\{u\}\)\|^\{2\}\\,d\\mathbf\{u\}=\\frac\{1\}\{\(2\\pi\)^\{2\}\}\\iint\|Y\(\\mathbf\{k\}\)\|^\{2\}\\,d\\mathbf\{k\},so spatial variance is exactly preserved and attributable spectrally\.

The power spectra of the target field and the model prediction are

\(5\)Py​\(𝐤\)=\|Y​\(𝐤\)\|2,Py^​\(𝐤\)=\|Y^​\(𝐤\)\|2,P\_\{y\}\(\\mathbf\{k\}\)=\|Y\(\\mathbf\{k\}\)\|^\{2\},\\qquad P\_\{\\hat\{y\}\}\(\\mathbf\{k\}\)=\|\\hat\{Y\}\(\\mathbf\{k\}\)\|^\{2\},whereY^​\(𝐤\)\\hat\{Y\}\(\\mathbf\{k\}\)denotes the transform of the centered prediction\. Ocean fields are scale\-organized rather than direction\-organized; consequently we radially average over the annulusΩk\\Omega\_\{k\}atk=kx2\+ky2k=\\sqrt\{k\_\{x\}^\{2\}\+k\_\{y\}^\{2\}\},

\(6\)P¯y​\(k\)=1\|Ωk\|​∑𝐤∈ΩkPy​\(𝐤\),\\bar\{P\}\_\{y\}\(k\)=\\frac\{1\}\{\|\\Omega\_\{k\}\|\}\\sum\_\{\\mathbf\{k\}\\in\\Omega\_\{k\}\}P\_\{y\}\(\\mathbf\{k\}\),and identically forP¯y^​\(k\)\\bar\{P\}\_\{\\hat\{y\}\}\(k\), giving the isotropic 1D PSD used in downscaling and turbulence diagnostics\(Mardani et al\.,[2025](https://arxiv.org/html/2608.04230#bib.bib10)\)\. We define the PSD ratio as:

\(7\)PSDR​\(k\)=P¯y^​\(k\)P¯y​\(k\)\+ε,\\mathrm\{PSDR\}\(k\)=\\frac\{\\bar\{P\}\_\{\\hat\{y\}\}\(k\)\}\{\\bar\{P\}\_\{y\}\(k\)\+\\varepsilon\},withε\\varepsilona small constant for numerical stability, evaluated between the coarse\-input scale and the kilometer\-scale Nyquist limit, reported over 5–50 km \(Section[5](https://arxiv.org/html/2608.04230#S5)\)\. A value ofPSDR​\(k\)≈1\\mathrm\{PSDR\}\(k\)\\approx 1indicates that variance is preserved at scalekk, whilePSDR​\(k\)<1\\mathrm\{PSDR\}\(k\)<1indicates over\-smoothing, the spectral signature of mean\-seeking collapse, andPSDR​\(k\)\>1\\mathrm\{PSDR\}\(k\)\>1indicates over\-amplified noise\. A model can post low RMSE andSS\>0\\mathrm\{SS\}\>0whilePSDR​\(k\)→0\\mathrm\{PSDR\}\(k\)\\to 0at highkk, exactly the failure mode invisible to pixel\-wise metrics, since eddies and fronts occupy bands ofkk; a faithful model must not collapse\.

### 4\.3\.Two Streams for Two Timescales

Atmospheric forcing and oceanic memory operate on separable timescales\. Synoptic weather systems driving short\-term SST change have lifetimes of two to seven days, while SST anomalies set by that forcing persist for weeks to months as the mixed layer relaxes toward equilibrium\(Bulgin et al\.,[2020](https://arxiv.org/html/2608.04230#bib.bib27)\)\. One encoder at one window and sampling rate cannot resolve both, soEddyFlowhas two input tensors,

\(8\)Xa∈ℝ28×23×20×48,Xo∈ℝ60×1×501×1201,X\_\{a\}\\in\\mathbb\{R\}^\{28\\times 23\\times 20\\times 48\},\\qquad X\_\{o\}\\in\\mathbb\{R\}^\{60\\times 1\\times 501\\times 1201\},with dimensions ordered as frames, channels, height, and width\.XaX\_\{a\}is a 7\-day, 6\-hourly window of 23 coarse atmospheric reanalysis forcing fields, spanning roughly two synoptic cycles, beyond which autocorrelation decays too quickly to add information\.XoX\_\{o\}is a 60\-day, daily window of fine\-resolution ocean surface temperature observations, chosen so that mesoscale eddies translating at 5–20 km/day\(Chelton et al\.,[2011](https://arxiv.org/html/2608.04230#bib.bib29)\)remain inside the receptive field even after several hundred kilometers, and long enough to span seasonal transitions such as summer heating giving way to autumn cooling\.

Because the two streams differ in resolution by roughly two orders of magnitude, they cannot share a patch embedding\(Dosovitskiy et al\.,[2021](https://arxiv.org/html/2608.04230#bib.bib30)\)\. The atmospheric stream uses a fixed patch size of 4, projecting each 20×\\times48 frame toN=5×12=60N=5\\times 12=60tokens via a strided convolution\. The ocean stream uses a dynamic patch embedding with patch size 64, keeping the token count tractable at 501×\\times1201 resolution and yieldingNo=8×19=152N\_\{o\}=8\\times 19=152tokens per frame\. The patch grid is computed at runtime from the input shape, so the same ocean encoder applies to differently sized domains without modification\.

### 4\.4\.Spatio\-Temporal Stream Encoders

Each stream factorizes attention into a spatial stage, mixing across theNNtokens per frame, and a temporal stage, mixing across frames per location\(Bertasius et al\.,[2021](https://arxiv.org/html/2608.04230#bib.bib31)\)\. For the atmospheric stream,

\(9\)ha,τ=SpatialAttn​\(Xa,τ\),τ=1,…,28,h\_\{a,\\tau\}=\\mathrm\{SpatialAttn\}\(X\_\{a,\\tau\}\),\\quad\\tau=1,\\dots,28,\(10\)a=TemporalAttn​\(ha,1,…,ha,28\),a=\\mathrm\{TemporalAttn\}\(h\_\{a,1\},\\dots,h\_\{a,28\}\),using 8 spatial layers then 4 temporal layers, with a learned frame embedding in place of a causal mask, since every frame in the window precedes the prediction time\. The spatial summaryssis kept separate from the temporal summaryaa:ssanswers where current forcing sits,aahow it evolved over the week, two physically distinct questions\.

The ocean stream’s spatial attention uses 2D Rotary Positional Embedding \(RoPE\)\(Su et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib32); Heo et al\.,[2024](https://arxiv.org/html/2608.04230#bib.bib33)\)rather than absolute position:

\(11\)ho,τ=SpatialAttnRoPE​\(Xo,τ\),τ=1,…,60,h\_\{o,\\tau\}=\\mathrm\{SpatialAttn\}\_\{\\mathrm\{RoPE\}\}\(X\_\{o,\\tau\}\),\\quad\\tau=1,\\dots,60,\(12\)o~=TemporalAttn​\(ho,1,…,ho,60\),\\tilde\{o\}=\\mathrm\{TemporalAttn\}\(h\_\{o,1\},\\dots,h\_\{o,60\}\),after which the ocean summary is pooled to match the atmospheric token count,o=AdaptiveAvgPool​\(o~\)∈ℝN×Do=\\mathrm\{AdaptiveAvgPool\}\(\\tilde\{o\}\)\\in\\mathbb\{R\}^\{N\\times D\}, since fusion requires a common grid despite the ocean stream’s finer resolution\.

![Refer to caption](https://arxiv.org/html/2608.04230v1/x1.png)Figure 1\.TheEddyFlowarchitecture\. An atmospheric \(ERA5\) encoder processes a 7\-day, 6\-hourly forcing window and an ocean \(MUR\) encoder a 60\-day daily SST history, each factorizing attention into spatial and temporal stages, with RoPE in the ocean stream\. Spatial \(ss\), temporal \(aa\), and pooled ocean \(oo\) summaries are fused by a two\-layer MLP into a latentzz, which the decoder upsamples under bathymetric guidance into the coarse tendencyxbasex\_\{\\mathrm\{base\}\}; a diffusion U\-Net refines the fine\-scale residual\.
### 4\.5\.Three\-Way Fusion

The atmospheric spatial summary, temporal summary, and pooled ocean summary are concatenated and fused with a two\-layer MLP rather than a single linear projection, since the mapping from three different summaries \(instantaneous forcing location, forcing history, ocean memory\) to a joint latent is unlikely to be linear:

\(13\)z=MLP​\(\[s;a;o\]\)∈ℝN×D\.z=\\mathrm\{MLP\}\\big\(\[\\,s\\,;\\,a\\,;\\,o\\,\]\\big\)\\in\\mathbb\{R\}^\{N\\times D\}\.
The concatenation keeps the spatial patternssdistinct from the temporal trendaa, since a forcing event’s location and persistence differently constrain where and how much SST will change\.

### 4\.6\.Decoder and Bathymetric Guidance

The latentzzis reshaped to\[D,nH,nW\]\[D,n\_\{H\},n\_\{W\}\]and passed through a deterministic decoder producing a coarse baseline prediction of the SST tendency,xbase=Rθ​\(z\)x\_\{\\mathrm\{base\}\}=R\_\{\\theta\}\(z\), via four bilinear upsampling stages\. The output size is derived at runtime from the bathymetry tensor, so the same decoder targets different grid sizes without modification\. Bathymetry is injected as a skip connection before the final convolution,cat​\[h,bathy\]\\mathrm\{cat\}\[h,\\mathrm\{bathy\}\]withhhthe penultimate feature map, so shelf breaks and banks shape the output where they are resolved, not only at the coarse encoder input where they modulate the forcing\.

### 4\.7\.Residual Formulation and Loss

As formalized in Section[3](https://arxiv.org/html/2608.04230#S3),EddyFlowpredicts the SST tendencyrt=yt−yt−1r\_\{t\}=y\_\{t\}\-y\_\{t\-1\}rather than the absolute field and recoversy^t=yt−1\+r^t\\hat\{y\}\_\{t\}=y\_\{t\-1\}\+\\hat\{r\}\_\{t\}\. The motivation is dynamic range\. Absolute SST in the training domain spans tens of degrees Celsius, while the residual diffusion component models have a standard deviation nearly two orders of magnitude smaller\. Learning the smaller signal concentrates capacity on the fine structure of daily change rather than the seasonal cycle, carried forward throughyt−1y\_\{t\-1\}\. The joint training objective, minimized in Phase 1 of the training schedule, is

\(14\)ℒ=ℒrec\+10​ℒdiff\+0\.1​ℒspec,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{rec\}\}\+10\\,\\mathcal\{L\}\_\{\\mathrm\{diff\}\}\+0\.1\\,\\mathcal\{L\}\_\{\\mathrm\{spec\}\},whereℒrec\\mathcal\{L\}\_\{\\mathrm\{rec\}\}is theL1L\_\{1\}loss between the deterministic baselinexbasex\_\{\\mathrm\{base\}\}and the tendencyrtr\_\{t\},ℒdiff\\mathcal\{L\}\_\{\\mathrm\{diff\}\}is the diffusion loss in Section[4\.8](https://arxiv.org/html/2608.04230#S4.SS8), and

\(15\)ℒspec=𝔼​\[w​\(k\)​\|\|Y^​\(k\)\|−\|Y​\(k\)\|\|\],w​\(k\)=1\+20​k2\.\\mathcal\{L\}\_\{\\mathrm\{spec\}\}=\\mathbb\{E\}\\\!\\left\[w\(k\)\\,\\bigl\|\\,\|\\hat\{Y\}\(k\)\|\-\|Y\(k\)\|\\,\\bigr\|\\right\],\\quad w\(k\)=\\sqrt\{1\+20k^\{2\}\}\.
The weightw​\(k\)w\(k\)grows approximately linearly in wavenumber, penalizing missing high\-wavenumber power more than low\-wavenumber error, which counteracts the tendency ofℒrec\\mathcal\{L\}\_\{\\mathrm\{rec\}\}to favor spatially smooth predictions\. The weights1010and0\.10\.1keep the initially small diffusion loss from being ignored and the large\-for\-textured\-fields spectral loss from dominating reconstruction\.

### 4\.8\.Diffusion Refinement with DDIM

The deterministic baselinexbasex\_\{\\mathrm\{base\}\}, even withℒspec\\mathcal\{L\}\_\{\\mathrm\{spec\}\}, still fails to represent mesoscale variance that is only partially determined by the coarse input, so a diffusion model refines the correction residual

\(16\)δt=rt−xbase,\\delta\_\{t\}=r\_\{t\}\-x\_\{\\mathrm\{base\}\},the part of the tendency the deterministic decoder fails to capture\. We writeδ\\deltaforδt\\delta\_\{t\}below, omitting the time index\. We use a standard linear\-schedule diffusion formulation with DDIM sampling, since the residual is modeled directly as a noise\-prediction problem under a DDPM\-style forward process\. The forward process uses a linear noise scheduleβt∈\[10−4,0\.02\]\\beta\_\{t\}\\in\[10^\{\-4\},0\.02\]overTdiffT\_\{\\text\{diff\}\}steps, withαt=1−βt\\alpha\_\{t\}=1\-\\beta\_\{t\}andα¯t=∏s≤tαs\\bar\{\\alpha\}\_\{t\}=\\prod\_\{s\\leq t\}\\alpha\_\{s\}\. For a clean residualδ\\deltaand Gaussian noiseϵ∼𝒩​\(0,I\)\\epsilon\\sim\\mathcal\{N\}\(0,I\), the noised sample is

\(17\)rt=α¯t​δ\+1−α¯t​ϵ\.r\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\delta\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\epsilon\.The denoiserϵθ​\(rt,xbase,z,bathy,t\)\\epsilon\_\{\\theta\}\(r\_\{t\},x\_\{\\text\{base\}\},z,\\text\{bathy\},t\)is trained to recover the injected noise directly:

\(18\)ℒdiff=𝔼t,ϵ​\[‖ϵθ​\(rt,xbase,z,bathy,t\)−ϵ‖2\]\.\\mathcal\{L\}\_\{\\mathrm\{diff\}\}=\\mathbb\{E\}\_\{t,\\epsilon\}\\left\[\\left\\\|\\epsilon\_\{\\theta\}\(r\_\{t\},x\_\{\\mathrm\{base\}\},z,\\mathrm\{bathy\},t\)\-\\epsilon\\right\\\|^\{2\}\\right\]\.At inference, DDIM sampling proceeds deterministically withη=0\\eta=0via thex0x\_\{0\}\-prediction form,

\(19\)δ^0=rt−1−α¯t​ϵθα¯t,rt−1=α¯t−1​δ^0\+1−α¯t−1​ϵθ,\\hat\{\\delta\}\_\{0\}=\\frac\{r\_\{t\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\epsilon\_\{\\theta\}\}\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\},\\quad r\_\{t\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\hat\{\\delta\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\epsilon\_\{\\theta\},run over a strided subsequence of the training timesteps for fast sampling\. As in the rest of the model, the full training objective combines baseline, diffusion, and spectral terms,

\(20\)ℒ=ℒbase\+10​ℒdiff\+0\.1​ℒspec\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{base\}\}\+10\\,\\mathcal\{L\}\_\{\\mathrm\{diff\}\}\+0\.1\\,\\mathcal\{L\}\_\{\\mathrm\{spec\}\}\.

### 4\.9\.RoPE and Domain Conditioning

Absolute positional encoding binds a token’s representation to its row and column in the training grid\. Trained with absolute 2D sinusoidal position, the ocean stream learned associations between fixed grid coordinates and features, associations that become meaningless once the same coordinates index a different basin\. We therefore replace absolute position with 2D RoPE in the ocean stream’s spatial attention, which rotates query and key vectors by an angle determined by token position so that the attention scoreqi⊤​kj∝cos⁡\(θ​\(Δ​ui​j,Δ​vi​j\)\)q\_\{i\}^\{\\top\}k\_\{j\}\\;\\propto\\;\\cos\\big\(\\theta\(\\Delta u\_\{ij\},\\,\\Delta v\_\{ij\}\)\\big\)depends only on the*relative*row–column displacement\(Δ​ui​j,Δ​vi​j\)\(\\Delta u\_\{ij\},\\,\\Delta v\_\{ij\}\)between tokensiiandjj, never their absolute position\. An eddy of a given radius therefore produces the same attention pattern wherever, and in whichever domain, it appears, a property absolute encoding cannot provide\.

In parallel, the diffusion head is conditioned on the per\-domain mean and standard deviation of the residual field\. At training time, these statistics default to\(0,1\)\(0,1\); at zero\-shot inference, the true domain statistics are substituted, letting the head rescale its output to an unseen basin without gradient\-based adaptation\. Together they target the two most basin\-specific quantities, spatial relational structure and the field’s statistical scale, leaving the atmospheric stream’s absolute positions as the component not yet transfer\-invariant\.

## 5\.Experiments and Analysis

### 5\.1\.Experimental Setup

Models are implemented in PyTorch\(Paszke et al\.,[2019](https://arxiv.org/html/2608.04230#bib.bib34)\)and trained on Digital Research Alliance of Canada infrastructure\. Each run uses 4 CPU cores and one 3g\.40 GB Multi\-Instance GPU \(MIG\) slice of an NVIDIA H100; one MIG slice per run, without distributed data parallelism\. This provides enough memory for all components while keeping runs reproducible and easy to schedule on a shared cluster\.

#### 5\.1\.1\.Dataset and Preprocessing

ERA5 is the fifth\-generation European Centre for Medium\-Range Weather Forecasts \(ECMWF\)222[https://www\.ecmwf\.int](https://www.ecmwf.int/)atmospheric reanalysis, combining short\-range forecasts with observations through data assimilation into a physically consistent, globally complete hourly atmospheric state\(Hersbach et al\.,[2020](https://arxiv.org/html/2608.04230#bib.bib8)\)\. We use ERA5 at 6\-hourly resolution \(00, 06, 12, 18 UTC\), cropped to 47∘N–52∘N, 68∘W–56∘W, and regridded to a coarse 20×\\times48 grid at 0\.25∘spacing\.

MUR \(Multi\-scale Ultra\-high Resolution\)333[https://podaac\.jpl\.nasa\.gov/dataset/MUR\-JPL\-L4\-GLOB\-v4\.1](https://podaac.jpl.nasa.gov/dataset/MUR-JPL-L4-GLOB-v4.1)SST is a satellite\-derived Level\-4 product that fuses infrared and microwave observations with in\-situ measurements into a gap\-filled daily SST analysis at approximately 1 km resolution\(JPL MUR MEaSUREs Project,[2015](https://arxiv.org/html/2608.04230#bib.bib35)\)\. We use MUR at its native 0\.01∘spacing over the same crop, giving a 501×\\times1201 fine grid\. The training target is the SST tendencyrt=yt−yt−1r\_\{t\}=y\_\{t\}\-y\_\{t\-1\}of Section[3](https://arxiv.org/html/2608.04230#S3), computed from the MUR field; the absolute field is retained only to reconstructy^t\\hat\{y\}\_\{t\}at evaluation\. The training domain is the GSL; the zero\-shot and few\-shot evaluation domains are BOF and GOM\. Data are split temporally, where training uses 2013–2020 \(8 years\), validation uses 2021, and test uses 2022–2023 \(732 valid samples\)\. Normalization statistics are computed over each stream’s training\-period data\.

ERA5’s SST channel is not an independently analyzed ocean product; it is a lower boundary condition prescribed from an external, coarser SST/sea\-ice product used only to force the atmospheric model rather than to resolve ocean structure\(Hirahara and Hersbach,[2016](https://arxiv.org/html/2608.04230#bib.bib36)\)\. ERA5 SST is thus unsuited to fine\-scale reconstruction and, in our crop, shows substantially weaker gradients than the MUR analysis\.

#### 5\.1\.2\.Evaluation Metrics

We report the three metrics of Sections[4\.1](https://arxiv.org/html/2608.04230#S4.SS1)and[4\.2](https://arxiv.org/html/2608.04230#S4.SS2), instantiated as follows\. First, RMSE is pixel\-and\-day\-pooled over the full test period, restricted to valid ocean pixels only:

\(21\)RMSE=∑\(y^−y\)2⋅m∑m,\\mathrm\{RMSE\}=\\sqrt\{\\frac\{\\sum\(\\hat\{y\}\-y\)^\{2\}\\cdot m\}\{\\sum m\}\},wheremmis the valid mask, excluding pixels with sea\-ice fraction above 0\.15 or non\-finite SST values\. RMSE is computed in normalized units and converted to Celsius for reporting, except on the GOM, where figures and tables state normalized units and∘C values follow by multiplying byσ=6\.3153\\sigma=6\.3153\. Second, we report the skill score of Eq\. \([1](https://arxiv.org/html/2608.04230#S4.E1)\), where the persistence reference is ERA5’s coarse SST channel, denormalized and bilinearly upsampled to the fine grid; this is the sole persistence denominator used throughout the paper\. Third, we report the PSD ratio of Section[4\.2](https://arxiv.org/html/2608.04230#S4.SS2), the mean azimuthally averaged 2D PSD of prediction over target in the 5–50 km band, plus the logarithmic biaslog10⁡\(PSDR\)\\log\_\{10\}\(\\mathrm\{PSDR\}\)\. Correlation is not a primary aggregate metric\. All evaluations provide the model with ERA5 inputs up to timettand MUR history up to timet−1t\-1\.

#### 5\.1\.3\.Models and Training Protocol

Table 1\.The six\-stage architectural ladder used for all experiments\. Stage 1, joint spatio\-temporal attention with a residual head, serves as the strong downscaling baseline; each later stage modifies one part of the design \(rightmost column\), so metric changes between adjacent stages localize, though do not perfectly isolate, each component’s contribution\.Table[1](https://arxiv.org/html/2608.04230#S5.T1)defines the six\-stage ladder used throughout the evaluation\. Stage 1, a single joint spatio\-temporal attention model with a residual head, serves as the strong downscaling baseline; each subsequent stage builds intoEddyFlow\.

Training proceeds in three phases; Phase 1 jointly trains the encoder, decoder, and diffusion module for 100 epochs on the GSL under the objective of Eq\. \([14](https://arxiv.org/html/2608.04230#S4.E14)\); Phase 2 fine\-tunes the diffusion module \(encoder and decoder frozen\) for 40 epochs on the BOF and the GOM; and Phase 3 performs few\-shot transfer for 20 epochs at 1\-, 7\-, and 30\-day target windows\. All zero\-shot results use the Phase 1 checkpoint; Phases 2 and 3 constitute the few\-shot experiments\. All phases use AdamW\(Loshchilov and Hutter,[2019](https://arxiv.org/html/2608.04230#bib.bib38)\)\(β1,β2=0\.9,0\.95\\beta\_\{1\},\\beta\_\{2\}=0\.9,0\.95, weight decay10−410^\{\-4\}\), with learning rates3×10−43\\times 10^\{\-4\}and10−510^\{\-5\}for joint training and fine\-tuning under cosine annealing to0\.01×0\.01\\times, bfloat16 mixed precision, gradient clipping at 1\.0 post\-accumulation, no early stopping, and non\-finite\-loss batches skipped rather than aborting\.

#### 5\.1\.4\.Baselines

We report two baselines\. The first is ERA5 coarse persistence, constructed as in the skill\-score definition above; it serves as the skill\-score denominator and represents no learned downscaling model beyond the reanalysis SST channel itself\. The second is MUR\-yesterday, which is reported only as an oracle upper\-bound reference and is not used to compute the main skill score, since the MUR history is an input to the model \(Section[4\.1](https://arxiv.org/html/2608.04230#S4.SS1)\)\. We include no separate classical interpolation or super\-resolution baseline; bilinear up\-sampling appears only inside the ERA5 persistence construction, not as an independent competitor\.

### 5\.2\.GSL Results

Table[2](https://arxiv.org/html/2608.04230#S5.T2)collects in\-domain GSL performance across all six stages for RMSE, skill, and PSD ratio\. Across all three metrics, performance is largely set by the early stages and, except for Stage 5’s latent\-domain variant, changes marginally thereafter, confirming that the ladder’s later stages are not evaluated on further in\-domain gains; the transfer results below should be read against this baseline\.

Table 2\.In\-domain GSL RMSE \(∘C\), persistence\-relative skill, and 5–50 km PSD ratio for all six stages, together with the zero\-shot PSD ratio on the transfer basins \(BOF, GOM\)\. PSD values closer to 1 indicate closer agreement with target spectral power\. In\-domain performance changes little across stages, so the ladder’s differences emerge under transfer\.
### 5\.3\.RMSE Results

![Refer to caption](https://arxiv.org/html/2608.04230v1/rmse.png)Figure 2\.RMSE across the six\-stage ladder on the \(a\) BOF and \(b\) GOM, versus the few\-shot budgetnn, the number of target\-domain days used for adaptation; lower is better\.Figure[2](https://arxiv.org/html/2608.04230#S5.F2)reports zero\-shot RMSE across the six\-stage ladder on GSL, BOF, and GOM\. Source\-domain error saturates almost immediately as, past the Stage 1→\\to2 transition, GSL RMSE stays within a narrow band, with Stage 5 as the sole outlier, so little of the ladder shows up in\-domain\. Transfer tells a different story as Stage 4 cuts BOF RMSE by roughly 22% and GOM RMSE by roughly 21% relative to the Stage 1 baseline, the largest transfer\-domain gap in the ladder and the lowest zero\-shot error on either transfer domain\. With no GSL counterpart, the gap reflects better generalization from the representation, not better fitting of the training distribution\.

Stages 5 and 6, despite more expressive generative refinement, do not improve on Stage 4’s zero\-shot RMSE; added diffusion\-head flexibility does not translate into transfer accuracy\. Few\-shot adaptation adds little further evidence either way as Stage 4 moves by under 1% at 1 day and is measurably worse by 30 days on both domains as the diffusion starts over\-fitting, so the zero\-shot checkpoint already sits near the transfer ceiling before any target supervision\.

### 5\.4\.Skill Score Results

![Refer to caption](https://arxiv.org/html/2608.04230v1/skill.png)Figure 3\.Skill relative to ERA5 persistence across the ladder stages on \(a\) BOF and \(b\) GOM versus the few\-shot budget; higher is better; zero corresponds to carrying the reference forward, and positive values indicate improvement over it\.Skill relative to ERA5 persistence \(Figure[3](https://arxiv.org/html/2608.04230#S5.F3)\) reveals a pattern RMSE alone does not capture\. Compared with Stage 1, Stage 2 leaves source\-domain skill essentially unchanged, yet transfer skill fails to improve, dipping on GOM while staying flat on BOF, indicating that optimizing further for the GSL training distribution does not necessarily improve generalization\. Stage 4 breaks this trend decisively as source\-domain skill remains nearly unchanged, while transfer skill increases on both BOF and GOM\. This matches the RMSE transition, now as an improvement over ERA5 persistence rather than absolute error\.

Stage 5 reduces transfer skill relative to Stage 4, and Stage 6 remains competitive without consistently surpassing it; despite substantially different diffusion formulations, neither improves on Stage 4’s transfer skill\. This suggests that the primary source of the performance gain lies in the learned encoder representation rather than in the downstream diffusion refinement\. Few\-shot experiments reinforce this on both domains as the ranking barely changes after fine\-tuning, so limited supervision mainly confirms the zero\-shot representation rather than producing a better one\.

### 5\.5\.PSD Ratio Results

PSD ratio in the 5–50 km band \(Table[2](https://arxiv.org/html/2608.04230#S5.T2)\) answers a different question, asking not how large the error is but whether fine\-scale structure survives\. The ladder settles even earlier than in RMSE or skill\. Stage 1 is the only model with a visible departure from spectral parity on GSL, and every later stage stays close to 1\.0 across all three domains\. That rules out the simplest alternative explanation as Stage 4 is not buying lower RMSE by smoothing away the variance PSD measures\.

What makes this metric distinctive is how little it moves across Stages 4–6 even though their diffusion heads differ substantially and their RMSE and skill numbers disagree\. A shared encoder producing near\-identical spectral fidelity under three refinement strategies is the clearest evidence that preserved mesoscale content is a property of the representation, not the generative head\. Few\-shot fine\-tuning moves this metric least as PSD stays near parity at every shot level, so adaptation recalibrates without altering retained spatial content\.

![Refer to caption](https://arxiv.org/html/2608.04230v1/x2.png)Figure 4\.Zero\-shot downscaling under geographic domain shift for two basins held out of training\. Each row shows a single date in an unseen basin: the Bay of Fundy \(top\) and the Gulf of Mexico \(bottom\), chosen to span contrasting thermal regimes\. From left to right, the panels present the coarse input, the MUR ground\-truth analysis, and the Stage 4 zero\-shot prediction\. The prediction is produced with no supervision from the target basin, showing that the learned representation reconstructs mesoscale sea\-surface\-temperature structure across basins rather than memorizing the training domain\.Figure[4](https://arxiv.org/html/2608.04230#S5.F4)illustrates the qualitative counterpart of this transition\. Together, the three metrics localize the gains in the encoder as Stage 4’s dual\-stream, RoPE\-based representation enables zero\-shot transfer across accuracy, skill, and spectral fidelity; Stages 5 and 6 mainly modulate adaptation and stability; and few\-shot supervision leaves the ranking intact\. Detailed ablations appear in the appendix\.

## 6\.Conclusion

This paper introducedEddyFlow, a physics\-informed framework for kilometer\-scale SST downscaling\. It encodes atmospheric forcing and ocean memory in separate streams at their native timescales, applies relative \(RoPE\) spatial encoding in the ocean pathway, injects bathymetry where it is resolved, and refines a deterministic tendency with a diffusion head under a wavenumber\-weighted spectral loss, with performance judged jointly by RMSE, persistence\-relative skill, and the PSD ratio\. Trained on the GSL, it transfers zero\-shot to the BOF and the GOM, cutting RMSE by up to 21% over a strong baseline, reaching up to 85\.6% skill over ERA5 persistence, and holding the 5–50 km PSD ratio near 1\.00\. These gains localize in the dual\-stream encoder, where RoPE\-based representation sets transfer performance, while richer diffusion heads alter adaptation and stability without improving zero\-shot accuracy, so transfer\-aligned inductive bias matters more than added generative complexity\.

Three limitations qualify these results\. \(i\) the atmospheric stream still relies on absolute positions, so one pathway remains bound to training\-domain geometry and plausibly contributes to the residual skill degradation observed under transfer; \(ii\) evaluation covers a single source basin and two transfer basins in the western North Atlantic, leaving eastern\-boundary upwelling systems, tropical basins, and marginal ice zones untested, with ice\-affected pixels masked rather than modeled; \(iii\) because the model predicts a one\-day tendency from observed high\-resolution history, multi\-day horizons would require auto\-regressive roll\-out whose stability remains untested, and few\-shot adaptation overfits by the 30\-day window, eroding rather than extending the zero\-shot advantage\.

The work following these results should continue by extending relative positions to the atmospheric stream, training across multiple basins or through meta\-learning to test whether generalization compounds, stabilizing auto\-regressive roll\-out for multi\-day forecasting, and calibrating probabilistic downscaling from the stochastic head under spread–skill evaluation\. Beyond methodology, coupling outputs to front\-aware fisheries management and coastal hazard forecasting would test whether preserved mesoscale structure yields decision\-relevant value, closing the loop between representation quality and the real\-world use of the proposal\.

Acknowledgments

This work was supported by the Natural Sciences and Engineering Research Council of Canada — NSERC, the Faculty of Computer Science of Dalhousie University, and the Conselho Nacional de Desenvolvimento Científico e Tecnológico — CNPq, Brasil\.

Data Availability

The ERA5 atmospheric reanalysis is publicly available from the Copernicus Climate Data Store, and the Multiscale Ultrahigh Resolution \(MUR\) sea\-surface temperature analysis from the NASA JPL Physical Oceanography Distributed Active Archive Center\. The code that reproduces pre\-processing, training, evaluation, and figure generation is available at an anonymized repository for peer review \([https://anonymous\.4open\.science/r/EddyFlow](https://anonymous.4open.science/r/EddyFlow)\)\. Because every input product is publicly available, all reported forecasting, skill\-score, and ablation results can be reproduced from open data\.

Responsible AI Use

Generative AI technologies supported initial manuscript drafting and the creation of preliminary code scaffolding\. The authors independently reviewed, verified, and revised all AI\-assisted material and retain full responsibility for the study’s methods, claims, analyses, interpretations, figures, and reported results\.

## References

- \(1\)
- Stock et al\.\(2015\)C\. Stock, K\. Pegion, G\. Vecchi, M\. Alexander, D\. Tommasi, N\. Bond, P\. Fratantoni, R\. Gudgel, T\. Kristiansen, T\. O’Brien, Y\. Xue, and X\. Yang\. 2015\.Seasonal Sea Surface Temperature Anomaly Prediction for Coastal Ecosystems\.*Prog\. Oceanogr\.*137 \(2015\), 219–236\.[doi:10\.1016/j\.pocean\.2015\.06\.007](https://doi.org/10.1016/j.pocean.2015.06.007)
- Xing et al\.\(2026\)Q\. Xing, Z\. Gao, S\. Ito, H\. Yu, W\. Yu, and X\. Chen\. 2026\.Human\-Induced Intensification of Sea Surface Temperature Regime Shifts Threatens Global Large Marine Ecosystems\.*Nat\. Commun\.*17 \(2026\), 4172\.[doi:10\.1038/s41467\-026\-70986\-z](https://doi.org/10.1038/s41467-026-70986-z)
- Saxena and Sharma \(2020\)N\. Saxena and A\. Sharma\. 2020\.Efficient Downscaling of Satellite Oceanographic Data with Convolutional Neural Networks\. In*Proc\. ACM SIGSPATIAL*\. 1–4\.[doi:10\.1145/3397536\.3429335](https://doi.org/10.1145/3397536.3429335)
- Thiria et al\.\(2023\)S\. Thiria, C\. Sorror, T\. Archambault, A\. Charantonis, D\. Béréziat, C\. Mejia, J\. Molines, and M\. Crépon\. 2023\.Downscaling of Ocean Fields by Fusion of Heterogeneous Observations Using Deep Learning Algorithms\.*Ocean Model\.*182 \(2023\), 102174\.[doi:10\.1016/j\.ocemod\.2023\.102174](https://doi.org/10.1016/j.ocemod.2023.102174)
- Kim et al\.\(2023\)J\. Kim, T\. Kim, and J\. Ryu\. 2023\.Multi\-Source Deep Data Fusion and Super\-Resolution for Downscaling Sea Surface Temperature Guided by Generative Adversarial Network\-Based Spatiotemporal Dependency Learning\.*Int\. J\. Appl\. Earth Obs\. Geoinf\.*119 \(2023\), 103312\.[doi:10\.1016/j\.jag\.2023\.103312](https://doi.org/10.1016/j.jag.2023.103312)
- Lavoie et al\.\(2021\)D\. Lavoie, N\. Lambert, M\. Starr, J\. Chassé, O\. Riche, Y\. Le Clainche, K\. Azetsu\-Scott, B\. Béjaoui, J\. Christian, and D\. Gilbert\. 2021\.The Gulf of St\. Lawrence Biogeochemical Model: A Modelling Tool for Fisheries and Ocean Management\.*Front\. Mar\. Sci\.*8 \(2021\), 732269\.[doi:10\.3389/fmars\.2021\.732269](https://doi.org/10.3389/fmars.2021.732269)
- Hersbach et al\.\(2020\)H\. Hersbach, B\. Bell, P\. Berrisford, S\. Hirahara, A\. Horányi, J\. Muñoz\-Sabater, J\. Nicolas, C\. Peubey, R\. Radu, D\. Schepers, A\. Simmons, C\. Soci, S\. Abdalla, X\. Abellan, G\. Balsamo, P\. Bechtold, G\. Biavati, J\. Bidlot, M\. Bonavita, G\. De Chiara, P\. Dahlgren, D\. Dee, M\. Diamantakis, R\. Dragani, J\. Flemming, R\. Forbes, M\. Fuentes, A\. Geer, L\. Haimberger, S\. Healy, R\. Hogan, E\. Hólm, M\. Janisková, S\. Keeley, P\. Laloyaux, P\. Lopez, C\. Lupu, G\. Radnoti, P\. de Rosnay, I\. Rozum, F\. Vamborg, S\. Villaume, and J\. Thépaut\. 2020\.The ERA5 Global Reanalysis\.*Q\. J\. R\. Meteorol\. Soc\.*146, 730 \(2020\), 1999–2049\.[doi:10\.1002/qj\.3803](https://doi.org/10.1002/qj.3803)
- Lupin\-Jimenez et al\.\(2025\)L\. Lupin\-Jimenez, M\. Darman, S\. Hazarika, T\. Wu, M\. Gray, R\. He, A\. Wong, and A\. Chattopadhyay\. 2025\.Simultaneous Emulation and Downscaling with Physically\-Consistent Deep Learning\-Based Regional Ocean Emulators\.*J\. Geophys\. Res\. Mach\. Learn\. Comput\.*2, 3 \(2025\), e2025JH000851\.[doi:10\.1029/2025JH000851](https://doi.org/10.1029/2025JH000851)
- Mardani et al\.\(2025\)M\. Mardani, N\. Brenowitz, Y\. Cohen, J\. Pathak, C\. Chen, C\. Liu, A\. Vahdat, M\. Nabian, T\. Ge, A\. Subramaniam, K\. Kashinath, J\. Kautz, and M\. Pritchard\. 2025\.Residual Corrective Diffusion Modeling for km\-scale Atmospheric Downscaling\.*Communications Earth & Environment*6, 1 \(2025\), 124\.[doi:10\.1038/s43247\-025\-02042\-5](https://doi.org/10.1038/s43247-025-02042-5)
- Wang et al\.\(2024\)S\. Wang, X\. Li, X\. Zhu, J\. Li, and S\. Guo\. 2024\.Spatial Downscaling of Sea Surface Temperature Using Diffusion Model\.*Remote Sens\.*16, 20 \(2024\), 3843\.[doi:10\.3390/rs16203843](https://doi.org/10.3390/rs16203843)
- Liu et al\.\(2025\)Y\. Liu, J\. Doss\-Gollin, Q\. Dai, A\. Veeraraghavan, and G\. Balakrishnan\. 2025\.Downscaling Extreme Precipitation with Wasserstein Regularized Diffusion\.*IEEE Trans\. Geosci\. Remote Sens\.*63 \(2025\)\.[doi:10\.1109/TGRS\.2025\.3611872](https://doi.org/10.1109/TGRS.2025.3611872)
- Hess et al\.\(2025\)P\. Hess, M\. Aich, B\. Pan, and N\. Boers\. 2025\.Fast, Scale\-Adaptive and Uncertainty\-Aware Downscaling of Earth System Model Fields with Generative Machine Learning\.*Nat\. Mach\. Intell\.*7, 3 \(2025\), 363–373\.[doi:10\.1038/s42256\-025\-00980\-5](https://doi.org/10.1038/s42256-025-00980-5)
- Lloyd et al\.\(2021\)D\. Lloyd, A\. Abela, R\. Farrugia, A\. Galea, and G\. Valentino\. 2021\.Optically Enhanced Super\-Resolution of Sea Surface Temperature Using Deep Learning\.*IEEE Trans\. Geosci\. Remote Sens\.*60 \(2021\), 1–14\.[doi:10\.1109/TGRS\.2021\.3094117](https://doi.org/10.1109/TGRS.2021.3094117)
- Ducournau and Fablet \(2016\)A\. Ducournau and R\. Fablet\. 2016\.Deep Learning for Ocean Remote Sensing: An Application of Convolutional Neural Networks for Super\-Resolution on Satellite\-Derived SST Data\. In*Proc\. IAPR Workshop Pattern Recognit\. Remote Sens\. \(PRRS\)*\. 1–6\.[doi:10\.1109/PRRS\.2016\.7867019](https://doi.org/10.1109/PRRS.2016.7867019)
- Fanelli et al\.\(2024\)C\. Fanelli, D\. Ciani, A\. Pisano, and B\. Buongiorno Nardelli\. 2024\.Deep Learning for the Super Resolution of Mediterranean Sea Surface Temperature Fields\.*Ocean Sci\.*20, 4 \(2024\), 1035–1050\.[doi:10\.5194/os\-20\-1035\-2024](https://doi.org/10.5194/os-20-1035-2024)
- Fanelli et al\.\(2025\)C\. Fanelli, T\. Li, L\. Biferale, B\. Buongiorno Nardelli, D\. Ciani, A\. Pisano, and M\. Buzzicotti\. 2025\.Super\-Resolution of Satellite\-Derived SST Data via Generative Adversarial Networks\.arXiv preprint\.arXiv:2511\.22610[doi:10\.48550/arXiv\.2511\.22610](https://doi.org/10.48550/arXiv.2511.22610)
- Krestenitis et al\.\(2024\)M\. Krestenitis, Y\. Androulidakis, and Y\. Krestenitis\. 2024\.Deep learning\-based forecasting of sea surface temperature in the interim future: application over the Aegean, Ionian, and Cretan Seas \(NE Mediterranean Sea\)\.*Ocean Dynamics*\.[doi:10\.1007/s10236\-023\-01595\-3](https://doi.org/10.1007/s10236-023-01595-3)
- Hao et al\.\(2024\)D\. Hao, L\. Famei, and W\. Guomei\. 2024\.Sea surface temperature prediction by stacked generalization ensemble of deep learning\.*Deep Sea Research Part I: Oceanographic Research Papers*209 \(2024\)\.[doi:10\.1016/j\.dsr\.2024\.104343](https://doi.org/10.1016/j.dsr.2024.104343)
- Cui et al\.\(2025\)Y\. Cui, R\. Wu, X\. Zhang, Z\. Zhu, B\. Liu, J\. Shi, J\. Chen, H\. Liu, S\. Zhou, L\. Su, Z\. Jing, H\. An, and L\. Wu\. 2025\.Forecasting the Eddying Ocean with a Deep Neural Network\.*Nat\. Commun\.*16 \(2025\), 2268\.[doi:10\.1038/s41467\-025\-57389\-2](https://doi.org/10.1038/s41467-025-57389-2)
- Lam et al\.\(2023\)R\. Lam, A\. Sanchez\-Gonzalez, M\. Willson, P\. Wirnsberger, M\. Fortunato, F\. Alet, S\. Ravuri, T\. Ewalds, Z\. Eaton\-Rosen, W\. Hu, A\. Merose, S\. Hoyer, G\. Holland, O\. Vinyals, J\. Stott, A\. Pritzel, S\. Mohamed, and P\. Battaglia\. 2023\.Learning Skillful Medium\-Range Global Weather Forecasting\.*Science*382, 6677 \(2023\), 1416–1421\.[doi:10\.1126/science\.adi2336](https://doi.org/10.1126/science.adi2336)
- Price et al\.\(2024\)I\. Price, A\. Sanchez\-Gonzalez, F\. Alet, T\. Andersson, A\. El\-Kadi, D\. Masters, T\. Ewalds, J\. Stott, S\. Mohamed, P\. Battaglia, R\. Lam, and M\. Willson\. 2024\.Probabilistic Weather Forecasting with Machine Learning\.*Nature*637 \(2024\), 84–90\.[doi:10\.1038/s41586\-024\-08252\-9](https://doi.org/10.1038/s41586-024-08252-9)
- Niu et al\.\(2026\)Y\. Niu, Q\. Huang, X\. Zhong, A\. Guo, L\. Chen, D\. Zhang, Z\. Zhou, X\. Jia, L\. Wu, R\. Zhang, J\. Qi, X\. Zhang, and H\. Li\. 2026\.A Deep Learning Global Ocean Forecasting Model with Sub\-Daily and Eddy\-Resolving Resolution\.*npj Clim Atmos Sci*\(2026\)\.[doi:10\.1038/s41612\-026\-01444\-2](https://doi.org/10.1038/s41612-026-01444-2)
- Michel et al\.\(2022\)M\. Michel, S\. Obakrim, N\. Raillard, P\. Ailliot, and V\. Monbet\. 2022\.Deep Learning for Statistical Downscaling of Sea States\.*Adv\. Stat\. Climatol\. Meteorol\. Oceanogr\.*8, 1 \(2022\), 83–95\.[doi:10\.5194/ascmo\-8\-83\-2022](https://doi.org/10.5194/ascmo-8-83-2022)
- Ho et al\.\(2020\)J\. Ho, A\. Jain, and P\. Abbeel\. 2020\.Denoising Diffusion Probabilistic Models\. In*Advances in Neural Information Processing Systems*, Vol\. 33\. 6840–6851\.
- Karras et al\.\(2022\)T\. Karras, M\. Aittala, T\. Aila, and S\. Laine\. 2022\.Elucidating the Design Space of Diffusion\-Based Generative Models\.35 \(2022\), 26565–26577\.
- Bulgin et al\.\(2020\)C\. Bulgin, C\. Merchant, and D\. Ferreira\. 2020\.Tendencies, Variability and Persistence of Sea Surface Temperature Anomalies\.*Sci\. Rep\.*10 \(2020\), 7986\.[doi:10\.1038/s41598\-020\-64785\-9](https://doi.org/10.1038/s41598-020-64785-9)
- Knaff and Landsea \(1997\)J\. Knaff and C\. Landsea\. 1997\.An El Niño–Southern Oscillation Climatology and Persistence \(CLIPER\) Forecasting Scheme\.*Weather Forecast\.*12, 3 \(1997\), 633–652\.[doi:10\.1175/1520\-0434\(1997\)012<0633:AENOSO\>2\.0\.CO;2](https://doi.org/10.1175/1520-0434(1997)012%3C0633:AENOSO%3E2.0.CO;2)
- Chelton et al\.\(2011\)D\. Chelton, M\. Schlax, and R\. Samelson\. 2011\.Global Observations of Nonlinear Mesoscale Eddies\.*Prog\. Oceanogr\.*91, 2 \(2011\), 167–216\.[doi:10\.1016/j\.pocean\.2011\.01\.002](https://doi.org/10.1016/j.pocean.2011.01.002)
- Dosovitskiy et al\.\(2021\)A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. Houlsby\. 2021\.An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale\. In*Proc\. Int\. Conf\. Learn\. Represent\. \(ICLR\)*\.
- Bertasius et al\.\(2021\)G\. Bertasius, H\. Wang, and L\. Torresani\. 2021\.Is Space\-Time Attention All You Need for Video Understanding?\. In*Proc\. Int\. Conf\. Mach\. Learn\. \(ICML\)*\. 813–824\.
- Su et al\.\(2024\)J\. Su, M\. Ahmed, Y\. Lu, S\. Pan, W\. Bo, and Y\. Liu\. 2024\.RoFormer: Enhanced Transformer with Rotary Position Embedding\.*Neurocomputing*568 \(2024\), 127063\.[doi:10\.1016/j\.neucom\.2023\.127063](https://doi.org/10.1016/j.neucom.2023.127063)
- Heo et al\.\(2024\)B\. Heo, S\. Park, D\. Han, and S\. Yun\. 2024\.Rotary Position Embedding for Vision Transformer\. In*Proc\. Eur\. Conf\. Comput\. Vis\. \(ECCV\)*\. 289–305\.[doi:10\.1007/978\-3\-031\-72684\-2\_17](https://doi.org/10.1007/978-3-031-72684-2_17)
- Paszke et al\.\(2019\)A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga, A\. Desmaison, A\. Köpf, E\. Yang, Z\. DeVito, M\. Raison, A\. Tejani, S\. Chilamkurthy, B\. Steiner, L\. Fang, J\. Bai, and S\. Chintala\. 2019\.PyTorch: An Imperative Style, High\-Performance Deep Learning Library\.*NeurIPS*32 \(2019\), 8024–8035\.
- JPL MUR MEaSUREs Project \(2015\)JPL MUR MEaSUREs Project\. 2015\.GHRSST Level 4 MUR Global Foundation Sea Surface Temperature Analysis \(v4\.1\)\.NASA Physical Oceanography Distributed Active Archive Center \(PO\.DAAC\)\.[doi:10\.5067/GHGMR\-4FJ04](https://doi.org/10.5067/GHGMR-4FJ04)
- Hirahara and Hersbach \(2016\)Shoji Hirahara and Hans Hersbach\. 2016\.Sea Surface Temperature and Sea Ice Concentration for ERA 5\.[https://api\.semanticscholar\.org/CorpusID:201692081](https://api.semanticscholar.org/CorpusID:201692081)
- Song et al\.\(2021\)J\. Song, C\. Meng, and S\. Ermon\. 2021\.Denoising Diffusion Implicit Models\. In*Proc\. Int\. Conf\. Learn\. Represent\. \(ICLR\)*\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. Hutter\. 2019\.Decoupled Weight Decay Regularization\. In*Proc\. Int\. Conf\. Learn\. Represent\. \(ICLR\)*\.

Appendix

## Appendix AList of Acronyms

Table 3\.Acronyms used throughout the paper\.
## Appendix BList of Notations

Table 4\.Mathematical notation used throughout the paper\.
## Appendix CDomain Characterization and OOD Analysis

This appendix characterizes the three ocean domains used in theEddyFlowevaluation: GSL \(training domain\), BOF, and GOM\. It summarizes grid specifications, target SST distributions, ERA5 atmospheric input statistics, bathymetric properties, the shared normalization pipeline, and the main out\-of\-distribution \(OOD\) differences between the source domain and the two transfer domains\.

### C\.1\.Domain overview

GSL covers 44\.0–52\.8∘N and 69\.5–56\.0∘W and is the only Phase 1 training domain\. It is a sub\-polar estuary with a large seasonal SST cycle, winter sea ice, strong freshwater influence, and mixed shelf/deep\-channel bathymetry\(Lavoie et al\.,[2021](https://arxiv.org/html/2608.04230#bib.bib7)\)\. BOF covers 44\.0–47\.0∘N and 67\.0–63\.0∘W and is a zero\-shot and few\-shot evaluation domain\. It overlaps GSL in latitude and has a very similar SST range, but its tidal forcing is much stronger\. GOM covers 18\.0–31\.0∘N and 98\.0–80\.0∘W and is likewise used for zero\-shot and few\-shot evaluation\. It is the most different domain, with much warmer SST, no ice, and circulation dominated by the Loop Current and warm\-core eddies\.

### C\.2\.Grid specifications

The MUR SST target is evaluated on a 0\.01∘grid, which corresponds to roughly 1 km resolution\. The fine\-grid sizes are 500×\\times1200 for GSL, 301×\\times401 for BOF, and 1301×\\times1801 for GOM\. ERA5 is provided at 0\.25∘resolution and is used as the coarse atmospheric input\. The resulting downscaling factor is approximately 25×\\timesin each spatial direction for GSL and GOM, and about 23–24×\\timesfor BOF\.

### C\.3\.Target SST statistics

Table[5](https://arxiv.org/html/2608.04230#A3.T5)summarizes the MUR SST target distribution for the 2022–2023 test period\. GSL and BOF have similar SST means and ranges, whereas GOM is shifted to much higher temperatures and lies mostly outside the GSL training SST distribution\.

Table 5\.Geography, MUR SST target statistics over the 2022–2023 test period, bathymetry, and OOD indicators for the three ocean basins; BOF and GOM SST values are converted to∘C before statistics are computed\. The indicators quantify the transfer\-difficulty ordering used in the analysis; BOF is a mild OOD case \(full SST overlap with the GSL training range, no shifted ERA5 channels\), while GOM is a strong one \(26% overlap, 14 of 21 channels shifted beyond 0\.5σ\\sigma\)\.The shared SST normalization uses the GSL training mean and standard deviation,μ=5\.9023∘​C\\mu=5\.9023\\,^\{\\circ\}\\mathrm\{C\}andσ=6\.3153∘​C\\sigma=6\.3153\\,^\{\\circ\}\\mathrm\{C\}, for all domains\. Under this normalization, BOF lies almost entirely within the GSL training SST range, whereas GOM lies predominantly above it\.

### C\.4\.ERA5 input statistics

ERA5 atmospheric channels are normalized using per\-channel mean and standard deviation computed on the GSL training period\. The same statistics are applied to BOF and GOM without domain\-specific rescaling\. Table[6](https://arxiv.org/html/2608.04230#A3.T6)shows the GSL training statistics together with the corresponding test\-period percentiles\. The GSL training and test distributions are close, suggesting that degradation on BOF and GOM is not due to a shift within the training domain\.

Table 6\.ERA5 per\-channel normalization statistics computed on the GSL training period \(2013–2020\), applied unchanged to BOF and GOM, alongside GSL test\-period 1st and 99th percentiles\. The proximity of train and test distributions indicates that transfer degradation reflects cross\-basin shift rather than temporal drift within the source domain\.
### C\.5\.Bathymetry

Bathymetry is a static auxiliary input coarsened to the ERA5 grid, transformed withlog⁡\(1\+x\)\\log\(1\+x\), and normalized with GSL training statistics before encoder input\. The GSL bathymetry contains the Laurentian Channel and a mixed shelf/deep\-channel structure; BOF is much shallower, and GOM contains a much deeper open\-ocean basin\. These differences matter because the model uses bathymetry as a spatial prior for where fine\-scale SST structure appears\.

### C\.6\.Normalization pipeline

All three domains use the same normalization statistics from the GSL 2013–2020 training period\. No domain\-specific rescaling is applied in the encoder or baseline decoder\. BOF and GOM are converted from Kelvin to degrees Celsius before the shared SST normalization is applied\. Where domain conditioning is enabled \(Section[4\.9](https://arxiv.org/html/2608.04230#S4.SS9); Stage 5 in Table[1](https://arxiv.org/html/2608.04230#S5.T1)\), the diffusion head additionally injects the target\-domain SST mean and standard deviation\.

The SST target is converted to a z\-score using the GSL training mean and standard deviation\. Physical RMSE is recovered from normalized RMSE by multiplying byσ=6\.3153​C∘\\sigma=6\.3153\\,\{\{\}^\{\\circ\}C\}\. Because the models predict SST changes rather than absolute SST levels, the constant GOM offset does not directly inflate the error; the day\-to\-day target is centered near zero in all domains\.

### C\.7\.OOD analysis

The OOD gap is small for BOF and large for GOM\. On SST, BOF sits inside the GSL training range, while GOM lies mostly above it\. On ERA5 forcing, BOF stays close to the GSL training distribution, whereas GOM shifts across many channels, especially moisture, temperature, and wind structure\. In bathymetry, BOF is shallow and coastal, whereas GOM features a deeper open\-ocean regime\.

Table[5](https://arxiv.org/html/2608.04230#A3.T5)summarizes the main domain\-level differences\. BOF is a mild OOD case, changing the local tidal and bathymetric context but not the broad SST scale\. GOM is a strong OOD case, changing SST level, atmospheric forcing regime, and circulation structure\.

## Appendix DAblations

We report two ablations: a hyperparameter and seed sweep summarized through the few\-shot transfer experiments, and an input ablation that removes the MUR history stream \(Appendix[D\.1](https://arxiv.org/html/2608.04230#A4.SS1)\)\. The few\-shot sweep covers all six stages, both transfer domains \(BOF and GOM\), three target\-data budgets \(1\-, 7\-, and 30\-day windows\), and three random seeds \(0, 42, 123\), for 108 configurations; Tables[7](https://arxiv.org/html/2608.04230#A4.T7)and[8](https://arxiv.org/html/2608.04230#A4.T8)report the best and mean result per stage and budget\.

Table 7\.Few\-shot transfer results on the BOF across 1\-, 7\-, and 30\-day adaptation budgets, reporting RMSE \(∘C\), skill score, and PSD ratio\. Best and Mean are computed over three random seeds \(0, 42, 123\): Best is the minimum for RMSE, the maximum for Skill, and the value closest to 1\.0 for PSD\. The ERA5 persistence baseline RMSE on BOF is≈\\approx4\.67∘C \(Table[5](https://arxiv.org/html/2608.04230#A3.T5)\)\. Stage 5 \(EDM diffusion\) shows high seed variance at the 7\-day budget \(σRMSE=0\.073\\sigma\_\{\\text\{RMSE\}\}=0\.073∘C\) due to bimodal convergence dynamics\.Table 8\.Few\-shot transfer results on the GOM under the same protocol and seed conventions as Table[7](https://arxiv.org/html/2608.04230#A4.T7)\. RMSE is reported in normalized units\. Stage 4 leads at the 1\- and 7\-day budgets, while by 30 days Stage 3 attains the best RMSE and skill and Stage 5 reaches spectral parity\.Table 9\.Percentage deltas for the no\-MUR ablation, in which the 60\-day MUR history is replaced by the z\-scored climatological average; positiveΔ\\DeltaRMSE andΔ\\DeltaSkill indicate degradation when MUR is removed, andΔ\\DeltaPSD reports its signed relative change \(positive when removal raises it\)\. Results average seeds\{0,42,123\}\\\{0,42,123\\\}; zero\-shot uses a single checkpoint\. Deltas are small almost everywhere, isolating Stage 5 on GOM at zero shot \(\+34\.2% RMSE, bold\) as the only configuration with MUR dependence\.### D\.1\.Evaluations Without MUR History

We assess how much the ocean stream depends on real target\-domain MUR SST history by comparing the standard setup \(with MUR\) to a no\-MUR variant in which the 60\-day MUR input is replaced by the z\-scored climatological average\. All other factors \(checkpoints, evaluation dates, seeds\) are held fixed\. Table[9](https://arxiv.org/html/2608.04230#A4.T9)reports percentage deltas, oriented so that positiveΔ\\DeltaRMSE andΔ\\DeltaSkill always indicate degradation when MUR is removed:

ΔRMSE\(%\)=100×RMSEno−RMSEwithRMSEwith,\\displaystyle\\Delta\\mathrm\{RMSE\}\\ \(\\%\)=100\\times\\frac\{\\mathrm\{RMSE\}\_\{\\text\{no\}\}\-\\mathrm\{RMSE\}\_\{\\text\{with\}\}\}\{\\mathrm\{RMSE\}\_\{\\text\{with\}\}\},ΔSkill\(%\)=100×Skillwith−SkillnoSkillwith,\\displaystyle\\Delta\\mathrm\{Skill\}\\ \(\\%\)=100\\times\\frac\{\\mathrm\{Skill\}\_\{\\text\{with\}\}\-\\mathrm\{Skill\}\_\{\\text\{no\}\}\}\{\\mathrm\{Skill\}\_\{\\text\{with\}\}\},whileΔPSD\(%\)=100×\(PSDRno−PSDRwith\)/PSDRwith\\Delta\\mathrm\{PSD\}\\ \(\\%\)=100\\times\(\\mathrm\{PSDR\}\_\{\\text\{no\}\}\-\\mathrm\{PSDR\}\_\{\\text\{with\}\}\)/\\mathrm\{PSDR\}\_\{\\text\{with\}\}reports its signed relative change, positive when removal raises it\.

Across most stages, domains, and shot counts, these percentage deltas are small\. RMSE changes are typically under 3%, and skill changes lie within±1\\pm 1% for the majority of configurations\. PSD ratios remain close to their original values, with most\|Δ​PSD\|\|\\Delta\\text\{PSD\}\|below 2%, indicating that spectral fidelity is largely preserved without real MUR history\. Stage 6 is the most robust as RMSE and skill deltas are consistently near zero or slightly favorable \(up to≈1\\approx 1% improvement in skill\), and PSD changes are negligible\. Stage 4 shows mild sensitivity on BOF atn=1n=1\(around 2–3% skill drop\) but is essentially neutral or slightly beneficial on GOM, especially atn=30n=30, where skill improves by about 3% without MUR\.

The clearest case of genuine MUR dependence appears forStage 5 on GOM in the zero\-shot setting, where removing MUR increases RMSE by roughly 34% and reduces skill by about 8%\. This is the largest degradation in the ablation and stands out against the otherwise mild pattern\. In few\-shot, Stage 5 on GOM recovers byn=7n=7–3030, RMSE and skill deltas shrink to within±3\\pm 3%, and PSD changes remain small\. Few\-shot adaptation thus mitigates most sensitivity asnnincreases from 1 to 30, skill and RMSE deltas contract, and PSD ratios stay stable\. Overall, real target\-domain MUR history is not uniformly essential; the framework is robust to its removal, with Stage 5 on GOM \(zero\-shot\) as the primary exception\.

Similar Articles

Domain-Adaptive Climate Downscaling Under Temporal Distribution Shift

arXiv cs.LG

This paper investigates temporal out-of-distribution shift in deep-learning-based climate downscaling and proposes a domain-adaptive framework that combines supervised reconstruction with domain alignment to improve high-resolution climate projections under non-stationary conditions.