SolarFlowRefiner: 感知细化流匹配用于地表太阳辐射降尺度
摘要
论文介绍了SolarFlowRefiner,一个感知细化的流匹配框架,用于将粗糙的ERA5数据降尺度为高分辨率的SolarCube场,展示了相对于独立生成和事后细化方法的一致改进。
arXiv:2609.22126v1 Announce Type: new
Abstract: High-resolution surface solar radiation (SSR) is important for solar forecasting and grid operation. However, physically consistent reanalysis products are too coarse to resolve localized cloud-driven variability. In this paper, we study a multisource downscaling task that reconstructs high-resolution SolarCube SSR fields from coarse ERA5 radiative variables and co-registered satellite channels. The task is challenging because a single ERA5 grid cell may contain both sunlit and cloud-shadowed regions. As a result, the missing high-resolution correction can be spatially sharp and inherently ambiguous. One-stage predictors often oversmooth these structures. Post-hoc refinement also introduces a stage-wise mismatch: the generator is optimized independently, even though its output determines the refiner's initial state. We introduce SolarFlowRefiner, a refinement-aware flow-matching framework for SSR downscaling. A conditional FlowMatch generator first predicts a normalized correction to an upsampled ERA5 baseline. The refiner is then trained on prediction-conditioned states between the current FlowMatch output and the target residual. This exposes the refiner to the structured errors produced by the generator. The refinement objective is also backpropagated through the FlowMatch sampler, allowing generation and correction to be jointly optimized for the final reconstruction. Experiments on a day-blocked ERA5--SolarCube benchmark show consistent improvements over standalone generation and post-hoc refinement. More broadly, SolarFlowRefiner provides a general strategy for coupling generative predictors with iterative correctors.
查看缓存全文
缓存时间: 2026/09/22 09:12
# Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling
Source: [https://arxiv.org/html/2609.22126](https://arxiv.org/html/2609.22126)
Antonita RachealYiheng ChenRunlong Yu\\correspondingXinyue Ye\\corresponding
###### Abstract
High\-resolution surface solar radiation \(SSR\) is important for solar forecasting and grid operation\. However, physically consistent reanalysis products are too coarse to resolve localized cloud\-driven variability\. In this paper, we study a multisource downscaling task that reconstructs high\-resolution SolarCube SSR fields from coarse ERA5 radiative variables and co\-registered satellite channels\. The task is challenging because a single ERA5 grid cell may contain both sunlit and cloud\-shadowed regions\. As a result, the missing high\-resolution correction can be spatially sharp and inherently ambiguous\. One\-stage predictors often oversmooth these structures\. Post\-hoc refinement also introduces a stage\-wise mismatch: the generator is optimized independently, even though its output determines the refiner’s initial state\. We introduce SolarFlowRefiner, a refinement\-aware flow\-matching framework for SSR downscaling\. A conditional FlowMatch generator first predicts a normalized correction to an upsampled ERA5 baseline\. The refiner is then trained on prediction\-conditioned states between the current FlowMatch output and the target residual\. This exposes the refiner to the structured errors produced by the generator\. The refinement objective is also backpropagated through the FlowMatch sampler, allowing generation and correction to be jointly optimized for the final reconstruction\. Experiments on a day\-blocked ERA5–SolarCube benchmark show consistent improvements over standalone generation and post\-hoc refinement\. More broadly, SolarFlowRefiner provides a general strategy for coupling generative predictors with iterative correctors\.
## 1Introduction
Solar radiation reaching the surface is quantified as downward shortwave irradiance\([22](https://arxiv.org/html/2609.22126#bib.bib18)\)\. This surface flux, composed of direct\-beam and diffuse sky components and strongly modulated by clouds and aerosols, is the primary driver of photovoltaic generation\([1](https://arxiv.org/html/2609.22126#bib.bib17)\)\. Ground pyranometers provide accurate point measurements but sparse spatial coverage; satellite retrievals provide spatially continuous estimates but rely on indirect observations; and atmospheric reanalysis provides physically consistent, gap\-free fields at coarse spatial resolution\([1](https://arxiv.org/html/2609.22126#bib.bib17);[13](https://arxiv.org/html/2609.22126#bib.bib14)\)\. In this work we use two complementary sources at opposite ends of this resolution spectrum\. ERA5 provides model\-derived radiative fields, including surface downward shortwave radiation, on a coarse global grid\([7](https://arxiv.org/html/2609.22126#bib.bib11)\)\. SolarCube provides high\-resolution surface solar radiation retrieved from geostationary satellite observations, together with visible, infrared, and cloud\-related channels that resolve fine\-scale cloud structure\([21](https://arxiv.org/html/2609.22126#bib.bib16);[13](https://arxiv.org/html/2609.22126#bib.bib14)\)\. Our task links these sources: recovering the fine\-scale surface solar radiation \(SSR\) field observed by SolarCube from coarse ERA5 radiative inputs\.
Accurate, spatially detailed radiation fields are increasingly important as solar generation is integrated into electricity grids at scale, because the combined output of distributed solar installations depends on radiation variations at spatial scales that coarse products cannot resolve\([1](https://arxiv.org/html/2609.22126#bib.bib17)\)\. Passing clouds create sharp, localized swings in surface irradiance, and these fine\-scale fluctuations drive the sudden ramps in power output that are difficult for grid operators to balance\. Yet the gap\-free, physically consistent estimates offered by reanalysis come at a coarse resolution, roughly0\.25∘0\.25^\{\\circ\}for ERA5\([7](https://arxiv.org/html/2609.22126#bib.bib11)\), at which a single grid cell averages over both sunlit and cloud\-shadowed ground\. The fine\-scale structure that governs local solar variability is therefore lost\. Reconstructing high\-resolution radiation fields from coarse inputs, i\.e\., spatial super\-resolution, is an important problem for solar forecasting, grid management, and climate applications\([23](https://arxiv.org/html/2609.22126#bib.bib12);[13](https://arxiv.org/html/2609.22126#bib.bib14)\)\.
Surface solar radiation is dominated by cloudcover, whose spatial structure is sharp, intermittent, and fast\-moving, and only weakly constrained by coarse reanalysis fields\([13](https://arxiv.org/html/2609.22126#bib.bib14);[21](https://arxiv.org/html/2609.22126#bib.bib16)\)\. Whereas clear\-sky irradiance varies smoothly with solar geometry, cloud cover imposes abrupt, high\-contrast gradients that shift on short timescales\. A single coarse ERA5 cell aggregates many distinct cloud and clear\-sky conditions, so many high\-resolution radiation fields may be consistent with the same coarse observation\. This one\-to\-many ambiguity makes solar super\-resolution a severely ill\-posed inverse problem\. It also explains a common limitation of deterministic models: pixel\-wise training tends to regress toward a conditional mean over plausible solutions, producing over\-smoothed fields that miss sharp radiation gradients near cloud edges\([2](https://arxiv.org/html/2609.22126#bib.bib15);[4](https://arxiv.org/html/2609.22126#bib.bib1)\)\. Recovering this structure requires a model that can represent residual uncertainty and a conditioning signal that carries the missing cloud information\.
Recent super\-resolution and generative restoration methods provide the starting point for this work\. Deep convolutional and transformer models have become strong deterministic predictors for natural\-image restoration\([15](https://arxiv.org/html/2609.22126#bib.bib4);[29](https://arxiv.org/html/2609.22126#bib.bib5);[14](https://arxiv.org/html/2609.22126#bib.bib8);[3](https://arxiv.org/html/2609.22126#bib.bib9)\), while generative methods such as residual\-shift\([27](https://arxiv.org/html/2609.22126#bib.bib10)\)models reconstruct high\-resolution images through iterative correction in a structured residual space\. More recently, flow matching\([16](https://arxiv.org/html/2609.22126#bib.bib19)\)has emerged as an efficient, simulation\-free generative formulation, and iterative refinement methods\([17](https://arxiv.org/html/2609.22126#bib.bib21);[5](https://arxiv.org/html/2609.22126#bib.bib22)\)improve an initial prediction through repeated correction rather than regenerating it from scratch; together these motivate the generation\-and\-refinement approach we adopt\. However, using generation and refinement as separately trained stages introduces a mismatch: the generator is optimized to produce the best standalone residual estimate, even though this estimate later serves as the initial state for a multi\-step correction process\. The best standalone prediction is therefore not necessarily the best starting point for refinement\. This motivates refinement\-aware generation, where the generator is trained not only to be accurate, but also to produce residual states that are easy for the refiner to improve\.
We propose SolarFlowRefiner, a refinement\-aware FlowMatch framework for ERA5\-to\-SolarCube SSR downscaling\. FlowMatch first generates an initial normalized residual conditioned on ERA5 radiative variables and SolarCube auxiliary channels\. A refiner then iteratively corrects the residual\. Our main model, SolarFlowRefiner, trains the FlowMatch generator and PDE\-style refiner end to end by backpropagating the refinement loss through the FlowMatch sampler\. In this way, the generator learns to be refinable\. We compare post\-hoc refinement variants that keep the base generator fixed: FlowRefiner\-ODE learns a continuous correction trajectory, while FlowRefiner\-PDE uses the same multilevel denoising refiner as the proposed model but blocks gradients into the FlowMatch generator\.
Our contributions are summarized as follows:
- •We formulate multisource SSR downscaling as residual reconstruction around a physically meaningful ERA5 baseline, focusing learning on fine\-scale cloud\-driven corrections\.
- •We introduce a prediction\-conditioned multilevel refinement scheme that trains the corrector on structured errors produced by the current generator\.
- •We develop a refinement\-aware end\-to\-end objective that jointly optimizes generation and correction for the final SSR reconstruction, and validate its effectiveness against standalone and post\-hoc refinement baselines\.
## 2Related Work
We review geophysical super\-resolution and generative refinement, highlighting the gap between post\-hoc correction and jointly optimized generation–refinement\.
### 2\.1Geophysical Super\-Resolution
Deep super\-resolution has progressed from convolutional models such as SRCNN and VDSR\([6](https://arxiv.org/html/2609.22126#bib.bib2);[9](https://arxiv.org/html/2609.22126#bib.bib3)\)to residual and attention\-based architectures including EDSR, RCAN, SwinIR, and HAT\([15](https://arxiv.org/html/2609.22126#bib.bib4);[29](https://arxiv.org/html/2609.22126#bib.bib5);[14](https://arxiv.org/html/2609.22126#bib.bib8);[3](https://arxiv.org/html/2609.22126#bib.bib9)\)\. Adversarial methods such as SRGAN and ESRGAN further target perceptual realism, while exposing the perception–distortion tradeoff\([11](https://arxiv.org/html/2609.22126#bib.bib6);[24](https://arxiv.org/html/2609.22126#bib.bib7);[2](https://arxiv.org/html/2609.22126#bib.bib15)\)\. These methods provide strong baselines, but their one\-stage predictions can oversmooth localized structures when multiple fine\-scale fields are compatible with the same coarse input\.
Learning\-based downscaling has also been applied to climatological wind and solar fields\([23](https://arxiv.org/html/2609.22126#bib.bib12)\), with WiSoSuper providing a benchmark for renewable\-energy super\-resolution\([10](https://arxiv.org/html/2609.22126#bib.bib13)\)\. SolarCube supplies co\-registered satellite observations and high\-resolution SSR across multiple sites\([13](https://arxiv.org/html/2609.22126#bib.bib14)\)\. We use ERA5 as a coarse physical baseline and SolarCube channels as cloud\-scale context, predicting and refining the normalized correction to the ERA5 field rather than directly regressing the full SSR field\.
### 2\.2Generative Refinement
Generative super\-resolution represents the ambiguity of recovering fine\-scale fields from coarse observations\. ResShift constructs a residual\-shifting process between degraded and high\-resolution images\([27](https://arxiv.org/html/2609.22126#bib.bib10)\), while normalizing\-flow and rectified\-flow methods learn conditional transport paths for restoration\([19](https://arxiv.org/html/2609.22126#bib.bib25);[16](https://arxiv.org/html/2609.22126#bib.bib19);[18](https://arxiv.org/html/2609.22126#bib.bib20);[26](https://arxiv.org/html/2609.22126#bib.bib23);[12](https://arxiv.org/html/2609.22126#bib.bib24);[30](https://arxiv.org/html/2609.22126#bib.bib26);[20](https://arxiv.org/html/2609.22126#bib.bib27)\)\. These approaches motivate our use of FlowMatch to generate an initial cloud\-conditioned SSR residual\.
Iterative methods instead improve an existing prediction\. PDE\-Refiner uses multilevel denoising to recover spatial\-frequency content missed by one\-shot predictors\([17](https://arxiv.org/html/2609.22126#bib.bib21)\), whereas FlowRefiner learns a continuous ODE\-style correction trajectory toward the target\([5](https://arxiv.org/html/2609.22126#bib.bib22)\)\. Both naturally support post\-hoc correction of a fixed prediction\. SolarFlowRefiner differs by constructing refinement states from the current FlowMatch output and backpropagating the refinement objective through the sampler, thereby optimizing the generator and refiner jointly for the final SSR reconstruction rather than as independent stages\.
Figure 1:Overview of SolarFlowRefiner\. ERA5 radiative variables and SolarCube auxiliary channels condition a FlowMatch residual generator, which predicts an initial normalized residual relative to the upsampled ERA5 SSR baseline\. A PDE\-style refiner then performs iterative correction updates overK=8K=8refinement levels before reconstructing the high\-resolution SSR field\. In SolarFlowRefiner, the refinement loss is backpropagated through both the refiner and the FlowMatch sampler, making generation refinement\-aware\.
## 3Problem Setup
We study supervised downscaling from coarse ERA5 radiative fields to high\-resolution SolarCube SSR\. Each sample contains conditioning inputsx∈ℝC×120×120x\\in\\mathbb\{R\}^\{C\\times 120\\times 120\}and a target SSR fieldy∈ℝ1×120×120y\\in\\mathbb\{R\}^\{1\\times 120\\times 120\}\. ERA5 variables are upsampled to the SolarCube grid\. Letbbdenote the upsampled ERA5 SSR baseline\. Rather than predictingyydirectly, we define the residual
We normalize this residual using training\-set statistics,
r~=r−μrσr,\\tilde\{r\}=\\frac\{r\-\\mu\_\{r\}\}\{\\sigma\_\{r\}\},\(2\)and reconstruct physical SSR from a predicted normalized residualr~^\\hat\{\\tilde\{r\}\}as
y^=b\+σrr~^\+μr\.\\hat\{y\}=b\+\\sigma\_\{r\}\\hat\{\\tilde\{r\}\}\+\\mu\_\{r\}\.\(3\)This formulation preserves the coarse physical estimate from ERA5 and focuses learning on high\-resolution cloud\-driven corrections\.
The conditioning variables include five ERA5 radiative channels:era\_ssrd,era\_ssr,era\_ssrdc,era\_fdir, andera\_cdir\. SolarCube auxiliary channels includevis047,vis086,ir133,sza, andcm\. Together, these channels provide both coarse radiative context and high\-resolution cloud information\. For input\-channel ablations, the ERA5 SSR baselinebbremains the residual reference in all settings, while the ablation changes only which channels are provided to the learned generator and refiner\.
We use the split 70\-10\-20 among train\-val\-test respectively\. Tiles 1–10, 12, and 14 pass the target\-valid quality\-control threshold and are used for training and evaluation\. Tiles 11, 13, and 15–19 are excluded because their target\-valid fraction is below 50%\. The split contains 32,159 training samples, 4,519 validation samples, and 8,994 test samples\. Samples are assigned by whole UTC\-day blocks within each tile/month, so all hours from the same tile\-month\-day remain in the same partition\. This reduces leakage from near\-duplicate cloud scenes across train, validation, and test sets\.
We report MAE and RMSE in physical irradiance units, SSIM\([25](https://arxiv.org/html/2609.22126#bib.bib28)\)for structural agreement, and LPIPS\([28](https://arxiv.org/html/2609.22126#bib.bib29)\)and FID\([8](https://arxiv.org/html/2609.22126#bib.bib30)\)for perceptual and distributional fidelity\.
## 4Method
SolarFlowRefiner is a learned predictor–corrector framework for high\-resolution surface solar radiation \(SSR\) reconstruction\. The predictor is a conditional FlowMatch model that generates an initial estimate of the normalized SSR residual\. The corrector is a multilevel denoising refiner that progressively removes the structured errors remaining in that estimate\. The central design choice is to optimize these two components jointly: the refinement objective is differentiated through the FlowMatch sampler, allowing the generator to adapt to the downstream correction process rather than being trained only as an independent predictor\.
Letxxdenote the conditioning variables,r~\\tilde\{r\}the normalized target residual, andbbthe upsampled ERA5 SSR baseline defined in Section[3](https://arxiv.org/html/2609.22126#S3)\. SolarFlowRefiner first generates a base residual estimater~b\\tilde\{r\}\_\{b\}and then appliesKKrefinement updates to obtain the final normalized residualr~^\\hat\{\\tilde\{r\}\}\. The physical SSR prediction is reconstructed as
y^=b\+σrr~^\+μr\.\\hat\{y\}=b\+\\sigma\_\{r\}\\hat\{\\tilde\{r\}\}\+\\mu\_\{r\}\.\(4\)
### 4\.1Conditional FlowMatch Predictor
The base predictor models the conditional distribution of normalized SSR residuals using rectified flow matching\. Letz0∼𝒩\(0,I\)z\_\{0\}\\sim\\mathcal\{N\}\(0,I\)denote a Gaussian source sample and letz1=r~z\_\{1\}=\\tilde\{r\}denote the target residual\. For a continuous flow timet∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\), we define the linear conditional probability path
zt=\(1−t\)z0\+tz1\.z\_\{t\}=\(1\-t\)z\_\{0\}\+tz\_\{1\}\.\(5\)Along this path, the target transport velocity is constant:
v⋆\(zt,z0,z1\)=z1−z0\.v^\{\\star\}\(z\_\{t\},z\_\{0\},z\_\{1\}\)=z\_\{1\}\-z\_\{0\}\.\(6\)A conditional velocity networkfθf\_\{\\theta\}receives the current stateztz\_\{t\}, the conditioning variablesxx, and the flow timett\. It is trained using
ℒFM\(θ\)=𝔼x,r~,z0,t\[‖fθ\(zt,x,t\)−\(r~−z0\)‖1\]\.\\mathcal\{L\}\_\{\\mathrm\{FM\}\}\(\\theta\)=\\mathbb\{E\}\_\{x,\\tilde\{r\},z\_\{0\},t\}\\left\[\\left\\\|f\_\{\\theta\}\(z\_\{t\},x,t\)\-\(\\tilde\{r\}\-z\_\{0\}\)\\right\\\|\_\{1\}\\right\]\.\(7\)
At inference, generation follows the learned ordinary differential equation
dz\(t\)dt=fθ\(z\(t\),x,t\),z\(0\)=z0\.\\frac\{dz\(t\)\}\{dt\}=f\_\{\\theta\}\(z\(t\),x,t\),\\qquad z\(0\)=z\_\{0\}\.\(8\)We numerically integrate Equation \([8](https://arxiv.org/html/2609.22126#S4.E8)\) fromt=0t=0tot=1t=1\. UsingMMintegration steps with time points0=t0<⋯<tM=10=t\_\{0\}<\\cdots<t\_\{M\}=1, an Euler update takes the form
z\(m\+1\)=z\(m\)\+\(tm\+1−tm\)fθ\(z\(m\),x,tm\)\.z^\{\(m\+1\)\}=z^\{\(m\)\}\+\(t\_\{m\+1\}\-t\_\{m\}\)f\_\{\\theta\}\(z^\{\(m\)\},x,t\_\{m\}\)\.\(9\)The resulting base residual is
r~b=Sθ\(z0,x\)=z\(M\),\\tilde\{r\}\_\{b\}=S\_\{\\theta\}\(z\_\{0\},x\)=z^\{\(M\)\},\(10\)whereSθS\_\{\\theta\}denotes the differentiable FlowMatch sampling procedure\.
The base prediction captures the overall residual field, but a finite\-capacity generator and few\-step sampler may leave localized errors around cloud boundaries, shadow transitions, and high\-gradient irradiance regions\. The refinement stage is designed to correct these candidate\-specific residual errors without regenerating the entire field from noise\.
### 4\.2Prediction\-Conditioned Refinement Path
A standard denoising model is commonly trained by perturbing the ground\-truth target directly\. Such perturbations do not necessarily resemble the structured errors produced by the FlowMatch generator\. We instead construct the refiner’s training states from the current generated residualr~b\\tilde\{r\}\_\{b\}\. This makes the training distribution explicitly dependent on the prediction that will be refined at inference\.
For refinement levelk∈\{0,…,K−1\}k\\in\\\{0,\\ldots,K\-1\\\}, define
αk=kK−1\.\\alpha\_\{k\}=\\frac\{k\}\{K\-1\}\.\(11\)We construct a prediction\-conditioned state along the path from the current FlowMatch residual to the target:
ck=\(1−αk\)r~b\+αkr~\.c\_\{k\}=\(1\-\\alpha\_\{k\}\)\\tilde\{r\}\_\{b\}\+\\alpha\_\{k\}\\tilde\{r\}\.\(12\)The path begins at the generator prediction,c0=r~bc\_\{0\}=\\tilde\{r\}\_\{b\}, and terminates at the target,cK−1=r~c\_\{K\-1\}=\\tilde\{r\}\. Writing the current generator error as
eθ=r~b−r~,e\_\{\\theta\}=\\tilde\{r\}\_\{b\}\-\\tilde\{r\},\(13\)Equation \([12](https://arxiv.org/html/2609.22126#S4.E12)\) gives
ck−r~=\(1−αk\)eθ\.c\_\{k\}\-\\tilde\{r\}=\(1\-\\alpha\_\{k\}\)e\_\{\\theta\}\.\(14\)Thus, the refinement levels expose the corrector to progressively smaller versions of the structured errors produced by the current generator, rather than to arbitrary interpolation errors\.
To improve robustness and train the corrector at multiple uncertainty levels, we perturb each intermediate state using Gaussian noise:
c¯k=ck\+σkϵ,ϵ∼𝒩\(0,I\)\.\\bar\{c\}\_\{k\}=c\_\{k\}\+\\sigma\_\{k\}\\epsilon,\\qquad\\epsilon\\sim\\mathcal\{N\}\(0,I\)\.\(15\)We use an exponentially decreasing noise schedule
σk=σmax\(σminσmax\)kK−1,\\sigma\_\{k\}=\\sigma\_\{\\max\}\\left\(\\frac\{\\sigma\_\{\\min\}\}\{\\sigma\_\{\\max\}\}\\right\)^\{\\frac\{k\}\{K\-1\}\},\(16\)so early refinement levels cover larger perturbations around the base prediction, while later levels focus on small corrections near the target\.
Combining Equations \([14](https://arxiv.org/html/2609.22126#S4.E14)\) and \([15](https://arxiv.org/html/2609.22126#S4.E15)\), the refiner observes
c¯k−r~=\(1−αk\)eθ\+σkϵ\.\\bar\{c\}\_\{k\}\-\\tilde\{r\}=\(1\-\\alpha\_\{k\}\)e\_\{\\theta\}\+\\sigma\_\{k\}\\epsilon\.\(17\)Its input therefore contains both the structured error of the current generator and a controlled stochastic perturbation\.
### 4\.3Multilevel Denoising Refiner
The PDE\-style refinergϕg\_\{\\phi\}receives the noisy intermediate state, the original conditioning variables, the base FlowMatch prediction, and an embedding of the refinement level:
r^k=gϕ\(c¯k,x,r~b,k\)\.\\hat\{r\}\_\{k\}=g\_\{\\phi\}\(\\bar\{c\}\_\{k\},x,\\tilde\{r\}\_\{b\},k\)\.\(18\)The base residual is supplied separately because it identifies the prediction being corrected and allows the refiner to distinguish candidate\-specific errors from the Gaussian perturbation introduced during training\.
The network predicts the clean target residual rather than the added noise\. For a sampled refinement levelkk, its loss is
ℒref\(k\)=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{ref\}\}^\{\(k\)\}=‖r^k−r~‖1\+λmse‖r^k−r~‖22\\displaystyle\\left\\\|\\hat\{r\}\_\{k\}\-\\tilde\{r\}\\right\\\|\_\{1\}\+\\lambda\_\{\\mathrm\{mse\}\}\\left\\\|\\hat\{r\}\_\{k\}\-\\tilde\{r\}\\right\\\|\_\{2\}^\{2\}\+λ∇‖∇r^k−∇r~‖1\.\\displaystyle\+\\lambda\_\{\\nabla\}\\left\\\|\\nabla\\hat\{r\}\_\{k\}\-\\nabla\\tilde\{r\}\\right\\\|\_\{1\}\.\(19\)The first term provides robust pixel\-level supervision, while the MSE term places additional emphasis on large residual errors\. The gradient\-consistency term encourages the reconstruction to preserve spatial transitions associated with cloud boundaries and localized irradiance changes\. The complete refinement objective averages over training samples, FlowMatch source noise, refinement levels, and perturbation noise:
ℒref\(θ,ϕ\)=𝔼x,r~,z0,k,ϵ\[ℒref\(k\)\]\.\\mathcal\{L\}\_\{\\mathrm\{ref\}\}\(\\theta,\\phi\)=\\mathbb\{E\}\_\{x,\\tilde\{r\},z\_\{0\},k,\\epsilon\}\\left\[\\mathcal\{L\}\_\{\\mathrm\{ref\}\}^\{\(k\)\}\\right\]\.\(20\)The dependence onθ\\thetaarises becauser~b=Sθ\(z0,x\)\\tilde\{r\}\_\{b\}=S\_\{\\theta\}\(z\_\{0\},x\)determines both the refinement path and one of the refiner inputs\.
### 4\.4Iterative Refinement at Inference
At inference, the refinement process begins from the generated residual
u0=r~b\.u\_\{0\}=\\tilde\{r\}\_\{b\}\.\(21\)Fork=0,…,K−1k=0,\\ldots,K\-1, the refiner predicts a clean residual estimate
r^k=gϕ\(uk,x,r~b,k\),\\hat\{r\}\_\{k\}=g\_\{\\phi\}\(u\_\{k\},x,\\tilde\{r\}\_\{b\},k\),\(22\)and the current state is moved toward this prediction:
uk\+1=uk\+ηk\(r^k−uk\)\.u\_\{k\+1\}=u\_\{k\}\+\\eta\_\{k\}\(\\hat\{r\}\_\{k\}\-u\_\{k\}\)\.\(23\)Here,ηk∈\(0,1\]\\eta\_\{k\}\\in\(0,1\]controls the correction strength\. We use a constantηk=η\\eta\_\{k\}=\\etain our experiments\. Unlike training, no ground\-truth residual is available and no prediction\-to\-target interpolation is constructed at inference\. The refiner instead follows the sequence of level\-conditioned correction operators learned from Equation \([12](https://arxiv.org/html/2609.22126#S4.E12)\)\. The final normalized residual is
r~^=uK\.\\hat\{\\tilde\{r\}\}=u\_\{K\}\.\(24\)
The distinction between Equations \([12](https://arxiv.org/html/2609.22126#S4.E12)\) and \([23](https://arxiv.org/html/2609.22126#S4.E23)\) is important\. The former defines the supervised training distribution using the known target, whereas the latter defines the test\-time correction trajectory using only the current estimate and available conditioning information\.
### 4\.5Refinement\-Aware Joint Optimization
A frozen generator–refiner cascade optimizes its two stages independently\. The generator is first trained with Equation \([7](https://arxiv.org/html/2609.22126#S4.E7)\), after which its output is treated as a fixed input when training the refiner\. In that setting, the refinement objective cannot influence the states produced by the generator\.
SolarFlowRefiner instead runs the differentiable samplerSθS\_\{\\theta\}inside the refinement training loop\. Its coupled objective is
ℒjoint\(θ,ϕ\)=λFMℒFM\(θ\)\+λrefℒref\(θ,ϕ\)\.\\mathcal\{L\}\_\{\\mathrm\{joint\}\}\(\\theta,\\phi\)=\\lambda\_\{\\mathrm\{FM\}\}\\mathcal\{L\}\_\{\\mathrm\{FM\}\}\(\\theta\)\+\\lambda\_\{\\mathrm\{ref\}\}\\mathcal\{L\}\_\{\\mathrm\{ref\}\}\(\\theta,\\phi\)\.\(25\)Because the base residual appears in both the noisy path state and the refiner conditioning, the refinement gradient with respect to the generator can be written schematically as
∇θℒref=\[\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{ref\}\}=\\Bigg\[\(1−αk\)∂ℒref∂c¯k\+∂ℒref∂r~b\]∂Sθ\(z0,x\)∂θ\.\\displaystyle\(1\-\\alpha\_\{k\}\)\\frac\{\\partial\\mathcal\{L\}\_\{\\mathrm\{ref\}\}\}\{\\partial\\bar\{c\}\_\{k\}\}\+\\frac\{\\partial\\mathcal\{L\}\_\{\\mathrm\{ref\}\}\}\{\\partial\\tilde\{r\}\_\{b\}\}\\Bigg\]\\frac\{\\partial S\_\{\\theta\}\(z\_\{0\},x\)\}\{\\partial\\theta\}\.\(26\)The first term propagates through the prediction\-conditioned refinement state, while the second propagates through the explicit base\-residual conditioning of the refiner\. Therefore,
∇θℒref≠0\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{ref\}\}\\neq 0\(27\)in SolarFlowRefiner\.
This coupling does not encourage the generator to produce a less accurate prediction\. Rather, it augments the standalone FlowMatch objective with information about how the generated state behaves under downstream correction\. The generator and refiner consequently co\-adapt: the generator supplies the states defining the refiner’s training distribution, while the refinement objective discourages generator outputs that lead to large final correction errors\.
### 4\.6Model Variants
We evaluate the following four flow\-based configurations\.
#### FlowMatch\.
This is the standalone conditional residual generator described in Section[4\.1](https://arxiv.org/html/2609.22126#S4.SS1)\. Its outputr~b\\tilde\{r\}\_\{b\}is used directly for SSR reconstruction, without downstream correction\.
#### FlowRefiner\-ODE\.
This post\-hoc variant freezes the pretrained FlowMatch generator and replaces the denoising corrector with an ODE\-style velocity model\. The velocity model learns a deterministic correction trajectory from the fixed FlowMatch residual toward the target residual\. At inference, this correction field is numerically integrated fromr~b\\tilde\{r\}\_\{b\}to obtain the final prediction\. Gradients from the ODE refiner are not propagated into the FlowMatch generator\.
#### FlowRefiner\-PDE\.
This variant uses the multilevel denoising refiner defined in Sections[4\.2](https://arxiv.org/html/2609.22126#S4.SS2)–[4\.4](https://arxiv.org/html/2609.22126#S4.SS4), but trains it on residuals produced by a frozen FlowMatch generator\. The base predictions may be precomputed and cached\. Formally, the refinement path is constructed using
r~b=sg\[Sθ\(z0,x\)\],\\tilde\{r\}\_\{b\}=\\operatorname\{sg\}\\left\[S\_\{\\theta\}\(z\_\{0\},x\)\\right\],\(28\)wheresg\[⋅\]\\operatorname\{sg\}\[\\cdot\]denotes the stop\-gradient operator\. Hence,
∇θℒref=0\.\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{ref\}\}=0\.\(29\)
#### SolarFlowRefiner\.
This is the proposed refinement\-aware model\. The FlowMatch prediction is generated online, the prediction\-conditioned refinement path is constructed from the current generator output, and the refinement loss is differentiated through the sampler\. SolarFlowRefiner therefore differs from FlowRefiner\-PDE in training coupling rather than in the inference\-time refinement architecture\.
Table 1:Test\-set SSR downscaling performance\. Lower is better for MAE, RMSE, LPIPS, and FID; higher is better for SSIM\. All metrics are computed after reconstructing the physical SSR field from the predicted normalized residual\.Table 2:SolarFlowRefiner input\-channel ablation\. The ERA5 SSR baseline is used for residual reconstruction in all settings; the ablation changes only the conditioning channels provided to the learned generator and refiner\.Figure 2:Qualitative SSR reconstruction and error maps for a held\-out test sample\. The top row shows reconstructed SSR fields for Bicubic/ERA5, deterministic baselines, generative baselines, refinement variants, and the SolarCube ground truth\. The bottom row shows the corresponding error maps computed as prediction minus ground truth, with per\-sample MAE and RMSE reported below each panel\.Figure 3:Performance gain versus scene complexity\. Scene complexity is computed as the mean Sobel gradient magnitude over z\-scored SolarCube auxiliary channelsvis047,vis086,ir133, andcm\. Each gray point is a held\-out test sample\. Red curves show binned mean MAE gain with standard\-error bars, where gain is baseline MAE minus SolarFlowRefiner MAE\. Positive values indicate that SolarFlowRefiner has lower error than the baseline\.
## 5Experimental Setup
We conduct extensive experiments to address the following research questions:
- •RQ1: Baseline performance\.How effectively do deterministic, residual\-shift, and flow\-based methods reconstruct high\-resolution SSR residuals?
- •RQ2: Post\-hoc refinement\.Does iterative refinement improve the standalone FlowMatch prediction, and how do ODE\-style and PDE\-style refinement dynamics compare?
- •RQ3: Refinement\-aware coupling\.Does end\-to\-end optimization of the generator and refiner improve over independently trained post\-hoc refinement?
- •RQ4: Conditioning information\.How do ERA5 radiative variables and SolarCube satellite channels contribute to SolarFlowRefiner’s downscaling performance?
Dataset and evaluation\.We evaluate on the ERA5–SolarCube split described in the problem setup\. All models are trained on the 32,159\-sample training set, selected using validation performance, and evaluated on the 8,994\-sample held\-out test set\. Metrics are computed after reconstructing physical SSR from the predicted normalized residual\.
Baselines\.We compare against Bicubic interpolation of ERA5 SSR, deterministic residual super\-resolution baselines EDSR and RCAN, attention\-based SwinIR and HAT baselines, ResShift, standalone FlowMatch, FlowRefiner\-ODE, FlowRefiner\-PDE, and SolarFlowRefiner\. EDSR, RCAN, SwinIR, and HAT are included as representative deterministic or attention\-based super\-resolution models\([15](https://arxiv.org/html/2609.22126#bib.bib4);[29](https://arxiv.org/html/2609.22126#bib.bib5);[14](https://arxiv.org/html/2609.22126#bib.bib8);[3](https://arxiv.org/html/2609.22126#bib.bib9)\)\. ResShift represents residual\-shift generative super\-resolution\([27](https://arxiv.org/html/2609.22126#bib.bib10)\)\. FlowMatch, FlowRefiner\-ODE, and FlowRefiner\-PDE isolate the effects of generation and post\-hoc refinement, while SolarFlowRefiner tests the value of refinement\-aware end\-to\-end coupling\.
Implementation details\.The generative models use a multiscale residual U\-Net backbone\. During end\-to\-end training, FlowMatch sampling uses an 8\-step Euler solver for memory and runtime efficiency\. The refiner learning rate is2×10−52\\times 10^\{\-5\}and the FlowMatch learning rate is2×10−62\\times 10^\{\-6\}\. The refiner usesK=8K=8refinement levels with a noise schedule decreasing from 0\.35 to 0\.01\. We use AdamW with batch size 4 and train for 100 epochs\. In the refinement loss, we setλmse=0\.1\\lambda\_\{\\mathrm\{mse\}\}=0\.1andλ∇=0\.05\\lambda\_\{\\nabla\}=0\.05, and use refinement strengthη=1\.0\\eta=1\.0\. All training is performed on Nvidia RTX 5060 GPU\.
Input\-channel ablation\.To answer RQ4, we evaluate SolarFlowRefiner under three conditioning settings: ERA5\-only, SolarCube\-only, and ERA5\+SolarCube\. The ERA5\-only setting includesera\_ssrd,era\_ssr,era\_ssrdc,era\_fdir, andera\_cdir\. The SolarCube\-only setting includesvis047,vis086,ir133,sza, andcm\. The combined setting uses all ten channels\. In every setting, the upsampled ERA5 SSR field remains the residual baselinebbused for reconstruction\.
## 6Results
Main quantitative comparison\.Table[1](https://arxiv.org/html/2609.22126#S4.T1)reports the held\-out test\-set comparison\. Bicubic upsampling of ERA5 SSR has high error because the coarse reanalysis field cannot resolve localized cloud\-shadow and cloud\-edge structure\. Deterministic super\-resolution models reduce this error substantially, but their gains saturate once the model must recover sharper, high\-frequency SSR residuals\. The generative and refinement\-based models provide a stronger reconstruction of these residual fields\.
SolarFlowRefiner achieves the best MAE, RMSE, SSIM, and LPIPS among the methods in Table[1](https://arxiv.org/html/2609.22126#S4.T1)\. Relative to the strongest PDE\-refinement baseline, FlowRefiner\-PDE, SolarFlowRefiner reduces MAE from 13\.58 to 11\.13 and RMSE from 20\.31 to 17\.11, corresponding to approximately 18% lower MAE and 16% lower RMSE\. It also improves SSIM from 0\.727 to 0\.827, indicating better structural agreement with SolarCube\. Compared with standalone FlowMatch, SolarFlowRefiner reduces MAE by about 25% and RMSE by about 27%, showing that the improvement is not due only to using a flow\-based generator, but to coupling generation with downstream refinement\. FID is comparable to the strongest generative baselines, while LPIPS is lowest for SolarFlowRefiner, suggesting that the proposed model improves local perceptual structure without degrading distribution\-level realism\.
Input\-channel ablation\.Table[2](https://arxiv.org/html/2609.22126#S4.T2)reports the SolarFlowRefiner input\-channel ablation\. We include this ablation only for the proposed model so that the analysis isolates the role of conditioning sources without conflating input selection with architecture changes\.
The ablation confirms that both sources of conditioning information are useful\. Using ERA5 channels alone gives substantially higher error, because the model receives physically meaningful radiative variables but little high\-resolution cloud structure\. SolarCube\-only conditioning improves over ERA5\-only conditioning, reducing MAE from 21\.82 to 17\.31, which indicates that the high\-resolution auxiliary channels carry important cloud\-boundary information\. The full ERA5\+SolarCube setting performs best across all metrics, reducing MAE by about 49% relative to ERA5\-only conditioning and about 36% relative to SolarCube\-only conditioning\. This supports the residual formulation: ERA5 provides the coarse physical baseline, while SolarCube auxiliary channels guide the fine\-scale correction\.
Qualitative reconstruction\.Figure[2](https://arxiv.org/html/2609.22126#S4.F2)shows a representative held\-out sample\. Bicubic/ERA5 preserves only broad radiative structure and misses the fine cloud\-driven texture visible in the ground truth\. Deterministic baselines recover more local detail but still leave structured errors near high\-gradient cloud boundaries\. FlowMatch and the post\-hoc refinement variants reduce these errors, but SolarFlowRefiner produces the lowest per\-sample MAE and RMSE in this example\. The error maps show that refinement\-aware training reduces both broad residual bias and localized cloud\-edge mistakes relative to the post\-hoc refiners\.
Scene complexity analysis\.Figure[3](https://arxiv.org/html/2609.22126#S4.F3)evaluates whether SolarFlowRefiner’s gains are concentrated in scenes with stronger spatial structure\. For each test sample, we measure scene complexity as the mean Sobel gradient magnitude over the SolarCube auxiliary channelsvis047,vis086,ir133, andcm\. We then compute gain as the baseline MAE minus the SolarFlowRefiner MAE, so positive values indicate that SolarFlowRefiner is better\.
SolarFlowRefiner has positive average gain against all four baselines: 2\.97 MAE over ResShift, 4\.24 over FlowMatch, 4\.11 over FlowRefiner\-ODE, and 2\.63 over FlowRefiner\-PDE\. The Pearson correlations between complexity and gain are weakly negative, ranging fromr=−0\.063r=\-0\.063tor=−0\.096r=\-0\.096\. This indicates that the improvement is broad rather than restricted to a narrow complexity regime\. The binned means remain above zero across most of the complexity range, including against FlowRefiner\-PDE, supporting the claim that refinement\-aware coupling provides a consistent advantage over post\-hoc refinement\.
## 7Conclusion and Future Work
In this paper, we introduced SolarFlowRefiner for multisource surface solar radiation downscaling\. The framework addresses the stage\-wise mismatch between generation and post\-hoc refinement by constructing prediction\-conditioned refinement states from the current generator output and propagating the refinement objective through the differentiable sampler\. This aligns generation and correction with the final SSR reconstruction objective\. Experiments on the day\-blocked ERA5–SolarCube benchmark show that SolarFlowRefiner consistently outperforms standalone generation and frozen refinement variants\. Relative to the strongest post\-hoc PDE\-style refiner, it reduces MAE and RMSE by approximately 18% and 16%, respectively, while improving structural and perceptual fidelity\. The conditioning ablation further confirms that coarse ERA5 radiative information and high\-resolution SolarCube observations provide complementary information for fine\-scale reconstruction\. These findings point to a broader class of geophysical inverse problems in which coarse, physically consistent products capture the large\-scale system state, while remote\-sensing observations reveal localized spatial heterogeneity\. Future work will investigate refinement\-aware generation as a transferable predictor–corrector principle for such problems, with generation recovering a plausible global field and refinement resolving observation\-informed fine\-scale structure\. We will first extend this framework to land\-surface temperature, precipitation, and soil\-moisture reconstruction, and then examine cross\-region and cross\-sensor transfer toward more general AI methods for multisource Earth observation and environmental decision\-making\.
## References
- Antonanzaset al\.\(2016\)J\. Antonanzas, N\. Osorio, R\. Escobar, R\. Urraca, F\. J\. Martinez\-de\-Pison, and F\. Antonanzas\-TorresReview of photovoltaic power forecasting\.Solar Energy136,pp\. 78–111\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p1.1),[§1](https://arxiv.org/html/2609.22126#S1.p2.1)\.
- Blau and Michaeli \(2018\)Y\. Blau and T\. MichaeliThe perception\-distortion tradeoff\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 6228–6237\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1)\.
- Chenet al\.\(2023\)X\. Chen, X\. Wang, J\. Zhou, Y\. Qiao, and C\. DongActivating more pixels in image super\-resolution transformer\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 22367–22377\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.22126#S5.p4.1)\.
- Chenet al\.\(2026\)Y\. Chen, Z\. Ma, P\. Jiang, Y\. Dai, Q\. Hu, X\. Ye, L\. Li, R\. Sousa, and R\. YuWhen earth foundation models meet diffusion: an application to land surface temperature super\-resolution\.External Links:2604\.16841Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p3.1)\.
- Daiet al\.\(2026\)Y\. Dai, Y\. Sun, Y\. Chen, S\. Chen, X\. Jia, and R\. YuFlowRefiner: flow matching\-based iterative refinement for 3d turbulent flow simulation\.External Links:2604\.17149Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p2.1)\.
- Donget al\.\(2016\)C\. Dong, C\. C\. Loy, K\. He, and X\. TangImage super\-resolution using deep convolutional networks\.IEEE Transactions on Pattern Analysis and Machine Intelligence38\(2\),pp\. 295–307\.Cited by:[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1)\.
- Hersbachet al\.\(2020\)H\. Hersbach, B\. Bell, P\. Berrisford, S\. Hirahara, A\. Horányi, J\. Muñoz\-Sabater, J\. Nicolas, C\. Peubey, R\. Radu, D\. Schepers,et al\.The era5 global reanalysis\.Quarterly Journal of the Royal Meteorological Society146\(730\),pp\. 1999–2049\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p1.1),[§1](https://arxiv.org/html/2609.22126#S1.p2.1)\.
- Heuselet al\.\(2017\)M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. HochreiterGANs trained by a two time\-scale update rule converge to a local nash equilibrium\.InProceedings of the 31st International Conference on Neural Information Processing Systems,NIPS’17,Red Hook, NY, USA,pp\. 6629–6640\.External Links:ISBN 9781510860964Cited by:[§3](https://arxiv.org/html/2609.22126#S3.p4.1)\.
- Kimet al\.\(2016\)J\. Kim, J\. K\. Lee, and K\. M\. LeeAccurate image super\-resolution using very deep convolutional networks\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 1646–1654\.Cited by:[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1)\.
- Kurinchi\-Vendhanet al\.\(2021\)R\. Kurinchi\-Vendhan, B\. Lütjens, R\. Gupta, L\. Werner, and D\. NewmanWiSoSuper: benchmarking super\-resolution methods on wind and solar data\.External Links:2109\.08770Cited by:[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p2.1)\.
- Lediget al\.\(2017\)C\. Ledig, L\. Theis, F\. Huszár, J\. Caballero, A\. Cunningham, A\. Acosta, A\. Aitken, A\. Tejani, J\. Totz, Z\. Wang, and W\. ShiPhoto\-realistic single image super\-resolution using a generative adversarial network\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 4681–4690\.Cited by:[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1)\.
- Liet al\.\(2025\)J\. Li, J\. Cao, Y\. Guo, W\. Li, and Y\. ZhangOne diffusion step to real\-world super\-resolution via flow trajectory distillation\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=riYSkLG0vt)Cited by:[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
- Liet al\.\(2024\)R\. Li, Y\. Xie, X\. Jia, D\. Wang, Y\. Li, Y\. Zhang, Z\. Wang, and Z\. LiSolarCube: an integrative benchmark dataset harnessing satellite and in\-situ observations for large\-scale solar energy forecasting\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,pp\. 3499–3513\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p1.1),[§1](https://arxiv.org/html/2609.22126#S1.p2.1),[§1](https://arxiv.org/html/2609.22126#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p2.1)\.
- Lianget al\.\(2021\)J\. Liang, J\. Cao, G\. Sun, K\. Zhang, L\. Van Gool, and R\. TimofteSwinIR: image restoration using swin transformer\.InProceedings of the IEEE/CVF International Conference on Computer Vision Workshops \(ICCVW\),pp\. 1833–1844\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.22126#S5.p4.1)\.
- Limet al\.\(2017\)B\. Lim, S\. Son, H\. Kim, S\. Nah, and K\. M\. LeeEnhanced deep residual networks for single image super\-resolution\.In2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops \(CVPRW\),Vol\.,pp\. 1132–1140\.External Links:[Document](https://dx.doi.org/10.1109/CVPRW.2017.151)Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.22126#S5.p4.1)\.
- Lipmanet al\.\(2023\)Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. LeFlow matching for generative modeling\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
- Lippeet al\.\(2023\)P\. Lippe, B\. S\. Veeling, P\. Perdikaris, R\. E\. Turner, and J\. BrandstetterPDE\-refiner: achieving accurate long rollouts with neural pde solvers\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.36,pp\. 67398–67433\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p2.1)\.
- Liuet al\.\(2023\)X\. Liu, C\. Gong, and Q\. LiuFlow straight and fast: learning to generate and transfer data with rectified flow\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
- Lugmayret al\.\(2020\)A\. Lugmayr, M\. Danelljan, L\. Van Gool, and R\. TimofteSRFlow: learning the super\-resolution space with normalizing flow\.InEuropean Conference on Computer Vision \(ECCV\),pp\. 715–732\.Cited by:[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
- Ohayonet al\.\(2025\)G\. Ohayon, T\. Michaeli, and M\. EladPosterior\-mean rectified flow: towards minimum mse photo\-realistic image restoration\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
- Schmitet al\.\(2017\)T\. J\. Schmit, P\. Griffith, M\. M\. Gunshor, J\. M\. Daniels, S\. J\. Goodman, and W\. J\. LebairA closer look at the abi on the goes\-r series\.Bulletin of the American Meteorological Society98\(4\),pp\. 681–698\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p1.1),[§1](https://arxiv.org/html/2609.22126#S1.p3.1)\.
- Senguptaet al\.\(2021\)M\. Sengupta, A\. Habte, S\. Wilbert, C\. Gueymard, and J\. RemundBest practices handbook for the collection and use of solar resource data for solar energy applications: third edition\.Technical reportTechnical ReportNREL/TP\-5D00\-77635,National Renewable Energy Laboratory \(NREL\)\.External Links:[Document](https://dx.doi.org/10.2172/1778700)Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p1.1)\.
- Stengelet al\.\(2020\)K\. Stengel, A\. Glaws, D\. Hettinger, and R\. N\. KingAdversarial super\-resolution of climatological wind and solar data\.Proceedings of the National Academy of Sciences \(PNAS\)117\(29\),pp\. 16805–16815\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p2.1)\.
- Wanget al\.\(2018\)X\. Wang, K\. Yu, S\. Wu, J\. Gu, Y\. Liu, C\. Dong, Y\. Qiao, and C\. C\. LoyESRGAN: enhanced super\-resolution generative adversarial networks\.InComputer Vision – ECCV 2018 Workshops: Munich, Germany, September 8\-14, 2018, Proceedings, Part V,Berlin, Heidelberg,pp\. 63–79\.External Links:ISBN 978\-3\-030\-11020\-8,[Link](https://doi.org/10.1007/978-3-030-11021-5_5),[Document](https://dx.doi.org/10.1007/978-3-030-11021-5%5F5)Cited by:[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1)\.
- Wanget al\.\(2004\)Z\. Wang, A\.C\. Bovik, H\.R\. Sheikh, and E\.P\. SimoncelliImage quality assessment: from error visibility to structural similarity\.IEEE Transactions on Image Processing13\(4\),pp\. 600–612\.External Links:[Document](https://dx.doi.org/10.1109/TIP.2003.819861)Cited by:[§3](https://arxiv.org/html/2609.22126#S3.p4.1)\.
- Xuet al\.\(2026\)J\. Xu, W\. Li, H\. Sun, F\. Li, Z\. Wang, L\. Peng, J\. Ren, H\. Yang, X\. Hu, R\. Pei, and P\. HengFast image super\-resolution via consistency rectified flow\.External Links:2605\.12377,[Link](https://arxiv.org/abs/2605.12377)Cited by:[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
- Yueet al\.\(2023\)Z\. Yue, J\. Wang, and C\. C\. LoyResShift: efficient diffusion model for image super\-resolution by residual shifting\.Advances in Neural Information Processing Systems \(NeurIPS\)36,pp\. 13294–13307\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1),[§5](https://arxiv.org/html/2609.22126#S5.p4.1)\.
- Zhanget al\.\(2018a\)R\. Zhang, P\. Isola, A\. A\. Efros, E\. Shechtman, and O\. WangThe unreasonable effectiveness of deep features as a perceptual metric\.In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition,Vol\.,pp\. 586–595\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2018.00068)Cited by:[§3](https://arxiv.org/html/2609.22126#S3.p4.1)\.
- Zhanget al\.\(2018b\)Y\. Zhang, K\. Li, K\. Li, L\. Wang, B\. Zhong, and Y\. FuImage super\-resolution using very deep residual channel attention networks\.InProceedings of the European Conference on Computer Vision \(ECCV\),pp\. 286–301\.Cited by:[§1](https://arxiv.org/html/2609.22126#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.22126#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.22126#S5.p4.1)\.
- Zhuet al\.\(2024\)Y\. Zhu, W\. Zhao, A\. Li, Y\. Tang, J\. Zhou, and J\. LuFlowIE: efficient image enhancement via rectified flow\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 13–22\.Cited by:[§2\.2](https://arxiv.org/html/2609.22126#S2.SS2.p1.1)\.
## Appendix ADataset Construction Details
Our benchmark is constructed for hourly ERA5\-to\-SolarCube surface solar radiation \(SSR\) downscaling\. Each sample contains a coarse reanalysis\-derived radiative state from ERA5 and a high\-resolution SolarCube target on a120×120120\\times 120grid\. ERA5 variables are spatially upsampled to the SolarCube grid before being passed to the model\. The high\-resolution target is the hourly mean SolarCube SSR field, computed from the available 15\-minute SolarCube frames within the hour\. We exclude nighttime samples using a minimum ERA5 SSRD mean threshold of 10\.0, and we discard samples with target\-valid fraction below the quality\-control threshold used in the main split\.
The conditioning tensor contains ten channels\. The five ERA5 radiative channels areera\_ssrd,era\_ssr,era\_ssrdc,era\_fdir, andera\_cdir\. The five SolarCube auxiliary channels arevis047,vis086,ir133,sza, andcm\. The prediction target is the SolarCube SSR fieldsolarcube\_ssr\_hourly\. All learned models predict the normalized SSR residual rather than the HR field directly\. The residual is the SolarCube target minus the upsampled ERA5 SSR baseline, standardized using training\-set residual statistics\. Physical SSR is reconstructed by adding the predicted residual correction back to the ERA5 baseline\. This residual formulation preserves the coarse physical estimate from ERA5 and focuses learning on high\-resolution cloud\-driven corrections\.
We use tiles 1–10, 12, and 14, which pass the target\-valid quality\-control threshold\. Tiles 11, 13, and 15–19 are excluded because their target\-valid fraction is below 50%\. Samples are split by whole UTC\-day blocks within each tile/month, so all hours from the same tile\-month\-day remain in the same partition\. This reduces leakage from near\-duplicate cloud scenes across the training, validation, and test sets\. The final split uses a 70/10/20 train/validation/test ratio and is summarized in Table[3](https://arxiv.org/html/2609.22126#A1.T3)\.
Table 3:Dataset split used for model selection and final evaluation\.
## Appendix BSite\-Wise Spatial Performance Analysis
We further analyze whether the relative benefit of SolarFlowRefiner is uniform across the spatial sites\. This analysis uses the same 8,994 held\-out test samples and the same per\-sample MAE values used for the main quantitative comparison\. We focus on spatial variation rather than UTC\-hour variation because the sites span multiple longitudes; the same UTC hour can correspond to different local solar times at different stations\. This makes a raw UTC morning/afternoon analysis difficult to interpret without additional solar\-time normalization\.
Figure[4](https://arxiv.org/html/2609.22126#A2.F4)shows the spatial distribution of the test sites\. Marker color indicates the per\-sample MAE win rate of SolarFlowRefiner at each site, where the competing methods are ResShift, FlowMatch, FlowRefiner\-ODE, and FlowRefiner\-PDE\. The eight North American sites are shown separately in an inset because they are geographically clustered\. Station codes, station names, coordinates, and test sample counts are listed in Table[4](https://arxiv.org/html/2609.22126#A2.T4)\. The map uses Natural Earth country boundaries for geographic context; station labels and coordinates follow the BSRN/PANGAEA station metadata conventions\.
Figure 4:Global distribution of the 12 test tiles/sites\. Marker color shows the fraction of test samples at each site for which SolarFlowRefiner has the lowest MAE among the learned comparison models\. The right panel expands the North American cluster\.Table 4:Test sites used in the spatial analysis\.Table[5](https://arxiv.org/html/2609.22126#A2.T5)reports site\-wise mean MAE\. SolarFlowRefiner has the lowest mean MAE at every site, indicating that the coupled FlowMatch–refiner model improves the average reconstruction error across the spatial domain rather than only at a small subset of stations\. The size of the benefit, however, is spatially heterogeneous\. The largest absolute reductions occur at LRC, DWN, and HOW, where the competing methods have substantially higher mean errors and SolarFlowRefiner also wins most individual samples\.
Table 5:Per\-site mean MAE inWm−2\\mathrm\{W\\,m^\{\-2\}\}for the learned models\. Lower is better\.Figure[5](https://arxiv.org/html/2609.22126#A2.F5)provides one qualitative example from each site\. The same models used in the spatial MAE analysis are shown next to the SolarCube ground truth, with per\-panel MAE and RMSE listed below each reconstruction\.
Figure 5:Qualitative comparison across test sites\. Each row shows one sample from a different tile/site, and columns compare ResShift, FlowMatch, FlowRefiner\-ODE, FlowRefiner\-PDE, SolarFlowRefiner, and the SolarCube ground truth\. MAE and RMSE are reported below each reconstruction in physical SSR units; the ground\-truth column has zero error by definition\.Overall, the spatial analysis supports two conclusions\. First, the coupled SolarFlowRefiner model improves mean reconstruction accuracy across all sites in this split\. Second, the strength of the improvement is site dependent: the largest reductions occur at LRC, DWN, and HOW, while GWN and several central U\.S\. stations show smaller gains\. This spatial heterogeneity is consistent with the interpretation that refinement is most helpful when the base generator leaves structured cloud\-related errors that can be corrected by a prediction\-conditioned refiner\.
## Appendix CImplementation and Training Details
All learned models operate on120×120120\\times 120samples and predict normalized residuals\. Unless otherwise stated, models use all ten conditioning channels, are selected using validation performance, and are evaluated on the test set after reconstructing physical SSR\. We use EMA checkpoints for the generative and refinement models, with decay 0\.9999, and use Adam or AdamW optimization with gradient clipping at 1\.0 where implemented\. All reported quantitative results use one final run for each model configuration\. The day\-block split is generated with random seed 42, and model training scripts use seed 42 for pseudorandom initialization and dataloader control where applicable\. Experiments were run in a Python/PyTorch CUDA environment on an Nvidia RTX 5060 GPU\. The anonymous code appendix includes the preprocessing scripts, split manifests, model source code, run scripts, and arequirements\.txtfile listing the required Python packages, including PyTorch, NumPy, pandas, xarray, h5py, Pillow, and scikit\-image\.
#### Deterministic baselines\.
EDSR is trained as a residual CNN with 64 feature channels and 16 residual blocks\. RCAN is trained with 64 channels, 4 residual groups, and 4 residual blocks per group\. Both use batch size 16, learning rate10−410^\{\-4\}, early stopping on validation performance, and 100 training epochs\. SwinIR and HAT are included as attention\-based super\-resolution baselines\. The SwinIR run uses embedding dimension 60, window size 7, four stages with depths\(6,6,6,6\)\(6,6,6,6\), heads\(6,6,6,6\)\(6,6,6,6\), batch size 8, learning rate2×10−52\\times 10^\{\-5\}, L1 loss, and gradient clipping at 1\.0\. The HAT run uses embedding dimension 48, window size 6, depths\(2,2,2\)\(2,2,2\), heads\(4,4,4\)\(4,4,4\), batch size 8, learning rate3×10−53\\times 10^\{\-5\}, L1 loss, and gradient clipping at 1\.0\.
#### ResShift\.
The ResShift baseline uses a U\-Net\-style residual backbone with inner channel 64, channel multipliers\(1,2,4,8,16\)\(1,2,4,8,16\), and two residual blocks per stage\. It is trained for 100 epochs with batch size 2, learning rate5×10−55\\times 10^\{\-5\}, weight decay10−410^\{\-4\}, and EMA decay 0\.9999\. The residual shift process usesT=15T=15transition steps withκ=1\.0\\kappa=1\.0\.
#### FlowMatch\.
The standalone FlowMatch generator uses a multiscale residual U\-Net backbone with base channel 64, channel multipliers\(1,2,4,8\)\(1,2,4,8\), and two residual blocks per stage\. It learns the rectified\-flow velocity between Gaussian source noise and the normalized SSR residual using anℓ1\\ell\_\{1\}velocity objective and uniform flow\-time sampling,t∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\)\. The standalone FlowMatch baseline is trained for 100 epochs with batch size 4, learning rate2×10−52\\times 10^\{\-5\}, and EMA decay 0\.9999\.
#### Post\-hoc refiners\.
FlowRefiner\-ODE keeps the pretrained FlowMatch base prediction fixed and trains an ODE\-style correction refiner\. The refiner uses base channel 64, channel multipliers\(1,2,4,8\)\(1,2,4,8\), two residual blocks per stage, 8 Heun refinement steps, batch size 8, learning rate2×10−52\\times 10^\{\-5\}, weight decay10−410^\{\-4\}, and EMA decay 0\.9999\. Its loss combines an L1 velocity term, endpoint supervision with weight 0\.5, and gradient consistency with weight 0\.05\.
FlowRefiner\-PDE is the non\-ODE post\-hoc refiner built on the same pretrained FlowMatch generator\. It keeps the FlowMatch generator frozen, uses cached FlowMatch base residual predictions as the initial state, and appends this base prediction as an additional conditioning channel\. Unlike FlowRefiner\-ODE, which learns a velocity field and integrates an ODE\-style update, FlowRefiner\-PDE uses an iterative denoising\-style correction\. At each refinement level, the model predicts the clean normalized residual from a state sampled along the base\-to\-target path, and inference repeatedly moves the current residual toward this predicted clean residual\. The refiner uses base channel 64, channel multipliers\(1,2,4,8\)\(1,2,4,8\), two residual blocks per stage, 8 refinement levels, an exponentially decreasing noise schedule fromσmax=0\.35\\sigma\_\{\\max\}=0\.35toσmin=0\.01\\sigma\_\{\\min\}=0\.01, refinement strengthη=1\.0\\eta=1\.0, batch size 4, learning rate2×10−52\\times 10^\{\-5\}, weight decay10−410^\{\-4\}, and EMA decay 0\.9999\. Its loss combines L1 denoising, an MSE term with weight 0\.1, and gradient consistency with weight 0\.05\.
#### SolarFlowRefiner\.
SolarFlowRefiner is trained jointly from random initialization\. The FlowMatch residual generator and PDE\-style refiner are optimized together from the start\. Each training batch applies the FlowMatch velocity objective, samples the current generator with an 8\-step Euler solver, and trains the refiner on prediction\-conditioned states along the base\-to\-target path\. The refinement loss combines L1 reconstruction, an MSE term with weight 0\.1, and gradient consistency with weight 0\.05, and is backpropagated through the differentiable FlowMatch sampler\. The generator and refiner both use base channel 64, channel multipliers\(1,2,4,8\)\(1,2,4,8\), and two residual blocks per stage\. The FlowMatch generator uses learning rate2×10−62\\times 10^\{\-6\}and the refiner uses learning rate2×10−52\\times 10^\{\-5\}as separate AdamW parameter groups\. We use batch size 4, weight decay10−410^\{\-4\}, EMA decay 0\.9999, 100 training epochs, uniform FlowMatch time samplingt∼𝒰\(0,1\)t\\sim\\mathcal\{U\}\(0,1\),K=8K=8refinement levels, a noise schedule fromσmax=0\.35\\sigma\_\{\\max\}=0\.35toσmin=0\.01\\sigma\_\{\\min\}=0\.01, and refinement strengthη=1\.0\\eta=1\.0at inference\.
#### Evaluation and statistical testing\.
Metrics are computed after converting each predicted normalized residual back to physical SSR units\. We report MAE and RMSE inWm−2\\mathrm\{W\\,m^\{\-2\}\}for irradiance error, SSIM for structural agreement, and LPIPS and FID for perceptual and distributional fidelity\. Lower values are better for MAE, RMSE, LPIPS, and FID; higher values are better for SSIM\. For scene\-complexity analysis in the main paper, gain is defined as baseline MAE minus SolarFlowRefiner MAE, so positive values indicate lower error for SolarFlowRefiner\.
We additionally test whether the SolarFlowRefiner MAE improvements are consistent across matched held\-out samples using two\-sided paired Wilcoxon signed\-rank tests on the same 8,994 test samples\. The tests use paired per\-sample MAE differences rather than aggregate table values\.
Table 6:Paired Wilcoxon signed\-rank tests comparing per\-sample MAE between each learned baseline and SolarFlowRefiner on the 8,994 held\-out test samples\.All four tests reject equality of paired MAE distributions at conventional significance levels\. Because multiple hours from the same site can still be correlated, these paired tests should be interpreted together with the day\-blocked split and site\-wise spatial analysis\.相似文章
面向中尺度保持的海面温度降尺度的可迁移双流表示
本文介绍了EddyFlow,一个用于公里级海面温度降尺度的深度学习框架,它在预测精度、尺度相关结构和区域泛化之间取得平衡。该框架在多个海洋区域实现了强大的零样本性能和近乎理想的频谱保真度。
ReFlowSET:用于SAR到EO图像翻译的表示对齐潜在流匹配
ReFlowSET是一个用于高保真SAR到EO图像翻译的潜在流匹配框架,通过编解码器选择和带有视觉模型对齐的条件DiT,实现了最先进的性能。
用于多尺度物理系统超分辨率数据同化的迭代精炼扩散
本文提出了一种用于数据同化的迭代精炼框架,该框架结合神经算子和扩散模型,以超分辨率处理多尺度物理系统,在Kraichnan湍流等基准测试中优于基线方法。
Surflo:具有全局状态的一致3D表面流模型
Surflo是一种前馈3D重建模型,它将未定姿的RGB视图压缩成潜在标记,并通过流匹配解码出一致的3D表面点,支持可变分辨率输出,在速度上优于现有方法。
FarSky:面向生成式小时内太阳预报的任务感知潜空间耦合
FarSky 是一种生成式预报框架,利用任务感知潜空间耦合和潜扩散模型,生成确定性和概率性的小时内太阳辐照度预报,预报技巧最多提升 11 个百分点,并具有更好的爬坡事件检测能力。