When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision
Summary
This arXiv paper studies the problem of factor-derived proxy supervision in scientific machine learning, using RUSLE-based soil-loss prediction as a case study. It introduces a diagnostic framework and proposes RASPL, a formula-preserving residual learning framework that outperforms direct prediction and improves robustness to degraded factor information.
View Cached Full Text
Cached at: 08/06/26, 07:49 AM
# When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision
Source: [https://arxiv.org/html/2608.04393](https://arxiv.org/html/2608.04393)
###### Abstract
Scientific machine learning often relies on proxy targets computed from known domain factors when direct observations are limited\. When those same factors are used as model inputs, however, high predictive accuracy may reflect reconstruction of the proxy\-generating equation rather than robustness to degraded factor information\. We study this problem in RUSLE\-derived soil\-loss proxy prediction under controlled degradation of the soil\-erodibility factorKK\. We introduce a diagnostic framework that combines degraded\-formula references, classical tree\-based baselines, matched direct and formula\-feature predictors, contextual ablations, tail\-error analysis, and degradation robustness scoring\. We then propose RASPL, a formula\-preserving residual framework that retains the degraded formula estimate as the prediction anchor and learns an adaptively gated contextual correction\. RASPL substantially outperforms matched direct prediction and provides stronger degradation and tail robustness than treating the formula estimate as an ordinary input feature\. Within RASPL, a compact statistical encoder achieves the highest macro\-averagedR2R^\{2\}and lowest computational cost, whereas a convolutional encoder achieves the strongest degradation robustness and lowest Tail95 mean absolute error \(MAE\)\. These results establish formula preservation as the central design principle for robust learning from factor\-derived proxy targets\.
## Introduction
In many scientific machine learning applications, direct measurements are sparse, expensive, or unavailable at scale\. Watershed\-scale soil\-loss modeling is one such setting: field measurements are valuable for physical validation, but spatially complete erosion labels are rarely available for large\-scale supervised learning\(Benavidezet al\.[2018](https://arxiv.org/html/2608.04393#bib.bib1); Borrelliet al\.[2017](https://arxiv.org/html/2608.04393#bib.bib27)\)\. Consequently, soil\-erosion assessments often rely on Revised Universal Soil Loss Equation \(RUSLE\)\-derived proxy estimates constructed from rainfall erosivity, soil erodibility, topographic, and cover\-management factors\(Panagoset al\.[2015](https://arxiv.org/html/2608.04393#bib.bib2); Borrelliet al\.[2017](https://arxiv.org/html/2608.04393#bib.bib27); Liet al\.[2023](https://arxiv.org/html/2608.04393#bib.bib3)\)\. When these same factors are used as model inputs, however, high predictive accuracy may reflect reconstruction of the proxy\-generating equation rather than robustness to degraded, missing, or uncertain factor information\.
We study this setting as*factor\-derived proxy supervision*, where the target is computed from known scientific factors rather than independently observed outcomes\. Using RUSLE\-derived soil\-loss proxy prediction as a case study\(Geet al\.[2023](https://arxiv.org/html/2608.04393#bib.bib4); Samarinaset al\.[2024](https://arxiv.org/html/2608.04393#bib.bib5); Hlalet al\.[2025](https://arxiv.org/html/2608.04393#bib.bib6)\), the target is
A=R×K×LS×C×P,A=R\\times K\\times LS\\times C\\times P,\(1\)whereRR,KK,LSLS,CC, andPPdenote rainfall erosivity, soil erodibility, topographic, cover\-management, and support\-practice factors, respectively\. Following prior work, we setP=1P=1in the study region\(İpek and Kahya[2026](https://arxiv.org/html/2608.04393#bib.bib38)\); therefore, the implemented target reduces toA=R×K×LS×CA=R\\times K\\times LS\\times C\. When the four varying factors are provided as inputs, the target can be reconstructed directly from the known equation\. Strong learned\-model performance under this setting may therefore reflect approximation of the proxy\-generating equation rather than robustness when factor information is degraded\.
To separate reconstruction from robust learning, we evaluate models under complete, noisy, coarsened, and masked factor regimes, together with a separate missing\-factor stress test\. The evaluation combines matched degraded\-formula references, classical center and neighborhood baselines, matched direct and formula\-feature predictors, contextual\-representation ablations, tail\-error analysis, and degradation robustness scoring\. These diagnostics shift evaluation from reproducing the proxy under complete information to remaining robust when an important factor is degraded\.
We further introduce RASPL, a formula\-preserving residual framework that retains the degraded proxy\-formula estimate as the prediction anchor and learns an adaptively gated contextual correction\. The framework supports correction from either compact neighborhood statistics or learned convolutional representations, enabling a controlled analysis of whether useful information arises from neighborhood summaries, raw local values, or their spatial arrangement\. Our contributions are as follows:
- •We formalize*equation reconstruction*as a failure mode in factor\-derived proxy supervision, where high predictive accuracy may not imply robustness to degraded factor information\.
- •We introduce a diagnostic protocol based on controlled factor degradation, matched formula references, classical and neural baselines, contextual\-representation ablations, tail metrics, and degradation robustness scoring\.
- •We propose RASPL, a formula\-preserving residual framework that retains the degraded proxy\-formula estimate as the prediction anchor and learns an adaptively gated contextual correction\. Controlled statistical and convolutional evaluations demonstrate the benefit of formula preservation and characterize the tradeoff among average accuracy, tail robustness, and computational complexity\.
## Related Work
Factor\-derived proxy targets and machine\-learning\-enhanced erosion modeling\.Empirical formulas are widely used to construct environmental proxy targets when direct measurements are sparse or unavailable\. In soil\-loss modeling, RUSLE derives erosion proxies from rainfall erosivity, soil erodibility, topographic, and cover\-management factors\(Benavidezet al\.[2018](https://arxiv.org/html/2608.04393#bib.bib1); Panagoset al\.[2015](https://arxiv.org/html/2608.04393#bib.bib2); Liet al\.[2023](https://arxiv.org/html/2608.04393#bib.bib3)\)\. Prior work combines RUSLE\-derived factors, remote\-sensing data, and machine learning to estimate soil loss, improve factor layers such as theKKfactor, refine erosion maps, and analyze erosion drivers\(Samarinaset al\.[2024](https://arxiv.org/html/2608.04393#bib.bib5); Zeghmaret al\.[2024](https://arxiv.org/html/2608.04393#bib.bib39); Abiye and Dengiz[2025](https://arxiv.org/html/2608.04393#bib.bib7); Geet al\.[2023](https://arxiv.org/html/2608.04393#bib.bib4); Hlalet al\.[2025](https://arxiv.org/html/2608.04393#bib.bib6)\)\. Related studies use environmental predictors to map erosion susceptibility from categorical labels or regional erosion inventories\(Bouamraneet al\.[2024](https://arxiv.org/html/2608.04393#bib.bib8); Geleteet al\.[2024](https://arxiv.org/html/2608.04393#bib.bib9); Momeni Damanehet al\.[2025](https://arxiv.org/html/2608.04393#bib.bib10); Tkeshelashvili[2024](https://arxiv.org/html/2608.04393#bib.bib11); Oliiet al\.[2025](https://arxiv.org/html/2608.04393#bib.bib12); V\.M\.et al\.[2025](https://arxiv.org/html/2608.04393#bib.bib13); Islamet al\.[2025](https://arxiv.org/html/2608.04393#bib.bib14)\)\. In this literature, RUSLE generally remains a fixed empirical formulation while machine learning improves its inputs, outputs, or downstream analysis\. Our work instead treats the RUSLE\-derived quantity itself as the supervision target and evaluates controlled degradation of its generating factors to distinguish proxy\-equation reconstruction from degraded\-factor robustness and to assess whether formula\-preserving corrections remain effective\.
Spatial deep learning, foundation models, and formula\-aware prediction\.Prior studies have used CNNs, Transformers, and geospatial representation\-learning methods to capture local context and multichannel spatial structure\(Zhaoet al\.[2023](https://arxiv.org/html/2608.04393#bib.bib15); Karraet al\.[2021](https://arxiv.org/html/2608.04393#bib.bib16)\)\. More recent geospatial foundation models, including SatMAE, CROMA, and GeoCLIP, learn transferable representations from large\-scale Earth\-observation data\(Conget al\.[2022](https://arxiv.org/html/2608.04393#bib.bib17); Fulleret al\.[2023](https://arxiv.org/html/2608.04393#bib.bib18); Cepedaet al\.[2023](https://arxiv.org/html/2608.04393#bib.bib19)\)\. These methods primarily improve feature representation and transfer across geospatial tasks, whereas our work addresses the distinct setting in which the supervision target is generated by a known scientific equation\. Related physics\-informed and structure\-aware methods incorporate scientific knowledge through model architectures, constraints, or training objectives\(Raissiet al\.[2019](https://arxiv.org/html/2608.04393#bib.bib20); Karniadakiset al\.[2021](https://arxiv.org/html/2608.04393#bib.bib22); Willardet al\.[2022](https://arxiv.org/html/2608.04393#bib.bib21)\)\. In our setting, the known equation directly defines the supervision target and is retained as the prediction anchor rather than serving only as an auxiliary constraint\. We further compare statistical and convolutional contextual encoders to determine whether predictive gains arise from neighborhood summaries, raw local values, or their spatial arrangement\.
Robustness under degraded features\.Prior work addresses missing, noisy, or corrupted inputs through imputation, augmentation, robust losses, and corruption\-based evaluation\(Yoonet al\.[2018](https://arxiv.org/html/2608.04393#bib.bib23); Shorten and Khoshgoftaar[2019](https://arxiv.org/html/2608.04393#bib.bib25); Huber[1964](https://arxiv.org/html/2608.04393#bib.bib26); Hendrycks and Dietterich[2019](https://arxiv.org/html/2608.04393#bib.bib24)\)\. We instead use controlled degradation as both a robustness evaluation and a diagnostic for proxy supervision\. We compare complete, noisy, coarsened, and masked factor regimes, together with a separate missing\-factor stress test, to distinguish equation reconstruction from degraded\-factor robustness\.
## Problem Setup and Diagnostics
### Proxy Supervision and Equation Reconstruction
We study supervised prediction of a*proxy*target constructed from known scientific factors rather than from independent field measurements\. For sampleii,
Ai=g\(fi,1,…,fi,m\),A\_\{i\}=g\(f\_\{i,1\},\\ldots,f\_\{i,m\}\),\(2\)whereg\(⋅\)g\(\\cdot\)is a fixed generating function andfi,jf\_\{i,j\}denotes the value of factorjjat sampleii\. In our RUSLE setting,
Ai=Ri×Ki×LSi×Ci,A\_\{i\}=R\_\{i\}\\times K\_\{i\}\\times LS\_\{i\}\\times C\_\{i\},\(3\)whereAiA\_\{i\}is taken from the precomputed soil\-loss raster\.
When the same factors used to defineAiA\_\{i\}are also provided as model inputs, high accuracy may reflect*complete\-factor proxy reconstruction*—that is, approximation ofg\(⋅\)g\(\\cdot\)itself—rather than robust use of contextual information under imperfect factors\. We therefore evaluate each predictor against aformula reference
A^formula,i=Ri×K~i×LSi×Ci,\\hat\{A\}\_\{\\mathrm\{formula\},i\}=R\_\{i\}\\times\\tilde\{K\}\_\{i\}\\times LS\_\{i\}\\times C\_\{i\},\(4\)whereK~i\\tilde\{K\}\_\{i\}is the degraded center\-pixel value, or a fixed training\-set fallback whenKKis absent as described in the following subsection\. UnderKfullK\_\{\\mathrm\{full\}\}, this reference nearly recoversAiA\_\{i\}and serves as a reconstruction benchmark\. Under degraded regimes, comparison with the corresponding degraded formula reference isolates gains beyond direct equation evaluation\. We do not assess validity against independently observed soil\-loss measurements\.
### Degraded\-Factor Regimes
We perturb only the erodibility factorKK, denoting byK~\\tilde\{K\}the version available to both the predictors and the formula reference\. The main matched evaluation uses five regimes:KfullK\_\{\\mathrm\{full\}\},Knoise\_020K\_\{\\mathrm\{noise\\\_020\}\},Kcoarse\_8K\_\{\\mathrm\{coarse\\\_8\}\},Kmask\_050K\_\{\\mathrm\{mask\\\_050\}\}, andKcenter\_mask\_050K\_\{\\mathrm\{center\\\_mask\\\_050\}\}\. Complete absence ofKKis evaluated separately throughKmissingK\_\{\\mathrm\{missing\}\}\.
Raster\-level regimes\.ForKfullK\_\{\\mathrm\{full\}\},K~=K\\tilde\{K\}=K\. ForKnoise\_020K\_\{\\mathrm\{noise\\\_020\}\},K~=Kexp\(ϵ\)\\tilde\{K\}=K\\exp\(\\epsilon\), whereϵ∼𝒩\(0,0\.202\)\\epsilon\\sim\\mathcal\{N\}\(0,0\.20^\{2\}\)\. ForKcoarse\_8K\_\{\\mathrm\{coarse\\\_8\}\}, we apply nodata\-aware8×88\\times 8block averaging followed by piecewise\-constant upsampling\. Invalid pixels are excluded from the block means, and the original nodata mask is restored\. Noise and coarsening are applied to the completeKKraster before window extraction\.
Sample\-level masking\.LetK¯train\\bar\{K\}\_\{\\mathrm\{train\}\}denote the mean center\-pixelKKvalue over the training samples\. InKmask\_050K\_\{\\mathrm\{mask\\\_050\}\}, approximately50%50\\%of the samples in each split are masked independently\. For each masked sample, the completeKKwindow and the center value used by the formula reference are replaced byK¯train\\bar\{K\}\_\{\\mathrm\{train\}\}\. InKcenter\_mask\_050K\_\{\\mathrm\{center\\\_mask\\\_050\}\}, the same masking rate is used, but only the center pixel and the corresponding formula\-reference value are replaced; the neighboringKKvalues remain available\. Unmasked samples retain their originalKKvalues\. The reliability vector𝐳\\mathbf\{z\}contains a mask or missingness indicator, a one\-hot degradation tag, sevenKK\-neighborhood statistics whenKKis observed, and the absolute center\-to\-neighborhood\-mean difference\|K~center−K~mean\|\\left\|\\tilde\{K\}\_\{\\mathrm\{center\}\}\-\\tilde\{K\}\_\{\\mathrm\{mean\}\}\\right\|\. AllKK\-derived entries are set to zero underKmissingK\_\{\\mathrm\{missing\}\}\. RASPL\-MLP\-STATS and RASPL\-CNN\-RAW\+STATS use the full reliability vector in both the gating and residual branches\. For RASPL\-MLP\-CENTER, the gate conditions on the full reliability vector, whereas the residual branch combines the four center\-pixel channels with a reduced eight\-dimensional reliability subset consisting of the mask or missingness indicator and the one\-hot degradation tag\. The residual branches of RASPL\-MLP\-FLAT and RASPL\-CNN\-RAW likewise use a reduced eight\-dimensional non\-statistical subset of the reliability vector together with their respective flattened\-window and convolutional representations\.
Missing\-KKstress test\.UnderKmissingK\_\{\\mathrm\{missing\}\}, observedKKis omitted from the spatial window, which becomes\[R,LS,C\]\[R,LS,C\]together with a constant missingness\-indicator plane\. The formula reference is
A^formula,i=Ri×K¯train×LSi×Ci\.\\hat\{A\}\_\{\\mathrm\{formula\},i\}=R\_\{i\}\\times\\bar\{K\}\_\{\\mathrm\{train\}\}\\times LS\_\{i\}\\times C\_\{i\}\.\(5\)Results forKmissingK\_\{\\mathrm\{missing\}\}are reported separately and excluded from the main degradation averages and rankings\.
### Diagnostic Criteria
We use five diagnostics\.Formula reconstructiontests whether strong performance underKfullK\_\{\\mathrm\{full\}\}is explained by direct recovery of the generating equation\.Formula preservation\(Q1\) compares matched direct, formula\-feature, and RASPL predictors using statistical MLP and convolutional CNN encoders\.Spatial\-context gain\(Q2\) compares center\-only, neighborhood\-statistics, flattened\-window, and convolutional\-window representations\.Tail robustnessis evaluated using Tail95 MAE and the Tail95 underprediction rate over positive samples whose target values are at or above the9595th percentile\. Underprediction is defined asA^<0\.8A\.\\hat\{A\}<0\.8A\.Degradation robustnessis summarized across regimes using the degradation robustness score \(DRS\):
DRS=0\.35Rall2~\+0\.25Rpos2~\+0\.25T95MAE~\+0\.15T95Under~\.\\begin\{array\}\[\]\{rl\}\\mathrm\{DRS\}=\{\}&0\.35\\,\\widetilde\{R^\{2\}\_\{\\mathrm\{all\}\}\}\+0\.25\\,\\widetilde\{R^\{2\}\_\{\\mathrm\{pos\}\}\}\\\\\[2\.0pt\] &\+0\.25\\,\\widetilde\{\\mathrm\{T95MAE\}\}\+0\.15\\,\\widetilde\{\\mathrm\{T95Under\}\}\.\\end\{array\}\(6\)Each component is min–max normalized within the stated comparison pool, separately for each regime across models\. The tail\-error components are direction\-reversed so that higher values are consistently better\. The component metrics are also reported separately\. For RASPL, stability is evaluated using the effective correction\|αδ\|\\left\|\\alpha\\delta\\right\|, rather than the raw residualδ\\deltaalone as defined in the RASPL Method section\.
## RASPL Method
### Motivation
Under complete factors, strong proxy\-prediction performance may reflect reconstruction of the generating equation rather than meaningful use of contextual information\. Under degraded factors, however, accurate prediction requires compensation for corrupted or missing inputs\. Direct predictors can use contextual information, but they treat the proxy as an ordinary regression target and do not preserve the known formula structure\. RASPL instead retains the available formula estimate as the prediction anchor and learns when contextual and reliability information should correct it\. To isolate the role of contextual representation, we compare statistical and convolutional residual encoders\. MLP\-STATS tests whether neighborhood summary statistics are sufficient for correction, whereas CNN\-RAW\+STATS tests whether the spatial arrangement of local values provides additional information\. Center\-only, flattened\-window, and raw\-window variants serve as controlled ablations\.
Figure 1:Overview of RASPL\. The degraded formula estimate is retained as the prediction anchor, and an adaptively gated contextual residual is added in log space before the prediction is mapped back to the proxy scale\.
### Formula\-Preserving Residual Prediction
As shown in Figure[1](https://arxiv.org/html/2608.04393#Sx4.F1), RASPL combines a fixed formula estimate, a contextual residual encoder, and an adaptive gate\. LetWiW\_\{i\}denote the regime\-specific local input and𝐳i\\mathbf\{z\}\_\{i\}the reliability vector defined in the Degraded\-Factor Regimes subsection\. The formula estimateAformula,iA\_\{\\mathrm\{formula\},i\}uses the degraded center\-pixel valueK~i\\tilde\{K\}\_\{i\}whenKKis observed and the training\-mean fallbackK¯train\\bar\{K\}\_\{\\mathrm\{train\}\}underKmissingK\_\{\\mathrm\{missing\}\}\.
Because the proxy is nonnegative and skewed, prediction is performed in log space:
yi=log\(1\+Ai\),yformula,i=log\(1\+Aformula,i\)\.y\_\{i\}=\\log\(1\+A\_\{i\}\),\\qquad y\_\{\\mathrm\{formula\},i\}=\\log\(1\+A\_\{\\mathrm\{formula\},i\}\)\.The model produces a raw residual proposalΔθ\(Wi,𝐳i\)\\Delta\_\{\\theta\}\(W\_\{i\},\\mathbf\{z\}\_\{i\}\)and a gate value
qθ\(𝐳i\)=σ\(ℓθ\(𝐳i\)\),αθ\(𝐳i\)=1−qθ\(𝐳i\),q\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\)=\\sigma\\\!\\left\(\\ell\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\)\\right\),\\qquad\\alpha\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\)=1\-q\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\),whereℓθ\\ell\_\{\\theta\}is a learned scalar logit computed from the reliability vector\. The predicted log value and proxy\-scale prediction are
y^i=clip\(yformula,i\+αθ\(𝐳i\)Δθ\(Wi,𝐳i\),−5,5\),A^i=max\{0,exp\(y^i\)−1\}\.\\begin\{array\}\[\]\{rcl\}\\hat\{y\}\_\{i\}&=&\\mathrm\{clip\}\\\!\\left\(y\_\{\\mathrm\{formula\},i\}\+\\alpha\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\)\\Delta\_\{\\theta\}\(W\_\{i\},\\mathbf\{z\}\_\{i\}\),\-5,5\\right\),\\\\\[4\.0pt\] \\hat\{A\}\_\{i\}&=&\\max\\left\\\{0,\\exp\(\\hat\{y\}\_\{i\}\)\-1\\right\\\}\.\\end\{array\}The effective correction applied to the formula estimate is thereforealphaθ\(𝐳i\)Δθ\(Wi,𝐳i\),alpha\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\)\\Delta\_\{\\theta\}\(W\_\{i\},\\mathbf\{z\}\_\{i\}\),rather than the raw residual alone\. A small value ofαθ\\alpha\_\{\\theta\}retains the formula estimate, whereas a large value permits a stronger contextual correction\. The gate is learned end\-to\-end without reliability labels and is interpreted as a residual\-allocation weight rather than calibrated uncertainty\.
### Contextual Residual Encoders
MLP\-STATS uses the full reliability vector𝐳i\\mathbf\{z\}\_\{i\}, including neighborhood summary statistics, in both the residual and gating branches\. This encoder tests whether compact statistical context is sufficient to correct the degraded formula estimate\.
CNN\-RAW\+STATS applies a compact two\-layer CNN to the raw local inputWiW\_\{i\}, pools the resulting representation, concatenates it with the full reliability vector𝐳i\\mathbf\{z\}\_\{i\}, and predicts the residual proposalΔθ\\Delta\_\{\\theta\}\. Its gate valueqθq\_\{\\theta\}also conditions on the full reliability vector\. This encoder tests whether the spatial arrangement of local values provides information beyond neighborhood summaries\.
The remaining variants use the same formula\-preserving prediction rule with different contextual inputs\. MLP\-CENTER uses the four center\-pixel channels in the residual branch; its gate conditions on the full reliability vector, whereas its residual branch uses a reduced eight\-dimensional non\-statistical reliability subset\. MLP\-FLAT and CNN\-RAW use flattened\-window and convolutional raw\-window encoders, respectively, together with reduced eight\-dimensional non\-statistical reliability inputs rather than the full reliability vector\.
### Training Objective
All four channels ofWWare standardized using statistics from the training split\. In𝐳\\mathbf\{z\}, only the continuousKK\-derived entries are standardized; the indicator and one\-hot degradation entries are left unchanged\. RASPL is trained against the unstandardized log targetsyiy\_\{i\}and formula estimatesyformula,iy\_\{\\mathrm\{formula\},i\}\. The matched direct baselines instead predict a training\-standardized log target, while the formula\-feature baseline receives a training\-standardized version oflog\(1\+Aformula\)\\log\(1\+A\_\{\\mathrm\{formula\}\}\)as an additional scalar input\. RASPL minimizes a weighted log\-space regression loss with
wi=1\+λpos𝟏\[Ai\>0\]\+λtail𝟏\[Ai≥Q95train\],w\_\{i\}=1\+\\lambda\_\{\\mathrm\{pos\}\}\\mathbf\{1\}\[A\_\{i\}\>0\]\+\\lambda\_\{\\mathrm\{tail\}\}\\mathbf\{1\}\[A\_\{i\}\\geq Q\_\{95\}^\{\\mathrm\{train\}\}\],whereQ95trainQ\_\{95\}^\{\\mathrm\{train\}\}is the9595th percentile of the positive training targets when more than ten positive samples are available and the9595th percentile of all training targets otherwise\. We setλpos=0\.5\\lambda\_\{\\mathrm\{pos\}\}=0\.5andλtail=1\.0\\lambda\_\{\\mathrm\{tail\}\}=1\.0\. The main loss is
ℒmain=1N∑i=1Nwi\(y^i−yi\)2\.\\mathcal\{L\}\_\{\\mathrm\{main\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}w\_\{i\}\(\\hat\{y\}\_\{i\}\-y\_\{i\}\)^\{2\}\.
A residual penalty discourages large raw corrections when the gate favors retention of the formula estimate:
ℒres=1N∑i=1Nqθ\(𝐳i\)Δθ\(Wi,𝐳i\)2\.\\mathcal\{L\}\_\{\\mathrm\{res\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}q\_\{\\theta\}\(\\mathbf\{z\}\_\{i\}\)\\,\\Delta\_\{\\theta\}\(W\_\{i\},\\mathbf\{z\}\_\{i\}\)^\{2\}\.The total objective isℒ=ℒmain\+λresℒres,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{main\}\}\+\\lambda\_\{\\mathrm\{res\}\}\\mathcal\{L\}\_\{\\mathrm\{res\}\},withλres=0\.01\\lambda\_\{\\mathrm\{res\}\}=0\.01\.
### Matched Direct and Formula\-Feature Baselines
The direct baselines remove the formula anchor, adaptive gate, and residual addition while retaining the same contextual encoders and regime\-matched inputs as their corresponding RASPL variants\. They are trained to predict the training\-standardized log target directly, and inverse standardization maps their outputs back to the proxy scale\.
The matched formula\-feature baseline additionally receives the training\-standardized log formula estimate as an ordinary input feature but still predicts the complete target directly\. It does not preserve the formula estimate through an explicit residual connection\. Comparing these baselines with RASPL isolates the effect of retaining the proxy\-generating equation as the prediction anchor rather than ignoring it or treating it as an ordinary input feature\.
## Experimental Setup
Study Setting and Proxy Target\.We evaluate the framework on the RUSLE\-derived proxy task defined in the Problem Setup and Diagnostics section\. Each sample contains rainfall erosivityRR, soil erodibilityKK, topographic factorLSLS, and cover\-management factorCC, derived from PRISM precipitation\(Dalyet al\.[2002](https://arxiv.org/html/2608.04393#bib.bib31)\), USDA SSURGO\(U\.S\. Department of Agriculture, Natural Resources Conservation Service[2023](https://arxiv.org/html/2608.04393#bib.bib32)\), USGS 3DEP DEM\(U\.S\. Geological Survey[2021](https://arxiv.org/html/2608.04393#bib.bib34)\), and Sentinel\-2 LULC\(Karraet al\.[2021](https://arxiv.org/html/2608.04393#bib.bib16)\), respectively\. The targetAAis therefore a factor\-derived proxy rather than an independently measured soil\-loss observation\. The main experiments use co\-registered3×33\\times 3factor windows, center\-level features, and the reliability vector𝐳i\\mathbf\{z\}\_\{i\}defined in the Degraded\-Factor Regimes subsection; the window\-size study replaces the spatial input with5×55\\times 5and7×77\\times 7windows while holding the remaining protocol fixed\.
Evaluation Regimes\.All matched comparisons use seeds\{42,123,999\}\\\{42,123,999\\\}\. The formula\-preservation comparison \(Q1\) includesKfullK\_\{\\mathrm\{full\}\},Knoise\_020K\_\{\\mathrm\{noise\\\_020\}\},Kcoarse\_8K\_\{\\mathrm\{coarse\\\_8\}\},Kmask\_050K\_\{\\mathrm\{mask\\\_050\}\}, andKcenter\_mask\_050K\_\{\\mathrm\{center\\\_mask\\\_050\}\}, withKfullK\_\{\\mathrm\{full\}\}also serving as the reconstruction diagnostic\. The contextual\-encoder comparison \(Q2\) usesKnoise\_020K\_\{\\mathrm\{noise\\\_020\}\},Kcoarse\_8K\_\{\\mathrm\{coarse\\\_8\}\}, andKcenter\_mask\_050K\_\{\\mathrm\{center\\\_mask\\\_050\}\}\. TheKmissingK\_\{\\mathrm\{missing\}\}regime is evaluated separately as the strongest stress test and is excluded from the shared ranking pools\.
### Models and Baselines
Formula and Tree\-Based References\.Formula references computeAformulaA\_\{\\mathrm\{formula\}\}using the regime\-specific available factors defined in the Degraded\-Factor Regimes subsection; underKmissingK\_\{\\mathrm\{missing\}\}, they use the training\-mean fallbackK¯train\\bar\{K\}\_\{\\mathrm\{train\}\}\. Random Forest and XGBoost references use either center\-level factors or engineered neighborhood statistics comprising the center value, mean, standard deviation, minimum, maximum, median, lower and upper quartiles, range, and local gradient features\. A two\-stage XGBoost model is additionally included in the complete\-factor reconstruction diagnostic\.
Matched Formula\-Preservation Models\.The matched CNN comparison evaluates direct prediction without the formula estimate, direct prediction with the standardized log formula estimate as an ordinary feature, and RASPL with the formula estimate retained as an adaptively gated residual anchor\. The same three\-way comparison is applied to statistics\-only MLP encoders over the Q2 regime subset\. Additional RASPL ablations vary only the contextual representation supplied to the residual and gating components\.
Window\-Size Variants\.The convolutional window\-size study compares RASPL\-CNN\-RAW\+STATS with3×33\\times 3,5×55\\times 5, and7×77\\times 7windows over the Q2 subset while holding the encoder family, channel definitions, seeds, optimization procedure, and checkpoint\-selection rule fixed\.
Training Protocol\.Neural models follow the transformations and objectives defined in the RASPL Method section\. RASPL uses the weighted residual objective on unstandardized log targets and formula estimates, whereas direct and formula\-feature baselines use weighted mean\-squared error on standardized log targets\. Formula\-feature models additionally receive the standardized log formula estimate, and all outputs are inverse\-transformed to the proxy scale\.
All neural models use AdamW with batch size512512, weight decay10−410^\{\-4\}, at most5050epochs, patience55, gradient clipping at1\.01\.0, and learning\-rate selection from\{10−4,3×10−4,10−3\}\\\{10^\{\-4\},3\\times 10^\{\-4\},10^\{\-3\}\\\}\. Checkpoints are selected using the lowest finite validation loss\.
Tree\-based and neural models use identical data partitions\. Matched aggregates include only shared, completed regime–seed cells\.
### Evaluation Metrics
We reportR2R^\{2\}, MAE, and RMSE over all test samples and over the positive\-target subsetAi\>0A\_\{i\}\>0, denoted by the subscriptsall\\mathrm\{all\}andpos\\mathrm\{pos\}, respectively\.
Let
𝒯95=\{i:Ai≥Q95train\},\\mathcal\{T\}\_\{95\}=\\left\\\{i:A\_\{i\}\\geq Q\_\{95\}^\{\\mathrm\{train\}\}\\right\\\},whereQ95trainQ\_\{95\}^\{\\mathrm\{train\}\}is the9595th percentile of the positive training targets when more than ten positive samples are available and the9595th percentile of all training targets otherwise\. The tail metrics are
Tail95MAE=1\|𝒯95\|∑i∈𝒯95\|Ai−A^i\|\\mathrm\{Tail95\\ MAE\}=\\frac\{1\}\{\|\\mathcal\{T\}\_\{95\}\|\}\\sum\_\{i\\in\\mathcal\{T\}\_\{95\}\}\|A\_\{i\}\-\\hat\{A\}\_\{i\}\|and
Tail95Under=1\|𝒯95\|∑i∈𝒯95𝟏\[A^i<0\.8Ai\]\.\\mathrm\{Tail95\\ Under\}=\\frac\{1\}\{\|\\mathcal\{T\}\_\{95\}\|\}\\sum\_\{i\\in\\mathcal\{T\}\_\{95\}\}\\mathbf\{1\}\\left\[\\hat\{A\}\_\{i\}<0\.8A\_\{i\}\\right\]\.
The degradation robustness score combines normalizedRall2R^\{2\}\_\{\\mathrm\{all\}\},Rpos2R^\{2\}\_\{\\mathrm\{pos\}\}, Tail95 MAE, and Tail95 underprediction using the weights defined in the Diagnostic Criteria subsection\. Normalization is performed separately within each regime and comparison pool, with the tail\-error components direction\-reversed so that higher DRS is better\. Consequently, DRS values from different normalization pools are not directly comparable\.
## Results and Analysis
Complete\-Factor Reconstruction and Degradation Sensitivity\.Table[1](https://arxiv.org/html/2608.04393#Sx6.T1)shows that the formula reference attains exactlyRall2=1\.0000R^\{2\}\_\{\\mathrm\{all\}\}=1\.0000underKfullK\_\{\\mathrm\{full\}\}, confirming that the proxy target is exactly reconstructable from the generating factors when they are fully available\. The matched tree baselines likewise achieve strong complete\-factor performance, with three\-seed meanRall2R^\{2\}\_\{\\mathrm\{all\}\}ranging from0\.83870\.8387for XGB\-center to0\.94710\.9471for two\-stage XGBoost\. Mild multiplicative noise inKKproduces only small reductions: the formula reference remains at0\.95600\.9560, while RF\-center, XGB\-center, and two\-stage XGBoost attain0\.87940\.8794,0\.83120\.8312, and0\.92520\.9252, respectively\. Spatial coarsening reducesRall2R^\{2\}\_\{\\mathrm\{all\}\}for the formula reference to0\.92590\.9259and for all three learned baselines to values between0\.75350\.7535and0\.82500\.8250\.
Broad masking produces the largest deterioration\. The formula reference decreases to a three\-seed mean of0\.51140\.5114, while RF\-center, XGB\-center, and two\-stage XGBoost decrease to0\.46220\.4622,0\.32560\.3256, and0\.42730\.4273, respectively\. RF\-center, XGB\-center, and two\-stage XGBoost use center\-level factor information and therefore receive the same effective centerKKinput underKmaskK\_\{\\mathrm\{mask\}\}andKcenter\-maskK\_\{\\mathrm\{center\\mbox\{\-\}mask\}\}, yielding identical results across the two regimes\. The formula reference, which is also computed from center\-pixel factors, exhibits the same behavior\. Together, these results demonstrate that strong complete\-factor accuracy can reflect exploitation of the deterministic proxy\-generating relationship and does not, by itself, establish robustness to degraded factor information\.
Table 1:Complete\-factor reconstruction andKK\-degradation sensitivity\. Values are meanRall2R^\{2\}\_\{\\mathrm\{all\}\}over seeds\{42,123,999\}\\\{42,123,999\\\}\. Columns correspond toKfullK\_\{\\mathrm\{full\}\},KnoiseK\_\{\\mathrm\{noise\}\},KcoarseK\_\{\\mathrm\{coarse\}\},KmaskK\_\{\\mathrm\{mask\}\}, andKcenter\-maskK\_\{\\mathrm\{center\\mbox\{\-\}mask\}\}, respectively\.Formula Preservation\.We first isolate the effect of formula preservation using matched statistical encoders evaluated overKnoise\_020K\_\{\\mathrm\{noise\\\_020\}\},Kcoarse\_8K\_\{\\mathrm\{coarse\\\_8\}\}, andKcenter\_mask\_050K\_\{\\mathrm\{center\\\_mask\\\_050\}\}with seeds\{42,123,999\}\\\{42,123,999\\\}\. The three models use the same neighborhood\-statistics representation but differ in how they incorporate the degraded formula estimate: the direct model does not receive it, the formula\-feature model treats it as an ordinary input feature, and RASPL retains it as the prediction anchor\.
Table 2:Matched statistical\-encoder comparison over three degraded\-factor regimes and three seeds\. DRS is normalized within this focused three\-model comparison pool\.Table[2](https://arxiv.org/html/2608.04393#Sx6.T2)shows that neighborhood statistics alone are insufficient for reliable prediction in this matched MLP setting\. Adding the degraded formula estimate as an input feature increases macroRall2R^\{2\}\_\{\\mathrm\{all\}\}from0\.10170\.1017to0\.81880\.8188and reduces Tail95 MAE from1\.15731\.1573to0\.29940\.2994, indicating that access to the formula estimate accounts for most of the recovery in average accuracy\. Explicitly preserving the formula estimate through RASPL provides a further increase inRall2R^\{2\}\_\{\\mathrm\{all\}\}to0\.83430\.8343and reduces Tail95 MAE to0\.23850\.2385\.
Across the nine matched regime–seed cells, RASPL\-MLP\-STATS exceeds Direct\-MLP\-STATS inRall2R^\{2\}\_\{\\mathrm\{all\}\}and Tail95 MAE in all nine cells, with mean improvements of0\.7330\.733inRall2R^\{2\}\_\{\\mathrm\{all\}\}and0\.9190\.919in Tail95 MAE\. Relative to the formula\-feature model, RASPL improves both metrics in eight of nine cells, with mean improvements of0\.0150\.015inRall2R^\{2\}\_\{\\mathrm\{all\}\}and0\.0610\.061in Tail95 MAE\. Tail95 underprediction is more mixed, with RASPL improving five of nine cells, indicating that tail\-error magnitude and underprediction frequency capture related but distinct behavior\.
Table 3:Matched statistical\-encoder differences\. PositiveΔRall2\\Delta R^\{2\}\_\{\\mathrm\{all\}\}and positive Tail95 reductions favor the first model\.The matched convolutional comparison shows the same gain over direct prediction, but a more nuanced tradeoff relative to the formula\-feature baseline\. Across the 15 matched regime–seed cells, RASPL\-CNN\-RAW\+STATS exceeds Direct\-CNN\-RAW\+STATS by0\.34270\.3427in meanRall2R^\{2\}\_\{\\mathrm\{all\}\}, reduces Tail95 MAE by0\.24820\.2482, and achieves a paired DRS improvement of0\.89500\.8950\. Relative to the formula\-feature CNN, the averageRall2R^\{2\}\_\{\\mathrm\{all\}\}is slightly lower for RASPL \(0\.80260\.8026versus0\.80600\.8060\), but RASPL yields lower Tail95 MAE \(0\.23290\.2329versus0\.27970\.2797\), lower Tail95 underprediction \(0\.20150\.2015versus0\.24120\.2412\), and stronger paired degradation robustness\. Thus, formula preservation does not uniformly maximize every individual metric, but it provides the more favorable robustness–tail\-error tradeoff in the matched CNN setting\.
Table 4:Contextual\-encoder and window\-size comparison over three degradedKKregimes and seeds\{42,123,999\}\\\{42,123,999\\\}\. DRS is normalized within the common four\-model comparison pool\.
Figure 2:Spatial comparison underKmissingK\_\{\\mathrm\{missing\}\}: \(a\) proxy target, \(b\) formula\-reference prediction usingK¯train\\bar\{K\}\_\{\\mathrm\{train\}\}, \(c\) RASPL\-CNN\-RAW\+STATS prediction, \(d\)–\(e\) absolute errors, and \(f\) formula\-reference error minus RASPL error\. Panels \(a\)–\(c\) use a sharedlog\(1\+x\)\\log\(1\+x\)scale, panels \(d\)–\(e\) uselog\(1\+\|error\|\)\\log\(1\+\|\\mathrm\{error\}\|\), and panel \(f\) is symmetrically clipped at the99\.599\.5thpercentile; positive \(warm\) values favor RASPL, whereas negative \(cool\) values favor the formula reference\.
Contextual Encoder Tradeoff\.Table[2](https://arxiv.org/html/2608.04393#Sx6.F2)compares the statistical RASPL encoder with convolutional encoders using3×33\\times 3,5×55\\times 5, and7×77\\times 7local windows\. All results are macro\-averaged over the same three degraded\-factor regimes and three seeds\. The reported DRS values are normalized within the common four\-model comparison pool and therefore differ from the focused DRS values in Table[2](https://arxiv.org/html/2608.04393#Sx6.T2)\.
RASPL\-MLP\-STATS attains the highest macroRall2R^\{2\}\_\{\\mathrm\{all\}\}and uses only3,4583\{,\}458parameters, compared with32,16232\{,\}162for each CNN variant\. Its mean historical training wall time is549\.97549\.97seconds, compared with834\.16834\.16seconds for the3×33\\times 3CNN\. These results identify the statistical encoder as the most favorable average\-accuracy and computational\-efficiency operating point\.
RASPL\-CNN\-RAW\+STATS with a3×33\\times 3window instead achieves the highestRpos2R^\{2\}\_\{\\mathrm\{pos\}\}, the lowest Tail95 MAE, the lowest Tail95 underprediction rate, and the highest DRS\. The convolutional encoder therefore provides the stronger degradation\- and tail\-robustness operating point, despite its slightly lower macroRall2R^\{2\}\_\{\\mathrm\{all\}\}and higher computational cost\.
Window\-Size Ablation\.Increasing the CNN window beyond3×33\\times 3does not improve aggregate performance\. Relative to3×33\\times 3, the5×55\\times 5model decreasesRall2R^\{2\}\_\{\\mathrm\{all\}\}by0\.0326360\.032636and increases Tail95 MAE by0\.0403270\.040327\. The7×77\\times 7model decreasesRall2R^\{2\}\_\{\\mathrm\{all\}\}by0\.0213330\.021333and increases Tail95 MAE by0\.0366090\.036609\. Both larger windows also produce lowerRpos2R^\{2\}\_\{\\mathrm\{pos\}\}, higher Tail95 underprediction, and lower DRS\.
The CNN parameter count remains fixed across window sizes, but the larger inputs increase computational cost\. Relative to3×33\\times 3, multiply–accumulate operations per sample increase by factors of2\.702\.70and5\.255\.25for5×55\\times 5and7×77\\times 7, respectively\. Because neither larger window improves accuracy or tail robustness,3×33\\times 3is retained as the convolutional configuration\.
Missing\-Factor Stress Test\.TheKmissingK\_\{\\mathrm\{missing\}\}regime is evaluated separately as the strongest stress test and is not included in the shared main\-ranking tables\. Figure[2](https://arxiv.org/html/2608.04393#Sx6.F2)visualizes the selected3×33\\times 3RASPL\-CNN\-RAW\+STATS configuration on the non\-overlapping256×256256\\times 256candidate tile containing the largest number of Tail95 samples, with ties resolved by the smallest top\-left raster index\. This predefined rule avoids selecting the visualization according to appearance or model performance\. On this tile, the formula reference yields regional MAE0\.20230\.2023and Tail95 MAE1\.16051\.1605, whereas RASPL\-CNN\-RAW\+STATS reduces these to0\.17890\.1789and1\.05741\.0574, respectively\. RASPL produces lower absolute error than the formula reference on59\.70%59\.70\\%of valid cells, with a mean absolute\-error reduction of0\.02340\.0234\. The Tail95 underprediction rate nevertheless remains high for both models \(1\.00001\.0000for the formula reference and0\.97320\.9732for RASPL\), showing that complete removal ofKKremains a severe stress test even when average error improves\.
## Conclusions and Future Work
Using a RUSLE\-derived soil\-loss proxy, we showed that high complete\-factor accuracy can reflect equation reconstruction rather than robustness to degraded factors\. We introduced a diagnostic framework and RASPL, which retains the available formula estimate as an adaptively gated prediction anchor\. Matched experiments show that formula\-aware models outperform direct prediction and that explicit formula preservation yields a stronger degradation–tail\-error tradeoff than treating the formula estimate as an ordinary feature\. RASPL\-MLP\-STATS provides the best accuracy–efficiency tradeoff, whereas RASPL\-CNN\-RAW\+STATS with a3×33\\times 3window provides the strongest degradation and tail robustness\. Future work will test RASPL across regions and other factor\-derived scientific tasks\.
## References
- W\. Abiye and O\. Dengiz \(2025\)Digital mapping of soil erodibility factor in response to land use change using machine learning models\.Environmental Systems Research14\.External Links:[Document](https://dx.doi.org/10.1186/s40068-025-00402-w)Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- R\. Benavidez, B\. Jackson, D\. Maxwell, and K\. Norton \(2018\)A review of the \(revised\) universal soil loss equation \(r/usle\): with a view to increasing its global applicability and improving soil loss estimates\.Hydrology and Earth System Sciences22,pp\. 6059–6086\.External Links:[Document](https://dx.doi.org/10.5194/hess-22-6059-2018)Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- P\. Borrelli, D\. A\. Robinson, L\. R\. Fleischer, E\. Lugato, C\. Ballabio, C\. Alewell, K\. Meusburger, S\. Modugno, B\. Schütt, V\. Ferro, V\. Bagarello, K\. Van Oost, L\. Montanarella, and P\. Panagos \(2017\)An assessment of the global impact of 21st century land use change on soil erosion\.Nature Communications8\(1\),pp\. 2013\.External Links:[Document](https://dx.doi.org/10.1038/s41467-017-02142-7)Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p1.1)\.
- A\. Bouamrane, H\. Boutaghane, A\. Bouamrane, N\. Dahri, H\. Abida, M\. Saber, S\. A\. Kantoush, and T\. Sumi \(2024\)Soil erosion susceptibility prediction using ensemble hybrid models with multicriteria decision\-making analysis: case study of the medjerda basin, northern africa\.International Journal of Sediment Research39\(6\),pp\. 998–1014\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- V\. V\. Cepeda, G\. K\. Nayak, and M\. Shah \(2023\)GeoCLIP: clip\-inspired alignment between locations and images for effective worldwide geo\-localization\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.
- Y\. Cong, S\. Khanna, C\. Meng, P\. Liu, E\. Rozi, Y\. He, M\. Burke, D\. B\. Lobell, and S\. Ermon \(2022\)SatMAE: pre\-training transformers for temporal and multi\-spectral satellite imagery\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.
- C\. Daly, W\. Gibson, G\. Taylor, G\. Johnson, and P\. Pasteris \(2002\)A knowledge\-based approach to the statistical mapping of climate\.Climate Research22,pp\. 99–113\.Cited by:[Experimental Setup](https://arxiv.org/html/2608.04393#Sx5.p1.9)\.
- A\. Fuller, K\. Millard, and J\. R\. Green \(2023\)CROMA: remote sensing representations with contrastive radar\-optical masked autoencoders\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.
- Y\. Ge, L\. Zhao, J\. Chen, X\. Li, H\. Li, Z\. Wang, and Y\. Ren \(2023\)Study on soil erosion driving forces by using \(r\)usle framework and machine learning: a case study in southwest china\.Land12\(3\)\.Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p2.8),[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- T\. B\. Gelete, P\. Pasala, N\. G\. Abay, G\. W\. Woldemariam, K\. H\. Yasin, E\. Kebede, and I\. Aliyi \(2024\)Integrated machine learning and geospatial analysis enhanced gully erosion susceptibility modeling in the erer watershed in eastern ethiopia\.Frontiers in Environmental ScienceVolume 12 \- 2024\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- D\. Hendrycks and T\. Dietterich \(2019\)Benchmarking neural network robustness to common corruptions and perturbations\.InInternational Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p3.1)\.
- M\. Hlal, B\. El Monhim, J\. Chenal, J\. B\. Munyaka, R\. Azmi, A\. Sbai, G\. Cwick, and B\. B\. Hichou \(2025\)Application of deep learning and geospatial analysis in soil loss risk in the moulouya watershed, morocco\.Water17\(9\)\.Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p2.8),[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- P\. J\. Huber \(1964\)Robust Estimation of a Location Parameter\.The Annals of Mathematical Statistics35\(1\),pp\. 73 – 101\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177703732)Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p3.1)\.
- A\. F\. İpek and E\. Kahya \(2026\)Spatiotemporal prioritization of soil erosion risk using the rusle model and cmip6 projections under future climate scenarios in a mediterranean watershed\.Frontiers in Environmental Science14\.External Links:[Document](https://dx.doi.org/10.3389/fenvs.2026.1760569)Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p2.7)\.
- F\. Islam, T\. Bibi, N\. U\. Rehman, J\. B\. Davis, R\. W\. Aslam, N\. Y\. Rebouh, H\. Elmannai, and A\. Tariq \(2025\)GIS\- and rs\-based models for gully erosion susceptibility mapping using machine learning and remote sensing data\.Earth Surface Processes and Landforms50\(15\)\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- G\. E\. Karniadakis, I\. G\. Kevrekidis, L\. Lu, P\. Perdikaris, S\. Wang, and L\. Yang \(2021\)Physics\-informed machine learning\.Nature Reviews Physics3\(6\),pp\. 422–440\.External Links:[Document](https://dx.doi.org/10.1038/s42254-021-00314-5)Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.
- K\. Karra, C\. Kontgis, Z\. Statman\-Weil, J\. C\. Mazzariello, M\. Mathis, and S\. P\. Brumby \(2021\)Global land use / land cover with sentinel 2 and deep learning\.In2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS,Vol\.,pp\. 4704–4707\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1),[Experimental Setup](https://arxiv.org/html/2608.04393#Sx5.p1.9)\.
- P\. Li, A\. Tariq, Q\. Li, B\. Ghaffar, M\. Farhan, A\. Jamil, W\. Soufan, A\. El Sabagh, and M\. Freeshah \(2023\)Soil erosion assessment by RUSLE model using remote sensing and GIS in an arid zone\.International Journal of Digital Earth16\(1\),pp\. 3105–3124\.External Links:[Document](https://dx.doi.org/10.1080/17538947.2023.2243916)Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- J\. Momeni Damaneh, A\. A\. Safdari, N\. Azarnejad, M\. Ghorbani, F\. Panahi, S\. F\. Afzali, and S\. Loppi \(2025\)Modeling soil erosion susceptibility using machine learning techniques: rud\-e\-faryab basin, iran\.Land Degradation & Development36\(18\),pp\. 6396–6409\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- M\. R\. Olii, A\. K\. Zailani Olii, A\. Olii, R\. A\. Djau, M\. A\. Mokoagow, B\. A\. Kironoto, B\. Bachtiar, R\. S\. N\. Olii, and R\. Pakaya \(2025\)Tree\-based machine learning algorithms for soil erosion vulnerability \(sev\) prediction in saddang watershed, south sulawesi, indonesia\.Journal of Water and Climate Change16\(4\),pp\. 1459–1476\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- P\. Panagos, P\. Borrelli, J\. Poesen, C\. Ballabio, E\. Lugato, K\. Meusburger, L\. Montanarella, and C\. Alewell \(2015\)The new assessment of soil loss by water erosion in europe\.Environmental Science & Policy54,pp\. 438–447\.External Links:[Document](https://dx.doi.org/10.1016/j.envsci.2015.08.012)Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- M\. Raissi, P\. Perdikaris, and G\.E\. Karniadakis \(2019\)Physics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational Physics378,pp\. 686–707\.External Links:ISSN 0021\-9991,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jcp.2018.10.045),[Link](https://www.sciencedirect.com/science/article/pii/S0021999118307125)Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.
- N\. Samarinas, N\. L\. Tsakiridis, E\. Kalopesa, and G\. C\. Zalidis \(2024\)Soil loss estimation by water erosion in agricultural areas introducing artificial intelligence geospatial layers into the rusle model\.Land13\(2\)\.Cited by:[Introduction](https://arxiv.org/html/2608.04393#Sx1.p2.8),[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- C\. Shorten and T\. M\. Khoshgoftaar \(2019\)A survey on image data augmentation for deep learning\.Journal of Big Data6\(60\)\.External Links:[Document](https://dx.doi.org/10.1186/s40537-019-0197-0)Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p3.1)\.
- N\. Tkeshelashvili \(2024\)Empirical and machine learning models for soil erosion risk assessment: a case study of tsageri municipality, georgia\.Journal of Geography, Environment and Earth Science International28\(11\),pp\. 148–162\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- U\.S\. Department of Agriculture, Natural Resources Conservation Service \(2023\)Soil survey geographic \(ssurgo\) database\.Note:https://www\.nrcs\.usda\.gov/resources/data\-and\-reports/soil\-survey\-geographic\-database\-ssurgoAccessed 2025\-12\-14Cited by:[Experimental Setup](https://arxiv.org/html/2608.04393#Sx5.p1.9)\.
- U\.S\. Geological Survey \(2021\)USGS 1/3 Arc Second n39w108 20210312\.Note:U\.S\. Geological Survey, The National MapAccessed 2025\-12\-20External Links:[Link](https://arxiv.org/html/2608.04393v1/PASTE_THE_EXACT_TILE_PAGE_URL_HERE)Cited by:[Experimental Setup](https://arxiv.org/html/2608.04393#Sx5.p1.9)\.
- P\. V\.M\., G\. Aldehim, N\. Negm, and S\. Subathradevi \(2025\)Integrating geospatial techniques and machine learning for assessing soil erosion and associated geomorphic risks\.Journal of South American Earth Sciences156,pp\. 105463\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- J\. Willard, X\. Jia, S\. Xu, M\. Steinbach, and V\. Kumar \(2022\)Integrating scientific knowledge with machine learning for engineering and environmental systems\.ACM Comput\. Surv\.55\(4\),pp\. 1–37\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3514228),[Document](https://dx.doi.org/10.1145/3514228)Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.
- J\. Yoon, J\. Jordon, and M\. van der Schaar \(2018\)GAIN: missing data imputation using generative adversarial nets\.InProceedings of the 35th International Conference on Machine Learning,Vol\.80,pp\. 5689–5698\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p3.1)\.
- A\. Zeghmar, E\. Mokhtari, and N\. Marouf \(2024\)A machine learning approach for RUSLE\-based soil erosion modeling in the Beni Haroun Dam Watershed, northeast algeria\.Earth Science Informatics17\(4\),pp\. 2921–2936\.External Links:[Document](https://dx.doi.org/10.1007/s12145-024-01305-7),ISSN 1865\-0481Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p1.1)\.
- S\. Zhao, K\. Tu, S\. Ye, H\. Tang, Y\. Hu, and C\. Xie \(2023\)Land use and land cover classification meets deep learning: a review\.Sensors23\(21\),pp\. 8966\.Cited by:[Related Work](https://arxiv.org/html/2608.04393#Sx2.p2.1)\.Similar Articles
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Proposes PUST, a novel LLM post-training framework that decouples reward exploration from distribution alignment using a lightweight proxy model, enabling reusable update signals and efficient weak-to-strong enhancement across models.
Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction
This paper proposes an iterative proxy correction framework to enhance robustness in multimodal sentiment analysis when dealing with incomplete or corrupted inputs by refining a language proxy for better sentiment prediction.
Forecasting Downstream Performance of LLMs With Proxy Metrics
This paper introduces proxy metrics based on token-level statistics from expert-written solutions to forecast downstream LLM performance, significantly outperforming loss-based methods in model selection, pretraining data selection, and training-time forecasting.
R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning
Proposes R2R2, a regularization method for self-predictive learning in reinforcement learning to mitigate overfitting under high update-to-data ratios, achieving significant improvements on continuous control tasks.
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Introduces Surrogate Latent Policy Optimization (SLPO) to apply outcome-reward RL to autoregressive latent reasoners, enabling test-time scaling and variable-horizon policies that improve accuracy on harder instances.