Networks are difficult to sustain at scale, radar coverage is geographically limited and needs calibration — heterogeneous, data-sparse regions stay imperfectly constrained.
Sampling, retrieval physics, surface and precipitation-type heterogeneity propagate into climate diagnostics and every workflow that uses precipitation as reference or forcing.
Inverting the soil-moisture water balance works — but skill is governed by (i) how infiltration/drainage physics is represented and (ii) the depth at which soil moisture is observed. Existing implementations fix literature parameters and use shallow layers only.
Fig. 1 of the paper maps global station coverage, colour-coded by how many of the four soil-moisture depths each station reports (1→4). Coverage is dense over North America and Europe, sparse elsewhere.
Depth × 4Land cover × 5Köppen climate × 8
The rainfall imprint in soil moisture is sharpest at the surface and becomes progressively attenuated, delayed and dispersed with depth — this is the paper's central control variable.
ET is assumed negligible during rainfall. Rainfall is inferred from the competition between storage tendency and loss terms — so errors in the drainage / infiltration term systematically bias retrieved rainfall and can make the inverse problem ill-conditioned.
| Z | depth-related soil parameter (soil layer depth) |
| θ(t) | relative / volumetric soil moisture |
| a, b, c | empirical drainage & runoff exponents (calibrated) |
Sadeghi et al.'s NWF model, embedded in SM2RAIN by Saeedi et al. It reconstructs the effective vertical water flux from SM by analytically inverting the 1-D linearized Richards equation.
| ΘN(Z) | normalised SM anomaly, (θ(t) − θ∞)/θ∞; θ∞ = background moisture |
| U(Z,T) | analytical unit response function |
| T | dimensionless time, T = k²t/Di (k = hydraulic conductivity, Di = diffusivity) |
| f(t) | drainage/flux term G(t) proxy inserted into the inversion |
| fmax, Di, kf, zf | infiltration parameters — fixed from literature in the original scheme |
Limitation attacked: fixed literature parameters neglect site-specific soil-hydraulic variability → systematic bias.
First application of a Green–Ampt infiltration calculation inside a bottom-up rainfall-inversion framework. A sharp wetting front separates a saturated zone (θs) above from the initial water content (θi) below and advances downward.
| F, F0 | cumulative infiltration depth at t and t0 (mm) |
| Ks | saturated hydraulic conductivity — constrained to [0.1, 200] mm h⁻¹ |
| Ψf | wetting-front suction head — constrained to [20, 1000] mm |
| Δθ | θs − θi, constant moisture deficit across the front |
Assumption & its cost: 1-D vertical flow with a sharp front is most defensible in shallow layers. Deeper, diffuse redistribution, lateral flow, root uptake, preferential flow and heterogeneity take over — so at 15–20 cm the GA term must be read as an effective first-order, physically constrained representation, not full mechanistic subsurface flow.
The fixed-literature-parameter assumption is relaxed: all infiltration- and SM2RAIN-specific parameters are estimated independently for every station.
ξ > 0 encourages exploratory sampling in high-posterior-uncertainty regions — balancing exploration against exploitation with a limited budget of expensive model evaluations. All infiltration- and SM2RAIN-specific parameters are jointly calibrated; the four reported against elevation in Fig. 8 are fmax, Di, kf, zf.
Isolates the marginal effect of each calibrated parameter: perturb one, hold the rest at their BO optima, keep the infiltration term fixed. 1000 realizations × 8 representative stations spanning a broad elevation and geographic range.
Reported per time step: mean, σ, 2.5th/97.5th percentiles (95% CI) and the 25–75% IQR. Result: at almost all stations the deterministic BO estimates for Z and a fall inside the 5–95% Monte-Carlo envelope and near the distribution median — the calibration is robust to parameter perturbation.
Together they separate "getting the totals right" from "getting the storms right".
Share of stations where each method ranks first on the joint rank-based score over all five metrics.
GA dominates at shallow depth, but the winner distribution shifts systematically downward: GA 54%→43%, NWF 30%→20%, while GPM climbs 16%→37%. Consistent with attenuation of the rainfall imprint in deeper soil.
Paired Wilcoxon signed-rank on station-wise metrics, FDR-adjusted q-values.
Read this way: FAR favours the bottom-up schemes at every depth for both methods — all eight cells green. NSE favours them everywhere except GA at 15–20 cm, where the difference is non-significant (q = 0.199). NWF loses correlation and POD in the two deepest layers; GA never loses POD anywhere — significantly better at all four depths.
Composite view: GA's advantage over GPM on event detection (POD ↑ and FAR ↓ simultaneously) is what separates it from NWF, and it survives deeper into the column.
Classes: 10 cropland (rainfed) · 70 tree cover, needle-leaved evergreen, closed-to-open (>15%) · 120 shrubland · 130 grassland · 190 urban. Mean improvement across classes — NSE +65% (NWF) / +53% (GA); FAR +20% vs +46%; POD +14% vs +49%. Replacing the drainage term with Green–Ampt is what buys transferability across surfaces.
Mean over classes — NSE +78% / +69%; R +14% / +17%; RMSE +29% / +24%; FAR +25% / +53%; POD +21% / +48%. Sign reversals matter: under NWF, POD collapses in Cfb (−55.3%) and Cfa (−7.6%); GA turns both positive (+11.7%, +10.5%). The one systematic exception is Cfa, where GA's RMSE change is slightly negative despite better FAR and POD — residual mismatch in infiltration timing.
Optimized fmax, Di, kf and zf span broad ranges with no monotonic dependence on elevation — elevation is not a primary control, and much of what BO captures is idiosyncratic rather than a large-scale driver.
Central tendency and spread of R and RMSE largely overlap between the two configurations; FAR changes are minor; the one systematic deviation is POD, which is higher under the fixed setting. Holding (fmax, Di, kf, zf) fixed does not materially degrade skill for most stations — and it removes the computational burden of site-wise BO.
1000 stochastic realizations per station, 8 stations. Box = 25–75% IQR, whiskers = 5th–95th percentile, dashed = mean ± 1σ, green dot = the deterministic BO estimate.
At almost all stations the deterministic calibration sits inside the 5–95% uncertainty envelope and close to the median of the distribution — the calibrated coefficients are relatively robust to the imposed perturbations. (Schematic reproduction of the paper's Fig. 9 layout; box widths/whiskers illustrate the reported structure.)
SM2RAIN estimates are validated against point-scale rain gauges while GPM IMERG represents areal rainfall at grid scale. To quantify that penalty the authors ran a multi-gauge pixel sensitivity analysis on GPM pixels containing several gauges (Fig. S1).
Verdict: scale mismatch explains only part of the observed performance difference — it does not overturn the main findings.
More defensible, observation-based rainfall constraints — with explicit uncertainty — for climate diagnostics, land–atmosphere coupling analysis, and benchmarking of gridded precipitation products in exactly the heterogeneous, data-scarce regions where gauges fail.