v3 Global CREST Parameters — 0.1°, daily, satellite-native
1,178,590 land cells × 15 parameters, learned end-to-end from 7,855
GRDC+Caravan stations through differentiable CREST physics, forced purely by IMERG.
Shipped as two fields (yardstick-optimal + physics-hardened) with a per-cell confidence
layer. Every number below is a paired-Wilcoxon, multi-seed verdict.
recipe CRESTPhys + tension store (ft) +
learnable Gamma UH · Snow-17 defaults · robust KGE loss (tail-capped + β penalty) ·
480 full-batch iterations · train 2012–23, evaluate 2003–11 · IMERG V07 daily,
ERA5 lapse-corrected T, Oudin PET — every input exists globally with no gauges.
Core numbers (all ungauged = never in training, independent window)
Skill
Metric
IMERG (product)
MSWEP (diagnostic)
daily KGE
0.514
0.605
monthly KGE
0.603
0.641
daily NSE
0.369
0.505
β (bias)
0.98
0.96
ceiling (per-station calib.)
0.579
0.650
MSWEP embeds rain-gauge information (+0.05 premium) that does not
exist in truly ungauged regions — IMERG is the deployable line, MSWEP bounds what
gauge-informed forcing would add.
Where it works (410-station holdout)
Slice
KGE
N. America / Europe
0.565 / 0.529
S. America / Africa / Oceania
0.485 / 0.438 / 0.394
area <500 → 10–50k km²
0.464 → 0.599
wet / mid / dry
0.558 / 0.555 / 0.407
Africa 0.12 → 0.44 across versions is the largest jump. Asia ships
as labeled extrapolation (MLP + hardened field) until CAMELS-IND ingestion.
Parameter table — meaning, production mode, boxes
Param
Meaning
Unit
Mode
Box
Provenance
wm
Max soil water capacity (VIC b-curve mean)
mm
mult ×⅓–3
1–2000
classic CREST
b
b-curve exponent (sub-cell heterogeneity)
–
mult ×⅓–3
0.01–6
classic CREST
im
Impervious fraction
–
mult ×0.2–5
0–0.5
classic CREST
fc
Infiltration capacity (fixed at prior)
mm/hr
fixed
0–150
classic CREST
ke
PET adjustment
–
direct
0.2–2.5
classic CREST
ksoil
Soil→aquifer drainage rate
1/day
mult ×0.2–5
0–0.5
CRESTPhys
hmaxaq
Aquifer capacity (fill-and-spill)
mm
mult ×⅓–3
50–4000
CRESTPhys
gwc / gwe
Baseflow coefficient / exponent
mm/d · –
direct
0.01–5 · 0.5–6
CRESTPhys
ft ★
Tension-store fraction of wm
–
mult ×⅓–3
0–0.7
this work
qc
Surface linear-reservoir release
1/day
direct
0.05–1
this work (operator-bound)
nshape / kmult
Gamma UH shape / time-scale multiplier
–
direct
1–8 · 0.1–4
this work (operator-bound)
pm
Precipitation multiplier (IMERG bias)
–
direct
0.7–1.4
this work
lf
Loss fraction (seepage/abstraction)
–
direct
0–0.5
this work
QC 13/15 parameters <1% bound-railing.
Disclosed: im sits at its near-zero prior outside cities (by construction); b hits the
×3 cap on 11% of tropical cells (box widened in v4). Strongly identified: ksoil
(seed rank-corr +0.97), pm, gwe; weakly identified (equifinality, see confidence
layer): wm, ft, qc. qc/nshape/kmult are operator-bound — re-estimate under sub-daily
step or explicit routing.
Parameter atlas — flip through all 15 maps in place
0.1° global fields (seed s1, the canonical
release), 50 m coastlines, robust 2–98% stretch. Use ← → keys, the arrows, or click a
parameter chip. Titles carry meaning, unit, box and global median.
1 / 15 · wm
Report deck — v3 in 8 slides
1 · At a glance
v3: satellite-native global parameters, honestly measured
Evaluation (never in training × never-trained window)
v3 release training set (7,855 train / 1,178 holdout)
CAMELS-IND — Peninsular India, CWC discharge
195 (of 242, five gates)
newly ingested: zero-shot v3 blind test done (KGE 0.118, β 0.71 — regime gap measured);
add-India training in progress (165 train / 30 holdout)
Asia goes from empty to anchored: the zero-shot number quantifies the risk the
confidence layer predicted; the in-training run will price what one monsoon observation
network is worth. Remaining gaps: China, SE Asia, Central Asia, interior Africa.
1c · Where the skill lands
Holdout KGE, station by station — v3 best
Every dot is a blind prediction (station never in training, window never
trained). Large = GRDC yardstick (410; med +0.514; 92% >0, 76% >0.3, 51% >0.5);
small = expansion-network holdout (749; med +0.568). The red pockets are physically
coherent — arid US Southwest, interior Spain, semi-arid NE Brazil, Western Australia:
the dry-regime mode shared by every global model, not random scatter.
2 · What moved the number
The ladder — every step paired, multi-seed
Lever
Δ KGE
Verdict note
Robust loss (tail-cap + β penalty)
+0.061
largest lever; kills mean-vs-median trap
CNN covariate head
+0.040
3 seeds; saturates at ~0.5° receptive field
Head width 128→256
+0.018
audited single-variable (p=4e-5)
Budget ×2 (→480 iters)
+0.024
960 & 2048 flat (paired p≥0.17): lever exhausted
Zero-effect club
0 / −
ensembling, attention, UNet, capacity-alone, lake covariate, NSE loss (−0.096, Gupta α*=r confirmed), learnable snow
0.343 (v1) → 0.405 → 0.445 → 0.484 → 0.514 (v3). Beyond 480 iters
train fit keeps climbing (0.61→0.68) while ungauged stays flat — pure memorization,
documented and stopped.
3 · Forcing decomposition
Why IMERG-native is an argued design, not a compromise
Increment (identical-harness twins)
Δ
ERA5-Land → IMERG (satellite beats reanalysis)
+0.087 (Africa +0.209)
+ merging algorithm (MSWEP no-gauge)
+0.004 (ns)
+ rain-gauge assimilation
+0.031*** → +0.05 at full budget; negative in Africa
The gauge premium exists only where gauges exist — precisely not the
deployment target of a global ungauged product. MSWEP twin (0.605) bounds it;
Africa flips its sign. Model structure is also not the bottleneck: CREST ≈ HBV
dead tie on the identical harness, both objectives × both forcings.
4 · Physics as weak supervision
Budyko lets ungauged continents train the model — ET cannot
Leave-continent-out (all 5 tested)
plain
+Budyko
paired d
South America (216)
+0.324
+0.377
+0.092 (p=1e-16)
Africa (152)
−0.070
−0.036
+0.049 (p=2e-5)
Oceania (417)
+0.126
+0.128
+0.005
N. America (996)
+0.310
+0.310
+0.020 (ns)
Europe (946)
+0.437
+0.402
−0.013
The law it traces: gain ∝ how badly transfer breaks the water
balance — large where β collapses (SA/AF), zero-to-negative where transfer is already
water-tight (EU). Placebo controls: a shuffled curve helps nowhere (levels must be
right); a global-constant target ties Budyko in mesic SA but is catastrophic in arid
AF (−0.31, β>2) — the curve's aridity structure is load-bearing exactly at climate
extremes. Satellite-ET supervision: −0.075 (refuted). Monsoon India: the curve itself
errs (seasonality) — its known boundary.
The boundary condition: an observation-free, low-dimensional
axiom (long-term Budyko water balance) transfers safely to unanchored regions;
a high-dimensional dynamic observation with its own errors (satellite monthly ET)
corrupts them — its bias gets absorbed wholesale where no discharge anchors the
balance. Ablation-clean (prior-reg zero, adaptive weighting negative), monotone
dose–response, two continents. In-distribution price −0.01…−0.02 → two fields.
5 · Transferability
Leave-continent-out map: risk is regime novelty, not distance
Left-out continent (holdout-subset protocol)
transfer loss
Europe
0.014 (near-free: NA supplies climate twins)
N. America / S. America / Africa
0.130 / 0.180 / 0.166
Oceania (no arid-regime twin exists)
0.214
CNN head off-continent: −0.49 (spatial texture is a fingerprint that
does not transfer) → dual-head rule: CNN where trained, MLP + hardened field where
not. Asia is assigned SA-class risk and shipped as labeled extrapolation.
6 · Against the published state of the art
Head-to-head with δHBV2-Globe at 116 common stations
Protocol
ours
δHBV2
verdict
Their TRAIN sites (n=72; in-sample for them, blind for us)
0.665
0.768
they win (p=8e-4) — in-sample + info advantages
Their UNGAUGED-TEST sites (n=43; symmetric protocol, both blind)
0.569
0.454
we lead (+0.059 paired, p=0.19)
Audited (units forced to m³/s→mm/day; per-station auto-detection bug fixed): at the only symmetric protocol — stations ungauged for BOTH models — we lead, not significantly. Their advantage concentrates entirely at their own training sites. β 0.958 vs 0.988, both unbiased. Their validation excludes Africa and
their "Asia" is essentially Japan+Korea; ours keeps Africa and South America with an
8-yr gate — the other end of the coverage-vs-completeness trade-off.
7 · The product
Two fields + a confidence layer
Field
yardstick KGE
use where
A · yardstick-optimal
0.514
gauged continents
B · Budyko-hardened
0.490
data-sparse regions (LRO +0.05…+0.09)
Per-cell confidence = 3-seed dispersion × covariate coverage ×
continent LRO prior. Blind-flood vignettes: Brazil 33k km² peak +1%; South
Africa 221 km² flash flood timing exact.
9 · Pure AI benchmark
Physics vs. a pure LSTM under the identical harness
Setting
LSTM
dPL
verdict
In-distribution spatial holdout
0.525-0.555
0.509-0.513
+0.01-0.03 for LSTM, not significant (p≥0.35)
Leave-continent-out (4 of 5)
0.01-0.24
0.13-0.44
physics +0.10 to +0.34 (p≤1e-13)
Africa (single arid South-African cluster)
+0.061
-0.070
dPL's own β collapses to 0.39 - physics fails there too
dPL + LSTM ensemble, in-distribution
0.594
free +0.059 over dPL (p=9e-11); harmful across continents (0.285 vs 0.337)
Mechanism, measured: the LSTM's transfer deficit is not timing
(r identical to dPL, p=0.25) but variability inflation (α median 1.82; 69% of its
error budget). Controlling for the LSTM's own water-balance error, every other predictor
- aridity, snow, area, and covariate novelty - loses significance; where the
LSTM's β lands in [0.8,1.25] (31% of stations) the two are statistically tied.
Physics buys failure avoidance, not a higher ceiling.
10 · What the physics constraint really is
An accurate obs-free anchor - the law itself is not the active ingredient
Placebo-controlled anchor test (South America held out)
Obs-free target
KGE
vs none
none
0.324
-
shuffled curve
0.364
ns
global constant runoff ratio
0.390
+0.057
Budyko-Fu (ω=2.6)
0.377
+0.092
empirical φ-anchor (training-basin lookup)
0.421
+0.164
Matched on per-basin target error the constant ties Budyko
(+0.110 vs +0.118): one monotone dose-response fits all three targets. Budyko wins
by landing inside tolerance most often and never being catastrophically wrong.
Replacing it with an empirical aridity lookup - same information, no theory, still
obs-free for holdout basins - nearly doubles the gain.
...and it regularizes, it does not fix mass
South America: 87% of the gain is hydrograph shape (r, α); the
bias term contributes +0.002 of +0.092;
dose 20→400 lifts KGE +0.034 while the distance to the true runoff
ratio stays put (0.091→0.087);
basins already inside the tolerance band - zero gradient - still gain +0.063,
through the shared parameter network;
the same constraint on the LSTM is statistically indistinguishable from
multiplying its hydrograph by a constant (p=0.13, seed-noise level),
whereas on dPL it redistributes water in time (baseflow +62%, peaks -10%,
p=3e-28). Physics must live in the model, not in the loss.
Known boundary: the anchor misfires where seasonality decouples P
from PET - Fennoscandian snow basins (Europe -0.013) and monsoon India - the same
failure class, now diagnosed.
11 · The physics-embedding spectrum (MC-LSTM matrix, 8/8)
Where must physics live? Loss < architecture < full structure
Leave-out
Free LSTM
MC-LSTM
dPL
conservation buys
South America
0.01
0.30 / 0.34 (2 seeds)
0.32
≈100% of the gap
Europe
0.23
0.35
0.44
≈half
North America
0.21-0.24
0.23
0.31
≈zero - all structure
Africa
+0.06
-0.05 (β0.57)
-0.07 (β0.39)
negative (under-delivery)
Oceania
+0.02
-0.26 (β1.64)
+0.13
negative (over-delivery)
Three closed verdicts. (1) The in-distribution cost of architectural
conservation (-0.11) is capacity-independent (32/64/128 cells: 0.42/0.42/0.42) - an intrinsic
tax, not a size problem. (2) Conservation's marginal value decays monotonically with how
water-anchored the target region already is, and turns negative in arid intermittent
regimes - in both directions of failure. (3) Stacking a Budyko loss anchor on top of the
conserving architecture adds nothing abroad (0.29 vs 0.30/0.34) and hurts at home (-0.06):
constraint gain = min(channel unanchoredness, reach) - anchor error.
8 · Limits & roadmap
Honesty list and what comes next
Known limits
Asia untrained (as all public global products) — labeled extrapolation;