CREST-Global · parameter set release

v3 Global CREST Parameters — 0.1°, daily, satellite-native

1,178,590 land cells × 15 parameters, learned end-to-end from 7,855 GRDC+Caravan stations through differentiable CREST physics, forced purely by IMERG. Shipped as two fields (yardstick-optimal + physics-hardened) with a per-cell confidence layer. Every number below is a paired-Wilcoxon, multi-seed verdict.

0.514ungauged daily KGE (410 st., 5-run plateau)
0.603ungauged monthly KGE
0.98β — near-zero bias
0.6184-init calibration envelope (gap ≈0.10)
2fields + confidence layer

▶ Report deck (in-page, 13 slides) 🗺 Parameter atlas (15 maps)

The release in one table

ComponentWhat it isEvidence
Field A — yardstick-optimal CNN head (two 3×3 conv, 256 ch) on 22 covariates; no physics constraint. For continents with training stations. ungauged 0.514 (runs 0.503/0.514/0.519/0.514/0.529, plateau ≥480 iters)
Field B — extrapolation-hardened Same + Budyko weak supervision (BUDW=100) applied to basins whose discharge never enters the loss. For data-sparse / regime-novel regions. leave-continent-out: SA +0.09, AF +0.05 (paired, p<1e-5); yardstick price −0.02 (0.490±0.003)
Confidence layer Per-cell [0,1]: 3-seed dispersion × covariate coverage × continent LRO prior tells users which field to trust where
Routing tier Gamma UH per basin <50k km²; differentiable Muskingum–Cunge above big-basin tier (213 st.): ungauged 0.324; mid-tier test: UH wins below 50k
Head network rule CNN in-distribution only; MLP head for continental extrapolation CNN off-continent collapses (LRO-SA −0.49); MLP loses only 0.014 (EU) – 0.21 (OC)

recipe CRESTPhys + tension store (ft) + learnable Gamma UH · Snow-17 defaults · robust KGE loss (tail-capped + β penalty) · 480 full-batch iterations · train 2012–23, evaluate 2003–11 · IMERG V07 daily, ERA5 lapse-corrected T, Oudin PET — every input exists globally with no gauges.

Core numbers (all ungauged = never in training, independent window)

Skill

MetricIMERG (product)MSWEP (diagnostic)
daily KGE0.5140.605
monthly KGE0.6030.641
daily NSE0.3690.505
β (bias)0.980.96
ceiling (per-station calib.)0.5790.650

MSWEP embeds rain-gauge information (+0.05 premium) that does not exist in truly ungauged regions — IMERG is the deployable line, MSWEP bounds what gauge-informed forcing would add.

Where it works (410-station holdout)

SliceKGE
N. America / Europe0.565 / 0.529
S. America / Africa / Oceania0.485 / 0.438 / 0.394
area <500 → 10–50k km²0.464 → 0.599
wet / mid / dry0.558 / 0.555 / 0.407

Africa 0.12 → 0.44 across versions is the largest jump. Asia ships as labeled extrapolation (MLP + hardened field) until CAMELS-IND ingestion.

Parameter table — meaning, production mode, boxes

ParamMeaningUnitModeBoxProvenance
wmMax soil water capacity (VIC b-curve mean)mmmult ×⅓–31–2000classic CREST
bb-curve exponent (sub-cell heterogeneity)–mult ×⅓–30.01–6classic CREST
imImpervious fraction–mult ×0.2–50–0.5classic CREST
fcInfiltration capacity (fixed at prior)mm/hrfixed0–150classic CREST
kePET adjustment–direct0.2–2.5classic CREST
ksoilSoil→aquifer drainage rate1/daymult ×0.2–50–0.5CRESTPhys
hmaxaqAquifer capacity (fill-and-spill)mmmult ×⅓–350–4000CRESTPhys
gwc / gweBaseflow coefficient / exponentmm/d · –direct0.01–5 · 0.5–6CRESTPhys
ft ★Tension-store fraction of wm–mult ×⅓–30–0.7this work
qcSurface linear-reservoir release1/daydirect0.05–1this work (operator-bound)
nshape / kmultGamma UH shape / time-scale multiplier–direct1–8 · 0.1–4this work (operator-bound)
pmPrecipitation multiplier (IMERG bias)–direct0.7–1.4this work
lfLoss fraction (seepage/abstraction)–direct0–0.5this work

QC 13/15 parameters <1% bound-railing. Disclosed: im sits at its near-zero prior outside cities (by construction); b hits the ×3 cap on 11% of tropical cells (box widened in v4). Strongly identified: ksoil (seed rank-corr +0.97), pm, gwe; weakly identified (equifinality, see confidence layer): wm, ft, qc. qc/nshape/kmult are operator-bound — re-estimate under sub-daily step or explicit routing.

Parameter atlas — flip through all 15 maps in place

0.1° global fields (seed s1, the canonical release), 50 m coastlines, robust 2–98% stretch. Use ← → keys, the arrows, or click a parameter chip. Titles carry meaning, unit, box and global median.

parameter map

Report deck — v3 in 8 slides

1 · At a glance

v3: satellite-native global parameters, honestly measured

Evaluation (never in training × never-trained window)nvalue
Ungauged daily KGE4100.514 (5-run plateau, ±0.013)
Ungauged monthly KGE / β / NSE4100.603 / 0.98 / 0.369
Expansion-network holdout7490.568
Per-station calibration ceiling (same physics/grid)410single-init 0.575–0.582; 4-init envelope 0.618 → gap ≈0.10; dPL beats envelope at 31% of stations
Temporal robustness (audited)—per-station calib loses 0.10 across windows; shared network loses 0.02

100% satellite inputs: IMERG V07 P, ERA5-lapse T → Snow-17/Oudin, 22 global covariates, SoilGrids priors. No rain gauges anywhere in the chain.

1b · Station network

The training network — now 8,460 basins with the monsoon anchor

station network
BlocknStatus
GRDC (global backbone) + Caravan expansion (HYSETS / LamaH / CAMELS-family) 2,989 + 5,276v3 release training set (7,855 train / 1,178 holdout)
CAMELS-IND — Peninsular India, CWC discharge195 (of 242, five gates) newly ingested: zero-shot v3 blind test done (KGE 0.118, β 0.71 — regime gap measured); add-India training in progress (165 train / 30 holdout)

Asia goes from empty to anchored: the zero-shot number quantifies the risk the confidence layer predicted; the in-training run will price what one monsoon observation network is worth. Remaining gaps: China, SE Asia, Central Asia, interior Africa.

1c · Where the skill lands

Holdout KGE, station by station — v3 best

holdout KGE map

Every dot is a blind prediction (station never in training, window never trained). Large = GRDC yardstick (410; med +0.514; 92% >0, 76% >0.3, 51% >0.5); small = expansion-network holdout (749; med +0.568). The red pockets are physically coherent — arid US Southwest, interior Spain, semi-arid NE Brazil, Western Australia: the dry-regime mode shared by every global model, not random scatter.

2 · What moved the number

The ladder — every step paired, multi-seed

LeverΔ KGEVerdict note
Robust loss (tail-cap + β penalty)+0.061largest lever; kills mean-vs-median trap
CNN covariate head+0.0403 seeds; saturates at ~0.5° receptive field
Head width 128→256+0.018audited single-variable (p=4e-5)
Budget ×2 (→480 iters)+0.024960 & 2048 flat (paired p≥0.17): lever exhausted
Zero-effect club0 / −ensembling, attention, UNet, capacity-alone, lake covariate, NSE loss (−0.096, Gupta α*=r confirmed), learnable snow

0.343 (v1) → 0.405 → 0.445 → 0.484 → 0.514 (v3). Beyond 480 iters train fit keeps climbing (0.61→0.68) while ungauged stays flat — pure memorization, documented and stopped.

3 · Forcing decomposition

Why IMERG-native is an argued design, not a compromise

Increment (identical-harness twins)Δ
ERA5-Land → IMERG (satellite beats reanalysis)+0.087 (Africa +0.209)
+ merging algorithm (MSWEP no-gauge)+0.004 (ns)
+ rain-gauge assimilation+0.031*** → +0.05 at full budget; negative in Africa

The gauge premium exists only where gauges exist — precisely not the deployment target of a global ungauged product. MSWEP twin (0.605) bounds it; Africa flips its sign. Model structure is also not the bottleneck: CREST ≈ HBV dead tie on the identical harness, both objectives × both forcings.

4 · Physics as weak supervision

Budyko lets ungauged continents train the model — ET cannot

Leave-continent-out (all 5 tested)plain+Budykopaired d
South America (216)+0.324+0.377+0.092 (p=1e-16)
Africa (152)−0.070−0.036+0.049 (p=2e-5)
Oceania (417)+0.126+0.128+0.005
N. America (996)+0.310+0.310+0.020 (ns)
Europe (946)+0.437+0.402−0.013

The law it traces: gain ∝ how badly transfer breaks the water balance — large where β collapses (SA/AF), zero-to-negative where transfer is already water-tight (EU). Placebo controls: a shuffled curve helps nowhere (levels must be right); a global-constant target ties Budyko in mesic SA but is catastrophic in arid AF (−0.31, β>2) — the curve's aridity structure is load-bearing exactly at climate extremes. Satellite-ET supervision: −0.075 (refuted). Monsoon India: the curve itself errs (seasonality) — its known boundary.

The boundary condition: an observation-free, low-dimensional axiom (long-term Budyko water balance) transfers safely to unanchored regions; a high-dimensional dynamic observation with its own errors (satellite monthly ET) corrupts them — its bias gets absorbed wholesale where no discharge anchors the balance. Ablation-clean (prior-reg zero, adaptive weighting negative), monotone dose–response, two continents. In-distribution price −0.01…−0.02 → two fields.

5 · Transferability

Leave-continent-out map: risk is regime novelty, not distance

Left-out continent (holdout-subset protocol)transfer loss
Europe0.014 (near-free: NA supplies climate twins)
N. America / S. America / Africa0.130 / 0.180 / 0.166
Oceania (no arid-regime twin exists)0.214

CNN head off-continent: −0.49 (spatial texture is a fingerprint that does not transfer) → dual-head rule: CNN where trained, MLP + hardened field where not. Asia is assigned SA-class risk and shipped as labeled extrapolation.

6 · Against the published state of the art

Head-to-head with δHBV2-Globe at 116 common stations

ProtocoloursδHBV2verdict
Their TRAIN sites (n=72; in-sample for them, blind for us) 0.6650.768they win (p=8e-4) — in-sample + info advantages
Their UNGAUGED-TEST sites (n=43; symmetric protocol, both blind) 0.5690.454we lead (+0.059 paired, p=0.19)

Audited (units forced to m³/s→mm/day; per-station auto-detection bug fixed): at the only symmetric protocol — stations ungauged for BOTH models — we lead, not significantly. Their advantage concentrates entirely at their own training sites. β 0.958 vs 0.988, both unbiased. Their validation excludes Africa and their "Asia" is essentially Japan+Korea; ours keeps Africa and South America with an 8-yr gate — the other end of the coverage-vs-completeness trade-off.

7 · The product

Two fields + a confidence layer

Fieldyardstick KGEuse where
A · yardstick-optimal0.514gauged continents
B · Budyko-hardened0.490data-sparse regions (LRO +0.05…+0.09)

Per-cell confidence = 3-seed dispersion × covariate coverage × continent LRO prior. Blind-flood vignettes: Brazil 33k km² peak +1%; South Africa 221 km² flash flood timing exact.

confidence
9 · Pure AI benchmark

Physics vs. a pure LSTM under the identical harness

mechanism figure
SettingLSTMdPLverdict
In-distribution spatial holdout0.525-0.5550.509-0.513+0.01-0.03 for LSTM, not significant (p≥0.35)
Leave-continent-out (4 of 5)0.01-0.240.13-0.44physics +0.10 to +0.34 (p≤1e-13)
Africa (single arid South-African cluster)+0.061-0.070dPL's own β collapses to 0.39 - physics fails there too
dPL + LSTM ensemble, in-distribution0.594free +0.059 over dPL (p=9e-11); harmful across continents (0.285 vs 0.337)

Mechanism, measured: the LSTM's transfer deficit is not timing (r identical to dPL, p=0.25) but variability inflation (α median 1.82; 69% of its error budget). Controlling for the LSTM's own water-balance error, every other predictor - aridity, snow, area, and covariate novelty - loses significance; where the LSTM's β lands in [0.8,1.25] (31% of stations) the two are statistically tied. Physics buys failure avoidance, not a higher ceiling.

10 · What the physics constraint really is

An accurate obs-free anchor - the law itself is not the active ingredient

Placebo-controlled anchor test (South America held out)

Obs-free targetKGEvs none
none0.324-
shuffled curve0.364ns
global constant runoff ratio0.390+0.057
Budyko-Fu (ω=2.6)0.377+0.092
empirical φ-anchor (training-basin lookup)0.421+0.164

Matched on per-basin target error the constant ties Budyko (+0.110 vs +0.118): one monotone dose-response fits all three targets. Budyko wins by landing inside tolerance most often and never being catastrophically wrong. Replacing it with an empirical aridity lookup - same information, no theory, still obs-free for holdout basins - nearly doubles the gain.

...and it regularizes, it does not fix mass

  • South America: 87% of the gain is hydrograph shape (r, α); the bias term contributes +0.002 of +0.092;
  • dose 20→400 lifts KGE +0.034 while the distance to the true runoff ratio stays put (0.091→0.087);
  • basins already inside the tolerance band - zero gradient - still gain +0.063, through the shared parameter network;
  • the same constraint on the LSTM is statistically indistinguishable from multiplying its hydrograph by a constant (p=0.13, seed-noise level), whereas on dPL it redistributes water in time (baseflow +62%, peaks -10%, p=3e-28). Physics must live in the model, not in the loss.

Known boundary: the anchor misfires where seasonality decouples P from PET - Fennoscandian snow basins (Europe -0.013) and monsoon India - the same failure class, now diagnosed.

11 · The physics-embedding spectrum (MC-LSTM matrix, 8/8)

Where must physics live? Loss < architecture < full structure

MC-LSTM matrix figure
Leave-outFree LSTMMC-LSTMdPLconservation buys
South America0.010.30 / 0.34 (2 seeds)0.32≈100% of the gap
Europe0.230.350.44≈half
North America0.21-0.240.230.31≈zero - all structure
Africa+0.06-0.05 (β0.57)-0.07 (β0.39)negative (under-delivery)
Oceania+0.02-0.26 (β1.64)+0.13negative (over-delivery)

Three closed verdicts. (1) The in-distribution cost of architectural conservation (-0.11) is capacity-independent (32/64/128 cells: 0.42/0.42/0.42) - an intrinsic tax, not a size problem. (2) Conservation's marginal value decays monotonically with how water-anchored the target region already is, and turns negative in arid intermittent regimes - in both directions of failure. (3) Stacking a Budyko loss anchor on top of the conserving architecture adds nothing abroad (0.29 vs 0.30/0.34) and hurts at home (-0.06): constraint gain = min(channel unanchoredness, reach) - anchor error.

8 · Limits & roadmap

Honesty list and what comes next

Known limits

  • Asia untrained (as all public global products) — labeled extrapolation;
  • dry regime 0.407 = weakest mode; residual error is sub-monthly recession timing;
  • equifinality on weak params — disclosed via confidence layer, not fixed;
  • Budyko fixes β only; ET weak supervision refuted (kept for audits);
  • b rails at ×3 on 11% tropical cells (v4 box widening).

Roadmap

  • CAMELS-IND monsoon anchor (access request pending);
  • MODIS snow cover to make snow identifiable;
  • GRACE TWS for the big-basin routing tier;
  • v4: wider b box, dual-field packaging, Zenodo DOI, re-audit.
1 / 8