0.1° · daily · differentiable parameter learning on 7,855 GRDC+Caravan stations, 5 continents · 410-station stratified ungauged GRDC holdout + 749-station expansion holdout · every step a paired-Wilcoxon verdict, multi-seed on the headline.

| Evaluation set | n | v1 | v2 | v3 |
|---|---|---|---|---|
| Ungauged GRDC (holdout, WB-gated) | 410 | +0.343 | +0.404 | +0.508 (β 0.98) |
| Ungauged, monthly | 410 | +0.425 | +0.504 | +0.603 |
| Seen basins (independent test window) | — | +0.338 | +0.410 | +0.590 |
| Ungauged expansion networks (never in GRDC) | 749 | — | — | +0.568 |
| Per-station free-calibration ceiling (identical gridded physics) | 410 | 0.579 → remaining regionalization cost 0.07 | ||
Medians; train 2012–23, test 2003–11, both spatial and temporal holdout. Headline replicated: seeds 0.502 / 0.514; NIT=960 run 0.513 ≈ flat (budget lever exhausted). Train fit 0.59 ≈ ceiling → the network is not under-fitted; what remains is parameter extrapolation, not optimization.
| Source | n kept | Note |
|---|---|---|
| GRDC (screened) | 2,989 | global backbone; NA 1,088 / EU 991 / OC 472 / SA 226 / AF 212 |
| HYSETS | 3,237 | N-America density + Canadian cold regimes |
| CAMELS-BR / CL | 587 + 237 | tropics + Andes (scarce regimes) |
| CAMELS-GB / LamaH | 550 + 521 | EU density + alpine snow |
| CAMELS-AUS / US | 97 + 47 | mostly de-dup'd vs GRDC/HYSETS |
Gates: area 50–50k km² · ≥8 yr valid obs in 2003–23 · <0.1% negatives · no 60-day flatline · mask/reported area ∈ [0.5, 2]. De-dup: 2 km + 15% area, ~1,600 duplicates removed across sources (prevents double-counting and train/test leakage).
Public global sets have no continental-monsoon interior: δHBV2-Globe's "Asia" is essentially Japan + South Korea (and it excludes Africa entirely via a ≥95%-completeness/41-yr gate). We keep Africa (212) and South America with an 8-yr gate instead — the other end of the coverage-vs-completeness trade-off — and harden information-poor regions with physics (slide 9). CAMELS-IND ingestion is queued: a monsoon observation set neither product has.
15% stratified holdout (region × aridity) — GRDC holdout (410 gated) is the fixed yardstick across all experiments; expansion holdout (749) reports skill on never-before-used networks. Observations of holdout stations never touch the loss; the 2003–11 window is never trained on.
| Stage | What | Acceptance |
|---|---|---|
| Covariates → params | CNN head (two 3×3 conv, 256 ch, ~0.5° receptive field) over the global covariate grid → 14 parameter maps; 6 heads are multipliers on pedotransfer priors, 8 direct | zero prior railing; LRO arbitration for any spatial-context gain |
| Physics | CRESTPhys + tension store (ft) per cell, Snow-17 melt, Oudin PET; unique-cell simulation + fractional-area gather (43% cheaper than basin-cell duplication) | mass balance closed; cell values spot-checked vs raw IMERG |
| Response | learnable Gamma UH per basin (nshape, kmult on geomorphic k) | operator-bound params flagged for re-estimation under other routing |
| Loss | robust KGE: tail-capped (1−KGE)≤2 + 2(β−1)², full-batch Adam, 480 iterations (23 yr × 7,855 basins per step) | best-iterate snapshot; holdout never in loss |
Everything backpropagates: ∂loss/∂(network weights) flows through 8,400 daily physics steps. One epoch = one full 23-year global simulation ≈ 60 s on a B200. Full-batch (state continuity forbids time-slicing), hence data growth requires budget growth — the optimization-budget rule below.
| Lever | Δ ungauged KGE | Evidence |
|---|---|---|
| Robust loss (cap catastrophic tail + β penalty) | +0.061 | largest single lever; the mean-loss-vs-median-metric trap: one −31 station hijacks the mean gradient |
| CNN covariate head | +0.040 | 3 seeds, p<1e-5; saturates at ~0.5° receptive field (UNet adds nothing); in-distribution only |
| Caravan expansion (+5,276 stations) | +0.035 | needs CNN capacity and matched budget; MLP-only +0.014 — data gain scales with architecture |
| Optimization budget (NIT 240→480) | +0.024 | 120→240 +0.036, 480→960 ≈ 0 (p=0.2) — a full dose–response curve, lever now exhausted |
| Ensembling / α-penalty / lake covariate / attention / UNet / learnable snow / prior-reg / NSE loss | 0 or negative | all priced and retired; NSE loss −0.096 confirms Gupta 2009 α*=r variability collapse |
Campaign rules that emerged: (1) when the dataset grows, scale iterations before interpreting any degradation — full-batch training gets no free updates from more data; (2) effects ~0.02 need ≥2 seeds before a verdict (one retraction taught us); (3) negative results are logged, not discarded — they define the frontier.
Every basin (including the 410 "holdout") gets its own freely-calibrated parameter set — same 0.1° grid, same physics, same robust loss, same optimizer, calibrated by backprop on the train window, scored on the test window. The only difference from the product is where parameters come from: per-station freedom vs a shared covariate network.
| ungauged 410 | |
|---|---|
| Per-station calibration (mode A) | 0.579 |
| Shared network, station never seen (v3) | 0.508 |
| Regionalization cost | 0.071 |
| Increment (identical harness twins) | Δ | Reading |
|---|---|---|
| ERA5-Land → IMERG | +0.087 | satellite beats reanalysis; Africa +0.209 |
| + merging algorithm (MSWEP no-gauge) | +0.004 ns | the algorithm itself is nearly worthless |
| + rain-gauge assimilation (MSWEP) | +0.031*** | the MSWEP premium ≈ gauge information; negative in Africa |
At v3 config the premium persists: MSWEP twin 0.546 vs IMERG 0.508 (+0.049, p=7.6e-9) — additive to all architecture/data gains.
Gauge assimilation is worth +0.03…+0.05 — where gauges exist. In the target deployment of a global product (ungauged, data-sparse regions) that information does not exist, and in Africa the gauge-informed product is actually worse than satellite-native. IMERG-native is therefore not a handicap we accept but the correct forcing for the question "what is achievable where only satellites exist?" — and the MSWEP twin quantifies exactly what is being given up elsewhere.
Ji et al. 2025 (Nat. Commun.): ungauged NSE ≈0.53 with gauge-fused forcing on ≥95%-complete stations, Africa excluded, "Asia" ≈ Japan+Korea. Different question, different data philosophy; our satellite-only v3 with Africa retained: KGE 0.508 / monthly 0.603.
Penalty: each basin's simulated long-term runoff ratio must sit within ±0.10 of the Budyko–Fu curve computed from its own P/PET climatology. Needs zero discharge observations → applied to basins whose obs never enter the loss.
| Leave-continent-out | plain | +Budyko | p |
|---|---|---|---|
| South America (216 st.) | +0.324 | +0.388 | <1e-14 |
| Africa (152 st.) | −0.070 | +0.002 | <1e-6 |
| Dose (BUDW) | SA gain |
|---|---|
| 5 | +0.018 |
| 20 | +0.040 |
| 100 | +0.09 (paired) |
| 100 (product) | SA near-peak; AF peaks 50–100, over-corrects at 200 |
| Left-out continent | transfer loss (MLP head) |
|---|---|
| Europe | −0.012 |
| North America | −0.094 |
| Oceania | −0.109 |
| Africa | −0.178 |
| South America | −0.219 |
Ordering follows regime redundancy — whether the training set contains climatic "twins" — not geographic distance. Budyko weak supervision claws back a third to half of the SA/AF loss (slide 9).
Asia's risk profile resembles SA (regime novelty: monsoon interior) more than EU. Priority: CAMELS-IND regime anchor; until then Asia ships MLP + Budyko-hardened parameters, explicitly labeled.
| Continent | n | KGE | v1 |
|---|---|---|---|
| N. America | 150 | +0.565 | +0.36 |
| Europe | 142 | +0.529 | +0.37 |
| S. America | 33 | +0.485 | +0.43 |
| Africa | 21 | +0.438 | +0.12 |
| Oceania | 64 | +0.394 | +0.20 |
| Area / humidity | KGE |
|---|---|
| <500 → 10–50k km² | 0.464 → 0.599 (monotone) |
| wet / mid / dry terciles | 0.558 / 0.555 / 0.407 |
Africa +0.12→+0.44 is the largest jump. Dry basins remain the hardest regime, as in every global model. β 0.98: the v1 IMERG under-catch signature (β 0.85) is fully absorbed by the robust loss + learned precip multiplier.
Blind predictions at ungauged stations, 2006–08: KGE 0.81–0.89 from Sweden/Canada snow to Brazilian tropics, 900–33,000 km².
| Identical harness, robust loss | ungauged |
|---|---|
| CREST+ft + Gamma UH (product) | 0.405 |
| HBV-96 + Gamma UH | 0.406 |
Tie replicated across both objectives (KGE/NSE) and both forcings (IMERG/MSWEP). The v1-era "HBV wins" verdict was an artifact of the fragile loss. With learned parameters at daily/0.1°, conceptual structure differences wash out — forcing and regionalization carry the skill. CREST keeps the product line (distributed field, EF5 deployability, sub-daily path).
Everything traceable: docs/camels_bench_log.md
(per-experiment verdicts) · docs/gg_experiments.csv (ledger) · results page:
global-v1.