Group meeting · Aug 2026 · CREST-Global

A globally-trained, IMERG-native CREST parameter set
v3: 0.508 ungauged — and physics as weak supervision

0.1° · daily · differentiable parameter learning on 7,855 GRDC+Caravan stations, 5 continents · 410-station stratified ungauged GRDC holdout + 749-station expansion holdout · every step a paired-Wilcoxon verdict, multi-seed on the headline.

0.508ungauged daily KGE (2 seeds ±0.006)
0.603ungauged monthly KGE
0.98β — near-zero bias
0.07gap to per-station ceiling (0.579)
1.18 Mland cells × 15 params
100%satellite forcing — no rain gauges
What is shipped

The product: a seamless global parameter field, not a set of basin fits

  • 1,178,590 land cells (0.1°, |lat| ≤ 60°) × 15 CREST parameters, inferred from 22 globally-available covariates (climate, SoilGrids, WorldCover, MERIT);
  • Dual head: CNN (two 3×3 convs — ~0.5° context) for trained continents; MLP for continental extrapolation (CNN fails catastrophically off-continent: LRO-SA −0.49);
  • Dual field: yardstick-optimal (no constraint) + extrapolation-hardened (Budyko weak supervision, BUDW=100) for data-sparse continents;
  • Tiered response: per-basin Gamma UH < 50,000 km²; differentiable Muskingum–Cunge river routing above (213 big-basin stations, ungauged 0.324);
  • QC: 13/15 parameters <1% bound-railing; disclosed exceptions (im at zero prior outside cities; b at ×3 cap on 11% of tropical cells).
v3 parameter maps
Headline

Ungauged skill across generations — and against the ceiling

Evaluation setnv1v2v3
Ungauged GRDC (holdout, WB-gated)410+0.343+0.404+0.508 (β 0.98)
Ungauged, monthly410+0.425+0.504+0.603
Seen basins (independent test window)—+0.338+0.410+0.590
Ungauged expansion networks (never in GRDC)749——+0.568
Per-station free-calibration ceiling (identical gridded physics)4100.579 → remaining regionalization cost 0.07

Medians; train 2012–23, test 2003–11, both spatial and temporal holdout. Headline replicated: seeds 0.502 / 0.514; NIT=960 run 0.513 ≈ flat (budget lever exhausted). Train fit 0.59 ≈ ceiling → the network is not under-fitted; what remains is parameter extrapolation, not optimization.

v3 map + CDF
Data

8,265 basins after five gates and cross-source de-duplication

Sourcen keptNote
GRDC (screened)2,989global backbone; NA 1,088 / EU 991 / OC 472 / SA 226 / AF 212
HYSETS3,237N-America density + Canadian cold regimes
CAMELS-BR / CL587 + 237tropics + Andes (scarce regimes)
CAMELS-GB / LamaH550 + 521EU density + alpine snow
CAMELS-AUS / US97 + 47mostly de-dup'd vs GRDC/HYSETS

Gates: area 50–50k km² · ≥8 yr valid obs in 2003–23 · <0.1% negatives · no 60-day flatline · mask/reported area ∈ [0.5, 2]. De-dup: 2 km + 15% area, ~1,600 duplicates removed across sources (prevents double-counting and train/test leakage).

The Asia hole — ours and everyone's

Public global sets have no continental-monsoon interior: δHBV2-Globe's "Asia" is essentially Japan + South Korea (and it excludes Africa entirely via a ≥95%-completeness/41-yr gate). We keep Africa (212) and South America with an 8-yr gate instead — the other end of the coverage-vs-completeness trade-off — and harden information-poor regions with physics (slide 9). CAMELS-IND ingestion is queued: a monsoon observation set neither product has.

Split design

15% stratified holdout (region × aridity) — GRDC holdout (410 gated) is the fixed yardstick across all experiments; expansion holdout (749) reports skill on never-before-used networks. Observations of holdout stations never touch the loss; the 2003–11 window is never trained on.

Method

Differentiable parameter learning: physics in the loop, end to end

StageWhatAcceptance
Covariates → paramsCNN head (two 3×3 conv, 256 ch, ~0.5° receptive field) over the global covariate grid → 14 parameter maps; 6 heads are multipliers on pedotransfer priors, 8 directzero prior railing; LRO arbitration for any spatial-context gain
PhysicsCRESTPhys + tension store (ft) per cell, Snow-17 melt, Oudin PET; unique-cell simulation + fractional-area gather (43% cheaper than basin-cell duplication)mass balance closed; cell values spot-checked vs raw IMERG
Responselearnable Gamma UH per basin (nshape, kmult on geomorphic k)operator-bound params flagged for re-estimation under other routing
Lossrobust KGE: tail-capped (1−KGE)≤2 + 2(β−1)², full-batch Adam, 480 iterations (23 yr × 7,855 basins per step)best-iterate snapshot; holdout never in loss

Everything backpropagates: ∂loss/∂(network weights) flows through 8,400 daily physics steps. One epoch = one full 23-year global simulation ≈ 60 s on a B200. Full-batch (state continuity forbids time-slicing), hence data growth requires budget growth — the optimization-budget rule below.

Attribution

What actually moved the number — every step priced, paired, replicated

LeverΔ ungauged KGEEvidence
Robust loss (cap catastrophic tail + β penalty)+0.061largest single lever; the mean-loss-vs-median-metric trap: one −31 station hijacks the mean gradient
CNN covariate head+0.0403 seeds, p<1e-5; saturates at ~0.5° receptive field (UNet adds nothing); in-distribution only
Caravan expansion (+5,276 stations)+0.035needs CNN capacity and matched budget; MLP-only +0.014 — data gain scales with architecture
Optimization budget (NIT 240→480)+0.024120→240 +0.036, 480→960 ≈ 0 (p=0.2) — a full dose–response curve, lever now exhausted
Ensembling / α-penalty / lake covariate / attention / UNet / learnable snow / prior-reg / NSE loss0 or negativeall priced and retired; NSE loss −0.096 confirms Gupta 2009 α*=r variability collapse

Campaign rules that emerged: (1) when the dataset grows, scale iterations before interpreting any degradation — full-batch training gets no free updates from more data; (2) effects ~0.02 need ≥2 seeds before a verdict (one retraction taught us); (3) negative results are logged, not discarded — they define the frontier.

Ceiling

The regionalization cost is now 0.07 — measured, not assumed

Mode A: the fair ceiling

Every basin (including the 410 "holdout") gets its own freely-calibrated parameter set — same 0.1° grid, same physics, same robust loss, same optimizer, calibrated by backprop on the train window, scored on the test window. The only difference from the product is where parameters come from: per-station freedom vs a shared covariate network.

ungauged 410
Per-station calibration (mode A)0.579
Shared network, station never seen (v3)0.508
Regionalization cost0.071

Why the gap is real information

  • Train fit 0.590 ≈ ceiling 0.579 → shared network expressiveness is not the bottleneck; the cost is purely "this station was never seen";
  • the same 0.07-class gap appears under MSWEP (0.650 ceiling vs 0.546 twin) — forcing-independent, so it is an information limit of regionalization, not precipitation quality;
  • part of it is irreducible station-specificity (reservoir operations, karst, abstractions) that no ungauged-legal information can recover;
  • remaining attack surfaces: satellite-ET multivariate calibration (B-track), CAMELS-IND, physics weak supervision (done, slide 9).
Forcing

Satellite-native is a design choice with a causal decomposition behind it

Increment (identical harness twins)ΔReading
ERA5-Land → IMERG+0.087satellite beats reanalysis; Africa +0.209
+ merging algorithm (MSWEP no-gauge)+0.004 nsthe algorithm itself is nearly worthless
+ rain-gauge assimilation (MSWEP)+0.031***the MSWEP premium ≈ gauge information; negative in Africa

At v3 config the premium persists: MSWEP twin 0.546 vs IMERG 0.508 (+0.049, p=7.6e-9) — additive to all architecture/data gains.

The argument

Gauge assimilation is worth +0.03…+0.05 — where gauges exist. In the target deployment of a global product (ungauged, data-sparse regions) that information does not exist, and in Africa the gauge-informed product is actually worse than satellite-native. IMERG-native is therefore not a handicap we accept but the correct forcing for the question "what is achievable where only satellites exist?" — and the MSWEP twin quantifies exactly what is being given up elsewhere.

Context: δHBV2-Globe

Ji et al. 2025 (Nat. Commun.): ungauged NSE ≈0.53 with gauge-fused forcing on ≥95%-complete stations, Africa excluded, "Asia" ≈ Japan+Korea. Different question, different data philosophy; our satellite-only v3 with Africa retained: KGE 0.508 / monthly 0.603.

Physics as weak supervision

Budyko constraint: ungauged regions become training signal

Penalty: each basin's simulated long-term runoff ratio must sit within ±0.10 of the Budyko–Fu curve computed from its own P/PET climatology. Needs zero discharge observations → applied to basins whose obs never enter the loss.

Leave-continent-outplain+Budykop
South America (216 st.)+0.324+0.388<1e-14
Africa (152 st.)−0.070+0.002<1e-6
Dose (BUDW)SA gain
5+0.018
20+0.040
100+0.09 (paired)
100 (product)SA near-peak; AF peaks 50–100, over-corrects at 200

Why it is credible

  • Ablation-clean: parameter-prior regularization = zero; novelty-gated weighting = negative (both retired) — the gain is 100% the weak-supervision term;
  • monotone dose–response, replicated on two continents;
  • mechanism matches diagnosis: continental extrapolation fails through β (water balance) — exactly what Budyko pins; it cannot teach timing;
  • honest scope: zero-to-negative on the regime-redundant yardstick (−0.01…−0.02 price) — gains live exactly where covariate twins don't exist;
  • lineage: PUB-era signature constraints (Yadav 2007, Zhang 2008), first made a gradient signal in a global differentiable framework.
Transferability

Leave-region-out: risk is regime novelty, not distance

Left-out continenttransfer loss (MLP head)
Europe−0.012
North America−0.094
Oceania−0.109
Africa−0.178
South America−0.219

Ordering follows regime redundancy — whether the training set contains climatic "twins" — not geographic distance. Budyko weak supervision claws back a third to half of the SA/AF loss (slide 9).

Architecture rule it produced

  • CNN head: +0.04 in-distribution, but −0.49 when South America is left out — spatial texture is a fingerprint that does not transfer;
  • MLP head: worse in-distribution, far safer off-continent;
  • → dual-head product: CNN where a continent has training stations, MLP (+ Budyko hardening) where it does not — e.g. Asia.

Implication for Asia

Asia's risk profile resembles SA (regime novelty: monsoon interior) more than EU. Priority: CAMELS-IND regime anchor; until then Asia ships MLP + Budyko-hardened parameters, explicitly labeled.

Where it works

v3 breakdowns — and blind case hydrographs

ContinentnKGEv1
N. America150+0.565+0.36
Europe142+0.529+0.37
S. America33+0.485+0.43
Africa21+0.438+0.12
Oceania64+0.394+0.20
Area / humidityKGE
<500 → 10–50k km²0.464 → 0.599 (monotone)
wet / mid / dry terciles0.558 / 0.555 / 0.407

Africa +0.12→+0.44 is the largest jump. Dry basins remain the hardest regime, as in every global model. β 0.98: the v1 IMERG under-catch signature (β 0.85) is fully absorbed by the robust loss + learned precip multiplier.

cases

Blind predictions at ungauged stations, 2006–08: KGE 0.81–0.89 from Sweden/Canada snow to Brazilian tropics, 900–33,000 km².

Structure & routing

Model structure is not the bottleneck; routing is tiered by evidence

CREST ≈ HBV — a structural tie

Identical harness, robust lossungauged
CREST+ft + Gamma UH (product)0.405
HBV-96 + Gamma UH0.406

Tie replicated across both objectives (KGE/NSE) and both forcings (IMERG/MSWEP). The v1-era "HBV wins" verdict was an artifact of the fragile loss. With learned parameters at daily/0.1°, conceptual structure differences wash out — forcing and regionalization carry the skill. CREST keeps the product line (distributed field, EF5 deployability, sub-daily path).

Routing: UH below 50k km², Muskingum–Cunge above

  • Differentiable MC on the 05′ MERIT network (1.89 M reaches, exact mass conservation, gradient-checked 6.7e-8);
  • impulse physics: routing matters from ~1,000 km channel length;
  • mid-tier test (<50k km²): MC loses to UH (−0.041) — UH is the right operator there;
  • big tier (≥50k km², 213 stations — new coverage, excluded from every number above): MC reaches ungauged 0.324 where UH is structurally invalid.
Limits

v3 honesty list

  • Asia untrained (no public monsoon-interior stations — true of all global products); ships as labeled extrapolation, Budyko-hardened;
  • dry regime 0.407 — weakest mode, as everywhere;
  • daily < monthly (0.508 vs 0.603): the residual error is sub-monthly timing of recessions/low flows, not volume;
  • equifinality: weakly-identifiable parameters differ between independent nets; prior-regularization priced at zero — disclosed, not fixed;
  • Budyko fixes β only; −0.01…−0.02 in-distribution price → two fields;
  • b-multiplier rails at ×3 on 11% of cells (tropics) — box widening queued for v4.

Verified-negative results (the frontier)

  • runoff gradients cannot identify snow parameters (→ needs MODIS SCA);
  • ensembling ≈ 0 under the robust loss (solution is seed-unique);
  • NSE loss −0.096 with α→r collapse — Gupta 2009, empirically confirmed;
  • attention / UNet / capacity-alone / lake covariate: zero (multi-seed);
  • novelty-adaptive constraint weighting: negative both ways.
Next

Outlook

Everything traceable: docs/camels_bench_log.md (per-experiment verdicts) · docs/gg_experiments.csv (ledger) · results page: global-v1.

← results page 1 / N ←/→ · Space · Click ‹ ›  ·  Ctrl+P to print all