CREST-GPM: from CAMELS certification to ungauged CONUS parameters

2026-07-23 · differentiable CREST · IMERG-native · vs the BAMS-2017 default global parameters

CAMELS: same-harness benchmark + optimizer-verified optimum structure upgraded from arid-strata diagnosis cell-level dPL → ungauged monthly KGE 0.55 2.3× the a-priori baseline
Three claims in this deck: (1) the differentiable harness is certified fair and reaches the model-structure optimum on CAMELS; (2) parameters transfer to never-gauged basins with quantified skill; (3) they more than double the uncalibrated global-default baseline (Clark et al. 2017 BAMS lineage) under identical forcing and evaluation.

Navigate: or corner buttons.

1 · CAMELS certification — protocol & fairness

ItemSetting
Scope673 CAMELS basins, lumped, daily; train 1999–2008, test 1989–1999 (Newman/Kratzert split)
CREST configSnow-17(9) + water balance + reservoirs + triangular UH; 20–22 params/basin, Adam, best-train selection
FairnessHBV & EXP-HYDRO re-implemented differentiably, identical harness; SAC-SMA = Newman's published SCE runs (Gamma-UH + calibrated PT-PET — component-equivalent to ours)
Certificationour same-harness HBV = 0.639, inside its published 0.63–0.68 range → harness neither cheats nor handicaps
Auditindependent leakage audit; obs-parity vs Newman r=1.0000; window/area bugs fixed pre-publication
Why lumped-daily CAMELS: it is HBV/SAC-SMA's home turf and the community's common yardstick — entering their band with a generic differentiable harness certifies the machinery before any regionalization claim.

2 · CAMELS results — CREST enters the conceptual band

Modelmed NSEmed KGEβ
SAC-SMA + Snow-17 (published) 0.6620.7170.94
HBV (ours, same harness) 0.6390.6711.00
CREST (CRESTPhys+ft, β-loss) 0.6180.6880.954
CREST+UH (single-layer, β-loss) 0.6080.6640.946
EXP-HYDRO (same harness) 0.5230.6211.03
CREST-dPL (regional, one net) 0.4620.4360.82
CREST-dPL (PUB, spatial holdout) 0.44–0.460.430.81
  • Every point attributed: UH +0.08 · β-loss fixes bias free · seeds/merge honestly rejected
  • Optimizer-verified: exhaustive SCE-UA (10,000 evals, 72 h) converges to the same score (0.597 vs 0.600) at 6× the compute → 0.60-band = structure ceiling, not an optimization artifact
NSE CDF
Test-decade KGE CDFs (673 basins, one metric code) — CREST sits right of HBV over most of the distribution.

3 · What CAMELS taught the structure — arid diagnosis → CRESTPhys+ft

  • Stratifying by aridity showed CREST's entire deficit was western: East of −100°, CREST = HBV = SAC-SMA (0.635/0.633/0.641)
  • Dry-basin forensics: calibrated params showed a "sewer strategy" (fgw 0.75, gwc 0.02, lf 0.36, β 0.69) — structural inability to hold small-storm water and to let streams dry
  • Fix = two complementary halves: tension store (ft, evap-only, fills first) + CRESTPhys dry-able exponential aquifer w/ GW-ET (from the lab's EF5 development fork, made differentiable)
  • Result: arid-strata NSE 0.28→0.43, west-humid +0.06, KGE 0.688 (above HBV's 0.671) — and scratch-training stability
aridity decomposition
Median NSE by aridity class and geographic block.

4 · Ungauged CONUS — cell-level dPL design

Pipeline

  • Per 0.1° cell: 21 covariates (IMERG/T climate ⊕ SoilGrids/KSAT soils ⊕ MERIT terrain ⊕ MODIS land cover) → shared MLP → 12 params (× hybrid physical priors)
  • Trained end-to-end through the hydrologic model on 364 USGS basins: monthly KGE + 0.25 × daily KGE (IMERG V07B forcing, 2010–2024)
  • 60 stratified basins held out entirely (never in any loss) = the honest ungauged verdict; 5-fold spatial PUB: fold medians 0.42–0.55, pooled cross-validated monthly KGE 0.497 (n=348)

Ablations (ungauged monthly / daily KGE)

Variantmonthlydaily (routed)
full (production)0.5530.187
− soil/landcover maps0.500<0
− daily loss term0.5720.100

Soil maps carry real transfer signal at cell level (daily ΔKGE ≈ 0.3); the daily loss pins fast-response parameters.

parameter fields
Production parameter fields. QC: boundary-artifact metric ≤ 1.18 everywhere (no basin blockiness); ranges physical; urban impervious & arid tension/loss patterns emerge unsupervised.

5 · Ungauged skill — where we stand

RegimeOursReference points (protocol differs — see notes)
CONUS never-trained basins, monthly KGE (IMERG forcing) 0.553 (β 0.97); 5-fold PUB pooled 0.497 (n=348, folds 0.42–0.55) per-basin calibrated upper ref 0.718; EF5 a-priori baseline 0.241
CONUS never-trained, daily KGE (geomorphic-UH routed) 0.187 δHBV PUB daily 0.67 (gauge-quality Daymet forcing, lumped)
CAMELS spatial-PUB daily NSE (older single-layer chain) 0.44–0.46 regional LSTM PUB ≈ 0.69 (Daymet); classical global regionalization KGE 0.46 (Beck 2020)
Honest read: monthly water-balance transfer is strong; daily ungauged skill carries three known costs — satellite forcing class (literature: 0.1–0.3 KGE vs gauge-quality), no calibrated routing, and IMERG-era records. An ERA5-forced twin (E5) is queued to separate forcing cost from method cost.
Unique among competitors: deliverable physical parameter fields (EF5-compatible GeoTIFFs + pedigree layer) with formal field QC (range / spatial-pattern / boundary-artifact audits) — no published regionalization work provides either.

6 · vs the BAMS-2017 default global parameters (EF5-Global-Parameters)

The prior art (Clark et al. 2017, BAMS; HyDROSLab/EF5-Global-Parameters): table-lookup a-priori grids — WM = porosity × fixed 100 cm depth (median 87 mm), b/im from lookup, uncalibrated (global daily correlation ≈ 0.63). We ran those exact grids through the same harness, same IMERG melt, same basins, same evaluation:

Configurationungauged monthly KGEβReading
a-priori params + standard Oudin PET−1.16 2.4887 mm bucket + realistic PET → budget cannot close
a-priori params + their FAO PET (verified 2–3× inflated) 0.2410.63 charitable variant: inflated PET compensates
CREST-GPM dPL (ours)0.553 0.9662.3× the charitable baseline; self-consistent water budget
Deeper finding: the default parameter set is configuration-locked — its water budget closes only through the inflated FAO PET, sub-daily stepping and kinematic-wave leak terms, not through the soil parameters themselves. Ported to any other setup it fails (β = 2.48). Our parameters carry their own closed water balance (β = 0.97 under standard PET) — that is what "observation-constrained" buys, and why this set can serve as a portable prior for the global stage.

7 · Takeaways & next

1 · Certified. Same-harness CAMELS puts CREST at 0.618 NSE / 0.688 KGE beside HBV 0.639/0.671 and SAC-SMA 0.662/0.717, with the optimum verified by two independent optimizers.
2 · Structure earned, not assumed. The arid-strata diagnosis drove the CRESTPhys+ft upgrade: arid NSE +0.15, snow-west +0.06, KGE past HBV.
3 · Ungauged delivered. Cell-level dPL transfers to 60 never-trained basins at monthly KGE 0.553 (β 0.97) — 2.3× the BAMS-2017 default baseline under identical forcing — with auditable, physically-patterned, blockiness-free parameter fields.
Open items (queued): 5-fold PUB CIs; b-clamp widening (railing over clay-rich humid soils is physical, box too tight); ERA5 twin to price the satellite-forcing cost; then M4: global 0.1° rollout + GRDC validation (1,818 stations ≥5,000 km², regulation-screened) against the 0.45/0.55 competitive lines.
Data/methods: IMERG V07B Final (2001–2024, verified); NWIS/USGS monthly+daily; SoilGrids 2.0 / Gupta KSAT / MERIT-IHU / MODIS MCD12C1; all numbers from a single evaluation harness; logs and audits in the project repository.