Streamflow model’s benchmark edge came with extra input
A Penn State study found that its hourly streamflow model, h-Diffusion, scored modestly higher than two LSTM benchmarks in one test across 516 U.S. river basins. But h-Diffusion used daily simulated hydrologic states as prior information, while the two comparison models did not. The authors say that difference may help explain the gain.
In the single-forcing test, the median Nash–Sutcliffe Efficiency (NSE) was 0.780 for h-Diffusion, compared with 0.763 for MTS-LSTM and 0.756 for MF-LSTM. The paper reports statistically significant differences among the results at p < 0.01. In an ablation that removed h-Diffusion’s daily simulation input, its median NSE was 0.728.
Results changed with the test setup
The research team includes Penn State researchers based at University Park, Pennsylvania. The study evaluated h-Diffusion and an inpainting-based data-assimilation version, h-Diffusion-DA, using the CAMELS-US dataset. Models were trained on data from Oct. 1, 1990, through Sept. 30, 2003. The single-forcing test covered Oct. 1, 2008, through Sept. 30, 2014; the multiple-forcing test covered Oct. 1, 2003, through Sept. 30, 2008.
| Model | Single-forcing median NSE | Multiple-forcing median NSE |
|---|---|---|
| h-Diffusion | 0.780 | 0.800 |
| MTS-LSTM | 0.763 | 0.812 |
| MF-LSTM | 0.756 | 0.805 |
The multiple-forcing test produced a different ranking: both LSTM models had higher median NSE scores than h-Diffusion. The paper reports no statistically significant differences among the three models in that comparison. The single-forcing result therefore did not hold across both test configurations.
Assimilating recent observations improved a separate test
A separate single-forcing experiment examined whether recent gauge observations could improve a 24-hour prediction window. When observations were assumed to be available five hours before that window, the median NSE rose from 0.780 for h-Diffusion to 0.832 for h-Diffusion-DA. The approach uses inpainting to incorporate observations without additional model training. This result compares the diffusion model with and without data assimilation; it is not a comparison with the LSTM benchmarks.
The paper also reports probabilistic scores for high flows in its inpainting comparison: a median Continuous Ranked Probability Skill Score (CRPSS) of 0.37 for the top 10% of flows and 0.21 for the top 1%. These figures describe the inpainting comparison, not performance against the LSTM models.
The study reports retrospective tests on U.S. basin data, not an operating public flood-warning service or Pennsylvania-specific forecast accuracy. The authors identify computational optimization and integration with physical process-based models as future research directions. The journal article first appeared in Water Resources Research on July 21, 2026; Penn State’s account of the research is marked last updated Oct. 6, 2026.
Sources
- Diffusion-Based Probabilistic Modeling for Hourly Streamflow Prediction and Assimilation, American Geophysical Union / Water Resources Research
- The techniques AI uses to create images can help predict floods, researchers say, Penn State University
Look for updates to this story
Discover more from Interactive News
Subscribe to get the latest posts sent to your email.