USGS ScienceSearch

USGS · 70218255

Determination of vadose zone and saturated zone nitrate lag times using long-term groundwater monitoring data and statistical machine learning

Abstract

In this study, we explored the use of statistical machine learning and long-term groundwater nitrate monitoring data to estimate vadose zone and saturated zone lag times in an irrigated alluvial agricultural setting. Unlike most previous statistical machine learning studies that sought to predict groundwater nitrate concentrations within aquifers, the focus of this study was to leverage available groundwater nitrate concentrations and other environmental variables to determine mean regional vertical velocities (transport rates) of water and solutes in the vadose zone and saturated zone (3.50 and 3.75 m yr −1 , respectively). The statistical machine learning results are consistent with two primary recharge processes in this western Nebraska aquifer, namely ( 1 ) diffuse recharge from irrigation and precipitation across the landscape and ( 2 ) focused recharge from leaking irrigation conveyance canals. The vadose zone mean velocity yielded a mean recharge rate (0.46 m yr −1 ) consistent with previous estimates from groundwater age dating in shallow wells (0.38 m yr −1 ). The saturated zone mean velocity yielded a recharge rate (1.31 m yr −1 ) that was more consistent with focused recharge from leaky irrigation canals, as indicated by previous results of groundwater age dating in intermediate-depth wells (1.22 m yr −1 ). Collectively, the statistical machine learning model results are consistent with previous observations of relatively high water fluxes and short transit times for water and nitrate in the primarily oxic aquifer. Partial dependence plots from the model indicate a sharp threshold in which high groundwater nitrate concentrations are mostly associated with total travel times of 7 years or less, possibly reflecting some combination of recent management practices and a tendency for nitrate concentrations to be higher in diffuse infiltration recharge than in canal leakage water. Limitations to the machine learning approach include the non-uniqueness of different transport rate combinations when comparing model performance and highlight the need to corroborate statistical model results with a robust conceptual model and complementary information such as groundwater age.

Explore related subjects

90° N90° S · 180° W ← longitude → 180° E
Source-reported bounding extent: 41.27367811566259° to 42.407234661551875° latitude; -104.03228759765625° to -102.39257812499999° longitude. This indicates report coverage, not an exact sampling location. View area on OpenStreetMap.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Martin J. Wells, Troy E. Gilmore, Natalie Nelson, Aaron Mittelstet, J.K. Bohlke. 2021-02-19. Determination of vadose zone and saturated zone nitrate lag times using long-term groundwater monitoring data and statistical machine learning. https://doi.org/10.5194/hess-25-811-2021

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Technical note: A low-cost approach to monitoring relative streamflow dynamics in small headwater streams using time lapse imagery and a deep learning model

Despite their ubiquity and importance as freshwater habitat, small headwater streams are under-monitored by existing stream gage networks. To address this gap, we describe a low-cost, non-contact, and low-effort method that enables organizations to monitor relative streamflow dynamics in small headwater streams. The method uses a camera to capture repeat images of the stream from a fixed position. A person then annotates pairs of images, in each case indicating which image has more apparent streamflow or indicating equal flow if no difference is discernible. A deep learning modeling framework called streamflow rank estimation (SRE) is then trained on the annotated image pairs and applied to rank all images from highest to lowest apparent streamflow. From this result a relative hydrograph can be derived. We found that our modeled relative hydrograph dynamics matched the observed hydrograph dynamics well for 11 cameras at 8 streamflow sites in western Massachusetts. Higher performance was observed during the annotation period (median Kendall's Tau rank correlation of 0.75, with a range of 0.6–0.83) than after it (median Kendall's Tau of 0.59, with range 0.34–0.74). We found that annotation performance was generally consistent across the 11 camera sites and 2 individual annotators and was positively correlated with streamflow variability at a site. A scaling simulation determined that model performance improvements were limited after 1000 annotation pairs. Our model's estimates of relative flow, while not equivalent to absolute flow, may still be useful for many applications, such as ecological modeling and calculating event-based hydrological statistics (e.g., the number of out-of-bank floods). We anticipate that this method will be a valuable tool to extend existing stream monitoring networks and provide new insights on dynamic headwater systems.

Massachusetts

Interrogating process deficiencies in large-scale hydrologic models with interpretable machine learning

Large-scale hydrologic models are increasingly being developed for operational use in the forecasting and planning of water resources. However, the predictive strength of such models depends on how well they resolve various functions of catchment hydrology, which are influenced by gradients in climate, topography, soils, and land use. Most assessments of hydrologic model uncertainty have been limited to traditional statistical methods. Here, we present a proof-of-concept approach that uses interpretable machine learning techniques to provide post hoc assessment of model sensitivity and process deficiency in hydrologic models. We train a random forest model to predict the Kling–Gupta efficiency (KGE) of National Water Model (NWM) and National Hydrologic Model (NHM) streamflow predictions for 4383 stream gauges in the conterminous United States. Thereafter, we explain the local and global controls that 48 catchment attributes exert on KGE prediction using interpretable Shapley values. Overall, we find that soil water content is the most impactful feature controlling successful model performance, suggesting that soil water storage is difficult for hydrologic models to resolve, particularly for arid locations. We identify nonlinear thresholds beyond which predictive performance decreases for NWM and NHM. For example, soil water content less than 210 mm, precipitation less than 900 mm yr −1 , road density greater than 5 km km −2 , and lake area percent greater than 10 % contributed to lower KGE values. These results suggest that improvements in how these influential processes are represented could result in the largest increases in NWM and NHM predictive performance. This study demonstrates the utility of interrogating process-based models using data-driven techniques, which has broad applicability and potential for improving the next generation of large-scale hydrologic models.

conterminous United States

Pluvial and potential compound flooding in a coupled coastal modeling framework: New York City during post-tropical Cyclone Ida (2021)

Many coastal urban areas are prone to extreme pluvial flooding due to limitations in stormwater system capacity, with the additional potential for flooding compounded by storm surge, tides, and waves. Understanding and simulating these processes can improve prediction and flood risk management. Here, we adapt the Coupled Ocean–Atmosphere–Wave–Sediment Transport modeling framework (COAWST) to simulate pluvial flooding from post-tropical Cyclone Ida (2021) in the Jamaica Bay watershed of New York City (NYC). We modify the model to capture the volumetric effects of rainfall and parameterize soil infiltration and a stormwater conveyance system as the drainage rate. We generate a spatially continuous flood map of Ida with a root-mean-square error (RMSE) of 20 cm when compared to high-water marks, useful for understanding Ida's impacts and subsequent mitigation planning. Results show that over 23 km 2 and 4621 buildings were flooded deeper than 0.3 m during Ida. Sensitivity analyses are used to study the broader risk from events like Ida (pluvial flooding) as well as potential compound (pluvial–coastal) flooding. Spatial shifting of the storm track within a typical 12 h forecast uncertainty reveals a worst-case scenario that increases this flooded area to 62 km 2 (5907 buildings). Shifting Ida's rainfall to coincide with high tide increases this flooded area by 1 km 2 , a relatively small change due to the lack of significant storm surge. The application of COAWST to this storm event addresses a broader goal of developing the capability to model compound pluvial–coastal flooding by simultaneously representing coastal storm processes such as rain, tide, waves, erosion, and atmosphere–wave–ocean interactions. The sensitivity analysis results underscore the need for detailed flood risk assessments, showing that Ida, already NYC's worst rain event, could have been even more devastating with slight shifts in the storm track.

New York