USGS ScienceSearch

Geology topics

Chad Babcock

Publications and source records attributed to Chad Babcock.

3 recordsLinked to original sources

How machine learning can improve predictions and provide insight into fluvial sediment transport in Minnesota

Understanding fluvial sediment transport is critical to addressing many environmental concerns such as exacerbated flooding, degradation of aquatic habitat, excess nutrients, and the economic challenges of restoring aquatic systems. However, fluvial sediment transport is difficult to understand because of the multitude of factors controlling the potential sources, delivery, mechanics, and storage of sediment in aquatic systems. While physical fluvial sediment samples are an integral part of developing solutions for these environmental concerns, samples cannot be collected at every river and time of interest. Therefore, accurate and cost-effective estimates of sediment loading are needed to manage riverine sediment transport at a multitude of scales (Ellison et al. 2016); also needed are methods to estimate sediment transport at sites where little or no physical samples have been collected (Gray & Simes 2008). The application of machine learning (ML) approaches to estimate sediment transport has grown over the past two decades (Afan et al. 2016). ML used in sediment transport research has shown multiple benefits over traditional approaches, such as increased prediction accuracy, the ability to learn complex linear and non-linear relations amongst the dataset and providing the ability to interpret these complex relations with important features used in the model (Cisty et al. 2021; Francke et al. 2008; Khan et al. 2021; Zounemat-Kermani et al. 2020; Cutler et al. 2007).

Minnesota

Using machine learning to improve predictions and provide insight into fluvial sediment transport

A thorough understanding of fluvial sediment transport is critical to addressing many environmental concerns such as exacerbated flooding, degradation of aquatic habitat, excess nutrients, and the economic challenges of restoring aquatic systems. Fluvial sediment samples are integral for addressing these environmental concerns but cannot be collected at every river and time of interest. Therefore, to gain a better understanding for rivers where direct measurements have not been made, extreme gradient boosting machine learning (ML) models were developed and trained to predict suspended sediment and bedload from sampling data collected in Minnesota, United States (U.S.), by the U.S. Geological Survey. Approximately 400 watershed (full upstream area), catchment (nearby landscape), near-channel, channel, and streamflow features were retrieved or developed from multiple sources, reduced to approximately 30 uncorrelated features, and used in the final ML models. The results indicate suspended sediment and bedload ML models explain approximately 70% of the variance in the datasets. Important features used in the models were interpreted with Shapley additive explanation (SHAP) plots, which provided insight into sediment transport processes. The most important features in the models were developed to normalize streamflow by the 2-year recurrence interval and quantify the rate of change in streamflow (slope), which helped account for sediment hysteresis. Generally, this study also showed a combination of mostly watershed and catchment geospatial features were important in ML models that predict sediment transport from physical samples. This study is a promising step forward in making fluvial sediment transport predictions using machine learning models trained by physically collected samples. The approach developed here can be used wherever similar datasets exists and will be useful for landscape and water management.

Minnesota

LiDAR based prediction of forest biomass using hierarchical models with spatially varying coefficients

Many studies and production inventory systems have shown the utility of coupling covariates derived from Light Detection and Ranging (LiDAR) data with forest variables measured on georeferenced inventory plots through regression models. The objective of this study was to propose and assess the use of a Bayesian hierarchical modeling framework that accommodates both residual spatial dependence and non-stationarity of model covariates through the introduction of spatial random effects. We explored this objective using four forest inventory datasets that are part of the North American Carbon Program, each comprising point-referenced measures of above-ground forest biomass and discrete LiDAR. For each dataset, we considered at least five regression model specifications of varying complexity. Models were assessed based on goodness of fit criteria and predictive performance using a 10-fold cross-validation procedure. Results showed that the addition of spatial random effects to the regression model intercept improved fit and predictive performance in the presence of substantial residual spatial dependence. Additionally, in some cases, allowing either some or all regression slope parameters to vary spatially, via the addition of spatial random effects, further improved model fit and predictive performance. In other instances, models showed improved fit but decreased predictive performance—indicating over-fitting and underscoring the need for cross-validation to assess predictive ability. The proposed Bayesian modeling framework provided access to pixel-level posterior predictive distributions that were useful for uncertainty mapping, diagnosing spatial extrapolation issues, revealing missing model covariates, and discovering locally significant parameters.

Colorado, Minnesota