USGS Science⌕ Search

SEARCH · USGS Science

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Regional flow duration curves: Geostatistical techniques versus multivariate regression

A period-of-record flow duration curve (FDC) represents the relationship between the magnitude and frequency of daily streamflows. Prediction of FDCs is of great importance for locations characterized by sparse or missing streamflow observations. We present a detailed comparison of two methods which are capable of predicting an FDC at ungauged basins: (1) an adaptation of the geostatistical method, Top-kriging, employing a linear weighted average of dimensionless empirical FDCs, standardised with a reference streamflow value; and (2) regional multiple linear regression of streamflow quantiles, perhaps the most common method for the prediction of FDCs at ungauged sites. In particular, Top-kriging relies on a metric for expressing the similarity between catchments computed as the negative deviation of the FDC from a reference streamflow value, which we termed total negative deviation (TND). Comparisons of these two methods are made in 182 largely unregulated river catchments in the southeastern U.S. using a three-fold cross-validation algorithm. Our results reveal that the two methods perform similarly throughout flow-regimes, with average Nash-Sutcliffe Efficiencies 0.566 and 0.662, (0.883 and 0.829 on log-transformed quantiles) for the geostatistical and the linear regression models, respectively. The differences between the reproduction of FDC's occurred mostly for low flows with exceedance probability (i.e. duration) above 0.98.

Advances in Water Resources↗

Operationalizing crop model data assimilation for improved on-farm situational awareness

The ability of ‘digital agriculture’ to support on-farm decision making is predicated on the real-time combination of observations and prior knowledge into an integrated digital environment. The mathematical discipline that seeks to provide this integration is known as model data assimilation (DA), with demonstrated benefits including improved predictive reliability, and the capacity to identify unexpected changes in field conditions and potential measurement errors. Despite routine adoption in other fields, the delayed adoption of DA in agriculture is due to the need to express end-of-season outcomes such as yield, update forecasts of these outcomes throughout the growing season as data become available, and enhance forecast reliability. To overcome these challenges, three guiding principles are introduced, providing a means to operationalize crop model DA for robust on-farm decision support. We apply the guiding principles using a South Australian viticulture case study. Our case study involves application of an iterative form of a widely used DA algorithm (ensemble Kalman filter) to dynamically update both static parameters and states associated with a grapevine simulation model. Daily weather data as well as fortnightly ground-based leaf area index (LAI) data are used for assimilation. It is shown how crop model DA can lead to not only significant improvements in forecasts of LAI but also to forecasts of end-of-season yield. The guiding principles also enable observations of greatest value to be identified throughout the season. This study highlights the role that formal crop model DA can play in agricultural decision support through enhancing situational awareness in real time.

Agricultural and Forest Meteorology↗

Characterizing recent and projecting future potential patterns of mountain pine beetle outbreaks in the Southern Rocky Mountains

The recent widespread mountain pine beetle (MPB) outbreak in the Southern Rocky Mountains presents an opportunity to investigate the relative influence of anthropogenic, biologic, and physical drivers that have shaped the spatiotemporal patterns of the outbreak. The aim of this study was to quantify the landscape-level drivers that explained the dynamic patterns of MPB mortality, and simulate areas with future potential MPB mortality under projected climate-change scenarios in Grand County, Colorado, USA. The outbreak patterns of MPB were characterized by analysis of a decade-long Landsat time-series stack, aided by automatic attribution of change detected by the Landsat-based Detection of Trends in Disturbance and Recovery algorithm (LandTrendr). The annual area of new MPB mortality was then related to a suite of anthropogenic, biologic, and physical predictor variables under a general linear model (GLM) framework. Data from years 2001–2005 were used to train the model and data from years 2006–2011 were retained for validation. After stepwise removal of non-significant predictors, the remaining predictors in the GLM indicated that neighborhood mortality, winter mean temperature anomaly, and residential housing density were positively associated with MPB mortality, whereas summer precipitation was negatively related. The final model had an average area under the curve (AUC) of a receiver operating characteristic plot value of 0.72 in predicting the annual area of new mortality for the independent validation years, and the mean deviation from the base maps in the MPB mortality areal estimates was around 5%. The extent of MPB mortality will likely expand under two climate-change scenarios (RCP 4.5 and 8.5) in Grand County, which implies that the impacts of MPB outbreaks on vegetation composition and structure, and ecosystem functioning are likely to increase in the future.

Colorado↗

Satellite remotely-sensed land surface parameters and their climatic effects for three metropolitan regions

By using both high-resolution orthoimagery and medium-resolution Landsat satellite imagery with other geospatial information, several land surface parameters including impervious surfaces and land surface temperatures for three geographically distinct urban areas in the United States – Seattle, Washington, Tampa Bay, Florida, and Las Vegas, Nevada, are obtained. Percent impervious surface is used to quantitatively define the spatial extent and development density of urban land use. Land surface temperatures were retrieved by using a single band algorithm that processes both thermal infrared satellite data and total atmospheric water vapor content. Land surface temperatures were analyzed for different land use and land cover categories in the three regions. The heterogeneity of urban land surface and associated spatial extents were shown to influence surface thermal conditions because of the removal of vegetative cover, the introduction of non-transpiring surfaces, and the reduction in evaporation over urban impervious surfaces. Fifty years of in situ climate data were integrated to assess regional climatic conditions. The spatial structure of surface heating influenced by landscape characteristics has a profound influence on regional climate conditions, especially through urban heat island effects.

Advances in Space Research↗

Optimal control of native predators

We apply decision theory in a structured decision-making framework to evaluate how control of raccoons ( Procyon lotor ), a native predator, can promote the conservation of a declining population of American Oystercatchers ( Haematopus palliatus ) on the Outer Banks of North Carolina. Our management objective was to maintain Oystercatcher productivity above a level deemed necessary for population recovery while minimizing raccoon removal. We evaluated several scenarios including no raccoon removal, and applied an adaptive optimization algorithm to account for parameter uncertainty. We show how adaptive optimization can be used to account for uncertainties about how raccoon control may affect Oystercatcher productivity. Adaptive management can reduce this type of uncertainty and is particularly well suited for addressing controversial management issues such as native predator control. The case study also offers several insights that may be relevant to the optimal control of other native predators. First, we found that stage-specific removal policies (e.g., yearling versus adult raccoon removals) were most efficient if the reproductive values among stage classes were very different. Second, we found that the optimal control of raccoons would result in higher Oystercatcher productivity than the minimum levels recommended for this species. Third, we found that removing more raccoons initially minimized the total number of removals necessary to meet long term management objectives. Finally, if for logistical reasons managers cannot sustain a removal program by removing a minimum number of raccoons annually, managers may run the risk of creating an ecological trap for Oystercatchers.

North Carolina↗

Prioritizing restoration areas to conserve multiple sagebrush-associated wildlife species

Strategic restoration of altered habitat is one method for addressing worldwide biodiversity declines. Within the sagebrush steppe of western North America, habitat degradation has been linked to declines in many species, making restoration a priority for managers; however, limited funding, spatiotemporal variation in restoration success, and the need to manage for diverse wildlife species make decision-making regarding restoration actions challenging. To address the challenge of spatial conservation prioritization, we developed the Prioritizing Restoration of Sagebrush Ecosystems Tool (PReSET). This decision support tool utilizes the prioritizr package in program R and an integer linear programming algorithm to select parcels representing both high biodiversity value and high probability of restoration success. We tested PReSET on a sagebrush steppe system within southwestern Wyoming using distributional data for six species with diverse life histories and a spatial layer of predicted sagebrush recovery times to identify restoration targets at both broad and local scales. While the broad-scale portion of our tool outputs can inform policy, the local-scale results can be applied directly to on-the-ground restoration. We identified restoration priority areas with greater precision than existing spatial prioritizations and incorporated range differences among species. We noted tradeoffs, including that restoring for habitat connectivity may require restoration actions in areas with lower probability of success. Future applications of PReSET will draw from emerging datasets, including spatially-varying economic costs of restoration, animal movement data, and additional species, to further improve our ability to target effective sagebrush restoration.

Wyoming↗

KGS-HighK: A Fortran 90 program for simulation of hydraulic tests in highly permeable aquifers

Slug and pumping tests (hydraulic tests) are frequently used by hydrogeologists to obtain in-situ estimates of the transmissive and storage properties of a formation (Streltsova, 1988; Kruseman and de Ridder, 1990; Butler, 1998). In aquifers of high hydraulic conductivity, hydraulic tests are affected by mechanisms that are not considered in the analysis of tests in less permeable media (Bredehoeft et al., 1966). Inertia-induced oscillations in hydraulic head are the most common manifestation of such mechanisms. Over the last three decades, a number of analytical solutions that incorporate these mechanisms have been developed for the analysis of hydraulic tests in highly permeable aquifers (see Butler and Zhan (2004) for a review of this previous work). These solutions, however, are restricted to a subset of the conditions commonly encountered in the field. Recently, a more general solution has been developed that builds on this previous work to remove many of the limitations imposed by these earlier approaches (Butler and Zhan, 2004). The purpose of this note is to present a Fortran 90 program, KGS-HighK, for the evaluation of this new solution. This note begins with a brief overview of the conceptual model that motivated the development of the solution of Butler and Zhan (2004) for pumping- and slug-induced flow to/from a central well. The major steps in the derivation of that solution are described, but no details are given. Instead, a Mathematica notebook is provided for those interested in the derivation details. The key algorithms used in KGS-HighK are then described and the program structure is briefly outlined. A field example is provided to demonstrate program performance. The note concludes with a short summary section. ?? 2005 Elsevier Ltd. All rights reserved.

Computers & Geosciences↗

A wetting and drying scheme for ROMS

The processes of wetting and drying have many important physical and biological impacts on shallow water systems. Inundation and dewatering effects on coastal mud flats and beaches occur on various time scales ranging from storm surge, periodic rise and fall of the tide, to infragravity wave motions. To correctly simulate these physical processes with a numerical model requires the capability of the computational cells to become inundated and dewatered. In this paper, we describe a method for wetting and drying based on an approach consistent with a cell-face blocking algorithm. The method allows water to always flow into any cell, but prevents outflow from a cell when the total depth in that cell is less than a user defined critical value. We describe the method, the implementation into the three-dimensional Regional Oceanographic Modeling System (ROMS), and exhibit the new capability under three scenarios: an analytical expression for shallow water flows, a dam break test case, and a realistic application to part of a wetland area along the Georgia Coast, USA.

Georgia↗

MTpy: A Python toolbox for magnetotellurics

We present the software package MTpy that allows handling, processing, and imaging of magnetotelluric (MT) data sets. Written in Python, the code is open source, containing sub-packages and modules for various tasks within the standard MT data processing and handling scheme. Besides the independent definition of classes and functions, MTpy provides wrappers and convenience scripts to call standard external data processing and modelling software. In its current state, modules and functions of MTpy work on raw and pre-processed MT data. However, opposite to providing a static compilation of software, we prefer to introduce MTpy as a flexible software toolbox, whose contents can be combined and utilised according to the respective needs of the user. Just as the overall functionality of a mechanical toolbox can be extended by adding new tools, MTpy is a flexible framework, which will be dynamically extended in the future. Furthermore, it can help to unify and extend existing codes and algorithms within the (academic) MT community. In this paper, we introduce the structure and concept of MTpy . Additionally, we show some examples from an everyday work-flow of MT data processing: the generation of standard EDI data files from raw electric ( E -) and magnetic flux density ( B -) field time series as input, the conversion into MiniSEED data format, as well as the generation of a graphical data representation in the form of a Phase Tensor pseudosection.

Computers & Geosciences↗

Determining on-fault earthquake magnitude distributions from integer programming

Earthquake magnitude distributions among faults within a fault system are determined from regional seismicity and fault slip rates using binary integer programming. A synthetic earthquake catalog (i.e., list of randomly sampled magnitudes) that spans millennia is first formed, assuming that regional seismicity follows a Gutenberg-Richter relation. Each earthquake in the synthetic catalog can occur on any fault and at any location. The objective is to minimize misfits in the target slip rate for each fault, where slip for each earthquake is scaled from its magnitude. The decision vector consists of binary variables indicating which locations are optimal among all possibilities. Uncertainty estimates in fault slip rates provide explicit upper and lower bounding constraints to the problem. An implicit constraint is that an earthquake can only be located on a fault if it is long enough to contain that earthquake. A general mixed-integer programming solver, consisting of a number of different algorithms, is used to determine the optimal decision vector. A case study is presented for the State of California, where a 4 kyr synthetic earthquake catalog is created and faults with slip ≥3 mm/yr are considered, resulting in >10 6 variables. The optimal magnitude distributions for each of the faults in the system span a rich diversity of shapes, ranging from characteristic to power-law distributions.

Computers & Geosciences↗

3D semantic mapping of surface geological features

Semantic mapping in 3D is fundamental to a wide range of geoscientific studies and applications, including geomorphology, hazard assessment, and environmental monitoring. However, automatically segmenting geological features from large-scale photogrammetric datasets remains a significant challenge. We present a methodology to address this gap. Using overlapping images collected over environments of interest, Structure-from-Motion (SfM) produces georeferenced point clouds and estimates camera poses. Existing large vision models, such as Segment Anything Model, segment objects in the images, generating pixel-segmentation associations. To produce pixel-point associations, we project the points back onto the camera image planes. As objects are independently segmented across multiple images with different perspectives, we develop a segmentation mosaicking algorithm to build probabilistic point-segmentation associations that combines the pixel-segmentation associations and pixel-point associations. Our methodology is validated using both synthetic data generated by Kubric and real-world UAV-SfM data. The implementation is designed to be compatible with existing SfM software, including Agisoft and OpenDroneMap, for photogrammetry mapping in geoscience studies. As a case study, we apply our method to the semantic mapping of precariously balanced rocks (PBRs), which provide upper-bound constraints on historical ground motion shaking intensity. To support object-level identification of PBRs, we additionally integrated Grounding DINO, enabling text-prompted segmentation of features of interest within UAV imagery. This case study demonstrates the effectiveness of our method in generating a 3D semantic map of PBRs, enabling spatial distribution of PBR fragility for earthquake hazard analysis.

Computers & Geosciences↗

Fingerprinting historical tributary contributions to floodplain sediment using bulk geochemistry

Sediment deposition on floodplains is essential for the development and maintenance of riparian ecosystems. Upstream erosion is known to influence downstream floodplain construction, but linking these disparate processes is challenging, especially over large spatial and temporal scales. Sediment fingerprinting is thus a robust tool to establish process linkages between downstream floodplain development and sediment production in distal headwater basins. Here we use sediment geochemistry to connect historical erosion in several tributaries of the Yampa River in Colorado and Wyoming, USA, to the construction of downstream floodplains on which extensive cottonwood forests established. Using a combination of conventional techniques and the relatively novel machine-learning random forest algorithm, we build multiple fingerprints of diagnostic geochemical tracers that are then input into a Bayesian mixing model to apportion provenance of floodplain sediment. Sediment samples for provenance analysis were collected from an excavated floodplain in Deerlodge Park on the Yampa River at the rooting surface of the surrounding cottonwood forest and dominantly comprised of very fine (4Φ) sand. Fingerprinting analysis of the 4Φ fraction of collected floodplain sink (n = 38) and tributary source (n = 218) samples revealed floodplain sediment to be dominantly sourced from the tributaries of Muddy Creek (45 ± 4%) and Sand Wash (42 ± 6%). Dendrochronology results moreover indicate the Deerlodge floodplain sediment was deposited in ∼1912, which falls squarely within the time (1880–1940) these tributaries were actively eroding. Taken together, study results indicate a demonstrable link between historical tributary erosion and downstream floodplain construction and concomitant forest establishment. Our findings suggest processes operating in tributary watersheds play an important role in the dynamics of large rivers and emphasize both the need for holistic, collaborative management of sediment as an essential resource and the potential to utilize sediment fingerprinting to inform and direct river ecosystem management.

Colorado, Wyoming↗

Machine learning and data augmentation approach for identification of rare earth element potential in Indiana Coals, USA

Rare earth elements and yttrium (REYs) are critical elements and valuable commodities due to their limited availability and high demand in a wide range of applications and especially in high-technology products. The increased demand and geopolitical pressures motivate the search for alternative sources of REYs, and coal, coal waste, and coal ash are considered as new sources for these critical elements. This research evaluates the REY potential of coals from Indiana (USA). However, although coal data revealed REY potential, it suffered from sparse samples with complete REY measurements. Therefore, we explore the applicability of machine learning (ML) models and data augmentation techniques to demonstrate their applicability to evaluate REY potential in Indiana, and other areas in coal basins, using selected coal parameters (Al2O3, Fe2O3, C, Ash, S, P, Mo, Zn, and As contents) as covariates (indicators). Due to the relatively small sample size with complete REY data in the Indiana Coal Database, two data augmentation techniques (Random Over-Sampling Examples and Synthetic Minority Over-Sampling Technique) were used. Four machine learning algorithms (linear discriminate analysis, support vector machine, random forest, and artificial neural networks) were applied for modeling REY potential as a classification problem. The results show that application of Synthetic Minority Over-Sampling Technique prior to development of the support vector machine (SVM) models generated the best REY classification with an accuracy of 95%. The encouraging results based on Indiana coal data may suggest that a similar approach can be used for other coal basins for screening the locations with REY potential. Those locations then can be targeted for more detailed geochemical surveys to identify most promising areas and evaluate overall REY resources.

Indiana↗

Merging machine learning and geostatistical approaches for spatial modeling of geoenergy resources

Geostatistics is the most commonly used probabilistic approach for modeling earth systems, including quality parameters of various geoenergy resources. In geostatistics, estimates, either on a point or block support, are generated as a spatially-weighted average of surrounding samples. The optimal weights are determined through the stationary variogram model which accounts for the spatial structure of the samples. Recently, efficient modeling workflows using various machine learning algorithms (MLAs) have been expanded to the spatial context for modeling geological heterogeneity. The flexible use of MLAs as a spatial estimation tool stems mainly from the fact that unlike kriging, they do not require any variogram, nor do they depend strongly on a prior stationarity assumption (i.e., second order stationarity). This study evaluates the performance of two MLAs (ensemble super learner and elliptical radial basis neural network), ordinary kriging, and hybrid spatial modeling approaches using ordinary intrinsic collocated cokriging. The aforementioned modeling techniques are compared for estimating resources for four coal variables (wash yield, ash yield, calorific value and thickness) as an example. The results suggest that MLAs, when implemented alone, do not outperform ordinary kriging, but the estimation accuracy of the final model, measured by the root mean squared error tends to subtly improve (

Virginia↗

Parameterization and simulation of near bed orbital velocities under irregular waves in shallow water

A set of empirical formulations is derived that describe important wave properties in shallow water as functions of commonly used parameters such as wave height, wave period, local water depth and local bed slope. These wave properties include time varying near-bed orbital velocities and statistical properties such as the distribution of wave height and wave period. Empirical expressions of characteristic wave parameters are derived on the basis of extensive analysis of field data using recently developed evolutionary algorithms. The field data covered a wide range of wave conditions, though there were few conditions with wave periods greater than 15 s. Comparison with field measurements showed good agreement both on a time scale of a single wave period as well as time averaged velocity moments.

Coastal Engineering↗

L-moments and TL-moments of the generalized lambda distribution

The 4-parameter generalized lambda distribution (GLD) is a flexible distribution capable of mimicking the shapes of many distributions and data samples including those with heavy tails. The method of L-moments and the recently developed method of trimmed L-moments (TL-moments) are attractive techniques for parameter estimation for heavy-tailed distributions for which the L- and TL-moments have been defined. Analytical solutions for the first five L- and TL-moments in terms of GLD parameters are derived. Unfortunately, numerical methods are needed to compute the parameters from the L- or TL-moments. Algorithms are suggested for parameter estimation. Application of the GLD using both L- and TL-moment parameter estimates from example data is demonstrated, and comparison of the L-moment fit of the 4-parameter kappa distribution is made. A small simulation study of the 98th percentile (far-right tail) is conducted for a heavy-tail GLD with high-outlier contamination. The simulations show, with respect to estimation of the 98th-percent quantile, that TL-moments are less biased (more robost) in the presence of high-outlier contamination. However, the robustness comes at the expense of considerably more sampling variability. ?? 2006 Elsevier B.V. All rights reserved.

Computational Statistics and Data Analysis↗

Debris-flow monitoring and warning: Review and examples

Debris flows represent one of the most dangerous types of mass movements, because of their high velocities, large impact forces and long runout distances. This review describes the available debris-flow monitoring techniques and proposes recommendations to inform the design of future monitoring and warning/alarm systems. The selection and application of these techniques is highly dependent on site and hazard characterization, which is illustrated through detailed descriptions of nine monitoring sites: five in Europe, three in Asia and one in the USA. Most of these monitored catchments cover less than ∼10 km 2 and are topographically rugged with Melton Indices greater than 0.5. Hourly rainfall intensities between 5 and 15 mm/h are sufficient to trigger debris flows at many of the sites, and observed debris-flow volumes range from a few hundred up to almost one million cubic meters. The sensors found in these monitoring systems can be separated into two classes: a class measuring the initiation mechanisms, and another class measuring the flow dynamics. The first class principally includes rain gauges, but also contains of soil moisture and pore-water pressure sensors. The second class involves a large variety of sensors focusing on flow stage or ground vibrations and commonly includes video cameras to validate and aid in the data interpretation. Given the sporadic nature of debris flows, an essential characteristic of the monitoring systems is the differentiation between a continuous mode that samples at low frequency (“non-event mode”) and another mode that records the measurements at high frequency (“event mode”). The event detection algorithm, used to switch into the “event mode” depends on a threshold that is typically based on rainfall or ground vibration. Identifying the correct definition of these thresholds is a fundamental task not only for monitoring purposes, but also for the implementation of warning and alarm systems.

Earth-Science Reviews↗

Evaluating a tandem human-machine approach to labelling of wildlife in remote camera monitoring

Remote cameras (“trail cameras”) are a popular tool for non-invasive, continuous wildlife monitoring, and as they become more prevalent in wildlife research, machine learning (ML) is increasingly used to automate or accelerate the labor-intensive process of labelling (i.e., tagging) photos. Human-machine hybrid tagging approaches have been shown to greatly increase tagging efficiency (i.e., time to tag a single image). However, those potential increases hinge on the extent to which an ML model makes correct vs. incorrect predictions. We performed an experiment using a ML model that produces bounding boxes around animals, people, and vehicles in remote camera imagery (MegaDetector) to consider the impact of a ML model’s performance on its ability to accelerate human labeling. Six participants tagged trail camera images collected from 12 sites in Vermont and Maine, USA (January–September 2022) using three tagging methods (one with ML bounding box assistance and two without assistance). We used a generalized linear mixed model to examine the influence of ML model performance and tagging method on tagging efficiency. We found that ML bounding boxes offer significant improvement in tagging efficiency when labelling data compared to unassisted tagging. Additionally, the time taken to label with bounding boxes was not statistically different from an unassisted tagging approach. However, we found that gains in efficiency are contingent on the ML algorithm’s performance and that incorrect ML predictions, particularly the 4.2% false positive and 3.6% false negative predictions, can slow the tagging process compared to a non-hybrid approach. These findings indicate that although practitioners usually forgo the production of bounding boxes when selecting a data labelling process due to the increased effort, ML bounding box-assisted tagging can offer an efficient method for labeling. More broadly, ML-assisted data labelling offers an opportunity to accelerate the analysis of trail camera imagery, but an assessment of the ML model’s performance can illuminate whether the hybrid-tagging approach is ultimately a help or hinderance.

Maine, Vermont↗