USGS ScienceSearch

SEARCH · USGS Science

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Beware of spatial autocorrelation when applying machine learning algorithms to borehole geophysical logs

Although many of the algorithms now considered to be machine learning algorithms (MLAs) have existed for nearly a century (e.g., Rosenblatt 1958 ), interest in MLAs has recently increased exponentially for solving data-driven problems across a variety of fields due to the expanded availability of large, complex datasets that may be difficult to interrogate using other methods, increases in computing power, and a growing library of easily implemented machine learning tools. While MLAs are often similar to statistical methods, there are key differences in the approach to problem solving. Namely, statistical methods are more concerned with generating informative models from “long” data (i.e., many more observations than explanatory variables), whereas MLAs are typically concerned with generating accurate predictions from “wide” data (i.e., a large number of variables with relatively fewer observations, Bzdok et al. 2018 ). In hydrogeologic studies, such wide datasets may be available from boreholes, where various types of geophysical, geochemical, and lithological information may exist. Borehole datasets are therefore a tempting target for MLAs to reveal hidden relations among gathered data and parameters of interest (e.g., contaminant concentration), and as a method of parameter reduction (e.g., reduce costs by collecting fewer datasets).

Groundwater

Comparison of an algebraic multigrid algorithm to two iterative solvers used for modeling ground water flow and transport

Numerical solution of large-scale ground water flow and transport problems is often constrained by the convergence behavior of the iterative solvers used to solve the resulting systems of equations. We demonstrate the ability of an algebraic multigrid algorithm (AMG) to efficiently solve the large, sparse systems of equations that result from computational models of ground water flow and transport in large and complex domains. Unlike geometric multigrid methods, this algorithm is applicable to problems in complex flow geometries, such as those encountered in pore-scale modeling of two-phase flow and transport. We integrated AMG into MODFLOW 2000 to compare two- and three-dimensional flow simulations using AMG to simulations using PCG2, a preconditioned conjugate gradient solver that uses the modified incomplete Cholesky preconditioner and is included with MODFLOW 2000. CPU times required for convergence with AMG were up to 140 times faster than those for PCG2. The cost of this increased speed was up to a nine-fold increase in required random access memory (RAM) for the three-dimensional problems and up to a four-fold increase in required RAM for the two-dimensional problems. We also compared two-dimensional numerical simulations of steady-state transport using AMG and the generalized minimum residual method with an incomplete LU-decomposition preconditioner. For these transport simulations, AMG yielded increased speeds of up to 17 times with only a 20% increase in required RAM. The ability of AMG to solve flow and transport problems in large, complex flow systems and its ready availability make it an ideal solver for use in both field-scale and pore-scale modeling.

Ground Water

Fast algorithm for automatically computing Strahler stream order

An efficient algorithm was developed to determine Strahler stream order for segments of stream networks represented in a Geographic Information System (GIS). The algorithm correctly assigns Strahler stream order in topologically complex situations such as braided streams and multiple drainage outlets. Execution time varies nearly linearly with the number of stream segments in the network. This technique is expected to be particularly useful for studying the topology of dense stream networks derived from digital elevation model data.

Water Resources Bulletin

A Double-difference Earthquake location algorithm: Method and application to the Northern Hayward Fault, California

We have developed an efficient method to determine high-resolution hypocenter locations over large distances. The location method incorporates ordinary absolute travel-time measurements and/or cross-correlation P-and S-wave differential travel-time measurements. Residuals between observed and theoretical travel-time differences (or double-differences) are minimized for pairs of earthquakes at each station while linking together all observed event-station pairs. A least-squares solution is found by iteratively adjusting the vector difference between hypocentral pairs. The double-difference algorithm minimizes errors due to unmodeled velocity structure without the use of station corrections. Because catalog and cross-correlation data are combined into one system of equations, interevent distances within multiplets are determined to the accuracy of the cross-correlation data, while the relative locations between multiplets and uncorrelated events are simultaneously determined to the accuracy of the absolute travel-time data. Statistical resampling methods are used to estimate data accuracy and location errors. Uncertainties in double-difference locations are improved by more than an order of magnitude compared to catalog locations. The algorithm is tested, and its performance is demonstrated on two clusters of earthquakes located on the northern Hayward fault, California. There it colapses the diffuse catalog locations into sharp images of seismicity and reveals horizontal lineations of hypocenter that define the narrow regions on the fault where stress is released by brittle failure.

Bulletin of the Seismological Society of America

Modeling seismic network detection thresholds using production picking algorithms

Estimating the detection threshold of a seismic network (the minimum magnitude earthquake that can be reliably located) is a critical part of network design and can drive network maintenance efforts. The ability of a station to detect an earthquake is often estimated by assuming the spectral amplitude for an earthquake of a given size, assuming an attenuation relationship, and comparing the predicted amplitude with the average station background noise level. This approach has significant uncertainty because of unknown regional attenuation and complications in computing small event power spectra, and it fails to account for the specific capabilities of the automatic seismic phase picker used in monitoring. We develop a data‐driven approach to determine network detection thresholds using a multiband phase picking algorithm that is currently in use at the U.S. Geological Survey National Earthquake Information Center. We apply this picking algorithm to cataloged earthquakes to determine an empirical relationship of the observability of earthquakes as a function of magnitude and distance. Using this relationship, we produce maps of detection threshold using station spatial configuration and station noise levels. We show that quiet, well‐sited stations significantly increase the detection capabilities of a network compared with a network composed of many noisy stations. Because our method is data driven, it has two distinct advantages: (1) it is less dependent on theoretical assumptions of source spectra and models of regional attenuation, and (2) it can easily be applied to any seismic network. This tool allows for an objective approach to the management of stations in regional seismic networks.

Seismological Research Letters

Landsat ecosystem disturbance adaptive processing system (LEDAPS) algorithm description

The Landsat Ecosystem Disturbance Adaptive Processing System (LEDAPS) software was originally developed by the National Aeronautics and Space Administration–Goddard Space Flight Center and the University of Maryland to produce top-of-atmosphere reflectance from LandsatThematic Mapper and Enhanced Thematic Mapper Plus Level 1 digital numbers and to apply atmospheric corrections to generate a surface-reflectance product.The U.S. Geological Survey (USGS) has adopted the LEDAPS algorithm for producing the Landsat Surface Reflectance Climate Data Record.This report discusses the LEDAPS algorithm, which was implemented by the USGS.

Open-File Report

Improved algorithms in the CE-QUAL-W2 water-quality model for blending dam releases to meet downstream water-temperature targets

Water-quality models allow water resource professionals to examine conditions under an almost unlimited variety of potential future scenarios. The two-dimensional (longitudinal, vertical) water-quality model CE-QUAL-W2, version 3.7, was enhanced and augmented with new features to help dam operators and managers explore and optimize potential solutions for temperature management downstream of thermally stratified reservoirs. Such temperature management often is accomplished by blending releases from multiple dam outlets that access water of different temperatures at different depths. The modified blending algorithm in version 3.7 of CE-QUAL-W2 allows the user to specify a time-series of target release temperatures, designate from 2 to 10 floating or fixed-elevation outlets for blending, impose minimum and maximum head and flow constraints for any blended outlet, and set priority designations for each outlet that allow the model to choose which outlets to use and how to balance releases among them. The modified model was tested with a variety of examples and against a previously calibrated model of Detroit Lake on the North Santiam River in northwestern Oregon, and the results compared well. These updates to the blending algorithms will allow more complicated dam-operation scenarios to be evaluated somewhat automatically with the model, with decreased need for multiple model runs or preprocessing of model inputs to fully characterize the operational constraints.

Open-File Report

Global cropland-extent product at 30-m resolution (GCEP30) derived from Landsat satellite time-series data for the year 2015 using multiple machine-learning algorithms on Google Earth Engine cloud

Executive Summary Global food and water security analysis and management require precise and accurate global cropland-extent maps. Existing maps have limitations, in that they are (1) mapped using coarse-resolution remote-sensing data, resulting in the lack of precise mapping location of croplands and their accuracies; (2) derived by collecting and collating national statistical data that are often subjective, leading to substantial uncertainties in cropland-area estimates, as well as their locations; and (3) extracted from one or more classes of a land use–land cover product in which cropland classes are not the focus of mapping, leading to their mixing with other classes and creating significant errors of omission and commission. These limitations can be overcome by producing high-resolution cropland-extent maps using satellite-sensor data, such as Landsat 30-m resolution or higher. The most fundamental cropland product is the high-resolution cropland-extent map because all higher level cropland products, such as crop-watering method (that is, whether crops are irrigated or rainfed), crop types, cropping intensities, cropland fallows, crop productivity, and crop-water productivity, are dependent on a precise and accurate cropland-extent product. Given these realities, the overarching goal of this study was to produce a Landsat satellite-derived global cropland-extent product at 30-m resolution. The work, which involved a paradigm shift in how global cropland-extent maps are produced, involved the following five key steps: (1) petabyte-scale computing that involved multiyear, 8- to 16-day, time-series Landsat 30-m resolution data for the global land surface; (2) composition of analysis-ready data (ARD) cubes; (3) creation of a large global-reference data hub for machine learning; (4) use of multiple machine-learning algorithms (MLAs) by writing software and computing in the cloud; and (5) Google Earth Engine (GEE) cloud computing. The five key steps involved nine distinct phases. First, the world was segmented into 74 agroecological zones (AEZs). Second, Landsat 8- to 16-day data were used to time-composite 10-band (blue, green, red, near-infrared, short-wave infrared band 1, short-wave infrared band 2, thermal infrared, enhanced vegetation index, normalized difference water index, and normalized difference vegetation index) Landsat 30-m resolution data cubes for every 2- to 4-month time period during 3- to 4-year periods (stated as nominal-year 2015 or, simply, 2015), along with two additional 30-m resolution bands (Shuttle Radar Topography Mission elevation, and slope) in each of the 74 AEZs. Third, more than 100,000 reference-training data samples were collected using ground data (some of which were collected using a mobile application), as well as submeter- to 5-m-resolution, very high-resolution imagery sourced from other reliable sources. Fourth, reference-training data were used to create a knowledge base for separating cropland from noncropland. Fifth, MLAs such as the pixel-based supervised random forest and support-vector machines were written on the GEE using Python and JavaScript. Sixth, object-based recursive hierarchical segmentation algorithm was used, in addition to MLAs, to overcome uncertainties. Seventh, MLAs used the knowledge base to classify and separate cropland from noncropland. Eighth, accuracy assessment was conducted by generating error matrices for each of the 74 AEZs using 19,171 independent validation-data samples. Ninth, cropland areas were computed for all countries of the world and compared with United Nation’s (UN’s) Food and Agricultural Organization (FAO) and other national statistics. The outcome was a Landsat-derived global cropland-extent product at 30-m resolution (GCEP30), which has an overall accuracy of 91.7 percent. For the cropland class, producer’s accuracy was 83.4 percent, and user’s accuracy was 78.3 percent. GCEP30 calculated (using direct pixel count) the global net-cropland area (GNCA) for the year 2015 as 1.873 billion hectares (~12.6 percent of the Earth’s terrestrial area). The continental cropland distribution as a percentage of GNCA was Asia, 33 percent; Europe, 25.5 percent; Africa, 16.7 percent; North America, 14.4 percent; South America, 8.1 percent; and Australia and Oceania, 2.4 percent. The worldwide cropland areas in GCEP30 for 2015 were higher by 236 to 299 million hectares (Mha) compared to national statistics reported elsewhere for the same year (for example, in Food and Agriculture Organization’s corporate statistical database [FAOSTAT] and in the monthly irrigated and rainfed crop areas [MIRCA] database). The global cropland area reported for 2015 increased by 344 Mha (22.5 percent), compared to the year 2000. During the same period (2000–2015), the world’s population increased by 20 percent. Whereas some of these areal increases are real increases in cropland areas, others are due to the types of data, methods, and approaches used. Using the highest known resolution (compared to previous coarse-resolution global products) enabled this study to capture fragmented croplands. Coarse-resolution data compute areas on the basis of subpixels, which, for a large proportion of certain land use–land cover classes, will show only a certain percentage of the total pixel area as actual area. Subpixel areas can lead to substantial uncertainties in area computation, as determining the exact fraction of cropland areas within a coarse-resolution pixel is resource intensive and subject to errors. Other innovations in GCEP30 include reference-data hubs, machine learning, and cloud computing. Cropland areas in 214 countries, territories, departments, and regions were calculated for the year 2015 using GCEP30, on the basis of UN’s global administrative unit layers (GAUL) boundaries. The 10 leading countries in terms of cropland area (as a percentage of the GNCA) were India (9.6 percent), United States (8.95 percent), China (8.82 percent), Russia (8.32 percent), Brazil (3.42 percent), Ukraine (2.32 percent), Canada (2.29 percent), Argentina (2.05 percent), Indonesia (2 percent), and Nigeria (1.91 percent). Together, these 10 countries occupy 50 percent of the global cropland, and they have 52 percent of the global population. Their combined cropland area increased by 2 percent between 2000 and 2015, compared to the substantial increase in population of 517 million (15.5 percent). Together, India, United States, China, and Russia encompass 36 percent of the total area. In the United States and Canada, from 2000 to 2015, cropland decreased by about 2 percent, whereas their populations increased by 14 and 13 percent, respectively. The additional food requirements in these 10 countries, which are caused by increased populations, as well as increasing nutritional demands, are met by production increases in existing cropland or through virtual food trade, or both. More than 18 countries, territories, departments, or regions had 60 percent or more of their geographic area as cropland: Republic of Moldova, San Marino, and Hungary had more than 80 percent of the country’s area as cropland; Denmark, Ukraine, Ireland, and Bangladesh, 70 to 80 percent; and Uruguay, Netherlands, United Kingdom, Spain, Lithuania, Poland, Gaza Strip, Czechia, Italy, India, and Azerbaijan, 60 to 70 percent. Europe and South Asia can be considered agricultural capitals of the world, on the basis of their percentages of geographic area as cropland. United States, China, and Russia, which all have high cropland areas, are ranked second, third, and fourth in the world; India is ranked first. However, the amount of cropland as a percentage of the country’s geographic area is relatively very low for United States (18.3 percent), China (17.7 percent), and Russia (9.5 percent), whereas it is 60.5 percent for India. Most African and South American countries, territories, departments, or regions have less than 15 percent of their geographic area as cropland. China and India together house 36 percent of the world’s population; however, between 2000 and 2015, the amount of China’s cropland area fell by 18.9 percent, owing to urban expansion and the abandonment of farmlands caused by demographic changes (that is, the movement of population from villages to cities). In contrast, China’s population grew by 10 percent. The amount of India’s cropland increased by 8.5 percent, whereas its population grew by 20 percent. This study showed that, out of the 10 leading cropland countries, Ukraine, Nigeria, Russia, and Indonesia showed an 18 to 31 percent increase in cropland areas, on the basis of GCEP30 by the year 2015, compared to 2000. Nigeria’s cropland area increased by 25 percent, and its population increased by 31 percent in the same period. In these countries, food security is maintained by cropland expansion, productivity increases, and virtual food trade. Nevertheless, this trend of increasing net-cropland area and productivity will likely become difficult to maintain, owing to diminishing arable lands and plateauing of 50 years of continual yield increases, requiring policymakers to explore novel and data-supported approaches to solving future food security issues. The GCEP30 product, which can be browsed at full resolution at www.croplands.org , has been released for public download and use through U.S. Geological Survey (USGS)–National Aeronautics and Space Administration (NASA) Land Processes Distributed Active Archive Center (see https://lpdaac.usgs.gov/news/release-of-gfsad-30-meter-cropland-extent-products/ ).

Professional Paper

Optimization of water-level monitoring networks in the eastern Snake River Plain aquifer using a kriging-based genetic algorithm method

Long-term groundwater monitoring networks can provide essential information for the planning and management of water resources. Budget constraints in water resource management agencies often mean a reduction in the number of observation wells included in a monitoring network. A network design tool, distributed as an R package, was developed to determine which wells to exclude from a monitoring network because they add little or no beneficial information. A kriging-based genetic algorithm method was used to optimize the monitoring network. The algorithm was used to find the set of wells whose removal leads to the smallest increase in the weighted sum of the (1) mean standard error at all nodes in the kriging grid where the water table is estimated, (2) root-mean-squared-error between the measured and estimated water-level elevation at the removed sites, (3) mean standard deviation of measurements across time at the removed sites, and (4) mean measurement error of wells in the reduced network. The solution to the optimization problem (the best wells to retain in the monitoring network) depends on the total number of wells removed; this number is a management decision. The network design tool was applied to optimize two observation well networks monitoring the water table of the eastern Snake River Plain aquifer, Idaho; these networks include the 2008 Federal-State Cooperative water-level monitoring network (Co-op network) with 166 observation wells, and the 2008 U.S. Geological Survey-Idaho National Laboratory water-level monitoring network (USGS-INL network) with 171 wells. Each water-level monitoring network was optimized five times: by removing (1) 10, (2) 20, (3) 40, (4) 60, and (5) 80 observation wells from the original network. An examination of the trade-offs associated with changes in the number of wells to remove indicates that 20 wells can be removed from the Co-op network with a relatively small degradation of the estimated water table map, and 40 wells can be removed from the USGS-INL network before the water table map degradation accelerates. The optimal network designs indicate the robustness of the network design tool. Observation wells were removed from high well-density areas of the network while retaining the spatial pattern of the existing water-table map.

Idaho

An algorithm for correction of atmospheric scattering dilution effects in volcanic gas emission measurements using skylight differential optical absorption spectroscopy

Differential Optical Absorption Spectroscopy (DOAS) is commonly used to measure gas emissions from volcanoes. DOAS instruments measure the absorption of solar ultraviolet (UV) radiation scattered in the atmosphere by sulfur dioxide (SO 2 ) and other trace gases contained in volcanic plumes. The standard spectral retrieval methods assume that all measured light comes from behind the plume and has passed through the plume along a straight line. However, a fraction of the light that reaches the instrument may have been scattered beneath the plume and thus has passed around it. Since this component does not contain the absorption signatures of gases in the plume, it effectively “dilutes” the measurements and causes underestimation of the gas abundance in the plume. This dilution effect is small for clean-air conditions and short distances between instrument and plume. However, plume measurements made at long distance and/or in conditions with significant atmospheric aerosol, haze, or clouds may be severely affected. Thus, light dilution is regarded as a major error source in DOAS measurements of volcanic degassing. Several attempts have been made to model the phenomena and the physical mechanisms are today relatively well understood. However, these models require knowledge of the local atmospheric aerosol composition and distribution, parameters that are almost always unknown. Thus, a practical algorithm to quantitatively correct for the dilution effect is still lacking. Here, we propose such an algorithm focused specifically on SO 2 measurements. The method relies on the fact that light absorption becomes non-linear for high SO 2 loads, and that strong and weak SO 2 absorption bands are unequally affected by the diluting signal. These differences can be used to identify when dilution is occurring. Moreover, if we assume that the spectral radiance of the diluting light is identical to the spectrum of light measured away from the plume, a measured clean air spectrum can be used to represent the dilution component. A correction can then be implemented by iteratively subtracting fractions of this clean air spectrum from the measured spectrum until the respective absorption signals on strong and weak SO 2 absorption bands are consistent with a single overhead SO 2 abundance. In this manner, we can quantify the magnitude of light dilution in each individual measurement spectrum as well as obtaining a dilution-corrected value for the SO 2 column density along the line of sight of the instrument. This paper first presents the theory behind the method, then discusses validation experiments using a radiative transfer model, as well as applications to field data obtained under different measurement conditions at three different locations; Fagradalsfjall located on the Reykjanaes peninsula in south Island, Manam located off the northeast coast of mainland Papua New Guinea and Holuhraun located in the inland of north east Island.

Frontiers in Earth Science

An empirical algorithm for estimating agricultural and riparian evapotranspiration using MODIS Enhanced Vegetation Index and ground measurements of ET. I. Description of method

We used the Enhanced Vegetation Index (EVI) from MODIS to scale evapotranspiration (ET actual ) over agricultural and riparian areas along the Lower Colorado River in the southwestern US. Ground measurements of ET actual by alfalfa, saltcedar, cottonwood and arrowweed were expressed as fraction of potential (reference crop) ET o (ET o F) then regressed against EVI scaled between bare soil (0) and full vegetation cover (1.0) (EVI*). EVI* values were calculated based on maximum and minimum EVI values from a large set of riparian values in a previous study. A satisfactory relationship was found between crop and riparian plant ET o F and EVI*, with an error or uncertainty of about 20% in the mean estimate (mean ET actual = 6.2 mm d −1 , RMSE = 1.2 mm d −1 ). The equation for ET actual was: ET actual = 1.22 × ET o-BC × EVI*, where ET o-BC is the Blaney Criddle formula for ET o . This single algorithm applies to all the vegetation types in the study, and offers an alternative to ET actual estimates that use crop coefficients set by expert opinion, by using an algorithm based on the actual state of the canopy as determined by time-series satellite images.

Remote Sensing

Implementation of the CCDC algorithm to produce the LCMAP Collection 1.0 annual land surface change product

The increasing availability of high-quality remote sensing data and advanced technologies have spurred land cover mapping to characterize land change from local to global scales. However, most land change datasets either span multiple decades at a local scale or cover limited time over a larger geographic extent. Here, we present a new land cover and land surface change dataset created by the Land Change Monitoring, Assessment, and Projection (LCMAP) program over the conterminous United States (CONUS). The LCMAP land cover change dataset consists of annual land cover and land cover change products over the period 1985-2017 at 30-meter resolution using Landsat and other ancillary data via the Continuous Change Detection and Classification (CCDC) algorithm. In this paper, we describe our novel approach to implement the CCDC algorithm to produce the LCMAP product suite composed of five land cover and five land surface change related products. The LCMAP land cover products were validated using a collection of ~ 25,000 reference samples collected independently across CONUS. The overall agreement for all years of the LCMAP primary land cover product reached 82.5%. The LCMAP products are produced through the LCMAP Information Warehouse and Data Store (IW+DS) and Shared Mesos Cluster systems that can process, store, and deliver all datasets for public access. To our knowledge, this is the first set of published 30m annual land cover and land cover change datasets that span from the 1980s to the present for the United States. The LCMAP product suite provides useful information for land resource management and facilitates studies to improve the understanding of terrestrial ecosystems and the complex dynamics of the Earth system. The LCMAP system could be implemented to produce global land change products in the future.

Earth System Science Data

Slope Unit Maker (SUMak): An efficient and parameter-free algorithm for delineating slope units to improve landslide modeling

Slope units are terrain partitions bounded by drainage and divide lines. In landslide modeling, including susceptibility modeling and event-specific modeling of landslide occurrence, slope units provide several advantages over gridded units, such as better capturing terrain geometry, improved incorporation of geospatial landslide-occurrence data in different formats (e.g., point and polygon), and better accommodating the varying data accuracy and precision in landslide inventories. However, the use of slope units in regional ( > 100 km 2 ) landslide studies remains limited due, in part, to the large computational costs and/or poor reproducibility with current delineation methods. We introduce a computationally efficient algorithm for the parameter-free delineation of slope units that leverages tools from within TauDEM and GRASS, using an R interface. The algorithm uses geomorphic laws to define the appropriate scaling of the slope units representative of hillslope processes, avoiding the often ambiguous determination of slope unit size. We then demonstrate how slope units enable more robust regional-scale landslide susceptibility and event-specific landslide occurrence maps.

Natural Hazards and Earth Systems Sciences (NHESS)

ALGORITHM DEVELOPMENT FOR SPATIAL OPERATORS.

An approach is given that develops spatial operators about the basic geometric elements common to spatial data structures. In this fashion, a single set of spatial operators may be accessed by any system that reduces its operands to such basic generic representations. Algorithms based on this premise have been formulated to perform operations such as separation, overlap, and intersection. Moreover, this generic approach is well suited for algorithms that exploit concurrent properties of spatial operators. The results may provide a framework for a geometry engine to support fundamental manipulations within a geographic information system.

Conference Paper

Modeling landscape evapotranspiration by integrating land surface phenology and a water balance algorithm

The main objective of this study is to present an improved modeling technique called Vegetation ET (VegET) that integrates commonly used water balance algorithms with remotely sensed Land Surface Phenology (LSP) parameter to conduct operational vegetation water balance modeling of rainfed systems at the LSP’s spatial scale using readily available global data sets. Evaluation of the VegET model was conducted using Flux Tower data and two-year simulation for the conterminous US. The VegET model is capable of estimating actual evapotranspiration (ETa) of rainfed crops and other vegetation types at the spatial resolution of the LSP on a daily basis, replacing the need to estimate crop- and region-specific crop coefficients.

Algorithms

Improving the accessibility and transferability of machine learning algorithms for identification of animals in camera trap images: MLWIC2

Motion‐activated wildlife cameras (or “camera traps”) are frequently used to remotely and noninvasively observe animals. The vast number of images collected from camera trap projects has prompted some biologists to employ machine learning algorithms to automatically recognize species in these images, or at least filter‐out images that do not contain animals. These approaches are often limited by model transferability, as a model trained to recognize species from one location might not work as well for the same species in different locations. Furthermore, these methods often require advanced computational skills, making them inaccessible to many biologists. We used 3 million camera trap images from 18 studies in 10 states across the United States of America to train two deep neural networks, one that recognizes 58 species, the “species model,” and one that determines if an image is empty or if it contains an animal, the “empty‐animal model.” Our species model and empty‐animal model had accuracies of 96.8% and 97.3%, respectively. Furthermore, the models performed well on some out‐of‐sample datasets, as the species model had 91% accuracy on species from Canada (accuracy range 36%–91% across all out‐of‐sample datasets) and the empty‐animal model achieved an accuracy of 91%–94% on out‐of‐sample datasets from different continents. Our software addresses some of the limitations of using machine learning to classify images from camera traps. By including many species from several locations, our species model is potentially applicable to many camera trap studies in North America. We also found that our empty‐animal model can facilitate removal of images without animals globally. We provide the trained models in an R package (MLWIC2: Machine Learning for Wildlife Image Classification in R), which contains Shiny Applications that allow scientists with minimal programming experience to use trained models and train new models in six neural network architectures with varying depths.

Ecology and Evolution

Calibration of an evapotranspiration algorithm in a semiarid sagebrush steppe using a 3-ha lysimeter and Landsat normalized difference vegetation index data

In arid and semiarid environments, evapotranspiration (ET) is the primary discharge component in the water balance, with potential ET exceeding precipitation. For this reason, reliable estimates of ET are needed to construct accurate water budgets in these environments. Remote sensing affords the ability to provide fast, accurate, field-scale ET estimates, but these methods have largely been restricted to deep rooted (phreatophytic) plant communities underlain by shallow groundwater. We used 13 years of data from a 3-ha drainage lysimeter in a semiarid sagebrush steppe and Landsat normalized difference vegetation index (NDVI) data to calibrate a generalized least squares model capable of predicting vadose zone ET in a high elevation upland ecosystem. Annual precipitation was the best predictor of annual ET, as they were nearly balanced every year analysed (mean difference = 3 mm). We incorporated reference crop ET and a linear combination of NDVI and precipitation to capably predict ET on a subannual, lag-determined interval of 48 days, with a mean error of only 9.92% across all observations. To our knowledge, this is the first vegetation index-ET algorithm calibrated in a semiarid upland plant community using field-scale lysimetry. Vadose zone ET is particularly important at waste disposal sites in the Desert Southwest, where accurate and spatially explicit ET estimates are needed for monitoring potential mobilization and transport of contaminants past the root zone into local aquifers and for monitoring and modelling effects of recharge on flow and transport of contaminants in underlying aquifers.

Utah

cBathy: A robust algorithm for estimating nearshore bathymetry

A three-part algorithm is described and tested to provide robust bathymetry maps based solely on long time series observations of surface wave motions. The first phase consists of frequency-dependent characterization of the wave field in which dominant frequencies are estimated by Fourier transform while corresponding wave numbers are derived from spatial gradients in cross-spectral phase over analysis tiles that can be small, allowing high-spatial resolution. Coherent spatial structures at each frequency are extracted by frequency-dependent empirical orthogonal function (EOF). In phase two, depths are found that best fit weighted sets of frequency-wave number pairs. These are subsequently smoothed in time in phase 3 using a Kalman filter that fills gaps in coverage and objectively averages new estimates of variable quality with prior estimates. Objective confidence intervals are returned. Tests at Duck, NC, using 16 surveys collected over 2 years showed a bias and root-mean-square (RMS) error of 0.19 and 0.51 m, respectively but were largest near the offshore limits of analysis (roughly 500 m from the camera) and near the steep shoreline where analysis tiles mix information from waves, swash and static dry sand. Performance was excellent for small waves but degraded somewhat with increasing wave height. Sand bars and their small-scale alongshore variability were well resolved. A single ground truth survey from a dissipative, low-sloping beach (Agate Beach, OR) showed similar errors over a region that extended several kilometers from the camera and reached depths of 14 m. Vector wave number estimates can also be incorporated into data assimilation models of nearshore dynamics.

North Carolina