USGS ScienceSearch

SEARCH · USGS Science

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Automated detection of clipping in broadband earthquake records

Because the amount of available ground‐motion data has increased over the last decades, the need for automated processing algorithms has also increased. One difficulty with automated processing is to screen clipped records. Clipping occurs when the ground‐motion amplitude exceeds the dynamic range of the linear response of the instrument. Clipped records in which the amplitude exceeds the dynamic range are relatively easy to identify visually yet challenging for automated algorithms. In this article, we seek to identify a reliable and fully automated clipping detection algorithm tailored to near‐real‐time earthquake response needs. We consider multiple alternative algorithms, including (1) an algorithm based on the percentage difference in adjacent data points, (2) the standard deviation of the data within a moving window, (3) the shape of the histogram of the recorded amplitudes, (4) the second derivative of the data, and (5) the amplitude of the data. To quantitatively compare these algorithms, we construct development and holdout datasets from earthquakes across a range of geographic regions, tectonic environments, and instrument types. We manually classify each record for the presence of clipping and use the classified records. We then develop an artificial neural network model that combines all the individual algorithms. Testing on the holdout dataset, the standard deviation and histogram approaches are the most accurate individual algorithms, with an overall accuracy of about 93%. The combined artificial neural network method yields an overall accuracy of 95%, and the choice of classification threshold can balance precision and recall.

Seismological Research Letters

Estimation of potential evapotranspiration from extraterrestrial radiation, air temperature and humidity to assess future climate change effects on the vegetation of the Northern Great Plains, USA

The potential evapotranspiration (PET) that would occur with unlimited plant access to water is a central driver of simulated plant growth in many ecological models. PET is influenced by solar and longwave radiation, temperature, wind speed, and humidity, but it is often modeled as a function of temperature alone. This approach can cause biases in projections of future climate impacts in part because it confounds the effects of warming due to increased greenhouse gases with that which would be caused by increased radiation from the sun. We developed an algorithm for linking PET to extraterrestrial solar radiation (incoming top-of atmosphere solar radiation), as well as temperature and atmospheric water vapor pressure, and incorporated this algorithm into the dynamic global vegetation model MC1. We tested the new algorithm for the Northern Great Plains, USA, whose remaining grasslands are threatened by continuing woody encroachment. Both the new and the standard temperature-dependent MC1 algorithm adequately simulated current PET, as compared to the more rigorous PenPan model of Rotstayn et al. (2006) . However, compared to the standard algorithm, the new algorithm projected a much more gradual increase in PET over the 21st century for three contrasting future climates. This difference led to lower simulated drought effects and hence greater woody encroachment with the new algorithm, illustrating the importance of more rigorous calculations of PET in ecological models dealing with climate change.

Montana, Nebraska, North Dakota, South Dakota, Wyo

Mapping forest change using stacked generalization: An ensemble approach

The ever-increasing volume and accessibility of remote sensing data has spawned many alternative approaches for mapping important environmental features and processes. For example, there are several viable but highly varied strategies for using time series of Landsat imagery to detect changes in forest cover. Performance among algorithms varies across complex natural systems, and it is reasonable to ask if aggregating the strengths of an ensemble of classifiers might result in increased overall accuracy. Relatively simple rules have been used in the past to aggregate classifications among remotely sensed maps (e.g. using majority predictions), and in other fields, empirical models have been used to create situationally specific algorithm weights. The latter process, called “stacked generalization” (or “stacking”), typically uses a parametric model for the fusion of algorithm outputs. We tested the performance of several leading forest disturbance detection algorithms against ensembles of the outputs of those same algorithms based upon stacking using both parametric and Random Forests-based fusion rules. Stacking using a Random Forests model cut omission and commission error rates in half in many cases in relation to individual change detection algorithms, and cut error rates by one quarter compared to more conventional parametric stacking. Stacking also offers two auxiliary benefits: alignment of outputs to the precise definitions built into a particular set of empirical calibration data; and, outputs which may be adjusted such that map class totals match independent estimates of change in each year. In general, ensemble predictions improve when new inputs are added that are both informative and uncorrelated with existing ensemble components. As increased use of cloud-based computing makes ensemble mapping methods more accessible, the most useful new algorithms may be those that specialize in providing spectral, temporal, or thematic information not already available through members of existing ensembles.

Remote Sensing of Environment

Incorporation of real-time earthquake magnitudes estimated via peak ground displacement scaling in the ShakeAlert Earthquake Early Warning system

The United States earthquake early warning (EEW) system, ShakeAlert®, currently employs two algorithms based on seismic data alone to characterize the earthquake source, reporting the weighted average of their magnitude estimates. Nonsaturating magnitude estimates derived in real time from Global Navigation Satellite System (GNSS) data using peak ground displacement (PGD) scaling relationships offer complementary information with the potential to improve EEW reliability for large earthquakes. We have adapted a method that estimates magnitude from PGD ( Crowell et al. , 2016 ) for possible production use by ShakeAlert. To evaluate the potential contribution of the modified algorithm, we installed it on the ShakeAlert development system for real‐time operation and for retrospective analyses using a suite of GNSS data that we compiled. Because of the colored noise structure of typical real‐time GNSS positions, observed PGD values drift over time periods relevant to EEW. To mitigate this effect, we implemented logic within the modified algorithm to control when it issues initial and updated PGD‐derived magnitude estimates ( ⁠ M PGD "> M PGD "> M PGD PGD ⁠ ), and to quantify M PGD "> M PGD PGD uncertainty for use in combining it with estimates from other ShakeAlert algorithms running in parallel. Our analysis suggests that, with these strategies, spuriously large M PGD "> M PGD PGD will seldom be incorporated in ShakeAlert’s magnitude estimate. Retrospective analysis of data from moderate‐to‐great earthquakes demonstrates that the modified algorithm can contribute to better magnitude estimates for M w > 7.0 "> M w > 7.0 w>7.0 events. GNSS station distribution throughout the ShakeAlert region limits how soon the modified algorithm can begin estimating magnitude in some locations. Furthermore, both the station density and the GNSS noise levels limit the minimum magnitude for which the modified algorithm is likely to contribute to the weighted average. This might be addressed by alternative GNSS processing strategies that reduce noise.

Bulletin of the Seismological Society of America

Parameter estimation for the 4-parameter Asymmetric Exponential Power distribution by the method of L-moments using R

The implementation characteristics of two method of L-moments (MLM) algorithms for parameter estimation of the 4-parameter Asymmetric Exponential Power (AEP4) distribution are studied using the R environment for statistical computing. The objective is to validate the algorithms for general application of the AEP4 using R. An algorithm was introduced in the original study of the L-moments for the AEP4. A second or alternative algorithm is shown to have a larger L-moment-parameter domain than the original. The alternative algorithm is shown to provide reliable parameter production and recovery of L-moments from fitted parameters. A proposal is made for AEP4 implementation in conjunction with the 4-parameter Kappa distribution to create a mixed-distribution framework encompassing the joint L-skew and L-kurtosis domains. The example application provides a demonstration of pertinent algorithms with L-moment statistics and two 4-parameter distributions (AEP4 and the Generalized Lambda) for MLM fitting to a modestly asymmetric and heavy-tailed dataset using R.

Computational Statistics and Data Analysis

Self-potential tomography preconditioned by particle swarm optimization— Application to monitoring hyporheic exchange in a bedrock river

A self-potential (SP) data-inversion algorithm was developed and tested on an analytical model of electrical-potential profile data attributed to single and multiple polarized electrical sources. The developed algorithm was then validated by an application to SP-monitoring field data measured on the floodplain of East Fork Poplar Creek, Oak Ridge, Tennessee, to image electrical sources in areas conducive to preferential flow into the flood plain from the bedrock-lined riverbed. The algorithm combined stochastic source-localization by particle-swarm-optimization (PSO) of electrical sources characterized by simplified geometries with source tomography by regularized weighted least-squares minimization of a quadratic objective function. Prior information was incorporated by preconditioning the tomography algorithm by PSO results. Variable percentages of random noise were added to analytical-model data to evaluate the algorithm performance. Results indicated that true parameters of single-source models were inverted and approximated with small residual error, whereas inversion of analytical-model data representing multiple electrical sources accurately approximated the locations of the sources but miscalculated some parameters because of the non-uniqueness of the inverse-model solution. Source tomography applied to analytical model data during testing produced a spatially continuous parameter field that identified the locations of point-scale synthetic dipole sources of electrical current flow with varying degrees of accuracy depending on the prior information incorporated into the tomography. When applied to SP-monitoring field data, the algorithm imaged electrical sources within a known fault that intersects the bedrock riverbed and flood plain of East Fork Poplar Creek and depicted dynamic electrical conditions attributed to hyporheic exchange.

Tennessee

Review of revised Klamath River Total Maximum Daily Load models from Link River Dam to Keno Dam, Oregon

Flow and water-quality models are being used to support the development of Total Maximum Daily Load (TMDL) plans for the Klamath River downstream of Upper Klamath Lake (UKL) in south-central Oregon. For riverine reaches, the RMA-2 and RMA-11 models were used, whereas the CE-QUAL-W2 model was used to simulate pooled reaches. The U.S. Geological Survey (USGS) was asked to review the most upstream of these models, from Link River Dam at the outlet of UKL downstream through the first pooled reach of the Klamath River from Lake Ewauna to Keno Dam. Previous versions of these models were reviewed in 2009 by USGS. Since that time, important revisions were made to correct several problems and address other issues. This review documents an assessment of the revised models, with emphasis on the model revisions and any remaining issues. The primary focus of this review is the 19.7-mile Lake Ewauna to Keno Dam reach of the Klamath River that was simulated with the CE-QUAL-W2 model. Water spends far more time in the Lake Ewauna to Keno Dam reach than in the 1-mile Link River reach that connects UKL to the Klamath River, and most of the critical reactions affecting water quality upstream of Keno Dam occur in that pooled reach. This model review includes assessments of years 2000 and 2002 current conditions scenarios, which were used to calibrate the model, as well as a natural conditions scenario that was used as the reference condition for the TMDL and was based on the 2000 flow conditions. The natural conditions scenario included the removal of Keno Dam, restoration of the Keno reef (a shallow spot that was removed when the dam was built), removal of all point-source inputs, and derivation of upstream boundary water-quality inputs from a previously developed UKL TMDL model. This review examined the details of the models, including model algorithms, parameter values, and boundary conditions; the review did not assess the draft Klamath River TMDL or the TMDL allocations. Attention to the details of a model is one of the best ways to identify potential problems, correct them if possible, and begin to assess the magnitude of potential model errors and uncertainty. Model users need to determine the level of acceptable uncertainty associated with their objectives, identify all sources of potential uncertainty (model uncertainty, data uncertainty, etc.), and assess their approach and results accordingly. In the draft Klamath River TMDL, the Oregon Department of Environmental Quality identified the upstream boundary conditions as the largest source of uncertainty for both the current and natural conditions scenarios, not the model algorithms or choice of model parameters. We agree that the upstream boundary conditions are one of the largest, if not the largest, source of model uncertainty; therefore, the derivation of upstream boundary conditions may be more important to the TMDL than some other model-related issues identified in this review. The revised models contain a number of changes, some of which were done to solve small problems and are largely inconsequential to model results, but others of which are important and affect model predictions of instream concentrations. A consistent version of the model is now applied to all scenarios, and an error in the source code was corrected that had inadvertently discarded 20 percent of the incoming solar radiation in the original model. The baseline light-extinction coefficient for water was decreased and set to a consistent and defensible value across all models of reservoir reaches. Inconsistencies among the values of certain parameters in the original models, such as the ammonia nitrification rate and the decomposition rates of organic matter, have been eliminated, although the reasoning behind the final selections was not documented. The dependence of the rate of sediment oxygen demand (SOD) on temperature was modified such that the SOD rate was substantially decreased at temperatures less than 20°C, causing the model to predict higher dissolved oxygen (DO) concentrations in spring, autumn, and winter. Although that change to the temperature dependence function was done to make the function more similar to the model’s default, this change was not accompanied by any documentation of recalibration or sensitivity exercises. The maximum SOD rate for the 2002 current conditions scenario was decreased from 3.0 grams per square meter per day (g/m 2 /d) in the original model to 2.0 g/m 2 /d in the revised model, a considerable adjustment that appears to have been needed to offset effects of a change to another variable (O2LIM) that would have resulted in a substantial increase in the effective SOD rate for 2002. A 50-percent decrease in the SOD rate over a 2-year period, however, is not likely to be mirrored by field measurements, so this change may be compensating for some process that is not represented correctly in the DO budget for the current conditions scenarios. Several important changes were made to the natural conditions scenario. First, the elevation of the Keno reef was corrected; the elevation specified in the original model was 1 foot too high, which affected the volume of the pooled reach and the travel time through it. The most important changes to this scenario were to the upstream boundary inputs of organic matter and algae, which affect incoming fluxes of nitrogen and phosphorus. Algal biomass inputs were increased by approximately 60 percent during summer because of a change in the way those inputs were derived from results of the UKL TMDL model. Non-algal organic matter inputs were decreased, particularly in summer to correct a problem attributed to double-counting of phosphorus in the original inputs. The distribution of non-algal organic matter was changed from 20 percent dissolved in the original model to 90 percent dissolved in the revised model in response to review comments and published data. The overall sum of algal biomass and non-living organic matter was decreased, which resulted in lower inputs of total phosphorus and nitrogen. Total phosphorus inputs were less than 0.03 mg/L, and although the inputs were derived from selected results of the UKL TMDL model, these concentrations seem too low to be representative of a historically eutrophic system surrounded by extensive wetlands, peat soils, and a groundwater system high in phosphorus. The draft TMDL states that the upstream boundary conditions are the greatest source of uncertainty, greater than any uncertainty associated with the models. Efforts to improve existing models of algal growth and nutrient cycling in UKL, therefore, would provide a substantial benefit to downstream modeling efforts on the Klamath River. Although many improvements were made in revising the Klamath River TMDL models, some issues and uncertainties remain. Several errors in the model source code remain, but do not affect model results for this application as long as certain options and rates are not changed; future users of these models should be aware of these issues. Although the distribution of dissolved and particulate organic matter was modified for the natural conditions scenario, that distribution was not changed for the current conditions scenarios. Recent data on that distribution and the likely rates of organic matter decomposition could be used to improve these models in the future. Nitrate predictions at Keno (Highway 66) still are too high for the current conditions scenarios; future efforts should re-evaluate the model’s denitrification rates and the release rate of ammonia from anoxic sediments. Possibly the most important of the remaining issues are tied to the two-state (healthy/unhealthy) hypothesis for the algae population that was coded into the model. Some of the rates and conversion functions could be refined to make them more acceptable; currently, the published literature does not support the concept of moderately low dissolved-oxygen concentrations as a stressor of algae in the ranges used by the model. More research is needed before these algorithms can be truly tested. The algorithms currently appear to help the model fit the patterns in the available data, and that is useful and perhaps sufficient for some purposes, but those algorithms are not truly predictive or reliable for certain purposes until they can be tested through well-designed experiments and research. In summary, the TMDL models used to simulate Link and Klamath Rivers from Link River Dam to Keno Dam were revised to fix several problems and address various issues. The resulting models are an improvement over those that were reviewed by USGS in 2009, and represent a useful advance in the simulation of a complex system that is difficult to model. However, several issues remain that cause increased uncertainty in the model results. Depending on the objectives of the modeling, now or in the future, these remaining issues could be more or less important. For the Klamath River TMDL, the upstream boundary conditions may be a larger source of uncertainty than the concerns with model algorithms and model parameters identified in this review. Efforts to re-evaluate the available models of algal growth and nutrient cycling in UKL would be highly beneficial to downstream modeling efforts in the Klamath River.

Oregon;California

Determination of water depth with high-resolution satellite imagery over variable bottom types

A standard algorithm for determining depth in clear water from passive sensors exists; but it requires tuning of five parameters and does not retrieve depths where the bottom has an extremely low albedo. To address these issues, we developed an empirical solution using a ratio of reflectances that has only two tunable parameters and can be applied to low-albedo features. The two algorithms--the standard linear transform and the new ratio transform--were compared through analysis of IKONOS satellite imagery against lidar bathymetry. The coefficients for the ratio algorithm were tuned manually to a few depths from a nautical chart, yet performed as well as the linear algorithm tuned using multiple linear regression against the lidar. Both algorithms compensate for variable bottom type and albedo (sand, pavement, algae, coral) and retrieve bathymetry in water depths of less than 10-15 m. However, the linear transform does not distinguish depths >15 m and is more subject to variability across the studied atolls. The ratio transform can, in clear water, retrieve depths in >25 m of water and shows greater stability between different areas. It also performs slightly better in scattering turbidity than the linear transform. The ratio algorithm is somewhat noisier and cannot always adequately resolve fine morphology (structures smaller than 4-5 pixels) in water depths >15-20 m. In general, the ratio transform is more robust than the linear transform.

Limnology and Oceanography

Semiautomated tremor detection using a combined cross-correlation and neural network approach

Despite observations of tectonic tremor in many locations around the globe, the emergent phase arrivals, low‒amplitude waveforms, and variable event durations make automatic detection a nontrivial task. In this study, we employ a new method to identify tremor in large data sets using a semiautomated technique. The method first reduces the data volume with an envelope cross‒correlation technique, followed by a Self‒Organizing Map (SOM) algorithm to identify and classify event types. The method detects tremor in an automated fashion after calibrating for a specific data set, hence we refer to it as being “semiautomated”. We apply the semiautomated detection algorithm to a newly acquired data set of waveforms from a temporary deployment of 13 seismometers near Cholame, California, from May 2010 to July 2011. We manually identify tremor events in a 3 week long test data set and compare to the SOM output and find a detection accuracy of 79.5%. Detection accuracy improves with increasing signal‒to‒noise ratios and number of available stations. We find detection completeness of 96% for tremor events with signal‒to‒noise ratios above 3 and optimal results when data from at least 10 stations are available. We compare the SOM algorithm to the envelope correlation method of Wech and Creager and find the SOM performs significantly better, at least for the data set examined here. Using the SOM algorithm, we detect 2606 tremor events with a cumulative signal duration of nearly 55 h during the 13 month deployment. Overall, the SOM algorithm is shown to be a flexible new method that utilizes characteristics of the waveforms to identify tremor from noise or other seismic signals.

California

Comparison of machine learning approaches used to identify the drivers of Bakken oil well productivity

Geologists and petroleum engineers have struggled to identify the mechanisms that drive productivity in horizontal hydraulically fractured oil wells. The machine learning algorithms of Random Forest (RF), gradient boosting trees (GBT) and extreme gradient boosting (XGBoost) were applied to a dataset containing 7311 horizontal hydraulically fractured wells drilled into the middle member of the Bakken Formation from 2010 through 2017. The initial goal is to use these data‐driven machine learning algorithms to identify the most important explanatory predictors of well productivity within nine subareas and the composite area. Predictor variables representing initial gas production, the initial 180‐day water cut, and vertical depth vary spatially and are identified with geologically favorable areas. Well‐completion predictors include the well lateral length, number of fracture stages, volume of proppant per stage, and the volume of injected fluids per stage. The performance of methods is compared based on a common test sample. The analysis then examines the comparative predictive performance of the three algorithms for 1330 wells that had initiated production after the initial 7311 well sample had been producing. The computations of predictor importance identified the initial 180‐day water cut and the 30‐day initial gas production predictors as having a dominant influence in most subareas and for the composite area. The relative importance of well completion predictor variables, that is, the number of fracture stages per well, volume of injected proppant per stage, volume of injected fluids per stage, and lateral length, varied considerably across the subareas. For the common test or holdout sample, the models calibrated with the XGBoost algorithm had superior predictive power. The predictive power of all the algorithms trained on the data from the original sample suffered some loss when tested with a sample of wells that had started production after the end of that period. Implications of the empirical findings and strategies to mitigate loss of predictive power are discussed in the concluding section.

Statistical Analysis and Data Mining

A strategy for recovering continuous behavioral telemetry data from Pacific walruses

Tracking animal behavior and movement with telemetry sensors can offer substantial insights required for conservation. Yet, the value of data collected by animal-borne telemetry systems is limited by bandwidth constraints. To understand the response of Pacific walruses ( Odobenus rosmarus divergens ) to rapid changes in sea ice availability, we required continuous geospatial chronologies of foraging behavior. Satellite telemetry offered the only practical means to systematically collect such data; however, data transmission constraints of satellite data-collection systems limited the data volume that could be acquired. Although algorithms exist for reducing sensor data volumes for efficient transmission, none could meet our requirements. Consequently, we developed an algorithm for classifying hourly foraging behavior status aboard a tag with limited processing power. We found a 98% correspondence of our algorithm's classification with a test classification based on time–depth data recovered and characterized through multivariate analysis in a separate study. We then applied our algorithm within a telemetry system that relied on remotely deployed satellite tags. Data collected by these tags from Pacific walruses across their range during 2007–2015 demonstrated the consistency of foraging behavior collected by this strategy with data collected by data logging tags; and demonstrated the ability to collect geospatial behavioral chronologies with minimal missing data where recovery of data logging tags is precluded. Our strategy for developing a telemetry system may be applicable to any study requiring intelligent algorithms to continuously monitor behavior, and then compress those data into meaningful information that can be efficiently transmitted.

Wildlife Society Bulletin

Estimating crustal heterogeneity from double-difference tomography

Seismic velocity parameters in limited, but heterogeneous volumes can be inferred using a double-difference tomographic algorithm, but to obtain meaningful results accuracy must be maintained at every step of the computation. MONTEILLER et al. (2005) have devised a double-difference tomographic algorithm that takes full advantage of the accuracy of cross-spectral time-delays of large correlated event sets. This algorithm performs an accurate computation of theoretical travel-time delays in heterogeneous media and applies a suitable inversion scheme based on optimization theory. When applied to Kilauea Volcano, in Hawaii, the double-difference tomography approach shows significant and coherent changes to the velocity model in the well-resolved volumes beneath the Kilauea caldera and the upper east rift. In this paper, we first compare the results obtained using MONTEILLER et al.'s algorithm with those obtained using the classic travel-time tomographic approach. Then, we evaluated the effect of using data series of different accuracies, such as handpicked arrival-time differences ("picking differences"), on the results produced by double-difference tomographic algorithms. We show that picking differences have a non-Gaussian probability density function (pdf). Using a hyperbolic secant pdf instead of a Gaussian pdf allows improvement of the double-difference tomographic result when using picking difference data. We completed our study by investigating the use of spatially discontinuous time-delay data. ?? Birkha??user Verlag, Basel, 2006.

Pure and Applied Geophysics

Aspects of numerical and representational methods related to the finite-difference simulation of advective and dispersive transport of freshwater in a thin brackish aquifer

The simulation of the transport of injected freshwater in a thin brackish aquifer, overlain and underlain by confining layers containing more saline water, is shown to be influenced by the choice of the finite-difference approximation method, the algorithm for representing vertical advective and dispersive fluxes, and the values assigned to parametric coefficients that specify the degree of vertical dispersion and molecular diffusion that occurs. Computed potable water recovery efficiencies will differ depending upon the choice of algorithm and approximation method, as will dispersion coefficients estimated based on the calibration of simulations to match measured data. A comparison of centered and backward finite-difference approximation methods shows that substantially different transition zones between injected and native waters are depicted by the different methods, and computed recovery efficiencies vary greatly. Standard and experimental algorithms and a variety of values for molecular diffusivity, transverse dispersivity, and vertical scaling factor were compared in simulations of freshwater storage in a thin brackish aquifer. Computed recovery efficiencies vary considerably, and appreciable differences are observed in the distribution of injected freshwater in the various cases tested. The results demonstrate both a qualitatively different description of transport using the experimental algorithms and the interrelated influences of molecular diffusion and transverse dispersion on simulated recovery efficiency. When simulating natural aquifer flow in cross-section, flushing of the aquifer occurred for all tested coefficient choices using both standard and experimental algorithms.

Journal of Hydrology

Performance metrics and variance partitioning reveal sources of uncertainty in species distribution models

Species distribution models (SDMs) are widely used in basic and applied ecology, making it important to understand sources and magnitudes of uncertainty in SDM performance and predictions. We analyzed SDM performance and partitioned variance among prediction maps for 15 rare vertebrate species in the southeastern USA using all possible combinations of seven potential sources of uncertainty in SDMs: algorithms, climate datasets, model domain, species presences, variable collinearity, CO 2 emissions scenarios, and general circulation models. The choice of modeling algorithm was the greatest source of uncertainty in SDM performance and prediction maps, with some additional variation in performance associated with the comprehensiveness of the species presences used for modeling. Other sources of uncertainty that have received attention in the SDM literature such as variable collinearity and model domain contributed little to differences in SDM performance or predictions in this study. Predictions from different algorithms tended to be more variable at northern range margins for species with more northern distributions, which may complicate conservation planning at the leading edge of species' geographic ranges. The clear message emerging from this work is that researchers should use multiple algorithms for modeling rather than relying on predictions from a single algorithm, invest resources in compiling a comprehensive set of species presences, and explicitly evaluate uncertainty in SDM predictions at leading range margins.

Ecological Modelling

Satellites for long-term monitoring of inland U.S. lakes: The MERIS time series and application for chlorophyll-a

Lakes and other surface fresh waterbodies provide drinking water, recreational and economic opportunities, food, and other critical support for humans, aquatic life, and ecosystem health. Lakes are also productive ecosystems that provide habitats and influence global cycles. Chlorophyll concentration provides a common metric of water quality, and is frequently used as a proxy for lake trophic state. Here, we document the generation and distribution of the complete MEdium Resolution Imaging Spectrometer (MERIS; Appendix A provides a complete list of abbreviations) radiometric time series for over 2300 satellite resolvable inland bodies of water across the contiguous United States (CONUS) and more than 5,000 in Alaska. This contribution greatly increases the ease of use of satellite remote sensing data for inland water quality monitoring, as well as highlights new horizons in inland water remote sensing algorithm development. We evaluate the performance of satellite remote sensing Cyanobacteria Index (CI)-based chlorophyll algorithms, the retrievals for which provide surrogate estimates of phytoplankton concentrations in cyanobacteria dominated lakes. Our analysis quantifies the algorithms' abilities to assess lake trophic state across the CONUS. As a case study, we apply a bootstrapping approach to derive a new CI-to-chlorophyll relationship, ChlBS, which performs relatively well with a multiplicative bias of 1.11 (11%) and mean absolute error of 1.60 (60%). While the primary contribution of this work is the distribution of the MERIS radiometric timeseries, we provide this case study as a roadmap for future stakeholders' algorithm development activities, as well as a tool to assess the strengths and weaknesses of applying a single algorithm across CONUS.

Alaska, Minnesota

How often can Earthquake Early Warning systems alert sites with high intensity ground motion?

Although numerous Earthquake Early Warning (EEW) algorithms have been developed we still lack a detailed understanding of how often and under what circumstances useful ground motion alerts can be provided to end-users. Here we analyze the alerting performance of the PLUM, EPIC and FinDer algorithms by running them retrospectively on the seismic strong motion data of the 219 earthquakes in Japan since 1996 that exceeded Modified Mercalli Intensity (MMI) of 4.5 on at least 10 sites (Mw 4.5-9.1). Our analysis suggests that, irrespective of the algorithm, EEW end-users should be prepared that EEW can often but not always provide useful ground motion alerts. A majority of sites with moderate-strong ground motion (MMI 5-6) can generally get at least a few seconds of warning time from all algorithms. If such shaking is caused by a shallow crustal event, around 50% of such sites receive alerts with warning times >5 s. Many sites with severe-extreme ground motion (MMI >=8) can be alerted successfully in the case of very large offshore earthquakes, but less than 20% can be alerted ahead of time if such shaking is caused by a shallow crustal event. Our results provide detailed quantitative insight into the expected alerting performance for EEW algorithms under realistic conditions. The main caveat is that the largest shallow crustal event in our data set has Mw7.0, i.e. the data set does not contain very large strike slip events.

Journal of Geophysical Research

A method to obtain remotely sensed grain size distributions from nonplanar granular deposits

Constraining the grain size distribution of granular deposits with complex surfaces is difficult with existing approaches. Field and laboratory techniques are time consuming and limited by the maximum grain size that laboratories can accommodate. In this study, we present a new method to identify the coarse fraction of the grain size distribution at a debris-flow fan deposit surveyed with terrestrial laser scanning (TLS) in Glenwood Canyon, Colorado, USA. This method is a novel grain segmentation algorithm developed for application to point cloud data of deposits with complex surfaces and angular grains ranging in size from centimeters to a meter. This approach combines an existing random forest machine learning method with a novel iterative clustering algorithm. We compared the grain size distribution from our algorithm with a Wolman pebble count conducted in the field, and found a root mean squared error of less than 2 cm from the 5th to 95th percentile of the grain size distribution of grains ranging from cobble to boulder sized (6.3–78 cm in our application). Finally, we compared our new algorithm with an existing open-source grain segregation algorithm, and our method outperformed the selected alternative when applied to the debris-flow deposit point cloud.

Colorado

Automated masking of cloud and cloud shadow for forest change analysis using Landsat images

Accurate masking of cloud and cloud shadow is a prerequisite for reliable mapping of land surface attributes. Cloud contamination is particularly a problem for land cover change analysis, because unflagged clouds may be mapped as false changes, and the level of such false changes can be comparable to or many times more than that of actual changes, even for images with small percentages of cloud cover. Here we develop an algorithm for automatically flagging clouds and their shadows in Landsat images. This algorithm uses clear view forest pixels as a reference to define cloud boundaries for separating cloud from clear view surfaces in a spectral-temperature space. Shadow locations are predicted according to cloud height estimates and sun illumination geometry, and actual shadow pixels are identified by searching the darkest pixels surrounding the predicted shadow locations. This algorithm produced omission errors of around 1% for the cloud class, although the errors were higher for an image that had very low cloud cover and one acquired in a semiarid environment. While higher values were reported for other error measures, most of the errors were found around the edges of detected clouds and shadows, and many were due to difficulties in flagging thin clouds and the shadow cast by them, both by the developed algorithm and by the image analyst in deriving the reference data. We concluded that this algorithm is especially suitable for forest change analysis, because the commission and omission errors of the derived masks are not likely to significantly bias change analysis results.

International Journal of Remote Sensing