USGS ScienceSearch

SEARCH · USGS Science

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Characterizing riverbed sediment using high-frequency acoustics 2: scattering signatures of Colorado River bed sediment in Marble and Grand Canyons

In this, the second of a pair of papers on the statistical signatures of riverbed sediment in high-frequency acoustic backscatter, spatially explicit maps of the stochastic geometries (length- and amplitude-scales) of backscatter are related to patches of riverbed surfaces composed of known sediment types, as determined by geo-referenced underwater video observations. Statistics of backscatter magnitudes alone are found to be poor discriminators between sediment types. However, the variance of the power spectrum, and the intercept and slope from a power-law spectral form (termed the spectral strength and exponent, respectively) successfully discriminate between sediment types. A decision-tree approach was able to classify spatially heterogeneous patches of homogeneous sands, gravels (and sand-gravel mixtures), and cobbles/boulders with 95, 88, and 91% accuracy, respectively. Application to sites outside the calibration, and surveys made at calibration sites at different times, were plausible based on observations from underwater video. Analysis of decision trees built with different training data sets suggested that the spectral exponent was consistently the most important variable in the classification. In the absence of theory concerning how spatially variable sediment surfaces scatter high-frequency sound, the primary advantage of this data-driven approach to classify bed sediment over alternatives is that spectral methods have well understood properties and make no assumptions about the distributional form of the fluctuating component of backscatter over small spatial scales.

Arizona

Implementation of the Next Generation Attenuation (NGA) ground-motion prediction equations in Fortran and R

This report presents two methods for implementing the earthquake ground-motion prediction equations released in 2008 as part of the Next Generation Attenuation of Ground Motions (NGA-West, or NGA) project coordinated by the Pacific Earthquake Engineering Research Center (PEER). These models were developed for predicting ground-motion parameters for shallow crustal earthquakes in active tectonic regions (such as California). Of the five ground-motion prediction equations (GMPEs) developed during the NGA project, four models are implemented: the GMPEs of Abrahamson and Silva (2008), Boore and Atkinson (2008), Campbell and Bozorgnia (2008), and Chiou and Youngs (2008a); these models are abbreviated as AS08, BA08, CB08, and CY08, respectively. Since site response is widely recognized as an important influence of ground motions, engineering applications typically require that such effects be modeled. The model of Idriss (2008) is not implemented in our programs because it does not explicitly include site response, whereas the other four models include site response and use the same variable to describe the site condition (VS30). We do not intend to discourage the use of the Idriss (2008) model, but we have chosen to implement the other four NGA models in our programs for those users who require ground-motion estimates for various site conditions. We have implemented the NGA models by using two separate programming languages: Fortran and R (R Development Core Team, 2010). Fortran, a compiled programming language, has been used in the scientific community for decades. R is an object-oriented language and environment for statistical computing that is gaining popularity in the statistical and scientific community. Derived from the S language and environment developed at Bell Laboratories, R is an open-source language that is freely available at http://www.r-project.org/ (last accessed 11 January 2011). In R, the functions for computing the NGA equations can be loaded as an add-on user-contributed code, which is referred to as a ?package? in R. The details of the nga package (Kaklamanos and Thompson, 2010) are presented in this report. In addition, differences between the R and Fortran implementations will be discussed later in this report. The NGA models have established a new baseline for seismic hazard assessments, and they have been incorporated into the most recent national seismic hazard maps published by the U.S. Geological Survey (Petersen and others, 2008). However, many of the new models are significantly more complicated than previous GMPEs and, therefore, require a substantial investment of time to implement and validate. We hope that the scientific and engineering communities find our implementations to be useful in research and practice. Our implementations may be considered as an alternate to the Microsoft Excel spreadsheet implementation available on the PEER NGA project Web site (http://peer.berkeley.edu/ngawest/, last accessed 11 January 2011). The implementations in Fortran and R are more appropriate for performing batch calculations than the implementation in Microsoft Excel. Spreadsheets and Fortran code for some of the individual models also are available on the PEER NGA project Web site; our programs implement the four GMPEs simultaneously. Our programs give the same results as the programs on the PEER NGA Web site, but we offer some additional flexibility of input, additional methods of estimating unknown input parameters, and additional options for output. Although these programs have been used by the U.S. Geological Survey (USGS), Tufts University, and others, no warranty, expressed or implied, is made by Tufts or the USGS as to the accuracy or functioning of the programs and related material, nor shall the fact of distribution constitute any such warranty, and no responsibility is assumed by Tufts or the USGS in connection therewith.

Open-File Report

Landslide initiation thresholds in data-sparse regions: Application to landslide early warning criteria in Sitka, Alaska, USA

Probabilistic models to inform landslide early warning systems often rely on rainfall totals observed during past events with landslides. However, these models are generally developed for broad regions using large catalogs, with dozens, hundreds, or even thousands of landslide occurrences. This study evaluates strategies for training landslide forecasting models with a scanty record of landslide-triggering events, which is a typical limitation in remote, sparsely populated regions. We evaluate 136 statistical models trained on a precipitation dataset with five landslide-triggering precipitation events recorded near Sitka, Alaska, USA, as well as > 6000 d of non-triggering rainfall (2002–2020). We also conduct extensive statistical evaluation for three primary purposes: (1) to select the best-fitting models, (2) to evaluate performance of the preferred models, and (3) to select and evaluate warning thresholds. We use Akaike, Bayesian, and leave-one-out information criteria to compare the 136 models, which are trained on different cumulative precipitation variables at time intervals ranging from 1 h to 2 weeks, using both frequentist and Bayesian methods to estimate the daily probability and intensity of potential landslide occurrence (logistic regression and Poisson regression). We evaluate the best-fit models using leave-one-out validation as well as by testing a subset of the data. Despite this sparse landslide inventory, we find that probabilistic models can effectively distinguish days with landslides from days without slide activity. Our statistical analyses show that 3 h precipitation totals are the best predictor of elevated landslide hazard, and adding antecedent precipitation (days to weeks) did not improve model performance. This relatively short timescale of precipitation combined with the limited role of antecedent conditions likely reflects the rapid draining of porous colluvial soils on the very steep hillslopes around Sitka. Although frequentist and Bayesian inferences produce similar estimates of landslide hazard, they do have different implications for use and interpretation: frequentist models are familiar and easy to implement, but Bayesian models capture the rare-events problem more explicitly and allow for better understanding of parameter uncertainty given the available data. We use the resulting estimates of daily landslide probability to establish two decision boundaries that define three levels of warning. With these decision boundaries, the frequentist logistic regression model incorporates National Weather Service quantitative precipitation forecasts into a real-time landslide early warning “dashboard” system ( https://sitkalandslide.org/ , last access: 9 October 2023). This dashboard provides accessible and data-driven situational awareness for community members and emergency managers.

Alaska

StreamStats: a U.S. geological survey web site for stream information

The U.S. Geological Survey has developed a Web application, named StreamStats, for providing streamflow statistics, such as the 100-year flood and the 7-day, 10-year low flow, to the public. Statistics can be obtained for data-collection stations and for ungaged sites. Streamflow statistics are needed for water-resources planning and management; for design of bridges, culverts, and flood-control structures; and for many other purposes. StreamStats users can point and click on data-collection stations shown on a map in their Web browser window to obtain previously determined streamflow statistics and other information for the stations. Users also can point and click on any stream shown on the map to get estimates of streamflow statistics for ungaged sites. StreamStats determines the watershed boundaries and measures physical and climatic characteristics of the watersheds for the ungaged sites by use of a Geographic Information System (GIS), and then it inserts the characteristics into previously determined regression equations to estimate the streamflow statistics. Compared to manual methods, StreamStats reduces the average time needed to estimate streamflow statistics for ungaged sites from several hours to several minutes.

Conference Paper

Arkansas StreamStats: A U.S. Geological Survey web map application for basin characteristics and streamflow statistics

The U.S. Geological Survey (USGS) provides streamflow and other related information needed by water-resource managers responsible for protecting people and property from floods, planning and managing water-resource activities, and protecting water quality. Streamflow statistics provided by the USGS, such as the 1-percent annual exceedance probability (100-year flood) and the 7-day 10-year low flow, are frequently used by engineers, flood forecasters, land managers, biologists, and others to guide their everyday decisions. Additionally, resource managers often need to know basin characteristics, the physical and climatic characteristics of a drainage basin, to help understand the mechanisms that control water availability, water quality, and aquatic habitats at various locations. Users of streamflow information often require streamflow statistics and basin characteristics at various locations along a stream. The USGS periodically calculates and publishes streamflow statistics and basin characteristics for streamflowgaging stations and partial-record stations, but these data commonly are scattered among many reports that may or may not be readily available to the public. The USGS also provides and periodically updates regional analyses of streamflow statistics that include regression equations and other prediction methods for estimating statistics for ungaged and unregulated streams across the State. Use of these regional predictions for a stream can be complex and often requires the user to determine a number of basin characteristics that may require interpretation. Basin characteristics may include drainage area, classifiers for physical properties, climatic characteristics, and other inputs. Obtaining these input values for gaged and ungaged locations traditionally has been time consuming, subjective, and can lead to inconsistent results.

Arkansas

Flood-frequency estimates for Ohio streamgages based on data through water year 2015 and techniques for estimating flood-frequency characteristics of rural, unregulated Ohio streams

Estimates of the magnitudes of annual peak streamflows with annual exceedance probabilities of 0.5, 0.2, 0.1, 0.04, 0.02, 0.01, and 0.002 (equivalent to recurrence intervals of 2-, 5-, 10-, 25-, 50-, 100-, and 500-years, respectively) were computed for 391 streamgages in Ohio and adjacent states based on data collected through the 2015 water year. The flood-frequency estimates were computed following guidance outlined in Bulletin 17C, developed by the Advisory Committee on Water Information. The Bulletin 17C guidelines retain the basic statistical framework of the superseded Bulletin 17B guidelines; however, the Bulletin 17C guidelines add several enhancements including an improved method of moments approach for fitting the log-Pearson Type III (LPIII) distribution to the flood peaks (called the expected moments algorithm), a generalization of the Grubbs Beck low-outlier test (called the Multiple Grubbs Beck test) that permits identification of multiple potentially influential low floods, and new methods for estimating regional skew and uncertainty. Equations for estimating flood-frequency characteristics at ungaged sites on rural, unregulated streams in Ohio were developed with a two-step process involving ordinary least-squares and generalized least-squares regression techniques. Data from 333 streamgages with 10 or more years of unregulated record were screened for redundancy and a regression dataset was selected that was composed of flood-frequency and basin-characteristic data for 275 streamgages in Ohio and adjacent states. Two sets of equations were developed—one set, referred to as the “simple model,” uses regression region and drainage area as regressor variables, and a second set, referred to as the “full model,” uses regression region, drainage area, main-channel slope, and the percentage of the watershed covered by water and wetlands as regressor variables. The average standard errors of prediction ranged from about 40.5 to 46.5 percent for the simple-model equations and from about 37.2 to 40.3 percent for the full-model equations. For sites meeting the rural, unregulated criteria, flood-frequency estimates determined by means of LPIII analyses are reported along with weighted flood-frequency estimates, computed as a function of the LPIII estimates and the regression estimates. For sites with homogenous periods of regulation, flood-frequency estimates determined by means of LPIII analyses are reported. Ninety-five percent confidence limits are reported for all estimates. Values of regressor variables were determined from digital spatial datasets by means of a geographic information system (GIS). The GIS datasets and the new full-model equations have been incorporated into Ohio’s StreamStats application, a web-based, GIS-backed system designed to facilitate the estimation of streamflow statistics at ungaged locations on streams. Seasonal patterns in peak flows were assessed for 295 streamgages in Ohio. Annual peak flows occurred most frequently between January and April, with March having the highest frequency of occurrence. The month with the fewest number of annual peaks was October. Peak-of-record flows occurred most frequently in March, followed by January (months in which two of Ohio’s most severe widespread floods in recent history occurred). None of the peak-of-record flows occurred in October and only two occurred in November. Temporal trend in annual peak flows were assessed for 133 streamgages on unregulated streams in Ohio with 30 or more years of systematic record. Trends were assessed by computing the rank correlation (as measured with the two-sided Kendall’s tau statistic) between time and annual peak flows. Weak but statistically significant trends were indicated at 15 of the 133 streamgages. Of the 15 streamgages with significant trend in annual peak flows, 12 had an upward trend (positive tau) and 3 had a downward trend (negative tau). All 12 streamgages with positive tau values were at latitudes north of 40°33', and streamgages with negative tau values were at latitudes south of 40°33'.

Ohio

A dynamic spatio-temporal model for spatial data

Analyzing spatial data often requires modeling dependencies created by a dynamic spatio-temporal data generating process. In many applications, a generalized linear mixed model (GLMM) is used with a random effect to account for spatial dependence and to provide optimal spatial predictions. Location-specific covariates are often included as fixed effects in a GLMM and may be collinear with the spatial random effect, which can negatively affect inference. We propose a dynamic approach to account for spatial dependence that incorporates scientific knowledge of the spatio-temporal data generating process. Our approach relies on a dynamic spatio-temporal model that explicitly incorporates location-specific covariates. We illustrate our approach with a spatially varying ecological diffusion model implemented using a computationally efficient homogenization technique. We apply our model to understand individual-level and location-specific risk factors associated with chronic wasting disease in white-tailed deer from Wisconsin, USA and estimate the location the disease was first introduced. We compare our approach to several existing methods that are commonly used in spatial statistics. Our spatio-temporal approach resulted in a higher predictive accuracy when compared to methods based on optimal spatial prediction, obviated confounding among the spatially indexed covariates and the spatial random effect, and provided additional information that will be important for containing disease outbreaks.

Wisconsin

Methodology for quantifying uncertainty in coal assessments with an application to a Texas lignite deposit

A common practice for characterizing uncertainty in coal resource assessments has been the itemization of tonnage at the mining unit level and the classification of such units according to distance to drilling holes. Distance criteria, such as those used in U.S. Geological Survey Circular 891, are still widely used for public disclosure. A major deficiency of distance methods is that they do not provide a quantitative measure of uncertainty. Additionally, relying on distance between data points alone does not take into consideration other factors known to have an influence on uncertainty, such as spatial correlation, type of probability distribution followed by the data, geological discontinuities, and boundary of the deposit. Several geostatistical methods have been combined to formulate a quantitative characterization for appraising uncertainty. Drill hole datasets ranging from widespread exploration drilling to detailed development drilling from a lignite deposit in Texas were used to illustrate the modeling. The results show that distance to the nearest drill hole is almost completely unrelated to uncertainty, which confirms the inadequacy of characterizing uncertainty based solely on a simple classification of resources by distance classes. The more complex statistical methods used in this study quantify uncertainty and show good agreement between confidence intervals in the uncertainty predictions and data from additional drilling.

International Journal of Coal Geology

Assessing the effectiveness of riparian restoration projects using Landsat and precipitation data from the cloud-computing application ClimateEngine.org

Riparian vegetation along streams provides a suite of ecosystem services in rangelands and thus is the target of restoration when degraded by over-grazing, erosion, incision, or other disturbances. Assessments of restoration effectiveness depend on defensible monitoring data, which can be both expensive and difficult to collect. We present a method and case study to evaluate the effectiveness of restoration of riparian vegetation using a web-based cloud-computing and visualization tool (ClimateEngine.org) to access and process remote sensing and climate data. Restoration efforts on an Eastern Oregon ranch were assessed by analyzing the riparian areas of four creeks that had in-stream restoration structures constructed between 2008 and 2011. Within each study area, we retrieved spatially and temporally aggregated values of summer (June, July, August) normalized difference vegetation index (NDVI) and total precipitation for each water year (October-September) from 1984 to 2017. We established a pre-restoration (1984–2007) linear regression between total water year precipitation and summer NDVI for each study area, and then compared the post-restoration (2012–2017) data to this pre-restoration relationship. In each study area, the post-restoration NDVI-precipitation relationship was statistically distinct from the pre-restoration relationship, suggesting a change in the fundamental relationship between precipitation and NDVI resulting from stream restoration. We infer that the in-stream structures, which raised the water table in the adjacent riparian areas, provided additional water to the streamside vegetation that was not available before restoration and reduced the dependence of riparian vegetation on precipitation. This approach provides a cost-effective, quantitative method for assessing the effects of stream restoration projects on riparian vegetation.

Oregon

Statistical analysis and evaluation of water-quality data for selected streams in the coal area of east-central Montana

To document and evaluate existing conditions of water quality prior to proposed coal development in east-central Montana, water-quality data were collected at 23 sites on selected streams from October 1975 through September 1981. The data were statistically summarized and regression equations were developed to define relationships between water-quality variables. Where applicable, measured water-quality conditions were compared to various water-use standards. Measured concentrations of dissolved solids ranged from 145 to 12,200 milligrams per liter. Concentrations commonly exceeded 1,000 milligrams per liter and thereby present a high to very high salinity hazard for irrigation. Streamflow of the area contains predominantly sodium and sulfate ions and generally constitutes a medium to very high sodium hazard for irrigation during base flow. The water in most streams is generally adequate for livestock consumption during base flow. Concentrations of suspended sediment were extremely variable and had a direct correlation to water discharge. Measured suspended-sediment concentrations ranged from 4 to 23 ,000 milligrams per liter. Sediment-transport curves were developed for 18 of the study sites. Mean annual suspended-sediment loads were determined at five sites using the flow-duration, sediment-transport curve method. Mean annual sediment loads ranged from 1,010 to 72,7000 tons. (USGS)

Water-Resources Investigations Report

Spatial heterogeneity in statistical power to detect changes in lake area in Alaskan National Wildlife Refuges

Over the past 50 years, the number and size of high-latitude lakes have decreased throughout many regions; however, individual lake trends have been variable in direction and magnitude. This spatial heterogeneity in lake change makes statistical detection of temporal trends challenging, particularly in small analysis areas where weak trends are difficult to separate from inter- and intra-annual variability. Factors affecting trend detection include inherent variability, trend magnitude, and sample size. In this paper, we investigated how the statistical power to detect average linear trends in lake size of 0.5, 1.0 and 2.0 %/year was affected by the size of the analysis area and the number of years of monitoring in National Wildlife Refuges in Alaska. We estimated power for large (930–4,560 sq km) study areas within refuges and for 2.6, 12.9, and 25.9 sq km cells nested within study areas over temporal extents of 4–50 years. We found that: (1) trends in study areas could be detected within 5–15 years, (2) trends smaller than 2.0 %/year would take >50 years to detect in cells within study areas, and (3) there was substantial spatial variation in the time required to detect change among cells. Power was particularly low in the smallest cells which typically had the fewest lakes. Because small but ecologically meaningful trends may take decades to detect, early establishment of long-term monitoring will enhance power to detect change. Our results have broad applicability and our method is useful for any study involving change detection among variable spatial and temporal extents.

Alaska

A model of strength

In her AAAS News & Notes piece "Can the Southwest manage its thirst?" (26 July, p. 362), K. Wren quotes Ajay Kalra, who advocates a particular method for predicting Colorado River streamflow "because it eschews complex physical climate models for a statistical data-driven modeling approach." A preference for data-driven models may be appropriate in this individual situation, but it is not so generally, Data-driven models often come with a warning against extrapolating beyond the range of the data used to develop the models. When the future is like the past, data-driven models can work well for prediction, but it is easy to over-model local or transient phenomena, often leading to predictive inaccuracy (1). Mechanistic models are built on established knowledge of the process that connects the response variables with the predictors, using information obtained outside of an extant data set. One may shy away from a mechanistic approach when the underlying process is judged to be too complicated, but good predictive models can be constructed with statistical components that account for ingredients missing in the mechanistic analysis. Models with sound mechanistic components are more generally applicable and robust than data-driven models.

Science

Combining particle-tracking and geochemical data to assess public supply well vulnerability to arsenic and uranium

Flow-model particle-tracking results and geochemical data from seven study areas across the United States were analyzed using three statistical methods to test the hypothesis that these variables can successfully be used to assess public supply well vulnerability to arsenic and uranium. Principal components analysis indicated that arsenic and uranium concentrations were associated with particle-tracking variables that simulate time of travel and water fluxes through aquifer systems and also through specific redox and pH zones within aquifers. Time-of-travel variables are important because many geochemical reactions are kinetically limited, and geochemical zonation can account for different modes of mobilization and fate. Spearman correlation analysis established statistical significance for correlations of arsenic and uranium concentrations with variables derived using the particle-tracking routines. Correlations between uranium concentrations and particle-tracking variables were generally strongest for variables computed for distinct redox zones. Classification tree analysis on arsenic concentrations yielded a quantitative categorical model using time-of-travel variables and solid-phase-arsenic concentrations. The classification tree model accuracy on the learning data subset was 70%, and on the testing data subset, 79%, demonstrating one application in which particle-tracking variables can be used predictively in a quantitative screening-level assessment of public supply well vulnerability. Ground-water management actions that are based on avoidance of young ground water, reflecting the premise that young ground water is more vulnerable to anthropogenic contaminants than is old ground water, may inadvertently lead to increased vulnerability to natural contaminants due to the tendency for concentrations of many natural contaminants to increase with increasing ground-water residence time.

Journal of Hydrology

Candidate soil indicators for monitoring the progress of constructed wetlands toward a natural state: a statistical approach

A persistent question among ecologists and environmental managers is whether constructed wetlands are structurally or functionally equivalent to naturally occurring wetlands. We examined 19 variables collected from 10 constructed and nine natural emergent wetlands in Ohio, USA. Our primary objective was to identify candidate indicators of wetland class (natural or constructed), based on measurements of soil properties and an index of vegetation integrity, that can be used to track the progress of constructed wetlands toward a natural state. The method of nearest shrunken centroids was used to find a subset of variables that would serve as the best classifiers of wetland class, and error rate was calculated using a five-fold cross-validation procedure. The shrunken differences of percent total organic carbon (% TOC) and percent dry weight of the soil exhibited the greatest distances from the overall centroid. Classification based on these two variables yielded a misclassification rate of 11% based on cross-validation. Our results indicate that % TOC and percent dry weight can be used as candidate indicators of the status of emergent, constructed wetlands in Ohio and for assessing the performance of mitigation. The method of nearest shrunken centroids has excellent potential for further applications in ecology.

Ohio

Standardized data quality acceptance criteria for a rapid Escherichia coli qPCR method (Draft Method C) for water quality monitoring at recreational beaches

There is growing interest in the application of rapid quantitative polymerase chain reaction (qPCR) and other PCR-based methods for recreational water quality monitoring and management programs. This interest has strengthened given the publication of U.S. Environmental Protection Agency (EPA)-validated qPCR methods for enterococci fecal indicator bacteria (FIB) and has extended to similar methods for Escherichia coli ( E. coli ) FIB. Implementation of qPCR-based methods in monitoring programs can be facilitated by confidence in the quality of the data produced by these methods. Data quality can be determined through the establishment of a series of specifications that should reflect good laboratory practice. Ideally, these specifications will also account for the typical variability of data coming from multiple users of the method. This study developed proposed standardized data quality acceptance criteria that were established for important calibration model parameters and/or controls from a new qPCR method for E. coli (EPA Draft Method C) based upon data that was generated by 21 laboratories. Each laboratory followed a standardized protocol utilizing the same prescribed reagents and reference and control materials. After removal of outliers, statistical modeling based on a hierarchical Bayesian method was used to establish metrics for assay standard curve slope, intercept and lower limit of quantification that included between-laboratory, replicate testing within laboratory, and random error variability. A nested analysis of variance (ANOVA) was used to establish metrics for calibrator/positive control, negative control, and replicate sample analysis data. These data acceptance criteria should help those who may evaluate the technical quality of future findings from the method, as well as those who might use the method in the future. Furthermore, these benchmarks and the approaches described for determining them may be helpful to method users seeking to establish comparable laboratory-specific criteria if changes in the reference and/or control materials must be made.

Water Research

Expanded target-chemical analysis reveals extensive mixed-organic-contaminant exposure in USA streams

Surface water from 38 streams nationwide was assessed using 14 target-organic methods (719 compounds). Designed-bioactive anthropogenic contaminants (biocides, pharmaceuticals) comprised 57% of 406 organics detected at least once. The 10 most-frequently detected anthropogenic-organics included eight pesticides (desulfinylfipronil, AMPA, chlorpyrifos, dieldrin, metolachlor, atrazine, CIAT, glyphosate) and two pharmaceuticals (caffeine, metformin) with detection frequencies ranging 66–84% of all sites. Detected contaminant concentrations varied from less than 1 ng L –1 to greater than 10 μg L –1 , with 77 and 278 having median detected concentrations greater than 100 ng L –1 and 10 ng L –1 , respectively. Cumulative detections and concentrations ranged 4–161 compounds (median 70) and 8.5–102 847 ng L –1 , respectively, and correlated significantly with wastewater discharge, watershed development, and toxic release inventory metrics. Log 10 concentrations of widely monitored HHCB, triclosan, and carbamazepine explained 71–82% of the variability in the total number of compounds detected (linear regression; p -values: < 0.001–0.012), providing a statistical inference tool for unmonitored contaminants. Due to multiple modes of action, high bioactivity, biorecalcitrance, and direct environment application (pesticides), designed-bioactive organics (median 41 per site at μg L –1 cumulative concentrations) in developed watersheds present aquatic health concerns, given their acknowledged potential for sublethal effects to sensitive species and lifecycle stages at low ng L –1 .

Environmental Science & Technology

Treed Gaussian processes for animal movement modeling

Wildlife telemetry data may be used to answer a diverse range of questions relevant to wildlife ecology and management. One challenge to modeling telemetry data is that animal movement often varies greatly in pattern over time, and current continuous-time modeling approaches to handle such nonstationarity require bespoke and often complex models that may pose barriers to practitioner implementation. We demonstrate a novel application of treed Gaussian process (TGP) modeling, a Bayesian machine learning approach that automatically captures the nonstationarity and abrupt transitions present in animal movement. The machine learning formulation of TGPs enables modeling to be nearly automated, while their Bayesian formulation allows for the derivation of movement descriptors with associated uncertainty measures. We demonstrate the use of an existing R package to implement TGPs using the familiar Markov chain Monte Carlo algorithm. We then use estimated movement trajectories to derive movement descriptors that can be compared across individuals and populations. We applied the TGP model to a case study of lesser prairie-chickens ( Tympanuchus pallidicinctus ) to demonstrate the benefits of TGP modeling and compared distance traveled and residence times across lesser prairie-chicken individuals and populations. For broad usability, we outline all steps necessary for practitioners to specify relevant movement descriptors (e.g., turn angles, speed, contact points) and apply TGP modeling and trajectory comparison to their own telemetry datasets. Combining the predictive power of machine learning and the statistical inference of Bayesian methods to model movement trajectories allows for the estimation of statistically comparable movement descriptors from telemetry studies. Our use of an accessible R package allows practitioners to model trajectories and estimate movement descriptors, facilitating the use of telemetry data to answer applied management questions.

Ecology and Evolution

A global search inversion for earthquake kinematic rupture history: Application to the 2000 western Tottori, Japan earthquake

[1] We present a two-stage nonlinear technique to invert strong motions records and geodetic data to retrieve the rupture history of an earthquake on a finite fault. To account for the actual rupture complexity, the fault parameters are spatially variable peak slip velocity, slip direction, rupture time and risetime. The unknown parameters are given at the nodes of the subfaults, whereas the parameters within a subfault are allowed to vary through a bilinear interpolation of the nodal values. The forward modeling is performed with a discrete wave number technique, whose Green's functions include the complete response of the vertically varying Earth structure. During the first stage, an algorithm based on the heat-bath simulated annealing generates an ensemble of models that efficiently sample the good data-fitting regions of parameter space. In the second stage (appraisal), the algorithm performs a statistical analysis of the model ensemble and computes a weighted mean model and its standard deviation. This technique, rather than simply looking at the best model, extracts the most stable features of the earthquake rupture that are consistent with the data and gives an estimate of the variability of each model parameter. We present some synthetic tests to show the effectiveness of the method and its robustness to uncertainty of the adopted crustal model. Finally, we apply this inverse technique to the well recorded 2000 western Tottori, Japan, earthquake ( Mw 6.6); we confirm that the rupture process is characterized by large slip (3-4 m) at very shallow depths but, differently from previous studies, we imaged a new slip patch (2-2.5 m) located deeper, between 14 and 18 km depth.

Tottori