USGS ScienceSearch

SEARCH · USGS Science

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Assessing conservation relevance of organism-environment relations using predicted changes in response variables

1. Organism–environment models are used widely in conservation. The degree to which they are useful for informing conservation decisions – the conservation relevance of these relations – is important because lack of relevance may lead to misapplication of scarce conservation resources or failure to resolve important conservation dilemmas. Even when models perform well based on model fit and predictive ability, conservation relevance of associations may not be clear without also knowing the magnitude and variability of predicted changes in response variables. 2. We introduce a method for evaluating the conservation relevance of organism–environment relations that employs confidence intervals for predicted changes in response variables. The confidence intervals are compared to a preselected magnitude of change that marks a threshold (trigger) for conservation action. To demonstrate the approach, we used a case study from the Chihuahuan Desert involving relations between avian richness and broad-scale patterns of shrubland. We considered relations for three winters and two spatial extents (1- and 2-km-radius areas) and compared predicted changes in richness to three thresholds (10%, 20% and 30% change). For each threshold, we examined 48 relations. 3. The method identified seven, four and zero conservation-relevant changes in mean richness for the 10%, 20% and 30% thresholds respectively. These changes were associated with major (20%) changes in shrubland cover, mean patch size, the coefficient of variation for patch size, or edge density but not with major changes in shrubland patch density. The relative rarity of conservation-relevant changes indicated that, overall, the relations had little practical value for informing conservation decisions about avian richness. 4. The approach we illustrate is appropriate for various response and predictor variables measured at any temporal or spatial scale. The method is broadly applicable across ecological environments, conservation objectives, types of statistical predictive models and levels of biological organization. By focusing on magnitudes of change that have practical significance, and by using the span of confidence intervals to incorporate uncertainty of predicted changes, the method can be used to help improve the effectiveness of conservation efforts.

Methods in Ecology and Evolution

δ 13 C and δ 18 O isotopic composition of CaCO 3 measured by continuous flow isotope ratio mass spectrometry: statistical evaluation and verification by application to Devils Hole core DH-11 calcite

A new method was developed to analyze the stable carbon and oxygen isotope ratios of small samples (400 ± 20 µg) of calcium carbonate. This new method streamlines the classical phosphoric acid/calcium carbonate (H 3 PO 4 /CaCO 3 ) reaction method by making use of a recently available Thermoquest-Finnigan GasBench II preparation device and a Delta Plus XL continuous flow isotope ratio mass spectrometer. Conditions for which the H 3 PO 4 /CaCO 3 reaction produced reproducible and accurate results with minimal error had to be determined. When the acid/carbonate reaction temperature was kept at 26 °C and the reaction time was between 24 and 54 h, the precision of the carbon and oxygen isotope ratios for pooled samples from three reference standard materials was ≤0.1 and ≤0.2 per mill or ‰, respectively, although later analysis showed that materials from one specific standard required reaction time between 34 and 54 h for δ 18 O to achieve this level of precision. Aliquot screening methods were shown to further minimize the total error. The accuracy and precision of the new method were analyzed and confirmed by statistical analysis. The utility of the method was verified by analyzing calcite from Devils Hole, Nevada, for which isotope-ratio values had previously been obtained by the classical method. Devils Hole core DH-11 recently had been re-cut and re-sampled, and isotope-ratio values were obtained using the new method. The results were comparable with those obtained by the classical method with correlation = +0.96 for both isotope ratios. The consistency of the isotopic results is such that an alignment offset could be identified in the re-sampled core material, and two cutting errors that occurred during re-sampling then were confirmed independently. This result indicates that the new method is a viable alternative to the classical reaction method. In particular, the new method requires less sample material permitting finer resolution and allows automation of some processes resulting in considerable time savings.

Rapid Communications in Mass Spectrometry

Modeling false positives

Many of the models we are concerned with included explicit descriptions of false negative errors. However, false positive errors can also be commin in practice, especially in citizen science applications where observer skill is highly variable. In addition, new methods which determine detection based on statistical classification or machine learning methods are also prone to false positive errors which must be accounted for. An early treatment of the false positive detection problem by Royle & Link (2006) recognized that false positive errors can be accommodated by a mixture model for detection probability: one value of detection at occupied sites and another non-zero value at unoccupied sites. This model has been extended greatly in recent years to include more informative data about false positives including validation or confirmation data (Miller et al. 2011) and multiple detection methods, among others. A new frontier for the application of false positives models lies in the use of modern technologies such as bioacoustics for efficient automated monitoring. For these technologies to realize their promise there must be improvements in automated processing of the vast quantities of output produced. Statistical classification methods (machine learning) are fallible and necessarily produce false positive detections. Therefore models which account for this process are necessary (Chambert et al. 2017). It stands to reason that false positives will need to be accounted for in other new technologies that rely on automated digital processing, including eDNA, genetic barcoding, and automated detection in remote camera studies. We devise a new occupancy model that integrates data from bioacoustics sampling with an occupancy model. This integrated model allows occupancy probability to inform species classification of samples and vice versa bioacoustics detection data inform occupancy. We provide a proof of concept for this new model in this chapter. As the core hierarchical model for the false positives models covered in this chapter are just ordinary occupancy models, extension of the ideas to open systems poses no technical challenges. We provide a suite of illustrations of these extensions. Perhaps the most prominent mechanism that leads to false positive errors it he mis-classification of species detections, or the confusion of one species for another. Very little work has been done on developing models based on this mechanistic understanding although Chambert et al. (2018) develop this idea as a 2-species occupancy model with error. We believe one important area of future research is to extend these ideas to truly multi-species systems.

Book chapter

Combining statistical inference and decisions in ecology

Statistical decision theory (SDT) is a sub-field of decision theory that formally incorporates statistical investigation into a decision-theoretic framework to account for uncertainties in a decision problem. SDT provides a unifying analysis of three types of information: statistical results from a data set, knowledge of the consequences of potential choices (i.e., loss), and prior beliefs about a system. SDT links the theoretical development of a large body of statistical methods including point estimation, hypothesis testing, and confidence interval estimation. The theory and application of SDT have mainly been developed and published in the fields of mathematics, statistics, operations research, and other decision sciences, but have had limited exposure in ecology. Thus, we provide an introduction to SDT for ecologists and describe its utility for linking the conventionally separate tasks of statistical investigation and decision making in a single framework. We describe the basic framework of both Bayesian and frequentist SDT, its traditional use in statistics, and discuss its application to decision problems that occur in ecology. We demonstrate SDT with two types of decisions: Bayesian point estimation, and an applied management problem of selecting a prescribed fire rotation for managing a grassland bird species. Central to SDT, and decision theory in general, are loss functions. Thus, we also provide basic guidance and references for constructing loss functions for an SDT problem.

Ecological Applications

Coastal loading and transport of Escherichia coli at an embayed beach in Lake Michigan

A Chicago beach in southwest Lake Michigan was revisited to determine the influence of nearshore hydrodynamic effects on the variability of Escherichia coli (E. coli) concentration in both knee-deep and offshore waters. Explanatory variables that could be used for identifying potential bacteria loading mechanisms, such as bed shear stress due to a combined wave-current boundary layer and wave runup on the beach surface, were derived from an existing wave and current database. The derived hydrodynamic variables, along with the actual observed E. coli concentrations in the submerged and foreshore sands, were expected to reveal bacteria loading through nearshore sediment resuspension and swash on the beach surface, respectively. Based on the observation that onshore waves tend to result in a more active hydrodynamic system at this embayed beach, multiple linear regression analysis of onshore-wave cases further indicated the significance of sediment resuspension and the interaction of swash with gull-droppings in explaining the variability of E. coli concentration in the knee-deep water. For cases with longshore currents, numerical simulations using the Princeton Ocean Model revealed current circulation patterns inside the embayment, which can effectively entrain bacteria from the swash zone into the central area of the embayed beach water and eventually release them out of the embayment. The embayed circulation patterns are consistent with the statistical results that identified that 1) the submerged sediment was an additional net source of E. coli to the offshore water and 2) variability of E. coli concentration in the knee-deep water contributed adversely to that in the offshore water for longshore-current cases. The embayed beach setting and the statistical and numerical methods used in the present study have wide applicability for analyzing recreational water quality at similar marine and freshwater sites. ?? 2010 American Chemical Society.

Environmental Science & Technology

Computing daily mean streamflow at ungaged locations in Iowa by using the Flow Anywhere and Flow Duration Curve Transfer statistical methods

The U.S. Geological Survey (USGS) maintains approximately 148 real-time streamgages in Iowa for which daily mean streamflow information is available, but daily mean streamflow data commonly are needed at locations where no streamgages are present. Therefore, the USGS conducted a study as part of a larger project in cooperation with the Iowa Department of Natural Resources to develop methods to estimate daily mean streamflow at locations in ungaged watersheds in Iowa by using two regression-based statistical methods. The regression equations for the statistical methods were developed from historical daily mean streamflow and basin characteristics from streamgages within the study area, which includes the entire State of Iowa and adjacent areas within a 50-mile buffer of Iowa in neighboring states. Results of this study can be used with other techniques to determine the best method for application in Iowa and can be used to produce a Web-based geographic information system tool to compute streamflow estimates automatically. The Flow Anywhere statistical method is a variation of the drainage-area-ratio method, which transfers same-day streamflow information from a reference streamgage to another location by using the daily mean streamflow at the reference streamgage and the drainage-area ratio of the two locations. The Flow Anywhere method modifies the drainage-area-ratio method in order to regionalize the equations for Iowa and determine the best reference streamgage from which to transfer same-day streamflow information to an ungaged location. Data used for the Flow Anywhere method were retrieved for 123 continuous-record streamgages located in Iowa and within a 50-mile buffer of Iowa. The final regression equations were computed by using either left-censored regression techniques with a low limit threshold set at 0.1 cubic feet per second (ft3/s) and the daily mean streamflow for the 15th day of every other month, or by using an ordinary-least-squares multiple linear regression method and the daily mean streamflow for the 15th day of every other month. The Flow Duration Curve Transfer method was used to estimate unregulated daily mean streamflow from the physical and climatic characteristics of gaged basins. For the Flow Duration Curve Transfer method, daily mean streamflow quantiles at the ungaged site were estimated with the parameter-based regression model, which results in a continuous daily flow-duration curve (the relation between exceedance probability and streamflow for each day of observed streamflow) at the ungaged site. By the use of a reference streamgage, the Flow Duration Curve Transfer is converted to a time series. Data used in the Flow Duration Curve Transfer method were retrieved for 113 continuous-record streamgages in Iowa and within a 50-mile buffer of Iowa. The final statewide regression equations for Iowa were computed by using a weighted-least-squares multiple linear regression method and were computed for the 0.01-, 0.05-, 0.10-, 0.15-, 0.20-, 0.30-, 0.40-, 0.50-, 0.60-, 0.70-, 0.80-, 0.85-, 0.90-, and 0.95-exceedance probability statistics determined from the daily mean streamflow with a reporting limit set at 0.1 ft 3 /s. The final statewide regression equation for Iowa computed by using left-censored regression techniques was computed for the 0.99-exceedance probability statistic determined from the daily mean streamflow with a low limit threshold and a reporting limit set at 0.1 ft 3 /s. For the Flow Anywhere method, results of the validation study conducted by using six streamgages show that differences between the root-mean-square error and the mean absolute error ranged from 1,016 to 138 ft 3 /s, with the larger value signifying a greater occurrence of outliers between observed and estimated streamflows. Root-mean-square-error values ranged from 1,690 to 237 ft 3 /s. Values of the percent root-mean-square error ranged from 115 percent to 26.2 percent. The logarithm (base 10) streamflow percent root-mean-square error ranged from 13.0 to 5.3 percent. Root-mean-square-error observations standard-deviation-ratio values ranged from 0.80 to 0.40. Percent-bias values ranged from 25.4 to 4.0 percent. Untransformed streamflow Nash-Sutcliffe efficiency values ranged from 0.84 to 0.35. The logarithm (base 10) streamflow Nash-Sutcliffe efficiency values ranged from 0.86 to 0.56. For the streamgage with the best agreement between observed and estimated streamflow, higher streamflows appear to be underestimated. For the streamgage with the worst agreement between observed and estimated streamflow, low flows appear to be overestimated whereas higher flows seem to be underestimated. Estimated cumulative streamflows for the period October 1, 2004, to September 30, 2009, are underestimated by -25.8 and -7.4 percent for the closest and poorest comparisons, respectively. For the Flow Duration Curve Transfer method, results of the validation study conducted by using the same six streamgages show that differences between the root-mean-square error and the mean absolute error ranged from 437 to 93.9 ft 3 /s, with the larger value signifying a greater occurrence of outliers between observed and estimated streamflows. Root-mean-square-error values ranged from 906 to 169 ft 3 /s. Values of the percent root-mean-square-error ranged from 67.0 to 25.6 percent. The logarithm (base 10) streamflow percent root-mean-square error ranged from 12.5 to 4.4 percent. Root-mean-square-error observations standard-deviation-ratio values ranged from 0.79 to 0.40. Percent-bias values ranged from 22.7 to 0.94 percent. Untransformed streamflow Nash-Sutcliffe efficiency values ranged from 0.84 to 0.38. The logarithm (base 10) streamflow Nash-Sutcliffe efficiency values ranged from 0.89 to 0.48. For the streamgage with the closest agreement between observed and estimated streamflow, there is relatively good agreement between observed and estimated streamflows. For the streamgage with the poorest agreement between observed and estimated streamflow, streamflows appear to be substantially underestimated for much of the time period. Estimated cumulative streamflow for the period October 1, 2004, to September 30, 2009, are underestimated by -9.3 and -22.7 percent for the closest and poorest comparisons, respectively.

Illinois;Iowa;Minnesota;Missouri;Nebraska;Wisconsi

Trends in groundwater levels in and near the Rosebud Indian Reservation, South Dakota, water years 1956–2017

The U.S. Geological Survey (USGS), in cooperation with the Rosebud Sioux Tribe, completed a study to characterize water-level fluctuations in observation wells to examine driving factors that affect water levels in and near the Rosebud Indian Reservation, which comprises all of Todd County. The study investigates concerns regarding potential effects of groundwater withdrawals and climate conditions on groundwater levels within an area that includes Todd County and a surrounding area that extends 10 miles north, east, and west of the county border. Characterization of water-level fluctuations in observation wells and relative driving factors was accomplished by statistical trend analysis. Two statistical methods were used for analysis of temporal trends for climatic and hydrologic data. To determine which trend analysis to use, applicable datasets were tested for statistically significant short-term persistence (STP). In the absence of significant STP, existence of statistical trends was determined using the standard Mann-Kendall test for probability values less than or equal to 0.10 (90-percent confidence level); however, a modified Mann-Kendall test was used for datasets where statistically significant STP was detected. Trend magnitudes were computed using the Sen’s slope estimator. Monthly data from the Parameter-elevation Regressions on Independent Slopes Model (PRISM) were aggregated to obtain annual and seasonal datasets for total precipitation, minimum air temperature ( T min ), and maximum air temperature ( T max ) for the study area and a surrounding buffer area. Trend tests for total precipitation, T min , and T max were completed for annual and seasonal time series for water years 1956–2017, which is about 2 years before the earliest available water-level measurements. A 2-year offset was arbitrarily selected because scrutiny of water-level and precipitation data indicated that responses of groundwater levels for many of the observation wells lagged major changes in precipitation patterns by about 2 years. Statistically significant upward trends were detected for annual precipitation and annual T min for almost all of the study area and the surrounding buffer area. Statistically significant downward trends in T max were detected for a very small part of the study area; however, the sparse spatial coverage reduces confidence that these are true trends. Spatial distributions of statistically significant trends in seasonal climate data were generally similar to the annual trends, but with substantial differences in the spatial density of the trends. Groundwater trends for 58 observation wells were analyzed for three separate water-level parameters (minimum, median, and maximum) because wells are measured sporadically and data are biased towards more frequent measurements during periods of heaviest irrigation demand. Trends in the time series of annual precipitation (from PRISM) starting 2 years earlier than for the associated water-level trend also were analyzed for the location of each individual observation well. Sen’s slope and Mann-Kendall probability values (p-values) were computed for the three water-level parameters and for the annual precipitation time series. Graphs showing results of trend analyses for each observation well also showed changes over time in the sum of licensed groundwater withdrawals within six specified radii (0.5, 1, 2, 3, 4, and 5 miles) of each well as a qualitative indicator of proximal groundwater demand. Of all 58 observation wells considered, 28 wells had significant upward trends for at least one of the three water-level parameters, 11 wells had significant downward trends for at least one water-level parameter, and 19 wells did not have any significant trends. Significant upward trends in annual precipitation were detected for 48 of the 58 wells. Results of trend analyses likely show the effects of groundwater withdrawals on water levels in the Ogallala aquifer in areas of substantial demand. Precipitation trends are significantly upward for 43 of the 48 wells completed in the Ogallala aquifer that were analyzed. Of the 48 Ogallala aquifer wells, 24 had significant upward trends for at least one water-level parameter (17 with all 3); however, 10 wells had statistically significant downward trends for at least one water-level parameter (8 with all 3 parameters). All but one of the wells with significant downward trends are located in the south-central part of the study area where licensed irrigation withdrawals are concentrated.

South Dakota

A comparison of methods to predict historical daily streamflow time series in the southeastern United States

Effective and responsible management of water resources relies on a thorough understanding of the quantity and quality of available water. Streamgages cannot be installed at every location where streamflow information is needed. As part of its National Water Census, the U.S. Geological Survey is planning to provide streamflow predictions for ungaged locations. In order to predict streamflow at a useful spatial and temporal resolution throughout the Nation, efficient methods need to be selected. This report examines several methods used for streamflow prediction in ungaged basins to determine the best methods for regional and national implementation. A pilot area in the southeastern United States was selected to apply 19 different streamflow prediction methods and evaluate each method by a wide set of performance metrics. Through these comparisons, two methods emerged as the most generally accurate streamflow prediction methods: the nearest-neighbor implementations of nonlinear spatial interpolation using flow duration curves (NN-QPPQ) and standardizing logarithms of streamflow by monthly means and standard deviations (NN-SMS12L). It was nearly impossible to distinguish between these two methods in terms of performance. Furthermore, neither of these methods requires significantly more parameterization in order to be applied: NN-SMS12L requires 24 regional regressions—12 for monthly means and 12 for monthly standard deviations. NN-QPPQ, in the application described in this study, required 27 regressions of particular quantiles along the flow duration curve. Despite this finding, the results suggest that an optimal streamflow prediction method depends on the intended application. Some methods are stronger overall, while some methods may be better at predicting particular statistics. The methods of analysis presented here reflect a possible framework for continued analysis and comprehensive multiple comparisons of methods of prediction in ungaged basins (PUB). Additional metrics of comparison can easily be incorporated into this type of analysis. By considering such a multifaceted approach, the top-performing models can easily be identified and considered for further research. The top-performing models can then provide a basis for future applications and explorations by scientists, engineers, managers, and practitioners to suit their own needs.

Scientific Investigations Report

Statewide analysis of the drainage-area ratio method for 34 streamflow percentile ranges in Texas

The drainage-area ratio method commonly is used to estimate streamflow for sites where no streamflow data are available using data from one or more nearby streamflow-gaging stations. The method is intuitive and straightforward to implement and is in widespread use by analysts and managers of surface-water resources. The method equates the ratio of streamflow at two stream locations to the ratio of the respective drainage areas. In practice, unity often is assumed as the exponent on the drainage-area ratio, and unity also is assumed as a multiplicative bias correction. These two assumptions are evaluated in this investigation through statewide analysis of daily mean streamflow in Texas. The investigation was made by the U.S. Geological Survey in cooperation with the Texas Commission on Environmental Quality. More than 7.8 million values of daily mean streamflow for 712 U.S. Geological Survey streamflow-gaging stations in Texas were analyzed. To account for the influence of streamflow probability on the drainage-area ratio method, 34 percentile ranges were considered. The 34 ranges are the 4 quartiles (0-25, 25-50, 50-75, and 75-100 percent), the 5 intervals of the lower tail of the streamflow distribution (0-1, 1-2, 2-3, 3-4, and 4-5 percent), the 20 quintiles of the 4 quartiles (0-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50, 50-55, 55-60, 60-65, 65-70, 70-75, 75-80, 80-85, 85-90, 90-95, and 95-100 percent), and the 5 intervals of the upper tail of the streamflow distribution (95-96, 96-97, 97-98, 98-99 and 99-100 percent). For each of the 253,116 (712X711/2) unique pairings of stations and for each of the 34 percentile ranges, the concurrent daily mean streamflow values available for the two stations provided for station-pair application of the drainage-area ratio method. For each station pair, specific statistical summarization (median, mean, and standard deviation) of both the exponent and bias-correction components of the drainage-area ratio method were computed. Statewide statistics (median, mean, and standard deviation) of the station-pair specific statistics subsequently were computed and are tabulated herein. A separate analysis considered conditioning station pairs to those stations within 100 miles of each other and with the absolute value of the logarithm (base-10) of the ratio of the drainage areas greater than or equal to 0.25. Statewide statistics of the conditional station-pair specific statistics were computed and are tabulated. The conditional analysis is preferable because of the anticipation that small separation distances reflect similar hydrologic conditions and the observation of large variation in exponent estimates for similar-sized drainage areas. The conditional analysis determined that the exponent is about 0.89 for streamflow percentiles from 0 to about 50 percent, is about 0.92 for percentiles from about 50 to about 65 percent, and is about 0.93 for percentiles from about 65 to about 85 percent. The exponent decreases rapidly to about 0.70 for percentiles nearing 100 percent. The computation of the bias-correction factor is sensitive to the range analysis interval (range of streamflow percentile); however, evidence suggests that in practice the drainage-area method can be considered unbiased. Finally, for general application, suggested values of the exponent are tabulated for 54 percentiles of daily mean streamflow in Texas; when these values are used, the bias correction is unity.

Scientific Investigations Report

Practical bias correction in aerial surveys of large mammals: Validation of hybrid double-observer with sightability method against known abundance of feral horse (Equus caballus) populations

Reliably estimating wildlife abundance is fundamental to effective management. Aerial surveys are one of the only spatially robust tools for estimating large mammal populations, but statistical sampling methods are required to address detection biases that affect accuracy and precision of the estimates. Although various methods for correcting aerial survey bias are employed on large mammal species around the world, these have rarely been rigorously validated. Several populations of feral horses ( Equus caballus ) in the western United States have been intensively studied, resulting in identification of all unique individuals. This provided a rare opportunity to test aerial survey bias correction on populations of known abundance. We hypothesized that a hybrid method combining simultaneous double-observer and sightability bias correction techniques would accurately estimate abundance. We validated this integrated technique on populations of known size and also on a pair of surveys before and after a known number was removed. Our analysis identified several covariates across the surveys that explained and corrected biases in the estimates. All six tests on known populations produced estimates with deviations from the known value ranging from -8.5% to +13.7% and <0.7 standard errors. Precision varied widely, from 6.1% CV to 25.0% CV. In contrast, the pair of surveys conducted around a known management removal produced an estimated change in population between the surveys that was significantly larger than the known reduction. Although the deviation between was only 9.1%, the precision estimate (CV = 1.6%) may have been artificially low. It was apparent that use of a helicopter in those surveys perturbed the horses, introducing detection error and heterogeneity in a manner that could not be corrected by our statistical models. Our results validate the hybrid method, highlight its potentially broad applicability, identify some limitations, and provide insight and guidance for improving survey designs.

Colorado, Nevada, Utah, Wyoming

Advancements in analytical approaches improve whitebark pine monitoring results

Long-term monitoring programs track the status and trends of species in increasingly vulnerable environments. These monitoring results provide critical information for evaluating, understanding, and managing natural resources. To accurately interpret if and how conditions may be changing for select ecological indicators, it is essential that monitoring programs adopt methods to ensure exceptional data quality. To do this and remain relevant, practitioners need to be flexible and willing to embrace a degree of adaptivity in their protocols. They must periodically re-evaluate their statistical methods and field data collection techniques to provide contemporary, significant, and applicable inferences.

Idaho, Montana, Wyoming

Using the U.S. Geological Survey National Water Quality Laboratory LT-MDL to Evaluate and Analyze Data

A long-term method detection level (LT-MDL) and laboratory reporting level (LRL) are used by the U.S. Geological Survey?s National Water Quality Laboratory (NWQL) when reporting results from most chemical analyses of water samples. Changing to this method provided data users with additional information about their data and often resulted in more reported values in the low concentration range. Before this method was implemented, many of these values would have been censored. The use of the LT-MDL and LRL presents some challenges for the data user. Interpreting data in the low concentration range increases the need for adequate quality assurance because even small contamination or recovery problems can be relatively large compared to concentrations near the LT-MDL and LRL. In addition, the definition of the LT-MDL, as well as the inclusion of low values, can result in complex data sets with multiple censoring levels and reported values that are less than a censoring level. Improper interpretation or statistical manipulation of low-range results in these data sets can result in bias and incorrect conclusions. This document is designed to help data users use and interpret data reported with the LTMDL/ LRL method. The calculation and application of the LT-MDL and LRL are described. This document shows how to extract statistical information from the LT-MDL and LRL and how to use that information in USGS investigations, such as assessing the quality of field data, interpreting field data, and planning data collection for new projects. A set of 19 detailed examples are included in this document to help data users think about their data and properly interpret lowrange data without introducing bias. Although this document is not meant to be a comprehensive resource of statistical methods, several useful methods of analyzing censored data are demonstrated, including Regression on Order Statistics and Kaplan-Meier Estimation. These two statistical methods handle complex censored data sets without resorting to substitution, thereby avoiding a common source of bias and inaccuracy.

Open-File Report

Chance-corrected classification for use in discriminant analysis: Ecological applications

A method for evaluating the classification table from a discriminant analysis is described. The statistic, kappa, is useful to ecologists in that it removes the effects of chance. It is useful even with equal group sample sizes although the need for a chance-corrected measure of prediction becomes greater with more dissimilar group sample sizes. Examples are presented.

American Midland Naturalist

Automating calibration, sensitivity and uncertainty analysis of complex models using the R package Flexible Modeling Environment (FME): SWAT as an example

Parameter optimization and uncertainty issues are a great challenge for the application of large environmental models like the Soil and Water Assessment Tool (SWAT), which is a physically-based hydrological model for simulating water and nutrient cycles at the watershed scale. In this study, we present a comprehensive modeling environment for SWAT, including automated calibration, and sensitivity and uncertainty analysis capabilities through integration with the R package Flexible Modeling Environment (FME). To address challenges (e.g., calling the model in R and transferring variables between Fortran and R) in developing such a two-language coupling framework, 1) we converted the Fortran-based SWAT model to an R function (R-SWAT) using the RFortran platform, and alternatively 2) we compiled SWAT as a Dynamic Link Library (DLL). We then wrapped SWAT (via R-SWAT) with FME to perform complex applications including parameter identifiability, inverse modeling, and sensitivity and uncertainty analysis in the R environment. The final R-SWAT-FME framework has the following key functionalities: automatic initialization of R, running Fortran-based SWAT and R commands in parallel, transferring parameters and model output between SWAT and R, and inverse modeling with visualization. To examine this framework and demonstrate how it works, a case study simulating streamflow in the Cedar River Basin in Iowa in the United Sates was used, and we compared it with the built-in auto-calibration tool of SWAT in parameter optimization. Results indicate that both methods performed well and similarly in searching a set of optimal parameters. Nonetheless, the R-SWAT-FME is more attractive due to its instant visualization, and potential to take advantage of other R packages (e.g., inverse modeling and statistical graphics). The methods presented in the paper are readily adaptable to other model applications that require capability for automated calibration, and sensitivity and uncertainty analysis.

Environmental Modelling and Software

Trends in plant cover derived from vegetation plot data using ordinal zero-augmented beta regression

Questions Plant cover values in vegetation plot data are bounded between 0 and 1, and cover is typically recorded in discrete classes with non-equal intervals. Consequently, cover data are skewed and heteroskedastic, which hampers the application of conventional regression methods. Recently developed ordinal beta regression models consider these statistical difficulties. Our primary question is whether we can detect species trends in vegetation plot time series data with this modelling approach. A second question is whether trends in cover have additional value compared to trends in occurrence, which are easier to assess for practitioners. Location The Netherlands, Western Europe. Methods We used vegetation plot data collected from 10,000 fixed plots which were surveyed once every four years during 1999–2022. We used the ordinal zero-augmented beta regression (OZAB) model, a hierarchical model consisting of a logistic regression for presence and an ordinal beta regression for cover. We adapted the OZAB model for longitudinal data and produced estimates of cover and occurrence for each four-year period. Thereafter we assessed trends in cover and in occurrence across all periods. Results We found evidence of a trend in cover in 318 out of the 721 species (44%) with sufficient data. Most species showed similar directional trends in occurrence and percent cover. No trend in occurrence was detected for 64 species that had evidence of a trend in cover. Declining species had stronger relative changes in cover than in occurrence. Conclusions Our model enables researchers to detect trends in cover using longitudinal vegetation plot data. Cover trends often corroborated trends in occurrence, but we also regularly found trends in cover even in the absence of evidence for trends in occurrence. Our approach thus contributes to a more complete picture of (changes in) vegetation composition based on large monitoring data sets.

Journal of Vegetation Science

Review of paleomagnetism

This review is an attempt to bring together and discuss relevant information concerning the magnetization of rocks, especially that having paleomagnetic significance. All paleomagnetic measurements available to the authors are here compiled and evaluated, with a key to the summary table and illustrations in English and Russian. The principles upon which the evaluation of paleomagnetic measurements is based are summarized, with special emphasis on statistical methods and on the evidence and tests for magnetic stability and paleomagnetic applicability. Evaluation of the data summarized leads to the following general conclusions: (1) The earth's average magnetic field, throughout Oligocene to Recent time, has very closely approximated that due to a dipole at the center of the earth oriented parallel to the present axis of rotation. (2) Paleomagnetic results for the Mesozoic and early Tertiary might be explained more plausibly by a relatively rapidly changing magnetic field, with or without wandering of the rotational pole, than by large-scale continental drift. (3) The Carboniferous and especially the Permian magnetic fields were relatively very “steady” and were vastly different from the present configuration of the field. (4) The Precambrian magnetic field was different from the present field configuration and, considering the time spanned, was remarkably consistent for all continents.

GSA Bulletin

Robust and resistant semivariogram modelling using a generalized bootstrap

The bootstrap is a computer-intensive resampling method for estimating the uncertainty of complex statistical models. We expand on an application of the bootstrap for inferring semivariogram parameters and their uncertainty. The model fitted to the median of the bootstrap distribution of the experimental semivariogram is proposed as an estimator of the semivariogram. The proposed application is not restricted to normal data and the estimator is resistant to outliers. Improvements are more significant for data-sets with less than 100 observations, which are those for which semivariogram model inference is the most difficult. The application is illustrated by using it to characterize a synthetic random field for which the true semivariogram type and parameters are known.

Journal of the Southern African Institute of Minin

Watershed Regressions for Pesticides (WARP) models for predicting stream concentrations of multiple pesticides

Watershed Regressions for Pesticides for multiple pesticides (WARP-MP) are statistical models developed to predict concentration statistics for a wide range of pesticides in unmonitored streams. The WARP-MP models use the national atrazine WARP models in conjunction with an adjustment factor for each additional pesticide. The WARP-MP models perform best for pesticides with application timing and methods similar to those used with atrazine. For other pesticides, WARP-MP models tend to overpredict concentration statistics for the model development sites. For WARP and WARP-MP, the less-than-ideal sampling frequency for the model development sites leads to underestimation of the shorter-duration concentration; hence, the WARP models tend to underpredict 4- and 21-d maximum moving-average concentrations, with median errors ranging from 9 to 38% As a result of this sampling bias, pesticides that performed well with the model development sites are expected to have predictions that are biased low for these shorter-duration concentration statistics. The overprediction by WARP-MP apparent for some of the pesticides is variably offset by underestimation of the model development concentration statistics. Of the 112 pesticides used in the WARP-MP application to stream segments nationwide, 25 were predicted to have concentration statistics with a 50% or greater probability of exceeding one or more aquatic life benchmarks in one or more stream segments. Geographically, many of the modeled streams in the Corn Belt Region were predicted to have one or more pesticides that exceeded an aquatic life benchmark during 2009, indicating the potential vulnerability of streams in this region.

Journal of Environmental Quality