USGS Science⌕ Search

SEARCH · USGS Science

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Quality of groundwater used for domestic supply in the Gilroy-Hollister basin and surrounding areas, California, 2022

More than 2 million Californians rely on groundwater from domestic wells for drinking-water supply. This report summarizes a 2022 California Groundwater Ambient Monitoring and Assessment Priority Basin Project (GAMA-PBP) water-quality survey of 33 domestic and small-system drinking-water supply wells in the Gilroy-Hollister Valley groundwater basin and the surrounding areas, where more than 20,000 residents are estimated to utilize privately owned domestic wells. The study area includes the Llagas subbasin in the north, the North San Benito subbasin in the south, and the surrounding uplands. The study was focused on groundwater resources used for domestic drinking-water supply, which are mostly drawn from shallower parts of aquifer systems rather than those of groundwater resources used for public drinking-water supply in the same area. This assessment characterized the quality of ambient groundwater in the aquifer before filtration or treatment, rather than the quality of drinking water delivered to the tap. To provide context, the measured concentrations of constituents in groundwater were compared to Federal and California State regulatory and non-regulatory benchmarks for drinking-water quality. A grid-based method was used to estimate the areal proportions of groundwater resources used for domestic drinking wells that have water-quality constituents present at high concentrations (above the benchmark), moderate concentrations (between one-half of the benchmark and the benchmark for inorganic constituents, or between one-tenth of the benchmark and the benchmark for organic constituents), and low concentrations (less than one-half or one-tenth the benchmark for inorganic and organic constituents, respectively). This method provides statistically representative results at the study-area scale and permits comparisons to other GAMA-PBP study areas. In the study area, inorganic constituents in groundwater were greater than regulatory benchmarks (U.S. Environmental Protection Agency [EPA] or State of California maximum contaminant levels [MCLs]) for public drinking-water quality in 24 percent of domestic groundwater resources. The inorganic constituents present at concentrations greater than MCLs for drinking water were nitrate (as nitrogen), barium, chromium, and selenium. Total dissolved solids (TDS) or manganese were present at concentrations greater than the secondary maximum contaminant levels (SMCLs) that the State of California uses as aesthetic-based benchmarks in 48 percent of domestic groundwater resources. No volatile organic compounds or pesticide constituents were present at concentrations greater than regulatory benchmarks. Total coliform bacteria and enterococci were detected in 4 percent of domestic groundwater resources. Per- and polyfluoroalkyl substances (PFAS) were detected in 19 percent of domestic groundwater resources, and 10 percent had concentrations greater than recently enacted (April 2024) EPA MCLs. Physical and chemical factors from natural and anthropogenic sources that could affect the groundwater quality were evaluated using results from statistical testing of associations between constituent concentrations and potential explanatory variables. In this study, relevant physical factors include well construction characteristics, groundwater age, site proximity to groundwater recharge or discharge zones, and potential sources of contamination. Relevant chemical factors include the initial chemistry of the recharge water, the mineralogy of the aquifer sediments, and the subsequent shifts in chemistry as biologic and geologic reactions alter groundwater in the subsurface. Nitrate concentrations were correlated to agricultural land use, distance from the boundary of the Gilroy-Hollister Valley groundwater basin, and the proportion of modern (post-1950s) water captured by the well. Denitrification under anoxic redox conditions can mitigate some nitrate derived from fertilizer application. Total dissolved solids primarily were derived from water-rock interactions with soils and aquifer materials in the study area, but there were high concentrations where agricultural practices contributed additional TDS. Mineralogy of aquifer sediments and rocks also affect barium, selenium, boron, and chromium concentrations in the Gilroy-Hollister Valley groundwater basin. PFAS were positively correlated with urban land use and the proportion of modern water captured by the well.

California↗

Demography of a reintroduced population: moving toward management models for an endangered species, the whooping crane

The reintroduction of threatened and endangered species is now a common method for reestablishing populations. Typically, a fundamental objective of reintroduction is to establish a self-sustaining population. Estimation of demographic parameters in reintroduced populations is critical, as these estimates serve multiple purposes. First, they support evaluation of progress toward the fundamental objective via construction of population viability analyses (PVAs) to predict metrics such as probability of persistence. Second, PVAs can be expanded to support evaluation of management actions, via management modeling. Third, the estimates themselves can support evaluation of the demographic performance of the reintroduced population, e.g., via comparison with wild populations. For each of these purposes, thorough treatment of uncertainties in the estimates is critical. Recently developed statistical methods - namely, hierarchical Bayesian implementations of state-space models - allow for effective integration of different types of uncertainty in estimation. We undertook a demographic estimation effort for a reintroduced population of endangered whooping cranes with the purpose of ultimately developing a Bayesian PVA for determining progress toward establishing a self-sustaining population, and for evaluating potential management actions via a Bayesian PVA-based management model. We evaluated individual and temporal variation in demographic parameters based upon a multi-state mark-recapture model. We found that survival was relatively high across time and varied little by sex. There was some indication that survival varied by release method. Survival was similar to that observed in the wild population. Although overall reproduction in this reintroduced population is poor, birds formed social pairs when relatively young, and once a bird was in a social pair, it had a nearly 50% chance of nesting the following breeding season. Also, once a bird had nested, it had a high probability of nesting again. These results are encouraging considering that survival and reproduction have been major challenges in past reintroductions of this species. The demographic estimates developed will support construction of a management model designed to facilitate exploration of management actions of interest, and will provide critical guidance in future planning for this reintroduction. An approach similar to what we describe could be usefully applied to many reintroduced populations.

Ecological Applications↗

Quick-start guide for version 3.0 of EMINERS - Economic Mineral Resource Simulator

Quantitative mineral resource assessment, as developed by the U.S. Geological Survey (USGS), consists of three parts: (1) development of grade and tonnage mineral deposit models; (2) delineation of tracts permissive for each deposit type; and (3) probabilistic estimation of the numbers of undiscovered deposits for each deposit type (Singer and Menzie, 2010). The estimate of the number of undiscovered deposits at different levels of probability is the input to the EMINERS (Economic Mineral Resource Simulator) program. EMINERS uses a Monte Carlo statistical process to combine probabilistic estimates of undiscovered mineral deposits with models of mineral deposit grade and tonnage to estimate mineral resources. It is based upon a simulation program developed by Root and others (1992), who discussed many of the methods and algorithms of the program. Various versions of the original program (called "MARK3" and developed by David H. Root, William A. Scott, and Lawrence J. Drew of the USGS) have been published (Root, Scott, and Selner, 1996; Duval, 2000, 2012). The current version (3.0) of the EMINERS program is available as USGS Open-File Report 2004-1344 (Duval, 2012). Changes from version 2.0 include updating 87 grade and tonnage models, designing new templates to produce graphs showing cumulative distribution and summary tables, and disabling economic filters. The economic filters were disabled because embedded data for costs of labor and materials, mining techniques, and beneficiation methods are out of date. However, the cost algorithms used in the disabled economic filters are still in the program and available for reference for mining methods and milling techniques included in Camm (1991). EMINERS is written in C++ and depends upon the Microsoft Visual C++ 6.0 programming environment. The code depends heavily on the use of Microsoft Foundation Classes (MFC) for implementation of the Windows interface. The program works only on Microsoft Windows XP or newer personal computers. It does not work on Macintosh computers. This report demonstrates how to execute EMINERS software using default settings and existing deposit models. Many options are available when setting up the simulation. Information and explanations addressing these optional parameters can be found in the EMINERS Help files. Help files are available during execution of EMINERS by selecting EMINERS Help from the pull-down menu under Help on the EMINERS menu bar. There are four sections in this report. Part I describes the installation, setup, and application of the EMINERS program, and Part II illustrates how to interpret the text file that is produced. Part III describes the creation of tables and graphs by use of the provided Excel templates. Part IV summarizes grade and tonnage models used in version 3.0 of EMINERS.

Open-File Report↗

Trends in environmental, anthropogenic, and water-quality characteristics in the upper White River Basin, Indiana

The U.S. Geological Survey (USGS), in cooperation with The Nature Conservancy, undertook a study to update and extend results from a previous study (Koltun, 2019, https://doi.org/10.3133/sir20195119 ), using data from 3 additional years and newer estimation methods. Koltun (2019) assessed trends in streamflow, precipitation, and estimated annual mean concentrations and flux of nitrate plus nitrite, total Kjeldahl nitrogen, total phosphorus, and total suspended solids (TSS) for USGS streamflow gages on the upper White River at Muncie, near Nora, and near Centerton, Indiana. Annual mean and maximum daily streamflows had statistically significant upward trends at all study gages between water years 1978 and 2020. An abrupt increase in streamflow occurred around water year 2001. Annual total precipitation at the Indianapolis International Airport increased between calendar years 1932 and 2020 at an average rate of 0.089 inches per year. The current study assessed the magnitude, direction, and likelihood of change in flow-normalized concentrations and flux of TSS, total phosphorus, nitrate plus nitrite, and total Kjeldahl nitrogen between water years 1997 and 2019. With two exceptions, concentration and flux changes that were statistically significant in Koltun (2019, https://doi.org/10.3133/sir20195119 ), which reported changes between water years 1997 and 2017, still have the same statistically significant change directions. The reliability of the current trend result for TSS is uncertain because of a large gap in the TSS record for the Centerton gage. For each constituent, spatial patterns were examined in the sampled distribution of nutrient and TSS concentration data from 20 mainstem, tributary, and distributary locations in the upper White River Basin. The largest median concentrations of TSS, total phosphorus, and total Kjeldahl nitrogen were associated with mainstem upper White River sites downstream from Indianapolis. The median total phosphorus and total Kjeldahl nitrogen concentrations were elevated relative to bracketing upstream/downstream mainstem sites at the upper White River site immediately downstream from Muncie. Data on several anthropogenic factors that could influence the concentrations and fluxes of nutrients and TSS were gathered and analyzed to better understand the factors’ spatial and temporal variations. Those anthropogenic factors included population, land cover, cropping and operational tillage practices, fertilizer application, and upgrades to wastewater treatment systems and delivery processes.

Indiana↗

Analysis of artificially matured shales with confocal laser scanning raman microscopy: Applications to organic matter characterization

Raman spectroscopy has been suggested as a method for characterizing the thermal maturity of rocks. The literature contains many empirical correlations between thermal maturity proxies, such as vitrinite reflectance (V Ro ) and pyrolysis-T max , with spectral metrics such as Raman peak-widths, peak-center positions, peak-areas and all manner of differences and ratios of these parameters. However, while these correlations may be convincing for small data sets from limited sample series, broader application of these metrics to disparate and heterogeneous samples proves difficult and there remains no consensus. In this extended abstract, Raman spectroscopy is introduced and the history of Raman analysis of carbonaceous material is briefly outlined, highlighting some of the latent difficulties and potential sources of bias. We suggest the organization of a community working group to establish terminology, guidelines, procedures and standards necessary for the successful development of this technique to characterize organic matter in an accessible, unbiased, and reproducible manner. For the present multi-phase study, immature shale samples from the Bakken and Duvernay formations were subjected to hydrous pyrolysis for 72 hours at temperatures from 280°C to 360°C. Rock residues were mounted and polished for analysis via confocal laser-scanning Raman microscopy and reflectance. The maturation series from the Bakken was randomized for the Phase-I single-blind study to be presented at this conference. For the Phase-II study, solid bitumen reflectance (B Ro ) values for the Duvernay series will be known. Multiple hyperspectral maps were collected from each Bakken sample, with each map consisting of a single diffraction-limited spot-size spectrum per 1 µm 2 in rectangular areas several hundred micrometers on a side. Initial attempts at using basic spectral metrics on small numbers of hand-selected spectra to sort the blind series produced inconclusive results: any number of possible correlations could be found. In an improved approach, the statistics of the full spectral datasets were leveraged to: 1) objectively identify organic carbon types (OCTs) in a given map based on Raman and fluorescence spectral characteristics, 2) identify those OCTs in other maps from the same sample and determine if the heterogeneity of the sample has been adequately characterized, and 3) identify the same OCTs in maps from other samples in the maturation series. In ongoing work, our goals are to: 1) use these analyses of the blind series to develop a hypothesis for a correlation to maturation, 2) test the hypothesis by applying the same analyses to the known Duvernay series (in Phase-II), 3) if necessary refine the hypothesis based on observations from the Duvernay analysis, and 4) finally reveal the true order of the Bakken series to verify if the hypothesized correlation accurately predicts the maturity order of the samples. In this document, we share progress to date. The analysis of one area of interest is detailed showing the differentiation of two OCTs based on Raman and fluorescence spectral features, including the use of 2-factor histograms, Principle Components Analysis (PCA), and Nonlinear Iterative Peak Fitting (NIPF).

Conference Paper↗

Geospatial Technology Applications and Infrastructure in the Biological Resources Division

Executive Summary -- Automated spatial processing technology such as geographic information systems (GIS), telemetry, and satellite-based remote sensing are some of the more recent developments in the long history of geographic inquiry. For millennia, humankind has endeavored to map the Earth's surface and identify spatial relationships. But the precision with which we can locate geographic features has increased exponentially with satellite positioning systems. Remote sensing, GIS, thematic mapping, telemetry, and satellite positioning systems such as the Global Positioning System (GPS) are tools that greatly enhance the quality and rapidity of analysis of biological resources. These technologies allow researchers, planners, and managers to more quickly and accurately determine appropriate strategies and actions. Researchers and managers can view information from new and varying perspectives using GIS and remote sensing, and GPS receivers allow the researcher or manager to identify the exact location of interest. These geospatial technologies support the mission of the U.S. Geological Survey (USGS) Biological Resources Division (BRD) and the Strategic Science Plan (BRD 1996) by providing a cost-effective and efficient method for collection, analysis, and display of information. The BRD mission is 'to work with others to provide the scientific understanding and technologies needed to support the sound management and conservation of our Nation's biological resources.' A major responsibility of the BRD is to develop and employ advanced technologies needed to synthesize, analyze, and disseminate biological and ecological information. As the Strategic Science Plan (BRD 1996) states, 'fulfilling this mission depends on effectively balancing the immediate need for information to guide management of biological resources with the need for technical assistance and long-range, strategic information to understand and predict emerging patterns and trends in ecological systems.' Information sharing plays a key role in nearly everything BRD does. The Strategic Science Plan discusses the need to (1) develop tools and standards for information transfer, (2) disseminate information, and (3) facilitate effective use of information. This effort centers around the National Biological Information Infrastructure (NBII) and the National Spatial Data Infrastructure (NSDI), components of the National Information Infrastructure. The NBII and NSDI are distributed electronic networks of biological and geographical data and information, as well as tools to help users around the world easily find and retrieve the biological and geographical data and information they need. The BRD is responsible for developing scientifically and statistically reliable methods and protocols to assess the status and trends of the Nation's biological resources. Scientists also conduct important inventory and monitoring studies to maintain baseline information on these same resources. Research on those species for which the Department of the Interior (DOI) has trust responsibilities (including endangered species and migratory species) involves laboratory and field studies of individual animals and the environments in which they live. Researchboth tactical and strategicis conducted at the BRD's 17 science centers and 81 field stations, 54 Cooperative Fish and Wildlife Research Units in 40 states, and at 11 former Cooperative Park Study Units. Studies encompass fish, birds, mammals, and plants, as well as their ecosystems and the surrounding landscape. Biological Resources Division researchers use a variety of scientific tools in their endeavors to understand the causes of biological and ecological trends. Research results are used by managers to predict environmental changes and to help them take appropriate measures to manage resources effectively. The BRD Geospatial Technology Program facilitates the collection, analysis, and dissemination of data and informat

Information and Technology Report↗

Development of regression equations for the estimation of the magnitude and frequency of floods at rural, unregulated gaged and ungaged streams in Puerto Rico through water year 2017

The methods of computation and estimates of the magnitude of flood flows were updated for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent chance exceedance levels for 91 streamgages on the main island of Puerto Rico by using annual peak-flow data through 2017. Since the previous flood frequency study in 1994, the U.S. Geological Survey has collected additional peak flows at additional streamgages, and Puerto Rico has experienced numerous flood events. This updated study was performed using longer annual peak-flow datasets from more stations to provide more representative equations to predict flood flows. Screening criteria for these streamgages included 10 or more years of annual peak-flow data, unregulated flow, and less than 10 percent impervious drainage area. The magnitude and frequency of floods at selected streamgages in Puerto Rico were estimated using updated methods outlined in Bulletin 17C. The new procedures include a regional skew analysis that incorporates Bayesian regression techniques, the Expected Moments Algorithm to better represent missing record and estimate parameters of the log-Pearson Type III distribution, and the Multiple Grubbs-Beck test for low outlier detection. Regional regression equations were developed to estimate peak-flow statistics at ungaged locations by using selected basin and climatic characteristics as explanatory variables. These variables were determined from digital spatial datasets and geographic information systems by using the most recent data available. Ordinary least-squares regression techniques were used to filter the basin characteristics and determine two separate regions, region 1 (west) and region 2 (east), based on residuals. A generalized least-squares procedure was used to account for cross-correlation of sites and develop the final set of equations that have drainage area as the only explanatory variable. The average standard errors of prediction ranged from 18.7 to 46.7 percent in region 1 and 33.4 to 57.6 percent in region 2 for all annual exceedance probabilities (AEPs) examined. The updated statistics showed a greater accuracy of prediction when compared to those from the previous study using drainage area as the only explanatory variable for all AEPs examined in region 1 and the 0.01 and 0.002 AEP flows for region 2. When compared to equations developed in the previous study that have drainage area, mean annual rainfall, and (or) depth-to-rock as explanatory variables, the updated statistics show a greater accuracy of prediction in region 1 at AEP flows of 0.02 and lower (that is, higher flows). Those developed for region 2 do not show a greater accuracy of prediction for any AEP flows when compared to the equations having multiple explanatory variables in the previous study. The calculated regression equations, basin characteristics, and at-site statistics will be incorporated into the U.S. Geological Survey web application, StreamStats ( https://streamstats.usgs.gov/ss/ ). This application allows users to select a location on a stream, whether gaged or ungaged, to obtain estimates of basin characteristics and flow statistics.

Puerto Rico↗

Evaluation of some random effects methodology applicable to bird ringing data

Existing models for ring recovery and recapture data analysis treat temporal variations in annual survival probability (S) as fixed effects. Often there is no explainable structure to the temporal variation in S1,..., Sk; random effects can then be a useful model: Si = E(S) + ??i. Here, the temporal variation in survival probability is treated as random with average value E(??2) = ??2. This random effects model can now be fit in program MARK. Resultant inferences include point and interval estimation for process variation, ??2, estimation of E(S) and var (E??(S)) where the latter includes a component for ??2 as well as the traditional component for v??ar(S??\S??). Furthermore, the random effects model leads to shrinkage estimates, Si, as improved (in mean square error) estimators of Si compared to the MLE, S??i, from the unrestricted time-effects model. Appropriate confidence intervals based on the Si are also provided. In addition, AIC has been generalized to random effects models. This paper presents results of a Monte Carlo evaluation of inference performance under the simple random effects model. Examined by simulation, under the simple one group Cormack-Jolly-Seber (CJS) model, are issues such as bias of ??s2, confidence interval coverage on ??2, coverage and mean square error comparisons for inference about Si based on shrinkage versus maximum likelihood estimators, and performance of AIC model selection over three models: Si ??? S (no effects), Si = E(S) + ??i (random effects), and S1,..., Sk (fixed effects). For the cases simulated, the random effects methods performed well and were uniformly better than fixed effects MLE for the Si.

Journal of Applied Statistics↗

Modeling participation duration, with application to the North American Breeding Bird Survey

We consider “participation histories,” binary sequences consisting of alternating finite sequences of 1s and 0s, ending with an infinite sequence of 0s. Our work is motivated by a study of observer tenure in the North American Breeding Bird Survey (BBS). In our analysis, j indexes an observer’s years of service and X j is an indicator of participation in the survey; 0s interspersed among 1s correspond to years when observers did not participate, but subsequently returned to service. Of interest is the observer’s duration D = max { j : X j = 1}. Because observed records X = ( X 1 , X 2 ,..., X n ) 1 are of finite length, all that we can directly infer about duration is that D ⩾ max { j ⩽ n : X j = 1}; model-based analysis is required for inference about D . We propose models in which lengths of 0s and 1s sequences have distributions determined by the index j at which they begin; 0s sequences are infinite with positive probability, an estimable parameter. We found that BBS observers’ lengths of service vary greatly, with 25.3% participating for only a single year, 49.5% serving for 4 or fewer years, and an average duration of 8.7 years, producing an average of 7.7 counts.

Communications in Statistics - Theory and Methods↗

Fire frequency in the Interior Columbia River Basin: Building regional models from fire history data

Fire frequency affects vegetation composition and successional pathways; thus it is essential to understand fire regimes in order to manage natural resources at broad spatial scales. Fire history data are lacking for many regions for which fire management decisions are being made, so models are needed to estimate past fire frequency where local data are not yet available. We developed multiple regression models and tree-based (classification and regression tree, or CART) models to predict fire return intervals across the interior Columbia River basin at 1-km resolution, using georeferenced fire history, potential vegetation, cover type, and precipitation databases. The models combined semiqualitative methods and rigorous statistics. The fire history data are of uneven quality; some estimates are based on only one tree, and many are not cross-dated. Therefore, we weighted the models based on data quality and performed a sensitivity analysis of the effects on the models of estimation errors that are due to lack of cross-dating. The regression models predict fire return intervals from 1 to 375 yr for forested areas, whereas the tree-based models predict a range of 8 to 150 yr. Both types of models predict latitudinal and elevational gradients of increasing fire return intervals. Examination of regional-scale output suggests that, although the tree-based models explain more of the variation in the original data, the regression models are less likely to produce extrapolation errors. Thus, the models serve complementary purposes in elucidating the relationships among fire frequency, the predictor variables, and spatial scale. The models can provide local managers with quantitative information and provide data to initialize coarse-scale fire-effects models, although predictions for individual sites should be treated with caution because of the varying quality and uneven spatial coverage of the fire history database. The models also demonstrate the integration of qualitative and quantitative methods when requisite data for fully quantitative models are unavailable. They can be tested by comparing new, independent fire history reconstructions against their predictions and can be continually updated, as better fire history data become available.

California, Idaho, Montana, Nevada, Oregon, Utah, ↗

Characterization of ambient groundwater quality within a statewide, fixed-station monitoring network in Pennsylvania, 2015–19

Pennsylvania leads the Nation in the number of individuals that use groundwater for private domestic water supply; more than 3 million rural and suburban Pennsylvania residents rely on private domestic supplies for drinking water. These supplies are not regulated nor routinely monitored; thus relevant groundwater-quality information is not widely available. The U.S. Geological Survey (USGS), in cooperation with the Pennsylvania Department of Environmental Protection (PaDEP) Safe Drinking Water Bureau, established a statewide, fixed-station ambient groundwater quality network in 2015. The goals for the Pennsylvania Groundwater Monitoring Network (GWMN) include characterizing ambient groundwater quality conditions in rural areas of the State and documenting potential changes in conditions over time. Seventeen wells were selected for monitoring at 6-month intervals beginning in 2015. Since then, several wells have been added to the GWMN, bringing the total number of wells sampled in the fall of 2019 to 28. Routinely monitored constituents included physical characteristics and chemical concentrations in filtered and unfiltered samples (major and trace elements, nutrients, and organic compounds). Samples for volatile organic compounds (VOCs), radionuclides, and dissolved hydrocarbon gases were collected during the first sampling event at each well. To offer insights on the quality of groundwater used for domestic supply in Pennsylvania, summary statistics for the 221 GWMN samples collected during 2015–19 are compared to U.S. Environmental Protection Agency (EPA) drinking-water standards, which are applicable to public water supplies. Results show that samples across the GWMN generally meet drinking-water standards for inorganic and organic constituents; however, a percentage of samples had concentrations that exceeded maximum contaminant level (MCL) thresholds for nitrate (3 percent) and secondary maximum contaminant level (SMCL) thresholds for iron (32 percent), manganese (36 percent), and aluminum (5 percent). Radon-222 activities, which were sampled only during the initial visit to a well, exceeded the lower proposed drinking water standard of 300 picocuries per liter (pCi/L) in 64 percent of wells in the GWMN; additionally, 7 percent of wells exceeded the higher proposed standard of 4,000 pCi/L. There were no exceedances for VOCs, but one well had a tribromomethane detection. Three wells had detectable concentrations of methane, with one sample exceeding the Pennsylvania action level of 7 milligrams per liter (mg/L). The pH and dissolved oxygen concentrations varied widely across the GWMN and were correlated with dissolved metal concentrations and other chemical characteristics of groundwater samples. Considering all samples collected for the study, the pH ranged from 4.2 to 8.3; 42 percent of pH values were either above or below the SMCL range of 6.5–8.5. The highest pH values resulted from contamination of loose grout used in the construction of one well and decreased to levels consistent with other wells in the vicinity after repeated sampling rounds. Dissolved oxygen (DO), which ranged from 0 to 13.9 mg/L, influences the mobility and prevalence of constituents with variable oxidation state, including iron, manganese, and nitrogen species. Samples with acidic pH (less than 6.5) and (or) low DO had the highest concentrations of manganese and iron, whereas those with neutral to alkaline pH values had the highest concentrations of calcium, magnesium, sodium, and other major ions. Analysis of major ions indicates that calcium/bicarbonate water types are the most common, with a few characterized as calcium/chloride or sodium/chloride, and most others as mixed water types including calcium-magnesium/bicarbonate, sodium-magnesium/bicarbonate, and sodium/bicarbonate-chloride. Nonparametric statistical methods were used to evaluate the data for spatial and temporal trends. A principal components analysis (PCA) model developed with ranked data values for the entire network resulted in three components, (1) dissolved solids, (2) redox, and (3) sodium-chloride, which explained 74.5 percent of variance in the dataset. On the basis of individual contributions to the PCA, certain wells were identified through hierarchical cluster analysis that shared relevant water-quality characteristics. The spatial distribution of sampling locations and the temporal trends of constituent concentrations indicate that hydrogeologic setting and topographic position as defined in the PCA model are important factors affecting the spatial and temporal patterns of groundwater quality in the GWMN.

Pennsylvania↗

Animal movement models for migratory individuals and groups

Animals often exhibit changes in their behaviour during migration. Telemetry data provide a way to observe geographic position of animals over time, but not necessarily changes in the dynamics of the movement process. Continuous‐time models allow for statistical predictions of the trajectory in the presence of measurement error and during periods when the telemetry device did not record the animal's position. However, continuous‐time models capable of mimicking realistic trajectories with sufficient detail are computationally challenging to fit to large datasets. Furthermore, basic continuous‐time model specifications (e.g. Brownian motion) lack realism in their ability to capture nonstationary dynamics. We present a unified class of animal movement models that are computationally efficient and provide a suite of approaches for accommodating nonstationarity in continuous trajectories due to migration and interactions among individuals. Our approach uses process convolutions to allow for flexibility in the movement process while facilitating implementation and incorporating location uncertainty. We show how to nest convolution models to incorporate interactions among migrating individuals to account for nonstationarity and provide inference about dynamic migratory networks. We demonstrate these approaches in two case studies involving migratory birds. Specifically, we used process convolution models with temporal deformation to account for heterogeneity in individual greater white‐fronted goose migrations in Europe and Iceland, and we used nested process convolutions to model dynamic migratory networks in sandhill cranes in North America. The approach we present accounts for various forms of temporal heterogeneity in animal movement and is not limited to migratory applications. Furthermore, our models rely on well‐established principles for modelling‐dependent data and leverage modern approaches for modelling dynamic networks to help explain animal movement and social interaction.

Methods in Ecology and Evolution↗

Total Phosphorus Loads for Selected Tributaries to Sebago Lake, Maine

The streamflow and water-quality datacollection networks of the Portland Water District (PWD) and the U.S. Geological Survey (USGS) as of February 2000 were analyzed in terms of their applicability for estimating total phosphorus loads for selected tributaries to Sebago Lake in southern Maine. The long-term unit-area mean annual flows for the Songo River and for small, ungaged tributaries are similar to the long-term unit-area mean annual flows for the Crooked River and other gaged tributaries to Sebago Lake, based on a regression equation that estimates mean annual streamflows in Maine. Unit-area peak streamflows of Sebago Lake tributaries can be quite different, based on a regression equation that estimates peak streamflows for Maine. Crooked River had a statistically significant positive relation (Kendall's Tau test, p=0.0004) between streamflow and total phosphorus concentration. Panther Run had a statistically significant negative relation (p=0.0015). Significant positive relations may indicate contributions from nonpoint sources or sediment resuspension, whereas significant negative relations may indicate dilution of point sources. Total phosphorus concentrations were significantly larger in the Crooked River than in the Songo River (Wilcoxon rank-sum test, p<0.0001). Evidence was insufficient, however, to indicate that phosphorus concentrations from medium-sized drainage basins, at a significance level of 0.05, were different from each other or that concentrations in small-sized drainage basins were different from each other (Kruskal-Wallis test, p= 0.0980, 0.1265). All large- and medium-sized drainage basins were sampled for total phosphorus approximately monthly. Although not all small drainage basins were sampled, they may be well represented by the small drainage basins that were sampled. If the tributaries gaged by PWD had adequate streamflow data, the current PWD tributary monitoring program would probably produce total phosphorus loading data that would represent all gaged and ungaged tributaries to Sebago Lake. Outside the PWD tributary-monitoring program, the largest ungaged tributary to Sebago Lake contains 1.5 percent of the area draining to the lake. In the absence of unique point or nonpoint sources of phosphorus, ungaged tributaries are unlikely to have total phosphorus concentrations that differ significantly from those in the small tributaries that have concentration data. The regression method, also known as the rating-curve method, was used to estimate the annual total phosphorus load for Crooked River, Northwest River, and Rich Mill Pond Outlet for water years 1996-98. The MOVE.1 method was used to estimate daily streamflows for the regression method at Northwest River and Rich Mill Pond Outlet, where streamflows were not continuously monitored. An averaging method also was used to compute annual loads at the three sites. The difference between the regression estimate and the averaging estimate for each of the three tributaries was consistent with what was expected from previous studies.

Water-Resources Investigations Report↗

Occam's shadow: levels of analysis in evolutionary ecology - where to next?

Evolutionary ecology is the study of evolutionary processes, and the ecological conditions that influence them. A fundamental paradigm underlying the study of evolution is natural selection. Although there are a variety of operational definitions for natural selection in the literature, perhaps the most general one is that which characterizes selection as the process whereby heritable variation in fitness associated with variation in one or more phenotypic traits leads to intergenerational change in the frequency distribution of those traits. The past 20 years have witnessed a marked increase in the precision and reliability of our ability to estimate one or more components of fitness and characterize natural selection in wild populations, owing particularly to significant advances in methods for analysis of data from marked individuals. In this paper, we focus on several issues that we believe are important considerations for the application and development of these methods in the context of addressing questions in evolutionary ecology. First, our traditional approach to estimation often rests upon analysis of aggregates of individuals, which in the wild may reflect increasingly non-random (selected) samples with respect to the trait(s) of interest. In some cases, analysis at the aggregate level, rather than the individual level, may obscure important patterns. While there are a growing number of analytical tools available to estimate parameters at the individual level, and which can cope (to varying degrees) with progressive selection of the sample, the advent of new methods does not reduce the need to consider carefully the appropriate level of analysis in the first place. Estimation should be motivated a priori by strong theoretical analysis. Doing so provides clear guidance, in terms of both (i) assisting in the identification of realistic and meaningful models to include in the candidate model set, and (ii) providing the appropriate context under which the results are interpreted. Second, while it is true that selection (as defined) operates at the level of the individual, the selection gradient is often (if not generally) conditional on the abundance of the population. As such, it may be important to consider estimating transition rates conditional on both the parameter values of the other individuals in the population (or at least their distribution), and population abundance. This will undoubtedly pose a considerable challenge, for both single- and multi-strata applications. It will also require renewed consideration of the estimation of abundance, especially for open populations. Thirdly, selection typically operates on dynamic, individually varying traits. Such estimation may require characterizing fitness in terms of individual plasticity in one or more state variables, constituting analysis of the norms of reaction of individuals to variable environments. This can be quite complex, especially for traits that are under facultative control. Recent work has indicated that the pattern of selection on such traits is conditional on the relative rates of movement among and frequency of spatially heterogeneous habitats, suggesting analyses of evolution of life histories in open populations can be misleading in some cases.

Journal of Applied Statistics↗

Methods for estimating annual exceedance-probability discharges and largest recorded floods for unregulated streams in rural Missouri

Regression analysis techniques were used to develop a set of equations for rural ungaged stream sites for estimating discharges with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities, which are equivalent to annual flood-frequency recurrence intervals of 2, 5, 10, 25, 50, 100, 200, and 500 years, respectively. Basin and climatic characteristics were computed using geographic information software and digital geospatial data. A total of 35 characteristics were computed for use in preliminary statewide and regional regression analyses. Annual exceedance-probability discharge estimates were computed for 278 streamgages by using the expected moments algorithm to fit a log-Pearson Type III distribution to the logarithms of annual peak discharges for each streamgage using annual peak-discharge data from water year 1844 to 2012. Low-outlier and historic information were incorporated into the annual exceedance-probability analyses, and a generalized multiple Grubbs-Beck test was used to detect potentially influential low floods. Annual peak flows less than a minimum recordable discharge at a streamgage were incorporated into the at-site station analyses. An updated regional skew coefficient was determined for the State of Missouri using Bayesian weighted least-squares/generalized least squares regression analyses. At-site skew estimates for 108 long-term streamgages with 30 or more years of record and the 35 basin characteristics defined for this study were used to estimate the regional variability in skew. However, a constant generalized-skew value of -0.30 and a mean square error of 0.14 were determined in this study. Previous flood studies indicated that the distinct physical features of the three physiographic provinces have a pronounced effect on the magnitude of flood peaks. Trends in the magnitudes of the residuals from preliminary statewide regression analyses from previous studies confirmed that regional analyses in this study were similar and related to three primary physiographic provinces. The final regional regression analyses resulted in three sets of equations. For Regions 1 and 2, the basin characteristics of drainage area and basin shape factor were statistically significant. For Region 3, because of the small amount of data from streamgages, only drainage area was statistically significant. Average standard errors of prediction ranged from 28.7 to 38.4 percent for flood region 1, 24.1 to 43.5 percent for flood region 2, and 25.8 to 30.5 percent for region 3. The regional regression equations are only applicable to stream sites in Missouri with flows not significantly affected by regulation, channelization, backwater, diversion, or urbanization. Basins with about 5 percent or less impervious area were considered to be rural. Applicability of the equations are limited to the basin characteristic values that range from 0.11 to 8,212.38 square miles (mi 2 ) and basin shape from 2.25 to 26.59 for Region 1, 0.17 to 4,008.92 mi 2 and basin shape 2.04 to 26.89 for Region 2, and 2.12 to 2,177.58 mi 2 for Region 3. Annual peak data from streamgages were used to qualitatively assess the largest floods recorded at streamgages in Missouri since the 1915 water year. Based on existing streamgage data, the 1983 flood event was the largest flood event on record since 1915. The next five largest flood events, in descending order, took place in 1993, 1973, 2008, 1994 and 1915. Since 1915, five of six of the largest floods on record occurred from 1973 to 2012.

Missouri↗

Spatially explicit models for inference about density in unmarked or partially marked populations

Recently developed spatial capture–recapture (SCR) models represent a major advance over traditional capture–recapture (CR) models because they yield explicit estimates of animal density instead of population size within an unknown area. Furthermore, unlike nonspatial CR methods, SCR models account for heterogeneity in capture probability arising from the juxtaposition of animal activity centers and sample locations. Although the utility of SCR methods is gaining recognition, the requirement that all individuals can be uniquely identified excludes their use in many contexts. In this paper, we develop models for situations in which individual recognition is not possible, thereby allowing SCR concepts to be applied in studies of unmarked or partially marked populations. The data required for our model are spatially referenced counts made on one or more sample occasions at a collection of closely spaced sample units such that individuals can be encountered at multiple locations. Our approach includes a spatial point process for the animal activity centers and uses the spatial correlation in counts as information about the number and location of the activity centers. Camera-traps, hair snares, track plates, sound recordings, and even point counts can yield spatially correlated count data, and thus our model is widely applicable. A simulation study demonstrated that while the posterior mean exhibits frequentist bias on the order of 5–10% in small samples, the posterior mode is an accurate point estimator as long as adequate spatial correlation is present. Marking a subset of the population substantially increases posterior precision and is recommended whenever possible. We applied our model to avian point count data collected on an unmarked population of the northern parula (Parula americana) and obtained a density estimate (posterior mode) of 0.38 (95% CI: 0.19–1.64) birds/ha. Our paper challenges sampling and analytical conventions in ecology by demonstrating that neither spatial independence nor individual recognition is needed to estimate population density—rather, spatial dependence can be informative about individual distribution and density.

Annals of Applied Statistics↗

Occupancy modeling species–environment relationships with non‐ignorable survey designs

Statistical models supporting inferences about species occurrence patterns in relation to environmental gradients are fundamental to ecology and conservation biology. A common implicit assumption is that the sampling design is ignorable and does not need to be formally accounted for in analyses. The analyst assumes data are representative of the desired population and statistical modeling proceeds. However, if data sets from probability and non‐probability surveys are combined or unequal selection probabilities are used, the design may be non‐ignorable. We outline the use of pseudo‐maximum likelihood estimation for site‐occupancy models to account for such non‐ignorable survey designs. This estimation method accounts for the survey design by properly weighting the pseudo‐likelihood equation. In our empirical example, legacy and newer randomly selected locations were surveyed for bats to bridge a historic statewide effort with an ongoing nationwide program. We provide a worked example using bat acoustic detection/non‐detection data and show how analysts can diagnose whether their design is ignorable. Using simulations we assessed whether our approach is viable for modeling data sets composed of sites contributed outside of a probability design. Pseudo‐maximum likelihood estimates differed from the usual maximum likelihood occupancy estimates for some bat species. Using simulations we show the maximum likelihood estimator of species–environment relationships with non‐ignorable sampling designs was biased, whereas the pseudo‐likelihood estimator was design unbiased. However, in our simulation study the designs composed of a large proportion of legacy or non‐probability sites resulted in estimation issues for standard errors. These issues were likely a result of highly variable weights confounded by small sample sizes (5% or 10% sampling intensity and four revisits). Aggregating data sets from multiple sources logically supports larger sample sizes and potentially increases spatial extents for statistical inferences. Our results suggest that ignoring the mechanism for how locations were selected for data collection (e.g., the sampling design) could result in erroneous model‐based conclusions. Therefore, in order to ensure robust and defensible recommendations for evidence‐based conservation decision‐making, the survey design information in addition to the data themselves must be available for analysts. Details for constructing the weights used in estimation and code for implementation are provided.

Ecological Applications↗

Factors Affecting Groundwater Quality Used for Domestic Supply in Marcellus Shale Region of North-Central and North-East Pennsylvania, USA

Factors affecting groundwater quality used for domestic supply within the Marcellus Shale footprint in north-central and north-east Pennsylvania are identified using a combination of spatial, statistical, and geochemical modeling. Untreated groundwater, sampled during 2011–2017 from 472 domestic wells within the study area, exhibited wide ranges in pH (4.5–9.3), total dissolved solids (TDS, 22–1960 mg/L), sodium (0.3–760 mg/L), chloride (0.3–1020 mg/L), bromide (<0.01–8.6 mg/L), and methane (<0.001–77 mg/L). The wells had depths ranging from 10 to 394 m; 69.5 percent were completed in sandstone bedrock, 19.3 percent in shale, 4.2 percent in siltstone, 4 percent in carbonate, and 3 percent in unconsolidated alluvial or glacial deposits. Groundwater quality in the Delaware River watershed, in the eastern part of the study area where Marcellus gas has not been developed, was similar to that in the Susquehanna, Allegheny, and Genesee River watersheds in the western part of the study area where natural gas production from Marcellus Shale has been ongoing since 2008. Most groundwaters were calcium/bicarbonate type with near-neutral pH; approximately 10 percent were sodium/bicarbonate and 1 percent were sodium/chloride types. Sodium-enriched waters, which were mostly from shale and siltstone aquifers, had the greatest frequency of elevated pH (>8.5) and elevated concentrations of TDS (>250 mg/L), bromide (>0.15 mg/L), methane (>7.0 mg/L), and lithium (>60 μg/L). Geochemical models indicate these characteristics could result from progressive mineral dissolution combined with cation exchange, plus mixing with locally important salinity sources, including as much as 0.7 percent Appalachian Basin brine and/or road-deicing salt. Multivariate correlation models suggest the observed variability in methane concentrations may be attributed to several environmental factors, such as geochemical evolution along groundwater flow paths, redox conditions, and/or mixing with saline groundwater or brine. Most samples having elevated methane were from shale aquifers, which were mainly in the Susquehanna River basin and had the greatest density of gas wells compared to other lithologies . Samples having elevated methane were also observed in the Delaware River watershed and other areas outside gas development. Isotopic compositions of methane for a subset of 39 samples (selected because of elevated methane) and relatively high ratios of methane to ethane in those samples indicated methane could be derived from microbial gas mixed with thermogenic gas that may have undergone degradation and/or fractionation during migration. The methods used in this study could be broadly applicable to understanding major factors affecting groundwater quality, particularly for explaining variations in ionic composition with pH and identifying sources of salinity and associated constituents (e.g. sodium, chloride, bromide, lithium, methane) that may have geogenic or anthropogenic origins.

Pennsylvania↗