USGS Science⌕ Search

SEARCH · USGS Science

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Methods for estimating annual exceedance-probability discharges and largest recorded floods for unregulated streams in rural Missouri

Regression analysis techniques were used to develop a set of equations for rural ungaged stream sites for estimating discharges with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities, which are equivalent to annual flood-frequency recurrence intervals of 2, 5, 10, 25, 50, 100, 200, and 500 years, respectively. Basin and climatic characteristics were computed using geographic information software and digital geospatial data. A total of 35 characteristics were computed for use in preliminary statewide and regional regression analyses. Annual exceedance-probability discharge estimates were computed for 278 streamgages by using the expected moments algorithm to fit a log-Pearson Type III distribution to the logarithms of annual peak discharges for each streamgage using annual peak-discharge data from water year 1844 to 2012. Low-outlier and historic information were incorporated into the annual exceedance-probability analyses, and a generalized multiple Grubbs-Beck test was used to detect potentially influential low floods. Annual peak flows less than a minimum recordable discharge at a streamgage were incorporated into the at-site station analyses. An updated regional skew coefficient was determined for the State of Missouri using Bayesian weighted least-squares/generalized least squares regression analyses. At-site skew estimates for 108 long-term streamgages with 30 or more years of record and the 35 basin characteristics defined for this study were used to estimate the regional variability in skew. However, a constant generalized-skew value of -0.30 and a mean square error of 0.14 were determined in this study. Previous flood studies indicated that the distinct physical features of the three physiographic provinces have a pronounced effect on the magnitude of flood peaks. Trends in the magnitudes of the residuals from preliminary statewide regression analyses from previous studies confirmed that regional analyses in this study were similar and related to three primary physiographic provinces. The final regional regression analyses resulted in three sets of equations. For Regions 1 and 2, the basin characteristics of drainage area and basin shape factor were statistically significant. For Region 3, because of the small amount of data from streamgages, only drainage area was statistically significant. Average standard errors of prediction ranged from 28.7 to 38.4 percent for flood region 1, 24.1 to 43.5 percent for flood region 2, and 25.8 to 30.5 percent for region 3. The regional regression equations are only applicable to stream sites in Missouri with flows not significantly affected by regulation, channelization, backwater, diversion, or urbanization. Basins with about 5 percent or less impervious area were considered to be rural. Applicability of the equations are limited to the basin characteristic values that range from 0.11 to 8,212.38 square miles (mi 2 ) and basin shape from 2.25 to 26.59 for Region 1, 0.17 to 4,008.92 mi 2 and basin shape 2.04 to 26.89 for Region 2, and 2.12 to 2,177.58 mi 2 for Region 3. Annual peak data from streamgages were used to qualitatively assess the largest floods recorded at streamgages in Missouri since the 1915 water year. Based on existing streamgage data, the 1983 flood event was the largest flood event on record since 1915. The next five largest flood events, in descending order, took place in 1993, 1973, 2008, 1994 and 1915. Since 1915, five of six of the largest floods on record occurred from 1973 to 2012.

Missouri↗

Spatially explicit models for inference about density in unmarked or partially marked populations

Recently developed spatial capture–recapture (SCR) models represent a major advance over traditional capture–recapture (CR) models because they yield explicit estimates of animal density instead of population size within an unknown area. Furthermore, unlike nonspatial CR methods, SCR models account for heterogeneity in capture probability arising from the juxtaposition of animal activity centers and sample locations. Although the utility of SCR methods is gaining recognition, the requirement that all individuals can be uniquely identified excludes their use in many contexts. In this paper, we develop models for situations in which individual recognition is not possible, thereby allowing SCR concepts to be applied in studies of unmarked or partially marked populations. The data required for our model are spatially referenced counts made on one or more sample occasions at a collection of closely spaced sample units such that individuals can be encountered at multiple locations. Our approach includes a spatial point process for the animal activity centers and uses the spatial correlation in counts as information about the number and location of the activity centers. Camera-traps, hair snares, track plates, sound recordings, and even point counts can yield spatially correlated count data, and thus our model is widely applicable. A simulation study demonstrated that while the posterior mean exhibits frequentist bias on the order of 5–10% in small samples, the posterior mode is an accurate point estimator as long as adequate spatial correlation is present. Marking a subset of the population substantially increases posterior precision and is recommended whenever possible. We applied our model to avian point count data collected on an unmarked population of the northern parula (Parula americana) and obtained a density estimate (posterior mode) of 0.38 (95% CI: 0.19–1.64) birds/ha. Our paper challenges sampling and analytical conventions in ecology by demonstrating that neither spatial independence nor individual recognition is needed to estimate population density—rather, spatial dependence can be informative about individual distribution and density.

Annals of Applied Statistics↗

Occupancy modeling species–environment relationships with non‐ignorable survey designs

Statistical models supporting inferences about species occurrence patterns in relation to environmental gradients are fundamental to ecology and conservation biology. A common implicit assumption is that the sampling design is ignorable and does not need to be formally accounted for in analyses. The analyst assumes data are representative of the desired population and statistical modeling proceeds. However, if data sets from probability and non‐probability surveys are combined or unequal selection probabilities are used, the design may be non‐ignorable. We outline the use of pseudo‐maximum likelihood estimation for site‐occupancy models to account for such non‐ignorable survey designs. This estimation method accounts for the survey design by properly weighting the pseudo‐likelihood equation. In our empirical example, legacy and newer randomly selected locations were surveyed for bats to bridge a historic statewide effort with an ongoing nationwide program. We provide a worked example using bat acoustic detection/non‐detection data and show how analysts can diagnose whether their design is ignorable. Using simulations we assessed whether our approach is viable for modeling data sets composed of sites contributed outside of a probability design. Pseudo‐maximum likelihood estimates differed from the usual maximum likelihood occupancy estimates for some bat species. Using simulations we show the maximum likelihood estimator of species–environment relationships with non‐ignorable sampling designs was biased, whereas the pseudo‐likelihood estimator was design unbiased. However, in our simulation study the designs composed of a large proportion of legacy or non‐probability sites resulted in estimation issues for standard errors. These issues were likely a result of highly variable weights confounded by small sample sizes (5% or 10% sampling intensity and four revisits). Aggregating data sets from multiple sources logically supports larger sample sizes and potentially increases spatial extents for statistical inferences. Our results suggest that ignoring the mechanism for how locations were selected for data collection (e.g., the sampling design) could result in erroneous model‐based conclusions. Therefore, in order to ensure robust and defensible recommendations for evidence‐based conservation decision‐making, the survey design information in addition to the data themselves must be available for analysts. Details for constructing the weights used in estimation and code for implementation are provided.

Ecological Applications↗

Factors Affecting Groundwater Quality Used for Domestic Supply in Marcellus Shale Region of North-Central and North-East Pennsylvania, USA

Factors affecting groundwater quality used for domestic supply within the Marcellus Shale footprint in north-central and north-east Pennsylvania are identified using a combination of spatial, statistical, and geochemical modeling. Untreated groundwater, sampled during 2011–2017 from 472 domestic wells within the study area, exhibited wide ranges in pH (4.5–9.3), total dissolved solids (TDS, 22–1960 mg/L), sodium (0.3–760 mg/L), chloride (0.3–1020 mg/L), bromide (<0.01–8.6 mg/L), and methane (<0.001–77 mg/L). The wells had depths ranging from 10 to 394 m; 69.5 percent were completed in sandstone bedrock, 19.3 percent in shale, 4.2 percent in siltstone, 4 percent in carbonate, and 3 percent in unconsolidated alluvial or glacial deposits. Groundwater quality in the Delaware River watershed, in the eastern part of the study area where Marcellus gas has not been developed, was similar to that in the Susquehanna, Allegheny, and Genesee River watersheds in the western part of the study area where natural gas production from Marcellus Shale has been ongoing since 2008. Most groundwaters were calcium/bicarbonate type with near-neutral pH; approximately 10 percent were sodium/bicarbonate and 1 percent were sodium/chloride types. Sodium-enriched waters, which were mostly from shale and siltstone aquifers, had the greatest frequency of elevated pH (>8.5) and elevated concentrations of TDS (>250 mg/L), bromide (>0.15 mg/L), methane (>7.0 mg/L), and lithium (>60 μg/L). Geochemical models indicate these characteristics could result from progressive mineral dissolution combined with cation exchange, plus mixing with locally important salinity sources, including as much as 0.7 percent Appalachian Basin brine and/or road-deicing salt. Multivariate correlation models suggest the observed variability in methane concentrations may be attributed to several environmental factors, such as geochemical evolution along groundwater flow paths, redox conditions, and/or mixing with saline groundwater or brine. Most samples having elevated methane were from shale aquifers, which were mainly in the Susquehanna River basin and had the greatest density of gas wells compared to other lithologies . Samples having elevated methane were also observed in the Delaware River watershed and other areas outside gas development. Isotopic compositions of methane for a subset of 39 samples (selected because of elevated methane) and relatively high ratios of methane to ethane in those samples indicated methane could be derived from microbial gas mixed with thermogenic gas that may have undergone degradation and/or fractionation during migration. The methods used in this study could be broadly applicable to understanding major factors affecting groundwater quality, particularly for explaining variations in ionic composition with pH and identifying sources of salinity and associated constituents (e.g. sodium, chloride, bromide, lithium, methane) that may have geogenic or anthropogenic origins.

Pennsylvania↗

Evaluation of some software measuring displacements using GPS in real-time

For the past decade, the USGS has been monitoring deformation at various locations in the western United States using continuous GPS. The main focus of these measurements are estimates of displacement averaged over one day. Essentially, these consist of recording at 30 seconds intervals the carrier-frequency phase-data (equivalent to travel-time) between a GPS receiver and the GPS satellite network. In turn, these observations, which are converted to pseudo—ranges, are processed using one of the “research grade” programs (GIPSY, Zumberge et al., or GAMIT, wwwgpsg.mit.edu/~simon/gtgk) to estimate the position of the GPS receiver averaged over 24 hours. However, it is possible and desirable to estimate the position of the receiver (actually the antenna) more frequently and to do this within a few seconds of the time actual measurement (known as real-time). A recent example, the 2004 Magnitude 6, Parkfield, California earthquake, demonstrated that having GPS estimates of position more frequently than simply a daily average is required if one requires discrimination between co-seismic and post-seismic deformation (Langbein et al., 2006). The high-rate estimates of position obtained at Parkfield show that post-seismic deformation started less than one-hour after the mainshock and that this deformation was roughly the same magnitude as the co-seismic deformation. The high-rate solutions for Parkfield were done by others including Yehuda Bock at UCSD and Kristine Larson at U. of Colorado, but not the USGS. The Parkfield experience points out the need for an in-house capability by the USGS to be able to accurately measure co-seismic displacements and other rapid, deformation signals using GPS. This applies to both the Earthquake and Volcano Hazard programs. Although at many locations where we monitor deformation, we have strainmeters and tiltmeters in addition to GPS which, in principle, are far more sensitive to rapid deformation over periods of less than a day (Langbein and Bock, 2004). But, not all locales include strain and tiltmeters. Thus, having the capability to extract signals with periods of less than a day is desirable since the distribution of GPS is more extensive than strain and tilt. At both Parkfield and Long Valley, the USGS has been using other software packages to process the GPS data at sub-daily intervals and in real-time. The underlying goal of these types of measurements is to detect any deformation event as it evolves; the 24 hour processing might not provide timely results if such a deformation event is precursory to a geologic hazard (an earthquake for Parkfield and either a volcanic event or an earthquake for Long Valley). In Long Valley, We use the software package called 3DTracker (http://www.3dtracker.com, http://www.condorearth.com) to estimate the changes of in position of a remote site relative to a “fixed” site. The 3DTracker software uses double difference GPS code measurements and receiversatellite-time triple differences from one epoch to the next of the GPS phase data (a proxy for travel-time measurements) and employs a Kalman filter to obtain stability in the estimate of position. That is, the estimate of the current position depends upon the estimate of the prior position. Hence, a time series of position looks fairly smooth depending upon the coefficient selected for the Kalman filter. With triple differences, the sometimes troublesome initial integer cycle ambiguity terms cancel (number of wavelengths between the receiver and each satellite), but only the incremental change in position is calculated. This triple difference Kalman filter solution is slow to converge and less accurate than a double difference (e.g., RTD, Track) solution, but it is robust and computationally efficient (Remondi and Brown, 2000). 3D-Tracker allows use of various single-frequency and dual-frequency GPS phase and code observables including the ionospheric-free combinations (known as LC or L3 and P(L3)) formed from an linear combination of the L1 and L2 carrier phase and code data. The lowest noise observable is the L1 carrier, but it is biased by ionospheric refraction that has amplitudes of about 1 to 10 ppm. This results in a systematic scale error in the relative positions. The L3 phase noise is about 3 times greater than the L1 phase noise, but it is generally used to solve for all but the shortest baselines (< 5 km). In addition, the software does output the position changes is a standard format that can be used for other analysis. At Parkfield, we use the software package called RTD (http://www.geodetics.com). The RTD software has been described in the literature (Bock et al., 2000) but basically, it estimates the position without the constraint of a Kalman filter. It uses double differences (in our studies the LC or ionospheric free observable is used) and the integer ambiguities are resolved independently for each 1-second measurement; Most GPS software that use double-differences require several epochs of measurements to resolve the integer ambiguities. The data files use a proprietary format and can not be read by me or others; rather, Yehuda Bock at UCSD (and author of RTD) translates these files into a standard format that can be read by me. Recently, Tom Herring of MIT has modified the GAMIT software to process kinematically GPS data (www-gpsg.mit.edu/~simon/gtgk/tutorial/Lecture_13.pdf). At this time, the software, known as TRACK, does not process the observations in real-time. Consequently, the latency between the time of the observation and the time when a position estimate is available depends upon the frequency that the data are downloaded and the speed of actually processing the observations; there could be a delay of an hour or two before the a position estimates are available. Unlike RTD and 3DTracker, TRACK comes with GAMIT (which is distributed freely) and is currently operating in a test mode at the USGS office in Pasadena. The LC or ionosphere free observable is used in our TRACK solutions. JPL has a version of their GIPSY software called “Real-time GIPSY (RTG)” (gipsy.jpl.nasa.gov/orms/rtg), which, like TRACK, can process the pseudo-range data “off—line”. However, this software is not freely distributed. Instead, at least one company, NAVCOM, has teamed with JPL to integrate RTG with GPS receivers and telemetry that yields positions in realtime. Kristine Larson of University of Colorado has modified the original GIPSY to estimate positions kinematically. Again, like TRACK, the positions are estimated off—line. Much of her research is described in Larson et al. (2003), and Choi et al. (2004). For Long Valley, out of the 17 GPS sites, we monitor 5 baselines within the caldera at 5 second intervals relative to the Bald Mountain site at the edge of the caldera using 3DTracker. The baseline measurement using 3DTracker consists of determination of the 3 dimensional positions of the 5 remote points (GPS receivers) relative to a GPS site at Bald. A second, independent system collects and downloads once a day the 30-second data used for the 24-hour solutions for the 12 sites not monitored with 3DTracker. For the sites monitored with 3DTracker, the pseudo—range data are decimated to 30 seconds and converted to a form used for the 24-hour solutions. Both sets of telemetry employ 900 MHz spread spectrum radios which require line of site between all of the links. The telemetry for the 3DTracker sites require a dedicated radios at each end and intermediate repeaters as needed, while the telemetry required for the other sites use a single master radio, repeaters as needed, and a radio at each remote site. (The 5 sites being monitored with 3DTracker require 13 radios.) At Parkfield, RTD is used to measure the position changes all 12 baselines at 1 second intervals relative to a site, Pomm, adjacent to the San Andreas Fault. The complete RTD package (hardware and software) collects all of the data and determines the position of each site relative to Pomm. In addition, the system stores both the 1-second and 30-second pseudo-range data for later downloading which are ultimately used in the 24-hour solutions. To do this, each site has a 2.4 GHz radio and a telemetry buffer. The telemetry buffer holds 24-hours of data (in the event that the telemetry link is broken) and converts the RS232 data stream from the GPS receiver into a form compatible with an IP (Internet protocol) network connection. In contrast with the Long Valley system, the telemetry link for GPS at Parkfield consists of a single radio at each remote sites and a single radio at the central site. Although position estimates are produced within 1-second of the observations, these results are not immediately available because there is no high speed Internet connection to Parkfield. Instead, the data are stored on a removable disk and sent to UCSD once per month. Below, I describe the results of a simple experiment to examine the response of some of these systems to simulated deformation that could be an analogue of a tectonic or volcanic event. In many engineering applications, the system response is tested by inputting a step to the system and measuring the output of the system. Essentially, this is what I've done. The experiment described below moves the GPS antenna from its original position to a new position within 1 second; the software tracks the translation. These measurements were conducted in August 2004 with the RTD software at Parkfield, and twice in Long Valley. The first Long Valley test was conducted in September 2004 using 3DTracker on a single baseline. The test was repeated in September 2005 using 3DTracker on two baselines and, importantly, saving the RINEX files of the data so that the data could be replayed through 3DTracker using other options in the program and, using other software packages including TRACK. In addition, we observed a short-term event at the Three Sisters volcano in Oregon. This event was snow melt at a remote GPS site which gave an apparent 15 cm displacement in vertical in less than one-day. 3DTracker is used to monitor this site, and the event was captured with this software. In addition, with the assistance of others, I got additional estimates of position using other software packages; those results are presented. Finally, the precision of both 3DTracker and RTD are compared using a power spectrum. Those results would suggest that 3DTracker using appropriate Kalman filter coefficients would have better precision than RTD; instead, the lower noise level from 3DTracker is a result of smoothing from the Kalman filter. Given the results described in this report, high-rate GPS is certainly capable of accurately measuring displacements of 1 centimeter with a high degree of statistical confidence. Plotting these results show that the time of the displacement can be visually determined to that of the sampling interval of the data. However, especially with small amplitude signals, any of the software packages can yield erroneous deformation “signals” that are either due excess travel-time of the GPS carrier frequency from multipath or a limitation in the software. Thus, the time series of displacements must be viewed with caution and knowledge of external circumstances that might cause a change in position. The casual reader should continue with the next section describing the methods then jump to the last two sections for the discussion and conclusions. I have made some recommendations there.

Open-File Report↗

Insights and strategic opportunities from the USGS 2024 Per- and Polyfluoroalkyl Substances (PFAS) Interagency Workshop

Introduction In 2021, the U.S. Geological Survey (USGS) published Circular 1490 titled, “Integrated Science for the Study of Perfluoroalkyl and Polyfluoroalkyl Substances (PFAS) in the Environment: A Strategic Science Vision for the U.S. Geological Survey” (Tokranov and others, 2021). Circular 1490 was created to be a resource for USGS scientists prioritizing and planning research related to per- and polyfluoroalkyl substances (PFAS) and to be a guide for developing partnerships with other scientists, State and Federal agencies, and stakeholders engaged in PFAS research and management and mitigation of the environmental and human-health effects of PFAS. This USGS PFAS Strategic Science Vision document was intended to be the foundation for a “living strategic vision,” periodically providing updates on the state of USGS PFAS research, emerging PFAS data gaps and needs, and progress on interagency and stakeholder PFAS partnerships and priorities. To meet this objective, the USGS planned to host an Interagency and Stakeholder PFAS Workshop every 2–3 years. During September 10–12, 2024, the USGS hosted the first Interagency and Stakeholder PFAS Workshop in Reston, Virginia. The Workshop brought together experts from other Federal agencies (U.S. Environmental Protection Agency, National Institute of Environmental Health Sciences, Food and Drug Administration, Department of Defense [Air Force, Army]), State agencies (Washington Fish and Wildlife, Virginia Department of Transportation), and academia (Harvard University, University of Maryland) to address key challenges relating to the measurement and modeling of PFAS and the implications for environmental health. Participants engaged in in-depth discussions centered around six pivotal topics related to PFAS: (1) sampling protocols, methods and interpretation; (2) environmental sources, source apportionment, and occurrence; (3) environmental fate and transport; (4) human and wildlife exposure routes and risk; (5) bioconcentration, bioaccumulation, and biomagnification; and (6) ecotoxicology and effects. Each topic had three breakout sessions. A recurrent theme of workshop discussions was how data on a nationwide scale for PFAS occurrence in various environmental matrices, including air, water, food crops, biota, soil, and streambed sediment could help to advance scientific understanding. Participants noted significant geospatial data gaps, particularly in the midwestern and southern United States and the Pacific Northwest. PFAS data collection tends to be more robust along the eastern seaboard and in California. Participants stressed how enhancing the integration of large and small datasets across various agencies could help to support national scale understanding of PFAS. To address these gaps, attendees suggested leveraging datasets from Federal entities like the USGS and the U.S. Department of Defense, State agencies, and municipal utility services to develop predictive contaminant detection and transport models. Improved coordination between water quality programs and USGS research could help to facilitate access to valuable data, leading to comprehensive databases that inform PFAS point (wastewater treatment plants and landfills) and nonpoint (runoff from land, atmospheric deposition, food packaging) sources, environmental transport mechanisms, environmental detection and concentrations, potential exposure routes, and health effects on different biota, including humans. A specific request was made to develop a map demarking the depth of modern (1953 or later) groundwater, which is susceptible to surface-derived anthropogenic (that is, human-made) contamination, based on tritium-age dating. Emphasis was placed on incorporation of hydrology, groundwater flow paths, groundwater–surface water interactions, and landscape factors in predictive statistical models as a step to improve contaminant source identification and tracking. Molecular fingerprinting approaches garnered attention as techniques to link specific PFAS mixtures detected in a sample to environmental sources and levels in biota (Dávila-Santiago and others, 2022). Integrating data from abiotic (that is, water, soil, and air) and biotic (that is, living organisms) systems identified as a research opportunity. For example, understanding the composition of soils and sediments, which include a mixture of mineral, plant, and animal components, could advance understanding of exposure pathways. The discussions highlighted opportunities to explore and understand the potential redistribution and biotic exposures of PFAS from biosolid and wastewater treatment plant effluent land application practices, in addition to atmospheric releases and discharges from landfill and wastewater treatment plants. Participants identified research gaps surrounding how these sources may contribute to contamination and may affect surrounding ecosystems, including a better definition of anthropogenic background concentrations. Moving forward, the collection of co-occurrence data was noted as a means to improve understanding of complex mixtures and to leverage companion modeling efforts focused on areas with high and low contamination levels to identify areas of concern and unaffected resources. Participants emphasized how centralized USGS databases and the establishment of sample-metadata archives can help to ensure that samples are preserved and accessible for future research. In conclusion, the workshop participants identified opportunities to bridge data gaps and improve measurement techniques, modeling frameworks, databases, and communication, to enhance the understanding of PFAS and their effects on environmental and human health. Upon completion of the workshop, participants indicated an interest in developing strategic data collection, modeling, and analytical approaches to address these challenges.

Open-File Report↗

Methods for estimating selected low-flow frequency statistics and mean annual flow for ungaged locations on streams in North Georgia

The U.S. Geological Survey, in cooperation with the Georgia Department of Natural Resources, Environmental Protection Division, developed regional regression equations for estimating selected low-flow frequency and mean annual flow statistics for ungaged streams in north Georgia that are not substantially affected by regulation, diversions, or urbanization. Selected low-flow frequency statistics and basin characteristics for 56 streamgage locations within north Georgia and 75 miles beyond the State’s borders in Alabama, Tennessee, North Carolina, and South Carolina were combined to form the final dataset used in the regional regression analysis. Because some of the streamgages in the study recorded zero flow, the final regression equations were developed using weighted left-censored regression analysis to analyze the flow data in an unbiased manner, with weights based on the number of years of record. The set of equations includes the annual minimum 1- and 7-day average streamflow with the 10-year recurrence interval (referred to as 1Q10 and 7Q10), monthly 7Q10, and mean annual flow. The final regional regression equations are functions of drainage area, mean annual precipitation, and relief ratio for the selected low-flow frequency statistics and drainage area and mean annual precipitation for mean annual flow. The average standard error of estimate was 13.7 percent for the mean annual flow regression equation and ranged from 26.1 to 91.6 percent for the selected low-flow frequency equations. The equations, which are based on data from streams with little to no flow alterations, can be used to provide estimates of the natural flows for selected ungaged stream locations in the area of Georgia north of the Fall Line. The regression equations are not to be used to estimate flows for streams that have been altered by the effects of major dams, surface-water withdrawals, groundwater withdrawals (pumping wells), diversions, or wastewater discharges. The regression equations should be used only for ungaged sites with drainage areas between 1.67 and 576 square miles, mean annual precipitation between 47.6 and 81.6 inches, and relief ratios between 0.146 and 0.607; these are the ranges of the explanatory variables used to develop the equations. An attempt was made to develop regional regression equations for the area of Georgia south of the Fall Line by using the same approach used during this study for north Georgia; however, the equations resulted with high average standard errors of estimates and poorly predicted flows below 0.5 cubic foot per second, which may be attributed to the karst topography common in that area. The final regression equations developed from this study are planned to be incorporated into the U.S. Geological Survey StreamStats program. StreamStats is a Web-based geographic information system that provides users with access to an assortment of analytical tools useful for water-resources planning and management, and for engineering design applications, such as the design of bridges. The StreamStats program provides streamflow statistics and basin characteristics for U.S. Geological Survey streamgage locations and ungaged sites of interest. StreamStats also can compute basin characteristics and provide estimates of streamflow statistics for ungaged sites when users select the location of a site along any stream in Georgia.

Georgia↗

Methods for determining magnitude and frequency of floods in California, based on data through water year 2006

Methods for estimating the magnitude and frequency of floods in California that are not substantially affected by regulation or diversions have been updated. Annual peak-flow data through water year 2006 were analyzed for 771 streamflow-gaging stations (streamgages) in California having 10 or more years of data. Flood-frequency estimates were computed for the streamgages by using the expected moments algorithm to fit a Pearson Type III distribution to logarithms of annual peak flows for each streamgage. Low-outlier and historic information were incorporated into the flood-frequency analysis, and a generalized Grubbs-Beck test was used to detect multiple potentially influential low outliers. Special methods for fitting the distribution were developed for streamgages in the desert region in southeastern California. Additionally, basin characteristics for the streamgages were computed by using a geographical information system. Regional regression analysis, using generalized least squares regression, was used to develop a set of equations for estimating flows with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities for ungaged basins in California that are outside of the southeastern desert region. Flood-frequency estimates and basin characteristics for 630 streamgages were combined to form the final database used in the regional regression analysis. Five hydrologic regions were developed for the area of California outside of the desert region. The final regional regression equations are functions of drainage area and mean annual precipitation for four of the five regions. In one region, the Sierra Nevada region, the final equations are functions of drainage area, mean basin elevation, and mean annual precipitation. Average standard errors of prediction for the regression equations in all five regions range from 42.7 to 161.9 percent. For the desert region of California, an analysis of 33 streamgages was used to develop regional estimates of all three parameters (mean, standard deviation, and skew) of the log-Pearson Type III distribution. The regional estimates were then used to develop a set of equations for estimating flows with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities for ungaged basins. The final regional regression equations are functions of drainage area. Average standard errors of prediction for these regression equations range from 214.2 to 856.2 percent. Annual peak-flow data through water year 2006 were analyzed for eight streamgages in California having 10 or more years of data considered to be affected by urbanization. Flood-frequency estimates were computed for the urban streamgages by fitting a Pearson Type III distribution to logarithms of annual peak flows for each streamgage. Regression analysis could not be used to develop flood-frequency estimation equations for urban streams because of the limited number of sites. Flood-frequency estimates for the eight urban sites were graphically compared to flood-frequency estimates for 630 non-urban sites. The regression equations developed from this study will be incorporated into the U.S. Geological Survey (USGS) StreamStats program. The StreamStats program is a Web-based application that provides streamflow statistics and basin characteristics for USGS streamgages and ungaged sites of interest. StreamStats can also compute basin characteristics and provide estimates of streamflow statistics for ungaged sites when users select the location of a site along any stream in California.

California↗

Analysis of the variability in ground-motion synthesis and inversion

In almost all past inversions of large-earthquake ground motions for rupture behavior, the goal of the inversion is to find the “best fitting” rupture model that predicts ground motions which optimize some function of the difference between predicted and observed ground motions. This type of inversion was pioneered in the linear-inverse sense by Olson and Apsel (1982), who minimized the square of the difference between observed and simulated motions (“least squares”) while simultaneously minimizing the rupture-model norm (by setting the null-space component of the rupture model to zero), and has been extended in many ways, one of which is the use of nonlinear inversion schemes such as simulated annealing algorithms that optimize some other misfit function. For example, the simulated annealing algorithm of Piatanesi and others (2007) finds the rupture model that minimizes a “cost” function which combines a least-squares and a waveform-correlation measure of misfit. All such inversions that look for a unique “best” model have at least three problems. (1) They have removed the null-space component of the rupture model—that is, an infinite family of rupture models that all fit the data equally well have been narrowed down to a single model. Some property of interest in the rupture model might have been discarded in this winnowing process. (2) Smoothing constraints are commonly used to yield a unique “best” model, in which case spatially rough rupture models will have been discarded, even if they provide a good fit to the data. (3) No estimate of confidence in the resulting rupture models can be given because the effects of unknown errors in the Green’s functions (“theory errors”) have not been assessed. In inversion for rupture behavior, these theory errors are generally larger than the data errors caused by ground noise and instrumental limitations, and so overfitting of the data is probably ubiquitous for such inversions. Recently, attention has turned to the inclusion of theory errors in the inversion process. Yagi and Fukahata (2011) made an important contribution by presenting a method to estimate the uncertainties in predicted large-earthquake ground motions due to uncertainties in the Green’s functions. Here we derive their result and compare it with the results of other recent studies that look at theory errors in a Bayesian inversion context particularly those by Bodin and others (2012), Duputel and others (2012), Dettmer and others (2014), and Minson and others (2014). Notably, in all these studies, the estimates of theory error were obtained from theoretical considerations alone; none of the investigators actually measured Green’s function errors. Large earthquakes typically have aftershocks, which, if their rupture surfaces are physically small enough, can be considered point evaluations of the real Green’s functions of the Earth. Here we simulate smallaftershock ground motions with (erroneous) theoretical Green’s functions. Taking differences between aftershock ground motions and simulated motions to be the “theory error,” we derive a statistical model of the sources of discrepancies between the theoretical and real Green’s functions. We use this model with an extended frequency-domain version of the time-domain theory of Yagi and Fukahata (2011) to determine the expected variance 2 τ caused by Green’s function error in ground motions from a larger (nonpoint) earthquake that we seek to model. We also differ from the above-mentioned Bayesian inversions in our handling of the nonuniqueness problem of seismic inversion. We follow the philosophy of Segall and Du (1993), who, instead of looking for a best-fitting model, looked for slip models that answered specific questions about the earthquakes they studied. In their Bayesian inversions, they inductively derived a posterior probability-density function (PDF) for every model parameter. We instead seek to find two extremal rupture models whose ground motions fit the data within the error bounds given by 2 τ , as quantified by using a chi-squared test described below. So, we can ask questions such as, “What are the rupture models with the highest and lowest average rupture speed consistent with the theory errors?” Having found those models, we can then say with confidence that the true rupture speed is somewhere between those values. Although the Bayesian approach gives a complete solution to the inverse problem, it is computationally demanding: Minson and others (2014) needed 1010 forward kinematic simulations to derive their posterior probability distribution. In our approach, only about107 simulations are needed. Moreover, in practical application, only a small set of rupture models may be needed to answer the relevant questions—for example, determining the maximum likelihood solution (achievable through standard inversion techniques) and the two rupture models bounding some property of interest. The specific property that we wish to investigate is the correlation between various rupturemodel parameters, such as peak slip velocity and rupture velocity, in models of real earthquakes. In some simulations of ground motions for hypothetical large earthquakes, such as those by Aagaard and others (2010) and the Southern California Earthquake Center Broadband Simulation Platform (Graves and Pitarka, 2015), rupture speed is assumed to correlate locally with peak slip, although there is evidence that rupture speed should correlate better with peak slip speed, owing to its dependence on local stress drop. We may be able to determine ways to modify Piatanesi and others’s (2007) inversion’s “cost” function to find rupture models with either high or low degrees of correlation between pairs of rupture parameters. We propose a cost function designed to find these two extremal models.

Open-File Report↗

Estimating abundance of an open population with an N-mixture model using auxiliary data on animal movements

Accurate assessment of abundance forms a central challenge in population ecology and wildlife management. Many statistical techniques have been developed to estimate population sizes because populations change over time and space and to correct for the bias resulting from animals that are present in a study area but not observed. The mobility of individuals makes it difficult to design sampling procedures that account for movement into and out of areas with fixed jurisdictional boundaries. Aerial surveys are the gold standard used to obtain data of large mobile species in geographic regions with harsh terrain, but these surveys can be prohibitively expensive and dangerous. Estimating abundance with ground‐based census methods have practical advantages, but it can be difficult to simultaneously account for temporary emigration and observer error to avoid biased results. Contemporary research in population ecology increasingly relies on telemetry observations of the states and locations of individuals to gain insight on vital rates, animal movements, and population abundance. Analytical models that use observations of movements to improve estimates of abundance have not been developed. Here we build upon existing multi‐state mark–recapture methods using a hierarchical N ‐mixture model with multiple sources of data, including telemetry data on locations of individuals, to improve estimates of population sizes. We used a state‐space approach to model animal movements to approximate the number of marked animals present within the study area at any observation period, thereby accounting for a frequently changing number of marked individuals. We illustrate the approach using data on a population of elk ( Cervus elaphus nelsoni ) in Northern Colorado, USA. We demonstrate substantial improvement compared to existing abundance estimation methods and corroborate our results from the ground based surveys with estimates from aerial surveys during the same seasons. We develop a hierarchical Bayesian N‐mixture model using multiple sources of data on abundance, movement and survival to estimate the population size of a mobile species that uses remote conservation areas. The model improves accuracy of inference relative to previous methods for estimating abundance of open populations.

Ecological Applications↗

Assessment of impacts of proposed coal-resource and related economic development on water resources, Yampa River basin, Colorado and Wyoming: A summary

Expanded mining and use of coal resources in the Rocky Mountain region of the western United States will have substantial impacts on water resources, environmental amenities, and social and economic conditions. The U.S. Geological Survey has completed a 3-year assessment of the Yampa River basin, Colorado and Wyoming, where increased coal-resource development has begun to affect the environment and quality of life. Economic projections of the overall effects of coal-resource development were used to estimate water use and the types and amounts of waste residuals that need to be assimilated into the environment. Based in part upon these projections, several physical-based models and other semiquantitative assessment methods were used to determine possible effects upon the basin's water resources. Depending on the magnitude of mining and use of coal resources in the basin, an estimated 0.7 to 2.7 million tons (0.6 to 2.4 million metric tons) of waste residuals may be discharged annually into the environment by coal-resource development and associated economic activities. If the assumed development of coal resources in the basin occurs, annual consumptive use of water, which was approximately 142,000 acre-feet (175 million cubic meters) during 1975, may almost double by 1990. In a related analysis of alternative cooling systems for coal-conversion facilities, four to five times as much water may be used consumptively in a wet-tower, cooling-pond recycling system as in once-through cooling. An equivalent amount of coal transported by slurry pipeline would require about one-third the water used consumptively by once-through cooling for in-basin conversion. Current conditions and a variety of possible changes in the water resources of the basin resulting from coal-resource development were assessed. Basin population may increase by as much as threefold between 1975 and 1990. Volumes of wastes requiring treatment will increase accordingly. Potential problems associated with ammonia-nitrogen concentrations in the Yampa River downstream from Steamboat Springs were evaluated using a waste-load assimilative-capacity model. Changes in sediment loads carried by streams due to increased coal mining and construction of roads and buildings may be apparent only locally; projected increases in sediment loads relative to historic loads from the basin are estimated to be 2 to 7 percent. Solid-waste residuals generated by coal-conversion processes and disposed of into old mine pits may cause widely dispersed ground-water contamination, based on simulation-modeling results. Projected increases in year-round water use will probably result in the construction of several proposed reservoirs. Current seasonal patterns of streamflow and of dissolvedsolids concentrations in streamflow will be altered appreciably by these reservoirs. Decreases in time-weighted mean-annual dissolved-solids concentrations of as much as 34 percent are anticipated, based upon model simulations of several configurations of proposed reservoirs. Detailed statistical analyses of water-quality conditions in the Yampa River basin were made. Regionalized maximum waterquality concentrations were estimated for possible comparison with future conditions. Using Landsat imagery and aerial photographs, potential remote-sensing applications were evaluated to monitor land-use changes and to assess both snow cover and turbidity levels in streams. The technical information provided by the several studies of the Yampa River basin assessment should be useful to regional planners and resource managers in evaluating the possible impacts of development on the basin's water resources.

Colorado, Wyoming↗

Computing and software

The reality is that the statistical methods used for analysis of data depend upon the availability of software. Analysis of marked animal data is no different than the rest of the statistical field. The methods used for analysis are those that are available in reliable software packages. Thus, the critical importance of having reliable, up–to–date software available to biologists is obvious. Statisticians have continued to develop more robust models, ever expanding the suite of potential analysis methods available. But without software to implement these newer methods, they will languish in the abstract, and not be applied to the problems deserving them. In the Computers and Software Session, two new software packages are described, a comparison of implementation of methods for the estimation of nest survival is provided, and a more speculative paper about how the next generation of software might be structured is presented. Rotella et al. (2004) compare nest survival estimation with different software packages: SAS logistic regression, SAS non–linear mixed models, and Program MARK. Nests are assumed to be visited at various, possibly infrequent, intervals. All of the approaches described compute nest survival with the same likelihood, and require that the age of the nest is known to account for nests that eventually hatch. However, each approach offers advantages and disadvantages, explored by Rotella et al. (2004). Efford et al. (2004) present a new software package called DENSITY. The package computes population abundance and density from trapping arrays and other detection methods with a new and unique approach. DENSITY represents the first major addition to the analysis of trapping arrays in 20 years. Barker & White (2004) discuss how existing software such as Program MARK require that each new model’s likelihood must be programmed specifically for that model. They wishfully think that future software might allow the user to combine pieces of likelihood functions together to generate estimates. The idea is interesting, and maybe some bright young statistician can work out the specifics to implement the procedure. Choquet et al. (2004) describe MSURGE, a software package that implements the multistate capture–recapture models. The unique feature of MSURGE is that the design matrix is constructed with an interpreted language called GEMACO. Because MSURGE is limited to just multistate models, the special requirements of these likelihoods can be provided. The software and methods presented in these papers gives biologists and wildlife managers an expanding range of possibilities for data analysis. Although ease–of–use is generally getting better, it does not replace the need for understanding of the requirements and structure of the models being computed. The internet provides access to many free software packages as well as user–discussion groups to share knowledge and ideas. (A starting point for wildlife–related applications is (http://www.phidot.org).

Animal Biodiversity and Conservation↗

An Automated Cropland Classification Algorithm (ACCA) for Tajikistan by combining Landsat, MODIS, and secondary data

The overarching goal of this research was to develop and demonstrate an automated Cropland Classification Algorithm (ACCA) that will rapidly, routinely, and accurately classify agricultural cropland extent, areas, and characteristics (e.g., irrigated vs. rainfed) over large areas such as a country or a region through combination of multi-sensor remote sensing and secondary data. In this research, a rule-based ACCA was conceptualized, developed, and demonstrated for the country of Tajikistan using mega file data cubes (MFDCs) involving data from Landsat Global Land Survey (GLS), Landsat Enhanced Thematic Mapper Plus (ETM+) 30 m, Moderate Resolution Imaging Spectroradiometer (MODIS) 250 m time-series, a suite of secondary data (e.g., elevation, slope, precipitation, temperature), and in situ data. First, the process involved producing an accurate reference (or truth) cropland layer (TCL), consisting of cropland extent, areas, and irrigated vs. rainfed cropland areas, for the entire country of Tajikistan based on MFDC of year 2005 (MFDC2005). The methods involved in producing TCL included using ISOCLASS clustering, Tasseled Cap bi-spectral plots, spectro-temporal characteristics from MODIS 250 m monthly normalized difference vegetation index (NDVI) maximum value composites (MVC) time-series, and textural characteristics of higher resolution imagery. The TCL statistics accurately matched with the national statistics of Tajikistan for irrigated and rainfed croplands, where about 70% of croplands were irrigated and the rest rainfed. Second, a rule-based ACCA was developed to replicate the TCL accurately (∼80% producer’s and user’s accuracies or within 20% quantity disagreement involving about 10 million Landsat 30 m sized cropland pixels of Tajikistan). Development of ACCA was an iterative process involving series of rules that are coded, refined, tweaked, and re-coded till ACCA derived croplands (ACLs) match accurately with TCLs. Third, the ACCA derived cropland layers of Tajikistan were produced for year 2005 (ACL2005), same year as the year used for developing ACCA, using MFDC2005. Fourth, TCL for year 2010 (TCL2010), an independent year, was produced using MFDC2010 using the same methods and approaches as the one used to produce TCL2005. Fifth, the ACCA was applied on MFDC2010 to derive ACL2010. The ACLs were then compared with TCLs (ACL2005 vs. TCL2005 and ACL2010 vs. TCL2010). The resulting accuracies and errors from error matrices involving about 152 million Landsat (30 m) pixels of the country of Tajikistan (of which about 10 million Landsat size, 30 m, cropland pixels) showed an overall accuracy of 99.6% (k hat = 0.97) for ACL2005 vs. TCL2005. For the 3 classes (irrigated, rainfed, and others) mapped in ACL2005, the producer’s accuracy was >86.4% and users accuracy was >93.6%. For ACL2010 vs. TCL2010, the error matrix showed an overall accuracy on 96.2% (k hat = 0.96). For the 3 classes (irrigated, rainfed, and others) mapped in ACL2010, the producer’s and user’s accuracies for the irrigated areas were ≥82.9%. Any intermixing was overwhelmingly between irrigated and rainfed croplands, indicating that croplands (irrigated plus rainfed areas) as well as irrigated areas were mapped with high levels of accuracies (∼90% or higher) even for the independent year. The ACL2005 and ACL2010, each, were produced using ACCA algorithm in ∼30 min using a Dell Precision desktop T7400 computer for the entire country of Tajikistan once the MFDCs for the years were ready. The ACCA algorithm for Tajikistan is made available through US Geological Survey’s ScienceBase: http://www.sciencebase.gov/catalog/folder/4f79f1b7e4b0009bd827f548 or at: https://powellcenter.usgs.gov/globalcroplandwater/content/models-algorithms . The research contributes to the efforts of global food security through research on global croplands and their water use (e.g., https://powellcenter.usgs.gov/globalcroplandwater/ ). The above results clearly demonstrated the ability of a rule-based ACCA to rapidly and accurately produce cropland data layer year after year (hindcast, nowcast, forecast) for the country it was developed using MFDCs that consist of combining multiple sensor data and secondary data. It needs to be noted that the ACCA is applicable to the area (e.g., country, region) for which it is developed. In this case, ACCA is applicable for the Country of Tajikistan to hindcast, nowcast, and forecast agricultural cropland extent, areas, and irrigated vs. rainfed. The same fundamental concept of ACCA applies to other areas of the World where ACCA codes need to be modified to suite the area/region of interest. ACCA can also be expanded to compute other crop characteristics such as crop types, cropping intensities, and phenologies.

Remote Sensing↗

Earthquake likelihood model testing

INTRODUCTION The Regional Earthquake Likelihood Models (RELM) project aims to produce and evaluate alternate models of earthquake potential (probability per unit volume, magnitude, and time) for California. Based on differing assumptions, these models are produced to test the validity of their assumptions and to explore which models should be incorporated in seismic hazard and risk evaluation. Tests based on physical and geological criteria are useful but we focus on statistical methods using future earthquake catalog data only. We envision two evaluations: a test of consistency with observed data and a comparison of all pairs of models for relative consistency. Both tests are based on the likelihood method, and both are fully prospective ( i.e. , the models are not adjusted to fit the test data). To be tested, each model must assign a probability to any possible event within a specified region of space, time, and magnitude. For our tests the models must use a common format: earthquake rates in specified “bins” with location, magnitude, time, and focal mechanism limits. Seismology cannot yet deterministically predict individual earthquakes; however, it should seek the best possible models for forecasting earthquake occurrence. This paper describes the statistical rules of an experiment to examine and test earthquake forecasts. The primary purposes of the tests described below are to evaluate physical models for earthquakes, assure that source models used in seismic hazard and risk studies are consistent with earthquake data, and provide quantitative measures by which models can be assigned weights in a consensus model or be judged as suitable for particular regions. In this paper we develop a statistical method for testing earthquake likelihood models. A companion paper ( Schorlemmer and Gerstenberger 2007 , this issue) discusses the actual implementation of these tests in the framework of the RELM initiative. Statistical testing of hypotheses is a common task and a wide range of possible testing procedures exist. Jolliffe and Stephenson ( 2003 ) present different forecast verifications from atmospheric science, among them likelihood testing of probability forecasts and testing the occurrence of binary events. Testing binary events requires that for each forecasted event, the spatial, temporal and magnitude limits be given. Although major earthquakes can be considered binary events, the models within the RELM project express their forecasts on a spatial grid and in 0.1 magnitude units; thus the results are a distribution of rates over space and magnitude. These forecasts can be tested with likelihood tests. In general, likelihood tests assume a valid null hypothesis against which a given hypothesis is tested. The outcome is either a rejection of the null hypothesis in favor of the test hypothesis or a nonrejection, meaning the test hypothesis cannot outperform the null hypothesis at a given significance level. Within RELM, there is no accepted null hypothesis and thus the likelihood test needs to be expanded to allow comparable testing of equipollent hypotheses. To test models against one another, we require that forecasts are expressed in a standard format: the average rate of earthquake occurrence within pre-specified limits of hypocentral latitude, longitude, depth, magnitude, time period, and focal mechanisms. Focal mechanisms should either be described as the inclination of P -axis, declination of P -axis, and inclination of the T -axis, or as strike, dip, and rake angles. Schorlemmer and Gerstenberger ( 2007 , this issue) designed classes of these parameters such that similar models will be tested against each other. These classes make the forecasts comparable between models. Additionally, we are limited to testing only what is precisely defined and consistently reported in earthquake catalogs. Therefore it is currently not possible to test such information as fault rupture length or area, asperity location, etc. Also, to account for data quality issues, we allow for location and magnitude uncertainties as well as the probability that an event is dependent on another event. As we mentioned above, only models with comparable forecasts can be tested against each other. Our current tests are designed to examine grid-based models. This requires that any fault-based model be adapted to a grid before testing is possible. While this is a limitation of the testing, it is an inherent difficulty in any such comparative testing. Please refer to appendix B for a statistical evaluation of the application of the Poisson hypothesis to fault-based models. The testing suite we present consists of three different tests: L-Test, N-Test, and R-Test. These tests are defined similarily to Kagan and Jackson ( 1995 ). The first two tests examine the consistency of the hypotheses with the observations while the last test compares the spatial performances of the models.

Seismological Research Letters↗

Western Mineral and Environmental Resources Science Center--providing comprehensive earth science for complex societal issues

Minerals in the environment and products manufactured from mineral materials are all around us and we use and come into contact with them every day. They impact our way of life and the health of all that lives. Minerals are critical to the Nation's economy and knowing where future mineral resources will come from is important for sustaining the Nation's economy and national security. The U.S. Geological Survey (USGS) Mineral Resources Program (MRP) provides scientific information for objective resource assessments and unbiased research results on mineral resource potential, production and consumption statistics, as well as environmental consequences of mining. The MRP conducts this research to provide information needed for land planners and decisionmakers about where mineral commodities are known and suspected in the earth's crust and about the environmental consequences of extracting those commodities. As part of the MRP scientists of the Western Mineral and Environmental Resources Science Center (WMERSC or 'Center' herein) coordinate the development of national, geologic, geochemical, geophysical, and mineral-resource databases and the migration of existing databases to standard models and formats that are available to both internal and external users. The unique expertise developed by Center scientists over many decades in response to mineral-resource-related issues is now in great demand to support applications such as public health research and remediation of environmental hazards that result from mining and mining-related activities. Western Mineral and Environmental Resources Science Center Results of WMERSC research provide timely and unbiased analyses of minerals and inorganic materials to (1) improve stewardship of public lands and resources; (2) support national and international economic and security policies; (3) sustain prosperity and improve our quality of life; and (4) protect and improve public health, safety, and environmental quality. The MRP supports approximately 40 USGS research specialists who utilize cooperative agreements with universities, industry, and other governmental agencies to support their collaborative research and information exchange. Scientists of the WMERSC study how and where non-fuel mineral resources form and are concentrated in the earth's crust, where mineral resources might be found in the future, and how mineral materials interact with the environment to affect human and ecosystem health. Natural systems (ecosystems) are complex - our understanding of how ecosystems operate requires collecting and synthesizing large amounts of geologic, geochemical, biologic, hydrologic, and meteorological information. Scientists in the Center strive to understand the interplay of various processes and how they affect the structure, composition, and health of ecosystems. Such understanding, which is then summarized in publicly available reports, is used to address and solve a wide variety of issues that are important to society and the economy. WMERSC scientists have extensive national and international experience in these scientific specialties and capabilities - they have collaborated with many Federal, State, and local agencies; with various private sector organizations; as well as with foreign countries and organizations. Nearly every scientific and societal challenge requires a different combination of scientific skills and capabilities. With their breadth of scientific specialties and capabilities, the scientists of the WMERSC can provide scientifically sound approaches to a wide range of societal challenges and issues. The following sections describe examples of important issues that have been addressed by scientists in the Center, the methods employed, and the relevant conclusions. New directions are inevitable as societal needs change over time. Scientists of the WMERSC have a diverse set of skills and capabilities and are proficient in the collection and integration of

Circular↗

Uncertainty in biological monitoring: a framework for data collection and analysis to account for multiple sources of sampling bias

Biological monitoring programmes are increasingly relying upon large volumes of citizen-science data to improve the scope and spatial coverage of information, challenging the scientific community to develop design and model-based approaches to improve inference. Recent statistical models in ecology have been developed to accommodate false-negative errors, although current work points to false-positive errors as equally important sources of bias. This is of particular concern for the success of any monitoring programme given that rates as small as 3% could lead to the overestimation of the occurrence of rare events by as much as 50%, and even small false-positive rates can severely bias estimates of occurrence dynamics. We present an integrated, computationally efficient Bayesian hierarchical model to correct for false-positive and false-negative errors in detection/non-detection data. Our model combines independent, auxiliary data sources with field observations to improve the estimation of false-positive rates, when a subset of field observations cannot be validated a posteriori or assumed as perfect. We evaluated the performance of the model across a range of occurrence rates, false-positive and false-negative errors, and quantity of auxiliary data. The model performed well under all simulated scenarios, and we were able to identify critical auxiliary data characteristics which resulted in improved inference. We applied our false-positive model to a large-scale, citizen-science monitoring programme for anurans in the north-eastern United States, using auxiliary data from an experiment designed to estimate false-positive error rates. Not correcting for false-positive rates resulted in biased estimates of occupancy in 4 of the 10 anuran species we analysed, leading to an overestimation of the average number of occupied survey routes by as much as 70%. The framework we present for data collection and analysis is able to efficiently provide reliable inference for occurrence patterns using data from a citizen-science monitoring programme. However, our approach is applicable to data generated by any type of research and monitoring programme, independent of skill level or scale, when effort is placed on obtaining auxiliary information on false-positive rates.

Methods in Ecology and Evolution↗

Ignoring species availability biases occupancy estimates in single-scale occupancy models

Most applications of single-scale occupancy models do not differentiate between availability and detectability, even though species availability is rarely equal to one. Species availability can be estimated using multi-scale occupancy models; however, for the practical application of multi-scale occupancy models, it can be unclear what a robust sampling design looks like and what the statistical properties of the multi-scale and single-scale occupancy models are when availability is less than one. Using simulations, we explore the following common questions asked by ecologists during the design phase of a field study: (Q1) what is a robust sampling design for the multi-scale occupancy model when there are a priori expectations of parameter estimates? (Q2) what is a robust sampling design when we have no expectations of parameter estimates? and (Q3) can a single-scale occupancy model with a random effects term adequately absorb the extra heterogeneity produced when availability is less than one and provide reliable estimates of occupancy probability? Our results show that there is a tradeoff between the number of sites and surveys needed to achieve a specified level of acceptable error for occupancy estimates using the multi-scale occupancy model. We also document that when species availability is low (<0.40 on the probability scale), then single-scale occupancy models underestimate occupancy by as much as 0.40 on the probability scale, produce overly precise estimates, and provide poor parameter coverage. This pattern was observed when a random effects term was and was not included in the single-scale occupancy model, suggesting that adding a random-effects term does not adequately absorb the extra heterogeneity produced by the availability process. In contrast, when species availability was high (>0.60), single-scale occupancy models performed similarly to the multi-scale occupancy model. Users can further explore our results and sampling designs across a number of different scenarios using the RShiny app https://gdirenzo.shinyapps.io/multi-scale-occ/ . Our results suggest that unaccounted for availability can lead to underestimating species distributions when using single-scale occupancy models, which can have large implications on inference and prediction, especially for those working in the fields of invasion ecology, disease emergence, and species conservation.

Methods in Ecology and Evolution↗

Estimation of annual agricultural pesticide use for counties of the conterminous United States, 1992-2009

A method was developed to calculate annual county level pesticide use for selected herbicides, insecticides, and fungicides applied to agricultural crops grown in the conterminous United States from 1992 through 2009. Pesticide-use data compiled by proprietary surveys of farm operations located within Crop Reporting Districts were used in conjunction with annual harvested-crop acreage reported by the U.S. Department of Agriculture National Agricultural Statistics Service (NASS) to calculate use rates per harvested crop acre, or an 'estimated pesticide use' (EPest) rate, for each crop by year. Pesticide-use data were not available for all Crop Reporting Districts and years. When data were unavailable for a Crop Reporting District in a particular year, EPest extrapolated rates were calculated from adjoining or nearby Crop Reporting Districts to ensure that pesticide use was estimated for all counties that reported harvested-crop acreage. EPest rates were applied to county harvested-crop acreage differently to obtain EPest-low and EPest-high estimates of pesticide-use for counties and states, with the exception of use estimates for California, which were taken from annual Department of Pesticide Regulation Pesticide Use Reports. Annual EPest-low and EPest-high use totals were compared with other published pesticide-use reports for selected pesticides, crops, and years. EPest-low and EPest-high national totals for five of seven herbicides were in close agreement with U.S. Environmental Protection Agency and National Pesticide Use Data estimates, but greater than most NASS national totals. A second set of analyses compared EPest and NASS annual state totals and state-by-crop totals for selected crops. Overall, EPest and NASS use totals were not significantly different for the majority of crop-stateyear combinations evaluated. Furthermore, comparisons of EPest and NASS use estimates for most pesticides had rank correlation coefficients greater than 0.75 and median relative errors of less than 15 percent. Of the 48 pesticide-by-crop combinations with 10 or more state-year combinations, 12 of the EPest-low and 17 of the EPest-high totals showed significant differences (p < 0.05) from NASS use estimates. The differences between EPest and NASS estimates did not follow consistent patterns related to particular crops, years, or states, and most correlation coefficients were greater than 0.75. EPest values from this study are suitable for making national, regional, and watershed assessments of annual pesticide use from 1992 to 2009. Although estimates are provided by county to facilitate estimation of watershed pesticide use for a wide variety of watersheds, there is a greater degree of uncertainty in individual county-level estimates when compared to Crop Reporting District or state-level estimates because (1) EPest crop-use rates were developed on the basis of pesticide use on harvested acres in multi-county areas (Crop Reporting Districts) and then allocated to county harvested cropland; (2) pesticide-by-crop use rates were not available for all Crop Reporting Districts in the conterminous United States, and extrapolation methods were used to estimate pesticide use for some counties; and (3) it is possible that surveyed pesticide-by-crop use rates do not reflect all agricultural use on all crops grown. The methods developed in this study also are applicable to other agricultural pesticides and years.

Scientific Investigations Report↗