USGS Science⌕ Search

SEARCH · USGS Science

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Methods for estimating selected low-flow frequency statistics for unregulated streams in Kentucky

This report provides estimates of, and presents methods for estimating, selected low-flow frequency statistics for unregulated streams in Kentucky including the 30-day mean low flows for recurrence intervals of 2 and 5 years (30Q 2 and 30Q 5 ) and the 7-day mean low flows for recurrence intervals of 5, 10, and 20 years (7Q 2 , 7Q 10 , and 7Q 20 ). Estimates of these statistics are provided for 121 U.S. Geological Survey streamflow-gaging stations with data through the 2006 climate year, which is the 12-month period ending March 31 of each year. Data were screened to identify the periods of homogeneous, unregulated flows for use in the analyses. Logistic-regression equations are presented for estimating the annual probability of the selected low-flow frequency statistics being equal to zero. Weighted-least-squares regression equations were developed for estimating the magnitude of the nonzero 30Q 2 , 30Q 5 , 7Q 2 , 7Q 10 , and 7Q 20 low flows. Three low-flow regions were defined for estimating the 7-day low-flow frequency statistics. The explicit explanatory variables in the regression equations include total drainage area and the mapped streamflow-variability index measured from a revised statewide coverage of this characteristic. The percentage of the station low-flow statistics correctly classified as zero or nonzero by use of the logistic-regression equations ranged from 87.5 to 93.8 percent. The average standard errors of prediction of the weighted-least-squares regression equations ranged from 108 to 226 percent. The 30Q 2 regression equations have the smallest standard errors of prediction, and the 7Q 20 regression equations have the largest standard errors of prediction. The regression equations are applicable only to stream sites with low flows unaffected by regulation from reservoirs and local diversions of flow and to drainage basins in specified ranges of basin characteristics. Caution is advised when applying the equations for basins with characteristics near the applicable limits and for basins with karst drainage features.

Scientific Investigations Report↗

Methods for estimating low-flow statistics for Massachusetts streams

Methods and computer software are described in this report for determining flow duration, low-flow frequency statistics, and August median flows. These low-flow statistics can be estimated for unregulated streams in Massachusetts using different methods depending on whether the location of interest is at a streamgaging station, a low-flow partial-record station, or an ungaged site where no data are available. Low-flow statistics for streamgaging stations can be estimated using standard U.S. Geological Survey methods described in the report. The MOVE.1 mathematical method and a graphical correlation method can be used to estimate low-flow statistics for low-flow partial-record stations. The MOVE.1 method is recommended when the relation between measured flows at a partial-record station and daily mean flows at a nearby, hydrologically similar streamgaging station is linear, and the graphical method is recommended when the relation is curved. Equations are presented for computing the variance and equivalent years of record for estimates of low-flow statistics for low-flow partial-record stations when either a single or multiple index stations are used to determine the estimates. The drainage-area ratio method or regression equations can be used to estimate low-flow statistics for ungaged sites where no data are available. The drainage-area ratio method is generally as accurate as or more accurate than regression estimates when the drainage-area ratio for an ungaged site is between 0.3 and 1.5 times the drainage area of the index data-collection site. Regression equations were developed to estimate the natural, long-term 99-, 98-, 95-, 90-, 85-, 80-, 75-, 70-, 60-, and 50-percent duration flows; the 7-day, 2-year and the 7-day, 10-year low flows; and the August median flow for ungaged sites in Massachusetts. Streamflow statistics and basin characteristics for 87 to 133 streamgaging stations and low-flow partial-record stations were used to develop the equations. The streamgaging stations had from 2 to 81 years of record, with a mean record length of 37 years. The low-flow partial-record stations had from 8 to 36 streamflow measurements, with a median of 14 measurements. All basin characteristics were determined from digital map data. The basin characteristics that were statistically significant in most of the final regression equations were drainage area, the area of stratified-drift deposits per unit of stream length plus 0.1, mean basin slope, and an indicator variable that was 0 in the eastern region and 1 in the western region of Massachusetts. The equations were developed by use of weighted-least-squares regression analyses, with weights assigned proportional to the years of record and inversely proportional to the variances of the streamflow statistics for the stations. Standard errors of prediction ranged from 70.7 to 17.5 percent for the equations to predict the 7-day, 10-year low flow and 50-percent duration flow, respectively. The equations are not applicable for use in the Southeast Coastal region of the State, or where basin characteristics for the selected ungaged site are outside the ranges of those for the stations used in the regression analyses. A World Wide Web application was developed that provides streamflow statistics for data collection stations from a data base and for ungaged sites by measuring the necessary basin characteristics for the site and solving the regression equations. Output provided by the Web application for ungaged sites includes a map of the drainage-basin boundary determined for the site, the measured basin characteristics, the estimated streamflow statistics, and 90-percent prediction intervals for the estimates. An equation is provided for combining regression and correlation estimates to obtain improved estimates of the streamflow statistics for low-flow partial-record stations. An equation is also provided for combining regression and drainage-area ratio estimates to obtain improved e

Massachusetts↗

Methods for estimating selected streamflow statistics at ungaged sites in Wyoming based on data through water year 2021

The U.S. Geological Survey, in cooperation with the Wyoming Water Development Office, developed regional regression equations based on basin characteristics and streamflow statistics for streamgages through water year 2021 (October 1, 2020, to September 30, 2021). The regression equations allow estimates of mean annual maximum, mean annual, mean seasonal, and mean monthly streamflows; frequency statistics for the 7-day mean low flows with 2-year and 10-year recurrence intervals, 14- and 30-day mean low flows with 5-year recurrence intervals, and 60- and 1-day mean high flow with 2-year and 5-year recurrence intervals, respectively; and the 0.1-, 0.2-, 0.5-, 1-, 2-, 4-, 5-, 10-, 20-, 25-, 30-, 50-, 60-, 70-, 75-, 80-, 90-, 95-, 98-, and 99-percent durations for annual streamflows and 0.1-, 0.5-, 10-, 15-, 20-, 25-, 30-, 40-, 50-, 60-, 70-, 75-, 80-, 85-, 90-, 95-, and 99-percent durations for monthly streamflows for most months for ungaged locations in Wyoming that are largely unaltered by diversions or upstream reservoirs. Regression equations were developed for 243 streamflow statistics. Best-subset selection was used to assess explanatory variables for respective streamflow statistics. Exploratory data analyses determined that, of the 81 basin characteristics evaluated as potential explanatory variables, characteristics such as drainage area and precipitation often produced models with the highest adjusted coefficient of determination and lowest mean squared error, as determined in the best-subset selection. To address heteroskedasticity of model residuals, model variables were regionalized using fixed-effects models; the percentages of the streamgage basins in selected ecoregions were defined as interaction terms, which represent the model slope for specific ecoregions. Most models were determined to be statistically significant for probability values less than or equal to 0.1 for one or more regional explanatory variables. The final regional regression equations defined in this report are available for use in the U.S. Geological Survey’s StreamStats web application at https://streamstats.usgs.gov/ss/ .

Colorado, Idaho, Montana, North Dakota, South Dako↗

Regression method for estimating long-term mean annual ground-water recharge rates from base flow in Pennsylvania

A method was developed for making estimates of long-term, mean annual ground-water recharge from streamflow data at 80 streamflow-gaging stations in Pennsylvania. The method relates mean annual base-flow yield derived from the streamflow data (as a proxy for recharge) to the climatic, geologic, hydrologic, and physiographic characteristics of the basins (basin characteristics) by use of a regression equation. Base-flow yield is the base flow of a stream divided by the drainage area of the basin, expressed in inches of water basinwide. Mean annual base-flow yield was computed for the period of available streamflow record at continuous streamflow-gaging stations by use of the computer program PART, which separates base flow from direct runoff on the streamflow hydrograph. Base flow provides a reasonable estimate of recharge for basins where streamflow is mostly unaffected by upstream regulation, diversion, or mining. Twenty-eight basin characteristics were included in the exploratory regression analysis as possible predictors of base-flow yield. Basin characteristics found to be statistically significant predictors of mean annual base-flow yield during 1971-2000 at the 95-percent confidence level were (1) mean annual precipitation, (2) average maximum daily temperature, (3) percentage of sand in the soil, (4) percentage of carbonate bedrock in the basin, and (5) stream channel slope. The equation for predicting recharge was developed using ordinary least-squares regression. The standard error of prediction for the equation on log-transformed data was 9.7 percent, and the coefficient of determination was 0.80. The equation can be used to predict long-term, mean annual recharge rates for ungaged basins, providing that the explanatory basin characteristics can be determined and that the underlying assumption is accepted that base-flow yield derived from PART is a reasonable estimate of ground-water recharge rates. For example, application of the equation for 370 hydrologic units in Pennsylvania predicted a range of ground-water recharge from about 6.0 to 22 inches per year. A map of the predicted recharge illustrates the general magnitude and variability of recharge throughout Pennsylvania.

Scientific Investigations Report↗

Testing and validating environmental models

Generally accepted standards for testing and validating ecosystem models would benefit both modellers and model users. Universally applicable test procedures are difficult to prescribe, given the diversity of modelling approaches and the many uses for models. However, the generally accepted scientific principles of documentation and disclosure provide a useful framework for devising general standards for model evaluation. Adequately documenting model tests requires explicit performance criteria, and explicit benchmarks against which model performance is compared. A model's validity, reliability, and accuracy can be most meaningfully judged by explicit comparison against the available alternatives. In contrast, current practice is often characterized by vague, subjective claims that model predictions show 'acceptable' agreement with data; such claims provide little basis for choosing among alternative models. Strict model tests (those that invalid models are unlikely to pass) are the only ones capable of convincing rational skeptics that a model is probably valid. However, 'false positive' rates as low as 10% can substantially erode the power of validation tests, making them insufficiently strict to convince rational skeptics. Validation tests are often undermined by excessive parameter calibration and overuse of ad hoc model features. Tests are often also divorced from the conditions under which a model will be used, particularly when it is designed to forecast beyond the range of historical experience. In such situations, data from laboratory and field manipulation experiments can provide particularly effective tests, because one can create experimental conditions quite different from historical data, and because experimental data can provide a more precisely defined 'target' for the model to hit. We present a simple demonstration showing that the two most common methods for comparing model predictions to environmental time series (plotting model time series against data time series, and plotting predicted versus observed values) have little diagnostic power. We propose that it may be more useful to statistically extract the relationships of primary interest from the time series, and test the model directly against them.

Science of the Total Environment↗

Standardization of reflectance measurements in dispersed organic matter: results of an exercise to improve interlaboratory agreement

Vitrinite reflectance generally is considered the most robust thermal maturity parameter available for application to hydrocarbon exploration and petroleum system evaluation. However, until 2011 there was no standardized methodology available to provide guidelines for vitrinite reflectance measurements in shale. Efforts to correct this deficiency resulted in publication of ASTM D7708: Standard test method for microscopical determination of the reflectance of vitrinite dispersed in sedimentary rocks . In 2012-2013, an interlaboratory exercise was conducted to establish precision limits for the D7708 measurement technique. Six samples, representing a wide variety of shale, were tested in duplicate by 28 analysts in 22 laboratories from 14 countries. Samples ranged from immature to overmature (0.31-1.53% R o ), from organic-lean to organic-rich (1-22 wt.% total organic carbon), and contained Type I (lacustrine), Type II (marine), and Type III (terrestrial) kerogens. Repeatability limits (maximum difference between valid repetitive results from same operator, same conditions) ranged from 0.03-0.11% absolute reflectance, whereas reproducibility limits (maximum difference between valid results obtained on same test material by different operators, different laboratories) ranged from 0.12-0.54% absolute reflectance. Repeatability and reproducibility limits degraded consistently with increasing maturity and decreasing organic content. However, samples with terrestrial kerogens (Type III) fell off this trend, showing improved levels of reproducibility due to higher vitrinite content and improved ease of identification. Operators did not consistently meet the reporting requirements of the test method, indicating that a common reporting template is required to improve data quality. The most difficult problem encountered was the petrographic distinction of solid bitumens and low-reflecting inert macerals from vitrinite when vitrinite occurred with reflectance ranges overlapping the other components. Discussion among participants suggested this problem could not be easily corrected via kerogen concentration or solvent extraction and is related to operator training and background. No statistical difference in mean reflectance was identified between participants reporting bitumen reflectance vs. vitrinite reflectance vs. a mixture of bitumen and vitrinite reflectance values, suggesting empirical conversion schemes should be treated with caution. Analysis of reproducibility limits obtained during this exercise in comparison to reproducibility limits from historical interlaboratory exercises suggests use of a common methodology (D7708) improves interlaboratory precision. Future work will investigate opportunities to improve reproducibility in high maturity, organic-lean shale varieties.

Marine and Petroleum Geology↗

A comprehensive plan for in-water sea turtle data collection in the US Gulf of Mexico

The Deepwater Horizon Open Ocean Trustee Implementation Group (OO TIG) released a Final Open Ocean Restoration Plan 2 in 2019, which included a project titled Developing a Gulf-wide Comprehensive Plan for In-water Sea Turtle Data Collection. This document, A Comprehensive Plan for In-water Sea Turtle Data Collection in the US Gulf of Mexico (Plan), is the culmination of that OO TIG project. This Plan serves as the OO TIG project’s technical report as well as a framework for a biologically and statistically-sound plan to support coordinated in-water sea turtle data collection in the United States (US) Gulf of Mexico (GoM) to determine sea turtle abundance and population trends. The purpose of this Plan is to act as a guide for collecting biologically and statistically robust, in-water sea turtle data in a comprehensive, coordinated, and standardized fashion in the US GoM. Several sea turtle in-water monitoring efforts are underway in the GoM; however, additional coordination and standardization of these efforts will benefit current restoration and recovery objectives. These efforts will aid in restoration project design, assess long-term effectiveness of restoration activities, and create abundance and distribution baselines across the GoM. This Plan provides guidance for researchers investigating sea turtle abundance and demographic questions, as well as for management agencies and restoration planners. A Steering Committee (SC) was assembled to develop this Plan and to recommend a coordinated approach to the formulation of an improved understanding of sea turtle population baselines in the GoM, from which determination of large-scale population changes, effects of specific threats (e.g., oil spills, anthropogenic hazards), and effects of changes in ocean conditions (e.g., climate change) can later be evaluated. In crafting this guidance, the SC considered species distribution and life history characteristics, spatial and logistical considerations, level of effort required to detect trends, methods available and the pros and cons of each, associated assumptions and biases with suggested monitoring methods, and standardization of data collection. Given the current level of data available, the SC has recommended species monitoring in two main phases in neritic and oceanic waters, with additional recommended sampling for surface pelagic drift communities. The two phases in this Plan focus on 1) monitoring a limited number of sites in the first 5 to 8 years, followed by 2) a refined monitoring design. To support implementation of this Plan, the SC also considered broader programmatic needs, including supplemental data collection, program and data management, potential international partnerships, program expansion, and applications including future technology.

Alabama, Florida, Louisiana, Mississippi, Texas↗

Occurrence of Pesticides in Ground Water of Wyoming, 1995-2006

Little existing information was available describing pesticide occurrence in ground water of Wyoming, so the U.S. Geological Survey, in cooperation with the Wyoming Department of Agriculture and the Wyoming Department of Environmental Quality on behalf of the Wyoming Ground-water and Pesticides Strategy Committee, collected ground-water samples twice (during late summer/early fall and spring) from 296 wells during 1995-2006 to characterize pesticide occurrence. Sampling focused on the State's ground water that was mapped as the most vulnerable to pesticide contamination because of either inherent hydrogeologic sensitivity (for example, shallow water table or highly permeable aquifer materials) or a combination of sensitivity and associated land use. Because of variations in reporting limits among different compounds and for the same compound during this study, pesticide detections were recensored to two different assessment levels to facilitate qualitative and quantitative examination of pesticide detection frequencies - a common assessment level (CAL) of 0.07 microgram per liter and an assessment level that differed by compound, referred to herein as a compound-specific assessment level (CSAL). Because of severe data censoring (fewer than 50 percent of the data are greater than laboratory reporting limits), categorical statistical methods were used exclusively for quantitative comparisons of pesticide detection frequencies between seasons and among various natural and anthropogenic (human-related) characteristics. One or more pesticides were detected at concentrations greater than the CAL in water from about 23 percent of wells sampled in the fall and from about 22 percent of wells sampled in the spring. Mixtures of two or more pesticides occurred at concentrations greater than the CAL in about 9 percent of wells sampled in the fall and in about 10 percent of wells sampled in the spring. At least 74 percent of pesticides detected were classified as herbicides. Considering only detections using the CAL, triazine pesticides were detected much more frequently than all other pesticide classes, and the number of different pesticides classified as triazines was the largest of all classes. More pesticides were detected at concentrations greater than the CSALs in water from wells sampled in the fall (28 different pesticides) than in the spring (21 different pesticides). Many pesticides were detected infrequently as nearly one-half of pesticides detected in the fall and spring at concentrations greater than the CSALs were detected only in one well. Using the CSALs for pesticides analyzed for in 11 or more wells, only five pesticides (atrazine, prometon, tebuthiuron, picloram, and 3,4-dichloroaniline, listed in order of decreasing detection frequency) were each detected in water from more than 5 percent of sampled wells. Atrazine was the pesticide detected most frequently at concentrations greater than the CSAL. Concentrations of detected pesticides generally were small (less than 1 microgram per liter), although many infrequent detections at larger concentrations were noted. All detected pesticide concentrations were smaller than U.S. Environmental Protection Agency (USEPA) drinking-water standards or applicable health advisories. Most concentrations were at least an order of magnitude smaller; however, many pesticides did not have standards or advisories. The largest percentage of pesticide detections and the largest number of different pesticides detected were in samples from wells located in the Bighorn Basin and High Plains/ Casper Arch geographic areas of north-central and southeastern Wyoming. Prometon was the only pesticide detected in all eight geographic areas of the State. Pesticides were detected much more frequently in samples from wells located in predominantly urban areas than in samples from wells located in predominantly agricultural or mixed areas. Pesticides were detected distinctly less often in sa

Scientific Investigations Report↗

The National Streamflow Statistics Program: A Computer Program for Estimating Streamflow Statistics for Ungaged Sites

The National Streamflow Statistics (NSS) Program is a computer program that should be useful to engineers, hydrologists, and others for planning, management, and design applications. NSS compiles all current U.S. Geological Survey (USGS) regional regression equations for estimating streamflow statistics at ungaged sites in an easy-to-use interface that operates on computers with Microsoft Windows operating systems. NSS expands on the functionality of the USGS National Flood Frequency Program, and replaces it. The regression equations included in NSS are used to transfer streamflow statistics from gaged to ungaged sites through the use of watershed and climatic characteristics as explanatory or predictor variables. Generally, the equations were developed on a statewide or metropolitan-area basis as part of cooperative study programs. Equations are available for estimating rural and urban flood-frequency statistics, such as the 1 00-year flood, for every state, for Puerto Rico, and for the island of Tutuila, American Samoa. Equations are available for estimating other statistics, such as the mean annual flow, monthly mean flows, flow-duration percentiles, and low-flow frequencies (such as the 7-day, 0-year low flow) for less than half of the states. All equations available for estimating streamflow statistics other than flood-frequency statistics assume rural (non-regulated, non-urbanized) conditions. The NSS output provides indicators of the accuracy of the estimated streamflow statistics. The indicators may include any combination of the standard error of estimate, the standard error of prediction, the equivalent years of record, or 90 percent prediction intervals, depending on what was provided by the authors of the equations. The program includes several other features that can be used only for flood-frequency estimation. These include the ability to generate flood-frequency plots, and plots of typical flood hydrographs for selected recurrence intervals, estimates of the probable maximum flood, extrapolation of the 500-year flood when an equation for estimating it is not available, and weighting techniques to improve flood-frequency estimates for gaging stations and ungaged sites on gaged streams. This report describes the regionalization techniques used to develop the equations in NSS and provides guidance on the applicability and limitations of the techniques. The report also includes a users manual and a summary of equations available for estimating basin lagtime, which is needed by the program to generate flood hydrographs. The NSS software and accompanying database, and the documentation for the regression equations included in NSS, are available on the Web at http://water.usgs.gov/software/.

Techniques and Methods↗

Proximate and landscape factors influence grassland bird distributions

Ecologists increasingly recognize that birds can respond to features well beyond their normal areas of activity, but little is known about the relative importance of landscapes and proximate factors or about the scales of landscapes that influence bird distributions. We examined the influences of tree cover at both proximate and landscape scales on grassland birds, a group of birds of high conservation concern, in the Sheyenne National Grassland in North Dakota, USA. The Grassland contains a diverse array of grassland and woodland habitats. We surveyed breeding birds on 2015 100 m long transect segments during 2002 and 2003. We modeled the occurrence of 19 species in relation to habitat features (percentages of grassland, woodland, shrubland, and wetland) within each 100-m segment and to tree cover within 200-1600 m of the segment. We used information-theoretic statistical methods to compare models and variables. At the proximate scales, tree cover was the most important variable, having negative influences on 13 species and positive influences on two species. In a comparison of multiple scales, models with only proximate variables were adequate for some species, but models combining proximate with landscape information were best for 17 of 19 species. Landscape-only models were rarely competitive. Combined models at the largest scales (800-1600 m) were best for 12 of 19 species. Seven species had best models including 1600-m landscapes plus proximate factors in at least one year. These were Wilson's Phalarope (Phalaropus tricolor), Sedge Wren (Cistothorus platensis), Field Sparrow (Spizella pusilla), Grasshopper Sparrow (Ammodramus savannarum), Bobolink (Dolychonix oryzivorus), Red-winged Blackbird (Agelaius phoeniceus), and Brown-headed Cowbird (Molothrus ater). These seven are small-bodied species; thus larger-bodied species do not necessarily respond most to the largest landscapes. Our findings suggest that birds respond to habitat features at a variety of scales. Models with only landscape-scale tree cover were rarely competitive, indicating that broad-scale modeling alone, such as that based solely on remotely sensed data, is likely to be inadequate in explaining species distributions. ?? 2006 by the Ecological Society of America.

Ecological Applications↗

Methods for estimating selected low-flow frequency statistics and mean annual flow for ungaged locations on streams in Alabama

Streamflow data and statistics are vitally important for proper protection and management of the water quality and water quantity of Alabama streams. Such data and statistics are generally available at U.S. Geological Survey streamflow-gaging stations, also referred to as streamgages or stations, but are often needed at ungaged stream locations. To address this need, the U.S. Geological Survey, in cooperation with numerous Alabama State agencies and organizations, developed regional regression equations for estimating selected low-flow frequency statistics and mean annual flow for ungaged locations on streams in Alabama that are not substantially affected by tides, regulation, diversions, or other anthropogenic influences. A small percentage of the streamgages included in this study experience zero flows during certain periods; thus, the final low-flow frequency regression equations were developed by using weighted left-censored regression analyses to analyze the flow data in an unbiased manner, with weights based on number of years of record. The equations developed include the annual minimum 1- and 7-day average streamflows with a 10-year recurrence interval (referred to as the 1Q10 and 7Q10 flows), the annual minimum 7-day average streamflow with a 2-year recurrence interval (referred to as the 7Q2 flow), and the mean annual flow using data from 174 streamgages from Alabama and surrounding States. For the 1Q10, 7Q2, and 7Q10 low-flow frequency statistics, the regional regression equations are functions of drainage area, streamflow-variability index, mean annual precipitation, and percentage of the drainage basin located in the Piedmont and Southeastern Plains ecoregions. The mean annual flow regression equation is a function of drainage area, mean annual precipitation, and percentage of the drainage basin located in the Southeastern Plains ecoregion. For the mean annual flow regression equation, the average standard error of estimate was 11.2 percent. For the selected low-flow frequency equations, the average standard errors of estimate ranged from 18.1 to 38.8 percent. The regional regression equations developed from this investigation have been incorporated into the U.S. Geological Survey StreamStats application for Alabama. StreamStats ( https://streamstats.usgs.gov/ss/ ) is a web-based geographic information system application that delineates drainage basins at selected stream locations and then generates the needed basin characteristics for available regional regression equations. Along with the low-flow frequency equations developed in this investigation, the StreamStats application also has regional regression equations for estimating flood-frequency statistics at locations on rural and urban streams in Alabama.

Alabama↗

Multivariate Statistical Models for Predicting Sediment Yields from Southern California Watersheds

Debris-retention basins in Southern California are frequently used to protect communities and infrastructure from the hazards of flooding and debris flow. Empirical models that predict sediment yields are used to determine the size of the basins. Such models have been developed using analyses of records of the amount of material removed from debris retention basins, associated rainfall amounts, measures of watershed characteristics, and wildfire extent and history. In this study we used multiple linear regression methods to develop two updated empirical models to predict sediment yields for watersheds located in Southern California. The models are based on both new and existing measures of volume of sediment removed from debris retention basins, measures of watershed morphology, and characterization of burn severity distributions for watersheds located in Ventura, Los Angeles, and San Bernardino Counties. The first model presented reflects conditions in watersheds located throughout the Transverse Ranges of Southern California and is based on volumes of sediment measured following single storm events with known rainfall conditions. The second model presented is specific to conditions in Ventura County watersheds and was developed using volumes of sediment measured following multiple storm events. To relate sediment volumes to triggering storm rainfall, a rainfall threshold was developed to identify storms likely to have caused sediment deposition. A measured volume of sediment deposited by numerous storms was parsed among the threshold-exceeding storms based on relative storm rainfall totals. The predictive strength of the two models developed here, and of previously-published models, was evaluated using a test dataset consisting of 65 volumes of sediment yields measured in Southern California. The evaluation indicated that the model developed using information from single storm events in the Transverse Ranges best predicted sediment yields for watersheds in San Bernardino, Los Angeles, and Ventura Counties. This model predicts sediment yield as a function of the peak 1-hour rainfall, the watershed area burned by the most recent fire (at all severities), the time since the most recent fire, watershed area, average gradient, and relief ratio. The model that reflects conditions specific to Ventura County watersheds consistently under-predicted sediment yields and is not recommended for application. Some previously-published models performed reasonably well, while others either under-predicted sediment yields or had a larger range of errors in the predicted sediment yields.

Open-File Report↗

A review of N-mixture models

N-mixture models were born in 2004 of the necessity to model animal population size from point counts with imperfect detection of individuals, where capture-recapture methods are infeasible. Initially developed for applications where population size was assumed constant, N-mixture models were extended in 2011 to include population dynamics, allowing application to populations whose size fluctuates during the study. A further extension in 2014 accommodates populations with multiple “states” such as age class or sex. More recent extensions model spatial movement of animals among habitat patches or the spatial spread of infectious disease in a human population. The core idea underlying this class of models is a hierarchical structure, where the observation model is defined conditional on the model for true abundance. This hierarchy allows researchers to incorporate information about observation and abundance processes, while permitting distinct inferences about elements affecting detection and those affecting abundance. Another benefit of the hierarchical approach is the ability to accommodate many existing sampling protocols such as removal sampling and distance sampling. One drawback to N-mixture models is that since they estimate both abundance and detection from replicated but unmarked counts, model parameters may not be clearly identifiable. A second drawback is that when observed counts are large, calculating the N-mixture likelihood is computationally infeasible. This difficulty motivated an approximate likelihood based on the normal approximation to the binomial. The normal approximation provides a diagnostic of parameter estimability based on the closed-form expression of the Fisher information matrix for a multivariate normal likelihood.

WIREs Computational Statistics↗

Simulation of groundwater flow and analysis of the effects of water-management options in the North Platte Natural Resources District, Nebraska

The North Platte Natural Resources District (NPNRD) has been actively collecting data and studying groundwater resources because of concerns about the future availability of the highly inter-connected surface-water and groundwater resources. This report, prepared by the U.S. Geological Survey in cooperation with the North Platte Natural Resources District, describes a groundwater-flow model of the North Platte River valley from Bridgeport, Nebraska, extending west to 6 miles into Wyoming. The model was built to improve the understanding of the interaction of surface-water and groundwater resources, and as an optimization tool, the model is able to analyze the effects of water-management options on the simulated stream base flow of the North Platte River. The groundwater system and related sources and sinks of water were simulated using a newton formulation of the U.S. Geological Survey modular three-dimensional groundwater model, referred to as MODFLOW–NWT, which provided an improved ability to solve nonlinear unconfined aquifer simulations with wetting and drying of cells. Using previously published aquifer-base-altitude contours in conjunction with newer test-hole and geophysical data, a new base-of-aquifer altitude map was generated because of the strong effect of the aquifer-base topography on groundwater-flow direction and magnitude. The largest inflow to groundwater is recharge originating from water leaking from canals, which is much larger than recharge originating from infiltration of precipitation. The largest component of groundwater discharge from the study area is to the North Platte River and its tributaries, with smaller amounts of discharge to evapotranspiration and groundwater withdrawals for irrigation. Recharge from infiltration of precipitation was estimated with a daily soil-water-balance model. Annual recharge from canal seepage was estimated using available records from the Bureau of Reclamation and then modified with canal-seepage potentials estimated using geophysical data. Groundwater withdrawals were estimated using land-cover data, precipitation data, and published crop water-use data. For fields irrigated with surface water and groundwater, surface-water deliveries were subtracted from the estimated net irrigation requirement, and groundwater withdrawal was assumed to be equal to any demand unmet by surface water. The groundwater-flow model was calibrated to measured groundwater levels and stream base flows estimated using the base-flow index method. The model was calibrated through automated adjustments using statistical techniques through parameter estimation using the parameter estimation suite of software (PEST). PEST was used to adjust 273 parameters, grouped as hydraulic conductivity of the aquifer, spatial multipliers to recharge, temporal multipliers to recharge, and two specific recharge parameters. Base flow of the North Platte River at Bridgeport, Nebraska, streamgage near the eastern, downstream end of the model was one of the primary calibration targets. Simulated base flow reasonably matched estimated base flow for this streamgage during 1950–2008, with an average difference of 15 percent. Overall, 1950–2008 simulated base flow followed the trend of the estimated base flow reasonably well, in cases with generally increasing or decreasing base flow from the start of the simulation to the end. Simulated base flow also matched estimated base flow reasonably well for most of the North Platte River tributaries with estimated base flow. Average simulated groundwater budgets during 1989–2008 were nearly three times larger for irrigation seasons than for non-irrigation seasons. The calibrated groundwater-flow model was used with the Groundwater-Management Process for the 2005 version of the U.S. Geological Survey modular three-dimensional groundwater model, MODFLOW–2005, to provide a tool for the NPNRD to better understand how water-management decisions could affect stream base flows of the North Platte River at Bridgeport, Nebr., streamgage in a future period from 2008 to 2019 under varying climatic conditions. The simulation-optimization model was constructed to analyze the maximum increase in simulated stream base flow that could be obtained with the minimum amount of reductions in groundwater withdrawals for irrigation. A second analysis extended the first to analyze the simulated base-flow benefit of groundwater withdrawals along with application of intentional recharge, that is, water from canals being released into rangeland areas with sandy soils. With optimized groundwater withdrawals and intentional recharge, the maximum simulated stream base flow was 15–23 cubic feet per second (ft 3 /s) greater than with no management at all, or 10–15 ft 3 /s larger than with managed groundwater withdrawals only. These results indicate not only the amount that simulated stream base flow can be increased by these management options, but also the locations where the management options provide the most or least benefit to the simulated stream base flow. For the analyses in this report, simulated base flow was best optimized by reductions in groundwater withdrawals north of the North Platte River and in the western half of the area. Intentional recharge sites selected by the optimization had a complex distribution but were more likely to be closer to the North Platte River or its tributaries. Future users of the simulation-optimization model will be able to modify the input files as to type, location, and timing of constraints, decision variables of groundwater withdrawals by zone, and other variables to explore other feasible management scenarios that may yield different increases in simulated future base flow of the North Platte River.

Nebraska↗

Methods for estimating annual exceedance-probability discharges for streams in Iowa, based on data through water year 2010

A statewide study was performed to develop regional regression equations for estimating selected annual exceedance-probability statistics for ungaged stream sites in Iowa. The study area comprises streamgages located within Iowa and 50 miles beyond the State’s borders. Annual exceedance-probability estimates were computed for 518 streamgages by using the expected moments algorithm to fit a Pearson Type III distribution to the logarithms of annual peak discharges for each streamgage using annual peak-discharge data through 2010. The estimation of the selected statistics included a Bayesian weighted least-squares/generalized least-squares regression analysis to update regional skew coefficients for the 518 streamgages. Low-outlier and historic information were incorporated into the annual exceedance-probability analyses, and a generalized Grubbs-Beck test was used to detect multiple potentially influential low flows. Also, geographic information system software was used to measure 59 selected basin characteristics for each streamgage. Regional regression analysis, using generalized least-squares regression, was used to develop a set of equations for each flood region in Iowa for estimating discharges for ungaged stream sites with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities, which are equivalent to annual flood-frequency recurrence intervals of 2, 5, 10, 25, 50, 100, 200, and 500 years, respectively. A total of 394 streamgages were included in the development of regional regression equations for three flood regions (regions 1, 2, and 3) that were defined for Iowa based on landform regions and soil regions. Average standard errors of prediction range from 31.8 to 45.2 percent for flood region 1, 19.4 to 46.8 percent for flood region 2, and 26.5 to 43.1 percent for flood region 3. The pseudo coefficients of determination for the generalized least-squares equations range from 90.8 to 96.2 percent for flood region 1, 91.5 to 97.9 percent for flood region 2, and 92.4 to 96.0 percent for flood region 3. The regression equations are applicable only to stream sites in Iowa with flows not significantly affected by regulation, diversion, channelization, backwater, or urbanization and with basin characteristics within the range of those used to develop the equations. These regression equations will be implemented within the U.S. Geological Survey StreamStats Web-based geographic information system tool. StreamStats allows users to click on any ungaged site on a river and compute estimates of the eight selected statistics; in addition, 90-percent prediction intervals and the measured basin characteristics for the ungaged sites also are provided by the Web-based tool. StreamStats also allows users to click on any streamgage in Iowa and estimates computed for these eight selected statistics are provided for the streamgage.

Iowa↗

Quality of groundwater used for domestic supply in the Gilroy-Hollister basin and surrounding areas, California, 2022

More than 2 million Californians rely on groundwater from domestic wells for drinking-water supply. This report summarizes a 2022 California Groundwater Ambient Monitoring and Assessment Priority Basin Project (GAMA-PBP) water-quality survey of 33 domestic and small-system drinking-water supply wells in the Gilroy-Hollister Valley groundwater basin and the surrounding areas, where more than 20,000 residents are estimated to utilize privately owned domestic wells. The study area includes the Llagas subbasin in the north, the North San Benito subbasin in the south, and the surrounding uplands. The study was focused on groundwater resources used for domestic drinking-water supply, which are mostly drawn from shallower parts of aquifer systems rather than those of groundwater resources used for public drinking-water supply in the same area. This assessment characterized the quality of ambient groundwater in the aquifer before filtration or treatment, rather than the quality of drinking water delivered to the tap. To provide context, the measured concentrations of constituents in groundwater were compared to Federal and California State regulatory and non-regulatory benchmarks for drinking-water quality. A grid-based method was used to estimate the areal proportions of groundwater resources used for domestic drinking wells that have water-quality constituents present at high concentrations (above the benchmark), moderate concentrations (between one-half of the benchmark and the benchmark for inorganic constituents, or between one-tenth of the benchmark and the benchmark for organic constituents), and low concentrations (less than one-half or one-tenth the benchmark for inorganic and organic constituents, respectively). This method provides statistically representative results at the study-area scale and permits comparisons to other GAMA-PBP study areas. In the study area, inorganic constituents in groundwater were greater than regulatory benchmarks (U.S. Environmental Protection Agency [EPA] or State of California maximum contaminant levels [MCLs]) for public drinking-water quality in 24 percent of domestic groundwater resources. The inorganic constituents present at concentrations greater than MCLs for drinking water were nitrate (as nitrogen), barium, chromium, and selenium. Total dissolved solids (TDS) or manganese were present at concentrations greater than the secondary maximum contaminant levels (SMCLs) that the State of California uses as aesthetic-based benchmarks in 48 percent of domestic groundwater resources. No volatile organic compounds or pesticide constituents were present at concentrations greater than regulatory benchmarks. Total coliform bacteria and enterococci were detected in 4 percent of domestic groundwater resources. Per- and polyfluoroalkyl substances (PFAS) were detected in 19 percent of domestic groundwater resources, and 10 percent had concentrations greater than recently enacted (April 2024) EPA MCLs. Physical and chemical factors from natural and anthropogenic sources that could affect the groundwater quality were evaluated using results from statistical testing of associations between constituent concentrations and potential explanatory variables. In this study, relevant physical factors include well construction characteristics, groundwater age, site proximity to groundwater recharge or discharge zones, and potential sources of contamination. Relevant chemical factors include the initial chemistry of the recharge water, the mineralogy of the aquifer sediments, and the subsequent shifts in chemistry as biologic and geologic reactions alter groundwater in the subsurface. Nitrate concentrations were correlated to agricultural land use, distance from the boundary of the Gilroy-Hollister Valley groundwater basin, and the proportion of modern (post-1950s) water captured by the well. Denitrification under anoxic redox conditions can mitigate some nitrate derived from fertilizer application. Total dissolved solids primarily were derived from water-rock interactions with soils and aquifer materials in the study area, but there were high concentrations where agricultural practices contributed additional TDS. Mineralogy of aquifer sediments and rocks also affect barium, selenium, boron, and chromium concentrations in the Gilroy-Hollister Valley groundwater basin. PFAS were positively correlated with urban land use and the proportion of modern water captured by the well.

California↗

Demography of a reintroduced population: moving toward management models for an endangered species, the whooping crane

The reintroduction of threatened and endangered species is now a common method for reestablishing populations. Typically, a fundamental objective of reintroduction is to establish a self-sustaining population. Estimation of demographic parameters in reintroduced populations is critical, as these estimates serve multiple purposes. First, they support evaluation of progress toward the fundamental objective via construction of population viability analyses (PVAs) to predict metrics such as probability of persistence. Second, PVAs can be expanded to support evaluation of management actions, via management modeling. Third, the estimates themselves can support evaluation of the demographic performance of the reintroduced population, e.g., via comparison with wild populations. For each of these purposes, thorough treatment of uncertainties in the estimates is critical. Recently developed statistical methods - namely, hierarchical Bayesian implementations of state-space models - allow for effective integration of different types of uncertainty in estimation. We undertook a demographic estimation effort for a reintroduced population of endangered whooping cranes with the purpose of ultimately developing a Bayesian PVA for determining progress toward establishing a self-sustaining population, and for evaluating potential management actions via a Bayesian PVA-based management model. We evaluated individual and temporal variation in demographic parameters based upon a multi-state mark-recapture model. We found that survival was relatively high across time and varied little by sex. There was some indication that survival varied by release method. Survival was similar to that observed in the wild population. Although overall reproduction in this reintroduced population is poor, birds formed social pairs when relatively young, and once a bird was in a social pair, it had a nearly 50% chance of nesting the following breeding season. Also, once a bird had nested, it had a high probability of nesting again. These results are encouraging considering that survival and reproduction have been major challenges in past reintroductions of this species. The demographic estimates developed will support construction of a management model designed to facilitate exploration of management actions of interest, and will provide critical guidance in future planning for this reintroduction. An approach similar to what we describe could be usefully applied to many reintroduced populations.

Ecological Applications↗

Quick-start guide for version 3.0 of EMINERS - Economic Mineral Resource Simulator

Quantitative mineral resource assessment, as developed by the U.S. Geological Survey (USGS), consists of three parts: (1) development of grade and tonnage mineral deposit models; (2) delineation of tracts permissive for each deposit type; and (3) probabilistic estimation of the numbers of undiscovered deposits for each deposit type (Singer and Menzie, 2010). The estimate of the number of undiscovered deposits at different levels of probability is the input to the EMINERS (Economic Mineral Resource Simulator) program. EMINERS uses a Monte Carlo statistical process to combine probabilistic estimates of undiscovered mineral deposits with models of mineral deposit grade and tonnage to estimate mineral resources. It is based upon a simulation program developed by Root and others (1992), who discussed many of the methods and algorithms of the program. Various versions of the original program (called "MARK3" and developed by David H. Root, William A. Scott, and Lawrence J. Drew of the USGS) have been published (Root, Scott, and Selner, 1996; Duval, 2000, 2012). The current version (3.0) of the EMINERS program is available as USGS Open-File Report 2004-1344 (Duval, 2012). Changes from version 2.0 include updating 87 grade and tonnage models, designing new templates to produce graphs showing cumulative distribution and summary tables, and disabling economic filters. The economic filters were disabled because embedded data for costs of labor and materials, mining techniques, and beneficiation methods are out of date. However, the cost algorithms used in the disabled economic filters are still in the program and available for reference for mining methods and milling techniques included in Camm (1991). EMINERS is written in C++ and depends upon the Microsoft Visual C++ 6.0 programming environment. The code depends heavily on the use of Microsoft Foundation Classes (MFC) for implementation of the Windows interface. The program works only on Microsoft Windows XP or newer personal computers. It does not work on Macintosh computers. This report demonstrates how to execute EMINERS software using default settings and existing deposit models. Many options are available when setting up the simulation. Information and explanations addressing these optional parameters can be found in the EMINERS Help files. Help files are available during execution of EMINERS by selecting EMINERS Help from the pull-down menu under Help on the EMINERS menu bar. There are four sections in this report. Part I describes the installation, setup, and application of the EMINERS program, and Part II illustrates how to interpret the text file that is produced. Part III describes the creation of tables and graphs by use of the provided Excel templates. Part IV summarizes grade and tonnage models used in version 3.0 of EMINERS.

Open-File Report↗