USGS ScienceSearch

SEARCH · USGS Science

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Predicting paleoclimate from compositional data using multivariate Gaussian process inverse prediction

Multivariate compositional count data arise in many applications including ecology, microbiology, genetics and paleoclimate. A frequent question in the analysis of multivariate compositional count data is what underlying values of a covariate(s) give rise to the observed composition. Learning the relationship between covariates and the compositional count allows for inverse prediction of unobserved covariates given compositional count observations. Gaussian processes provide a flexible framework for modeling functional responses with respect to a covariate without assuming a functional form. Many scientific disciplines use Gaussian process approximations to improve prediction and make inference on latent processes and parameters. When prediction is desired on unobserved covariates given realizations of the response variable, this is called inverse prediction. Because inverse prediction is often mathematically and computationally challenging, predicting unobserved covariates often requires fitting models that are different from the hypothesized generative model. We present a novel computational framework that allows for efficient inverse prediction using a Gaussian process approximation to generative models. Our framework enables scientific learning about how the latent processes co-vary with respect to covariates while simultaneously providing predictions of missing covariates. The proposed framework is capable of efficiently exploring the high dimensional, multi-modal latent spaces that arise in the inverse problem. To demonstrate flexibility, we apply our method in a generalized linear model framework to predict latent climate states given multivariate count data. Based on cross-validation, our model has predictive skill competitive with current methods while simultaneously providing formal, statistical inference on the underlying community dynamics of the biological system previously not available.

Annals of Applied Statistics

Scale-dependent approaches to modeling spatial epidemiology of chronic wasting disease.

This e-book is the product of a second workshop that was funded and promoted by the United States Geological Survey to enhance cooperation between states for the management of chronic wasting disease (CWD). The first workshop addressed issues surrounding the statistical design and collection of surveillance data for CWD. The second workshop, from which this document arose, followed logically from the first workshop and focused on appropriate methods for analysis, interpretation, and use of CWD surveillance and related epidemiology data. Consequently, the emphasis of this e-book is on modeling approaches to describe and gain insight of the spatial epidemiology of CWD. We designed this e-book for wildlife managers and biologists who are responsible for the surveillance of CWD in their state or agency. We chose spatial methods that are popular or common in the spatial epidemiology literature and evaluated them for their relevance to modeling CWD. Our opinion of the usefulness and relevance of each method was based on the type of field data commonly collected as part of CWD surveillance programs and what we know about CWD biology, ecology, and epidemiology. Specifically, we expected the field data to consist primarily of the infection status of a harvested or culled sample along with its date of collection (not date of infection), location, and demographic status. We evaluated methods in light of the fact that CWD does not appear to spread rapidly through wild populations, relative to more highly contagious viruses, and can be spread directly from animal to animal or indirectly through environmental contamination. We discovered that many of the wellpublished methods were developed for fast-spreading human diseases, such as influenza and measles. While these methods are applicable to fast spreading wildlife diseases, such as foot-and-mouth disease or West Nile virus, many are not likely to work well for CWD. Only limited data exist to evaluate geographic and spatial spread because many locations where we find CWD tend to be locations where samples have just been taken or sample sizes have just become large enough to have a high probability of detecting a low prevalence. Consequently, methods that work well to describe or predict the spread of foot-and-mouth disease throughout England, which occurred within a year, do not work well for describing or predicting CWD spread. We did not exclude methods that we regarded as inappropriate; rather, we included methods that are commonly used for disease epidemiology and then discussed their applicability for modeling the spatial epidemiology of CWD. We hope including inappropriate methods with an explanation of why they are ill-suited for CWD will make it easier to drop them from consideration and explain to others why they were not recommended for spatial modeling of CWD. We organized the three chapters by scale and extent for which each method was developed or best suited. The first chapter covers methods appropriate to multi-jurisdictional or multi-state modeling, which we call “regional” scale. The second chapter covers methods appropriate for within state areas such as wildlife management units or metapopulations, which we call “landscape” scale. The third chapter covers methods appropriate for population or individual-based modeling, which we call “fine” scale. We know this rubric is somewhat artificial because many methods work at multiple scales. We hope, however, that this structure addresses some of the challenges faced by managers that work at local, regional, state, and national scales. Further, the resolution of empirical data often changes with spatial scale, which affects the utility of different modeling approaches. For example, individual-based models work best at modeling spread within populations, while risk analysis is most useful for summarizing data over larger scales such as a region. Because some methods are applicable at several scales, however, we included a graphic at the beginning of each method that indicates the range of scales for which it applies. For example, the graphic to the right indicates that the method is most applicable for regional-scale modeling. There is also a question of resolution as well as scale and extent for each method. CWD surveillance data have been collected over large areas, such as a wildlife management unit or state, but the resolution of the data may be fine scale with GPS locations for many samples. For each method, we described the required resolution of the data and describe the type of data required, as well as what questions the method could answer and how useful the method is, given typical CWD data. For each scale, we presented a focal approach that would be useful for understanding the spatial pattern and epidemiology of CWD, as well as being a useful tool for CWD management. The focal approaches include risk analysis and micromaps for the regional scale, cluster analysis for the landscape scale, and individual based modeling for the fine scale of within population. For each of these methods, we used simulated data and walked through the method step by step to fully illustrate the “how to”, with specifics about what is input and output, as well as what questions the method addresses. We also provided a summary table to, at a glance, describe the scale, questions that can be addressed, and general data required for each method described in this e-book. We hope that this review will be helpful to biologists and managers by increasing the utility of their surveillance data, and ultimately be useful for increasing our understanding of CWD and allowing wildlife biologists and managers to move beyond retroactive fire-fighting to proactive preventative action.

Book

Experimental evaluation of spatial capture–recapture study design

A principal challenge impeding strong inference in analyses of wild populations is the lack of robust and long-term data sets. Recent advancements in analytical tools used in wildlife science may increase our ability to integrate smaller data sets and enhance the statistical power of population estimates. One such advancement, the development of spatial capture–recapture (SCR) methods, explicitly accounts for differences in spatial study designs, making it possible to equate multiple study designs in one analysis. SCR has been shown to be robust to variation in design as long as minimal sampling guidance is adhered to. However, these expectations are based on simulation and have yet to be evaluated in wild populations. Here we conduct a rigorously designed field experiment by manipulating the arrangement of artificial cover objects (ACOs) used to collect data on red-backed salamanders ( Plethodon cinereus ) to empirically evaluate the effects of design configuration on inference made using SCR. Our results suggest that, using SCR, estimates of space use and detectability are sensitive to study design configuration, namely the spacing and extent of the array, and that caution is warranted when assigning biological interpretation to these parameters. However, estimates of population density remain robust to design except when the configuration of detectors grossly violates existing recommendations.

Massachusetts

Methods for estimating annual exceedance probability discharges for streams in Arkansas, based on data through water year 2013

In 2013, the U.S. Geological Survey initiated a study to update regional skew, annual exceedance probability discharges, and regional regression equations used to estimate annual exceedance probability discharges for ungaged locations on streams in the study area with the use of recent geospatial data, new analytical methods, and available annual peak-discharge data through the 2013 water year. An analysis of regional skew using Bayesian weighted least-squares/Bayesian generalized-least squares regression was performed for Arkansas, Louisiana, and parts of Missouri and Oklahoma. The newly developed constant regional skew of -0.17 was used in the computation of annual exceedance probability discharges for 281 streamgages used in the regional regression analysis. Based on analysis of covariance, four flood regions were identified for use in the generation of regional regression models. Thirty-nine basin characteristics were considered as potential explanatory variables, and ordinary least-squares regression techniques were used to determine the optimum combinations of basin characteristics for each of the four regions. Basin characteristics in candidate models were evaluated based on multicollinearity with other basin characteristics (variance inflation factor < 2.5) and statistical significance at the 95-percent confidence level ( p ≤ 0.05). Generalized least-squares regression was used to develop the final regression models for each flood region. Average standard errors of prediction of the generalized least-squares models ranged from 32.76 to 59.53 percent, with the largest range in flood region D. Pseudo coefficients of determination of the generalized least-squares models ranged from 90.29 to 97.28 percent, with the largest range also in flood region D. The regional regression equations apply only to locations on streams in Arkansas where annual peak discharges are not substantially affected by regulation, diversion, channelization, backwater, or urbanization. The applicability and accuracy of the regional regression equations depend on the basin characteristics measured for an ungaged location on a stream being within range of those used to develop the equations.

Arkansas

Multiscale sagebrush rangeland habitat modeling in the Gunnison Basin of Colorado

North American sagebrush-steppe ecosystems have decreased by about 50 percent since European settlement. As a result, sagebrush-steppe dependent species, such as the Gunnison sage-grouse, have experienced drastic range contractions and population declines. Coordinated ecosystem-wide research, integrated with monitoring and management activities, is needed to help maintain existing sagebrush habitats; however, products that accurately model and map sagebrush habitats in detail over the Gunnison Basin in Colorado are still unavailable. The goal of this project is to provide a rigorous large-area sagebrush habitat classification and inventory with statistically validated products and estimates of precision across the Gunnison Basin. This research employs a combination of methods, including (1) modeling sagebrush rangeland as a series of independent objective components that can be combined and customized by any user at multiple spatial scales; (2) collecting ground measured plot data on 2.4-meter QuickBird satellite imagery in the same season the imagery is acquired; (3) modeling of ground measured data on 2.4-meter imagery to maximize subsequent extrapolation; (4) acquiring multiple seasons (spring, summer, and fall) of Landsat Thematic Mapper imagery (30-meter) for optimal modeling; (5) using regression tree classification technology that optimizes data mining of multiple image dates, ratios, and bands with ancillary data to extrapolate ground training data to coarser resolution Landsat Thematic Mapper; and 6) employing accuracy assessment of model predictions to enable users to understand their dependencies. Results include the prediction of four primary components including percent bare ground, percent herbaceous, percent shrub, and percent litter, and four secondary components including percent sagebrush (Artemisia spp.), percent big sagebrush (Artemisia tridentata), percent Wyoming sagebrush (Artemisia tridentata wyomingensis), and shrub height (centimeters). Results were validated with an independent accuracy assessment, with root mean square error values ranging from 3.5 (percent big sagebrush) to 10.8 (percent bare ground) at the QuickBird scale, and from 4.5 (percent Wyoming sagebrush) to 12.4 (percent herbaceous) at the full Landsat scale. These results offer significant improvement in sagebrush ecosystem quantification across the Gunnison Basin, and also provide maximum flexibility to users to employ for a wide variety of applications. Further refinement of these remote sensing component predictions in the future will be most likely achieved by focusing on more extensive ground plot sampling, employing new high and moderate-resolution satellite sensors that offer additional spectral bands for vegetation discrimination, and capturing more dates of satellite imagery to better represent phenological variation.

Colorado

Spatial capture–recapture with partial identity: An application to camera traps

Camera trapping surveys frequently capture individuals whose identity is only known from a single flank. The most widely used methods for incorporating these partial identity individuals into density analyses discard some of the partial identity capture histories, reducing precision, and, while not previously recognized, introducing bias. Here, we present the spatial partial identity model (SPIM), which uses the spatial location where partial identity samples are captured to probabilistically resolve their complete identities, allowing all partial identity samples to be used in the analysis. We show that the SPIM outperforms other analytical alternatives. We then apply the SPIM to an ocelot data set collected on a trapping array with double-camera stations and a bobcat data set collected on a trapping array with single-camera stations. The SPIM improves inference in both cases and, in the ocelot example, individual sex is determined from photographs used to further resolve partial identities—one of which is resolved to near certainty. The SPIM opens the door for the investigation of trapping designs that deviate from the standard two camera design, the combination of other data types between which identities cannot be deterministically linked, and can be extended to the problem of partial genotypes.

Annals of Applied Statistics

Whole-coal versus ash basis in coal geochemistry: a mathematical approach to consistent interpretations

Several standard methods require coal to be ashed prior to geochemical analysis. Researchers, however, are commonly interested in the compositional nature of the whole-coal, not its ash. Coal geochemical data for any given sample can, therefore, be reported in the ash basis on which it is analyzed or the whole-coal basis to which the ash basis data are back calculated. Basic univariate (mean, variance, distribution, etc.) and bivariate (correlation coefficients, etc.) measures of the same suite of samples can be very different depending which reporting basis the researcher uses. These differences are not real, but an artifact resulting from the compositional nature of most geochemical data. The technical term for this artifact is subcompositional incoherence. Since compositional data are forced to a constant sum, such as 100% or 1,000,000 ppm, they possess curvilinear properties which make the Euclidean principles on which most statistical tests rely inappropriate, leading to erroneous results. Applying the isometric logratio (ilr) transformation to compositional data allows them to be represented in Euclidean space and evaluated using traditional tests without fear of producing mathematically inconsistent results. When applied to coal geochemical data, the issues related to differences between the two reporting bases are resolved as demonstrated in this paper using major oxide and trace metal data from the Pennsylvanian-age Pond Creek coal of eastern Kentucky, USA. Following ilr transformation, univariate statistics, such as mean and variance, still differ between the ash basis and whole-coal basis, but in predictable and calculated manners. Further, the stability between two different components, a bivariate measure, is identical, regardless of the reporting basis. The application of ilr transformations addresses both the erroneous results of Euclidean-based measurements on compositional data as well as the inconsistencies observed on coal geochemical data reported on different bases.

International Journal of Coal Geology

Building a state-space life cycle model for naturally produced Snake River fall Chinook salmon

In 1992, Snake River basin fall Chinook salmon (Oncorhynchus tshawytscha) were listed for protection under the U.S. Endangered Species Act (NMFS 1992) and the population remained below 1000 individuals until 2000. Since then, returns from natural production has rebounded to over 20,000 spawners owing to a host of factors including reduced harvest (Peters et al. 2001), stable minimum spawning flows (Groves and Chandler 1999), summer flow augmentation (Connor et al. 2003), predator control (Beamesderfer et al. 1996), hatchery supplementation (Rosenberger et al. 2017), improved juvenile passage structures (Adams et al. 2014), summer spill operations (Perry et al. 2006; Adams et al. 2008), and periods of favorable ocean conditions and food availability (Logerwell et al. 2003; Peterson et al. 2014). Given this change in abundance coincident with numerous management actions and fluctuation in environmental drivers, quantifying which factors contributed to the observed rebound in natural production can provide critical insights into future management actions for this at-risk population. Multistage life cycle models provide a powerful analytical framework for understating how each life stage of a population contributes to population growth rate (Moussalli and Hilborn 1986; Greene and Beechie 2004). Multistage models may also be used as an analytical framework to explicitly estimate demographic parameters of a population model. This approach has an advantage over single-stage stock-recruitment models by allowing population growth rates to be partitioned among life stages rather than aggregated over an entire life cycle. Such partitioning allows for estimating 1) stage-specific density dependence, and 2) stage-specific effects of environmental factors or management actions. For example, Zabel et al. (2006) estimated parameters of a multistage model used in the context of a population viability analysis for spring/summer Chinook salmon in the Snake River, but such an approach has yet to be applied to fall Chinook salmon in the Snake River basin. Typically, data informing estimates of abundance at particular “check points” in the life cycle determines the complexity of the multistage model that can be fit to the data. For fall Chinook salmon, we are developing a two-stage model that encompasses: 1) upstream passage of spawners at Lower Granite Dam (LGR) to the subsequent downstream passage of their progeny at the dam, and 2) downstream passage of juveniles at LGR to their subsequent return from the ocean and passage at the Dam 2‒6 years later. This approach partitions the life cycle of fall Chinook salmon both spatially and temporally, which allows us to fit and compare alternative models with covariates specific to each stage. Our previous report to the ISAB (Zabel et al. 2013) detailed methods for estimating abundance of naturally produced adults and juveniles passing Lower Granite Dam, which provides the requisite data for fitting a two-stage model. The intent of this report is to describe the structure of the two-stage life cycle model, present preliminary results from fitting the model to data, and outline future directions and developments. As is clear from the diversity of models presented in this report, “life cycle models” range from very simple theoretically based population models (e.g., the Beverton-Holt stock- recruitment model) to very complex spatially explicit simulation models linked to hydrosystem hydrodynamic models (e.g., the COMPASS model for a single transition in a life cycle model, Zabel et al. 2008). We chose to develop a model of intermediate complexity that casts the two- stage life cycle model in a state-space framework (Newman et al. 2014). We chose to use a state-space framework implemented in a Bayesian framework because: • It provides both a statistical estimation framework for retrospective statistical analysis and a stochastic simulation framework for prospective analysis to evaluate alternative management actions. • Abundance estimates are uncertain. A state-space framework accounts for observation uncertainty in the abundance estimates and other data (e.g., age structure) while simultaneously estimating process uncertainty. • It allows for missing data. By drawing missing data from an appropriate probability model, uncertainty owing to missing data can be propagated without having to omit data or assume fixed values for missing data. Thus, a two-stage state-space life cycle model for fall Chinook salmon strikes an appropriate balance between model complexity, tractability, and applicability given the goals of performing both retrospective and prospective analysis to guide future management of this population.

Idaho, Oregon, Washington, Wyoming

Magnitude and frequency of floods on Kauaʻi, Oʻahu, Molokaʻi, Maui, and Hawaiʻi, State of Hawaiʻi, based on data through water year 2020

Accurate estimates of flood magnitude and frequency are needed to (1) optimize the design and location of infrastructure, including dams, culverts, bridges, industrial buildings, and highways, and (2) inform flood-zoning and flood-insurance studies. The U.S. Geological Survey (USGS), in cooperation with the State of Hawaiʻi Department of Transportation, estimated flood magnitudes for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities (AEP) for unregulated streamgages in Kauaʻi, Oʻahu, Molokaʻi, Maui, and Hawaiʻi, State of Hawaiʻi, using data through water year 2020. Regression equations were developed to estimate flood magnitude and associated frequency at ungaged streams. This study improves upon a previous USGS flood-frequency report (Oki and others, 2010) by including more peak-flow data, implementing new statistical methods in flood-frequency analysis, and using updated techniques to estimate the regional-skewness coefficient (regional skew). Flood magnitude and frequency at 238 streamgages were estimated—following national guidelines established in Bulletin 17C (England and others, 2019)—by fitting annual peak-flow data to the Log-Pearson Type III distribution using the expected moments algorithm and the PeakFQ flood-frequency software. Potentially influential low outliers in the data were identified and removed using the Multiple Grubbs-Beck Test. An updated regional skew for Hawaiʻi was estimated using the Bayesian weighted least squares/Bayesian generalized least squares method. The updated regional skew employs a constant model for the five islands in the study area and has a value of −0.157 (mean square error of 0.212). Multiple linear regression techniques were used to develop regression equations that relate basin and climatic characteristics to peak flows at streamgages. The regression equations can be applied to estimate flood magnitude and frequency at ungaged sites. The study area was split into 10 regions—2 regions per island, generally following a leeward/windward division—containing from 9 to 49 streamgages each. The final regression equations for each region were determined with generalized least-squares analysis using the USGS weighted-multiple-linear regression (WREG) program. The standard error of prediction at the 1-percent AEP for the regression equations ranged from 18 to 164 percent; the pseudo coefficient of determination (pseudo-R2) at the 1-percent AEP ranged from 46 to 100 percent. The regression equations performed well for all regions except leeward Molokaʻi and southern Island of Hawaiʻi; for all other regions, the pseudo-R2 values ranged from about 75 to 100 percent. Compared to the regression equations developed by Oki and others (2010), the regression equations in this study generally showed modest improvements, although the magnitude of differences varied for each region. Peak-flow estimates at the 238 streamgages included in this study are improved by weighting the at-site statistics computed with PeakFQ and the predicted flows based on the regression equations. Results of this study—including the final peak-flow estimates at streamgages and the regional regression equations—are implemented in the USGS StreamStats web application (U.S. Geological Survey, 2023, StreamStats: https://streamstats.usgs.gov/ss/ ). StreamStats provides a consistent approach for obtaining peak-flow estimates at streamgages and for applying the regional regression equations for estimating peak flows at ungaged locations.

Hawaii

Modeling vegetation heights from high resolution stereo aerial photography: an application for broad-scale rangeland monitoring

Vertical vegetation structure in rangeland ecosystems can be a valuable indicator for assessing rangeland health and monitoring riparian areas, post-fire recovery, available forage for livestock, and wildlife habitat. Federal land management agencies are directed to monitor and manage rangelands at landscapes scales, but traditional field methods for measuring vegetation heights are often too costly and time consuming to apply at these broad scales. Most emerging remote sensing techniques capable of measuring surface and vegetation height (e.g., LiDAR or synthetic aperture radar) are often too expensive, and require specialized sensors. An alternative remote sensing approach that is potentially more practical for managers is to measure vegetation heights from digital stereo aerial photographs. As aerial photography is already commonly used for rangeland monitoring, acquiring it in stereo enables three-dimensional modeling and estimation of vegetation height. The purpose of this study was to test the feasibility and accuracy of estimating shrub heights from high-resolution (HR, 3-cm ground sampling distance) digital stereo-pair aerial images. Overlapping HR imagery was taken in March 2009 near Lake Mead, Nevada and 5-cm resolution digital surface models (DSMs) were created by photogrammetric methods (aerial triangulation, digital image matching) for twenty-six test plots. We compared the heights of individual shrubs and plot averages derived from the DSMs to field measurements. We found strong positive correlations between field and image measurements for several metrics. Individual shrub heights tended to be underestimated in the imagery, however, accuracy was higher for dense, compact shrubs compared with shrubs with thin branches. Plot averages of shrub height from DSMs were also strongly correlated to field measurements but consistently underestimated. Grasses and forbs were generally too small to be detected with the resolution of the DSMs. Estimates of vertical structure will be more accurate in plots having low herbaceous cover and high amounts of dense shrubs. Through the use of statistically derived correction factors or choosing field methods that better correlate with the imagery, vegetation heights from HR DSMs could be a valuable technique for broad-scale rangeland monitoring needs.

California;Nevada

An updated genetic marker for detection of Lake Sinai Virus and metagenetic applications

Background Lake Sinai Viruses (LSV) are common RNA viruses of honey bees ( Apis mellifera ) that frequently reach high abundance but are not linked to overt disease. LSVs are genetically heterogeneous and collectively widespread, but despite frequent detection in surveys, the ecological and geographic factors structuring their distribution in A. mellifera are not understood. Even less is known about their distribution in other species. Better understanding of LSV prevalence and ecology have been hampered by high sequence diversity within the LSV clade. Methods Here we report a new polymerase chain reaction (PCR) assay that is compatible with currently known lineages with minimal primer degeneracy, producing an expected 365 bp amplicon suitable for end-point PCR and metagenetic sequencing. Using the Illumina MiSeq platform, we performed pilot metagenetic assessments of three sample sets, each representing a distinct variable that might structure LSV diversity (geography, tissue, and species). Results The first sample set in our pilot assessment compared cDNA pools from managed A. mellifera hives in California ( n = 8) and Maryland ( n = 6) that had previously been evaluated for LSV2, confirming that the primers co-amplify divergent lineages in real-world samples. The second sample set included cDNA pools derived from different tissues (thorax vs. abdomen, n = 24 paired samples), collected from managed A. mellifera hives in North Dakota. End-point detection of LSV frequently differed between the two tissue types; LSV metagenetic composition was similar in one pair of sequenced samples but divergent in a second pair. Overall, LSV1 and intermediate lineages were common in these samples whereas variants clustering with LSV2 were rare. The third sample set included cDNA from individual pollinator specimens collected from diverse landscapes in the vicinity of Lincoln, Nebraska. We detected LSV in the bee Halictus ligatus (four of 63 specimens tested, 6.3%) at a similar rate as A. mellifera (nine of 115 specimens, 7.8%), but only one H. ligatus sequencing library yielded sufficient data for compositional analysis. Sequenced samples often contained multiple divergent LSV lineages, including individual specimens. While these studies were exploratory rather than statistically powerful tests of hypotheses, they illustrate the utility of high-throughput sequencing for understanding LSV transmission within and among species.

California, Maryland, Nebraska

Parameter estimation at the conterminous United States scale and streamflow routing enhancements for the National Hydrologic Model infrastructure application of the Precipitation-Runoff Modeling System (NHM-PRMS)

This report documents a three-part continental-scale calibration procedure and a new streamflow routing algorithm using the U.S. Geological Survey National Hydrologic Model (NHM) infrastructure along with an application of the Precipitation-Runoff Modeling System (PRMS). The traditional approach to hydrologic model calibration and evaluation, which relies on comparing observed and simulated streamflow, is not sufficient for accurately representing the non-streamflow parts of the water budget. If intermediate process variables computed by the hydrologic model are not examined, the variables could be characterized by parameter values that do not replicate those hydrological processes present in the physical system. In answer to this potential problem, alternative hydrologic process variables from the model (in addition to streamflow) are included in a calibration procedure applied to the conterminous United States (CONUS) domain. The three-part calibration procedure presented in this report considers volume (calibration by hydrologic response unit [byHRU]), timing (calibration by headwater watershed [byHW]), and measured streamflow [byHWobs]). The first part, byHRU, is considered a water-balance volume calibration that uses five alternative (non-streamflow) hydrologic quantities (runoff, actual evapotranspiration, recharge, soil moisture, and snow-covered area) as calibration targets for each hydrologic response unit (HRU). These alternative data products were derived, with error bounds, from multiple sources for each of the 109,951 HRUs in the NHM on time scales varying from annual to daily. The second part of the calibration, byHW, is considered a streamflow timing calibration that uses statistically based streamflow simulations developed using ordinary kriging for 7,265 headwater watersheds that had drainage areas of less than 3,000 square kilometers (1,158 square miles) across the CONUS. Two streamflow routing algorithms were tested in this byHW calibration: (1) continuity without attenuation of the flood pulse and (2) a new formulation of the Muskingum routing method, which was added to the PRMS as part of this study. The third part of the calibration, byHWobs, refines the model parameters using available measured streamflow using 1,417 streamgage locations. A multiple-objective, stepwise, automated calibration procedure was used to identify the optimal set of parameters for each calibration procedure. Using a variety of alternative datasets for calibration of the water budget provides users of the NHM-PRMS with improved initial parameters and helps alleviate the equifinality problem (getting the right answer for the wrong reason). Through a community effort, these alternative data products, with error bounds, can be used to improve and expand our understanding of hydrologic-process representation in models. The broader modeling community can use these data products, with error bounds, to calibrate and evaluate hydrologic models using more than streamflow.

Conterminous United States

Procedures for using the Horiba Scientific Aqualog ® fluorometer to measure absorbance and fluorescence from dissolved organic matter

Advances in spectroscopic techniques have led to an increase in the use of optical measurements (absorbance and fluorescence) to assess dissolved organic matter composition and infer sources and processing. Although optical measurements are easy to make, they can be affected by many variables rendering them less comparable, including by inconsistencies in sample collection (for example, filter pore size, preservation), the application of corrections for interferences (for example, inner-filtering corrections), differences in holding times, and instrument drift (for example, lamp intensity). A documented, standardized procedure to address these variables ensures that the optical (absorbance and fluorescence) measurements collected by U.S. Geological Survey researchers are useful and widely comparable. Rigorous and quantifiable quality assurance and quality control are essential for making these data comparable, particularly because there is no published guideline for the measurement of dissolved organic matter absorbance and fluorescence, and especially because there is no National Institute of Standards and Technology standard for dissolved organic matter. Validation and quality-control samples are analyzed on a monthly basis to determine laboratory and instrument precision and daily (that is, each day samples are run) to ensure repeatability. Data are not considered acceptable unless they meet laboratory criteria: All standards should be within 10 percent of the target value, laboratory replicates should be within 5 percent relative percent difference, and laboratory blanks (that is, laboratory reagent-grade water) should be less than one-tenth of the long-term method detection limit. Finally, for data to be useful, they must be accessible to users in a format that can be easily analyzed and interpreted. The Organic Matter Research Laboratory staff has developed a processing routine that extracts a subset of the data, which is made available to the public through the USGS National Water Quality Information System ( http://nwis.waterdata.usgs.gov/usa/nwis/qwdata ), and organizes the full datasets (that is, complete absorbance spectra and fluorescence excitation-emission matrices) in different forms that allow for these data to be analyzed using multi-parameter and multi-way statistical approaches.

Open-File Report

Methods for estimating selected low-flow frequency statistics for unregulated streams in Kentucky

This report provides estimates of, and presents methods for estimating, selected low-flow frequency statistics for unregulated streams in Kentucky including the 30-day mean low flows for recurrence intervals of 2 and 5 years (30Q 2 and 30Q 5 ) and the 7-day mean low flows for recurrence intervals of 5, 10, and 20 years (7Q 2 , 7Q 10 , and 7Q 20 ). Estimates of these statistics are provided for 121 U.S. Geological Survey streamflow-gaging stations with data through the 2006 climate year, which is the 12-month period ending March 31 of each year. Data were screened to identify the periods of homogeneous, unregulated flows for use in the analyses. Logistic-regression equations are presented for estimating the annual probability of the selected low-flow frequency statistics being equal to zero. Weighted-least-squares regression equations were developed for estimating the magnitude of the nonzero 30Q 2 , 30Q 5 , 7Q 2 , 7Q 10 , and 7Q 20 low flows. Three low-flow regions were defined for estimating the 7-day low-flow frequency statistics. The explicit explanatory variables in the regression equations include total drainage area and the mapped streamflow-variability index measured from a revised statewide coverage of this characteristic. The percentage of the station low-flow statistics correctly classified as zero or nonzero by use of the logistic-regression equations ranged from 87.5 to 93.8 percent. The average standard errors of prediction of the weighted-least-squares regression equations ranged from 108 to 226 percent. The 30Q 2 regression equations have the smallest standard errors of prediction, and the 7Q 20 regression equations have the largest standard errors of prediction. The regression equations are applicable only to stream sites with low flows unaffected by regulation from reservoirs and local diversions of flow and to drainage basins in specified ranges of basin characteristics. Caution is advised when applying the equations for basins with characteristics near the applicable limits and for basins with karst drainage features.

Scientific Investigations Report

Methods for estimating low-flow statistics for Massachusetts streams

Methods and computer software are described in this report for determining flow duration, low-flow frequency statistics, and August median flows. These low-flow statistics can be estimated for unregulated streams in Massachusetts using different methods depending on whether the location of interest is at a streamgaging station, a low-flow partial-record station, or an ungaged site where no data are available. Low-flow statistics for streamgaging stations can be estimated using standard U.S. Geological Survey methods described in the report. The MOVE.1 mathematical method and a graphical correlation method can be used to estimate low-flow statistics for low-flow partial-record stations. The MOVE.1 method is recommended when the relation between measured flows at a partial-record station and daily mean flows at a nearby, hydrologically similar streamgaging station is linear, and the graphical method is recommended when the relation is curved. Equations are presented for computing the variance and equivalent years of record for estimates of low-flow statistics for low-flow partial-record stations when either a single or multiple index stations are used to determine the estimates. The drainage-area ratio method or regression equations can be used to estimate low-flow statistics for ungaged sites where no data are available. The drainage-area ratio method is generally as accurate as or more accurate than regression estimates when the drainage-area ratio for an ungaged site is between 0.3 and 1.5 times the drainage area of the index data-collection site. Regression equations were developed to estimate the natural, long-term 99-, 98-, 95-, 90-, 85-, 80-, 75-, 70-, 60-, and 50-percent duration flows; the 7-day, 2-year and the 7-day, 10-year low flows; and the August median flow for ungaged sites in Massachusetts. Streamflow statistics and basin characteristics for 87 to 133 streamgaging stations and low-flow partial-record stations were used to develop the equations. The streamgaging stations had from 2 to 81 years of record, with a mean record length of 37 years. The low-flow partial-record stations had from 8 to 36 streamflow measurements, with a median of 14 measurements. All basin characteristics were determined from digital map data. The basin characteristics that were statistically significant in most of the final regression equations were drainage area, the area of stratified-drift deposits per unit of stream length plus 0.1, mean basin slope, and an indicator variable that was 0 in the eastern region and 1 in the western region of Massachusetts. The equations were developed by use of weighted-least-squares regression analyses, with weights assigned proportional to the years of record and inversely proportional to the variances of the streamflow statistics for the stations. Standard errors of prediction ranged from 70.7 to 17.5 percent for the equations to predict the 7-day, 10-year low flow and 50-percent duration flow, respectively. The equations are not applicable for use in the Southeast Coastal region of the State, or where basin characteristics for the selected ungaged site are outside the ranges of those for the stations used in the regression analyses. A World Wide Web application was developed that provides streamflow statistics for data collection stations from a data base and for ungaged sites by measuring the necessary basin characteristics for the site and solving the regression equations. Output provided by the Web application for ungaged sites includes a map of the drainage-basin boundary determined for the site, the measured basin characteristics, the estimated streamflow statistics, and 90-percent prediction intervals for the estimates. An equation is provided for combining regression and correlation estimates to obtain improved estimates of the streamflow statistics for low-flow partial-record stations. An equation is also provided for combining regression and drainage-area ratio estimates to obtain improved e

Massachusetts

Methods for estimating selected streamflow statistics at ungaged sites in Wyoming based on data through water year 2021

The U.S. Geological Survey, in cooperation with the Wyoming Water Development Office, developed regional regression equations based on basin characteristics and streamflow statistics for streamgages through water year 2021 (October 1, 2020, to September 30, 2021). The regression equations allow estimates of mean annual maximum, mean annual, mean seasonal, and mean monthly streamflows; frequency statistics for the 7-day mean low flows with 2-year and 10-year recurrence intervals, 14- and 30-day mean low flows with 5-year recurrence intervals, and 60- and 1-day mean high flow with 2-year and 5-year recurrence intervals, respectively; and the 0.1-, 0.2-, 0.5-, 1-, 2-, 4-, 5-, 10-, 20-, 25-, 30-, 50-, 60-, 70-, 75-, 80-, 90-, 95-, 98-, and 99-percent durations for annual streamflows and 0.1-, 0.5-, 10-, 15-, 20-, 25-, 30-, 40-, 50-, 60-, 70-, 75-, 80-, 85-, 90-, 95-, and 99-percent durations for monthly streamflows for most months for ungaged locations in Wyoming that are largely unaltered by diversions or upstream reservoirs. Regression equations were developed for 243 streamflow statistics. Best-subset selection was used to assess explanatory variables for respective streamflow statistics. Exploratory data analyses determined that, of the 81 basin characteristics evaluated as potential explanatory variables, characteristics such as drainage area and precipitation often produced models with the highest adjusted coefficient of determination and lowest mean squared error, as determined in the best-subset selection. To address heteroskedasticity of model residuals, model variables were regionalized using fixed-effects models; the percentages of the streamgage basins in selected ecoregions were defined as interaction terms, which represent the model slope for specific ecoregions. Most models were determined to be statistically significant for probability values less than or equal to 0.1 for one or more regional explanatory variables. The final regional regression equations defined in this report are available for use in the U.S. Geological Survey’s StreamStats web application at https://streamstats.usgs.gov/ss/ .

Colorado, Idaho, Montana, North Dakota, South Dako

Regression method for estimating long-term mean annual ground-water recharge rates from base flow in Pennsylvania

A method was developed for making estimates of long-term, mean annual ground-water recharge from streamflow data at 80 streamflow-gaging stations in Pennsylvania. The method relates mean annual base-flow yield derived from the streamflow data (as a proxy for recharge) to the climatic, geologic, hydrologic, and physiographic characteristics of the basins (basin characteristics) by use of a regression equation. Base-flow yield is the base flow of a stream divided by the drainage area of the basin, expressed in inches of water basinwide. Mean annual base-flow yield was computed for the period of available streamflow record at continuous streamflow-gaging stations by use of the computer program PART, which separates base flow from direct runoff on the streamflow hydrograph. Base flow provides a reasonable estimate of recharge for basins where streamflow is mostly unaffected by upstream regulation, diversion, or mining. Twenty-eight basin characteristics were included in the exploratory regression analysis as possible predictors of base-flow yield. Basin characteristics found to be statistically significant predictors of mean annual base-flow yield during 1971-2000 at the 95-percent confidence level were (1) mean annual precipitation, (2) average maximum daily temperature, (3) percentage of sand in the soil, (4) percentage of carbonate bedrock in the basin, and (5) stream channel slope. The equation for predicting recharge was developed using ordinary least-squares regression. The standard error of prediction for the equation on log-transformed data was 9.7 percent, and the coefficient of determination was 0.80. The equation can be used to predict long-term, mean annual recharge rates for ungaged basins, providing that the explanatory basin characteristics can be determined and that the underlying assumption is accepted that base-flow yield derived from PART is a reasonable estimate of ground-water recharge rates. For example, application of the equation for 370 hydrologic units in Pennsylvania predicted a range of ground-water recharge from about 6.0 to 22 inches per year. A map of the predicted recharge illustrates the general magnitude and variability of recharge throughout Pennsylvania.

Scientific Investigations Report

Testing and validating environmental models

Generally accepted standards for testing and validating ecosystem models would benefit both modellers and model users. Universally applicable test procedures are difficult to prescribe, given the diversity of modelling approaches and the many uses for models. However, the generally accepted scientific principles of documentation and disclosure provide a useful framework for devising general standards for model evaluation. Adequately documenting model tests requires explicit performance criteria, and explicit benchmarks against which model performance is compared. A model's validity, reliability, and accuracy can be most meaningfully judged by explicit comparison against the available alternatives. In contrast, current practice is often characterized by vague, subjective claims that model predictions show 'acceptable' agreement with data; such claims provide little basis for choosing among alternative models. Strict model tests (those that invalid models are unlikely to pass) are the only ones capable of convincing rational skeptics that a model is probably valid. However, 'false positive' rates as low as 10% can substantially erode the power of validation tests, making them insufficiently strict to convince rational skeptics. Validation tests are often undermined by excessive parameter calibration and overuse of ad hoc model features. Tests are often also divorced from the conditions under which a model will be used, particularly when it is designed to forecast beyond the range of historical experience. In such situations, data from laboratory and field manipulation experiments can provide particularly effective tests, because one can create experimental conditions quite different from historical data, and because experimental data can provide a more precisely defined 'target' for the model to hit. We present a simple demonstration showing that the two most common methods for comparing model predictions to environmental time series (plotting model time series against data time series, and plotting predicted versus observed values) have little diagnostic power. We propose that it may be more useful to statistically extract the relationships of primary interest from the time series, and test the model directly against them.

Science of the Total Environment