USGS Science⌕ Search

SEARCH · USGS Science

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Spatial data reduction through element -of-interest (EOI) extraction

Any large, multifaceted data collection that is challenging to handle with traditional management practices can be branded ‘Big Data.’ Any big data containing geo-referenced attributes can be considered big geospatial data. The increased proliferation of big geospatial data is currently reforming the geospatial industry into a data-driven enterprise. Challenges in the big spatial data domain can be summarized as the ‘Big Vs’ – variety, volume, velocity, veracity and value. Big spatial data sources can be considered in two broad classes, active and passive, as each is impacted to varying degrees. Some of these challenges may be alleviated by reducing unprocessed, or minimally processed, (raw) data to features, which we refer to as the extraction of Elements of Interest (EOI). In fact, many applications require EOI extraction from raw data to enable their basic employment. This chapter presents current state-of-the-art methods to create EOI from some types of georeferenced big data. We classify the data types into two realms: active and passive. Active data are those collected specifically for the purpose to which they are applied. Passive data are those collected for purposes other than those for which they are utilized, included those ‘collected’ for no particular purpose at all. The chapter then presents use cases from both the active and passive spatial realms, including the active applications of terrain feature extraction from digital elevation models and vegetation mapping from remotely-sensed imagery and passive applications like building identification from VGI and point-of-interest data mining from social networks for land use classification. Finally, the chapter concludes with future research needs.

Book chapter↗

Data standardization and management to facilitate large-scale and interdisciplinary approaches access

Bringing data related to recreational fishers and fisheries together across large scales can provide tremendous insight. Methods for collecting, analysing, and storing data can vary dramatically, which can have significant implications for the use of these data. Efforts to standardise data within organisations often increase the ability to compare datasets from different areas, monitor changes over time, and increase the utility of the data for management and research. Though employing standardised methodology and data architecture results in the most straightforward and robust opportunities for data integration, doing so may not be possible due to variation in data collection objectives, continuity with historical programs, and resource limitations. Additionally, managing data according to FAIR principles (findability, accessibility, interoperability, and reusability) helps support data sharing and large-scale research efforts. Key elements of integrable data include appropriate data structures, adequate documentation of methodology, interpretable and complete metadata, and accessible storage formats. We document the potential benefits of standardising data and offer example approaches. We also explore best practices regarding the formatting, storage, and transmission of recreational fisher data. Efforts to increase the level of standardisation and integrability of recreational fisher data can create opportunities to better understand fisher behaviour, needs, and fulfilment.

Book chapter↗

Sharing FAIR monitoring program data improves discoverability and reuse

Data resulting from environmental monitoring programs are valuable assets for natural resource managers, decision-makers, and researchers. These data are often collected to inform specific reporting needs or decisions with a specific timeframe. While program-oriented data and related publications are effective for meeting program goals, sharing well-documented data and metadata allows users to research aspects outside initial program intentions. As part of an effort to integrate data from four long-term large-scale US aquatic monitoring programs, we evaluated the original datasets against the FAIR (Findable, Accessible, Interoperable, Reusable) data principles and offer recommendations and lessons learned. Differences in data governance across these programs resulted in considerable effort to access and reuse the original datasets. Requirements, guidance, and resources available to support data publishing and documentation are inconsistent across agencies and monitoring programs, resulting in various data formats and storage locations that are not easily found, accessed, or reused. Making monitoring data FAIR will reduce barriers to data discovery and reuse. Programs are continuously striving to improve data management, data products, and metadata; however, provision of related tools, consistent guidelines and standards, and more resources to do this work is needed. Given the value of these data and the significant effort required to access and reuse them, actions and steps intended on improving data documentation and accessibility are described.

Enviornmental Monitoring and Assessment↗

Organization of marine phenology data in support of planning and conservation in ocean and coastal ecosystems

Among the many effects of climate change is its influence on the phenology of biota. In marine and coastal ecosystems, phenological shifts have been documented for multiple life forms; however, biological data related to marine species' phenology remain difficult to access and is under-used. We conducted an assessment of potential sources of biological data for marine species and their availability for use in phenological analyses and assessments. Our evaluations showed that data potentially related to understanding marine species' phenology are available through online resources of governmental, academic, and non-governmental organizations, but appropriate datasets are often difficult to discover and access, presenting opportunities for scientific infrastructure improvement. The developing Federal Marine Data Architecture when fully implemented will improve data flow and standardization for marine data within major federal repositories and provide an archival repository for collaborating academic and public data contributors. Another opportunity, largely untapped, is the engagement of citizen scientists in standardized collection of marine phenology data and contribution of these data to established data flows. Use of metadata with marine phenology related keywords could improve discovery and access to appropriate datasets. When data originators choose to self-publish, publication of research datasets with a digital object identifier, linked to metadata, will also improve subsequent discovery and access. Phenological changes in the marine environment will affect human economics, food systems, and recreation. No one source of data will be sufficient to understand these changes. The collective attention of marine data collectors is needed—whether with an agency, an educational institution, or a citizen scientist group—toward adopting the data management processes and standards needed to ensure availability of sufficient and useable marine data to understand marine phenology.

Ecological Informatics↗

Challenges with secondary use of multi-source water-quality data in the United States

Combining water-quality data from multiple sources can help counterbalance diminishing resources for stream monitoring in the United States and lead to important regional and national insights that would not otherwise be possible. Individual monitoring organizations understand their own data very well, but issues can arise when their data are combined with data from other organizations that have used different methods for reporting the same common metadata elements. Such use of multi-source data is termed “secondary use”—the use of data beyond the original intent determined by the organization that collected the data. In this study, we surveyed more than 25 million nutrient records collected by 488 organizations in the United States since 1899 to identify major inconsistencies in metadata elements that limit the secondary use of multi-source data. Nearly 14.5 million of these records had missing or ambiguous information for one or more key metadata elements, including (in decreasing order of records affected) sample fraction, chemical form, parameter name, units of measurement, precise numerical value, and remark codes. As a result, metadata harmonization to make secondary use of these multi-source data will be time consuming, expensive, and inexact. Different data users may make different assumptions about the same ambiguous data, potentially resulting in different conclusions about important environmental issues. The value of these ambiguous data is estimated at \$US12 billion, a substantial collective investment by water-resource organizations in the United States. By comparison, the value of unambiguous data is estimated at \$US8.2 billion. The ambiguous data could be preserved for uses beyond the original intent by developing and implementing standardized metadata practices for future and legacy water-quality data throughout the United States.

Water Research↗

Stakeholder engagement to guide decision-relevant water data delivery

Water resources management and policy making require access to reliable scientific data. However, water managers may need to overcome various obstacles to accessing data. For example, insufficient technological infrastructures, low data literacy, and data format complexities often inhibit data user access. Thus, it is imperative to include stakeholders in the design of data delivery systems. The United States Geological Survey's Water Resources Mission Area is currently developing Integrated Water Availability Assessments (IWAAs) — multi-extent, stakeholder driven, near real-time water availability census and prediction for human and ecological uses. To provide appropriate user accessibility to data delivery systems developed for IWAAs, a user-centered design process including stakeholder focus groups was used to determine potential water data user needs and preferences. Focus groups identified five types of potential users: Public sector water resources managers, Public sector water resources manager data analysts, Industry and private companies, Tribal Nations, and Nonprofit organizations. Different water data user types depended on diverse spatial and temporal scale data. Public sector water resources managers benefitted most from data synthesized into user-friendly platforms and Public sector water resources data analysts preferred easy access to raw data. These findings can support the development of a water data delivery platform that meets a variety of user needs.

Journal of the American Water Resources Associatio↗

What you should know about land-cover data

Wildlife biologists are using land-characteristics data sets for a variety of applications. Many kinds of landscape variables have been characterized and the resultant data sets or maps are readily accessible. Often, too little consideration is given to the accuracy or traits of these data sets, most likely because biologists do not know how such data are compiled and rendered, or the potential pitfalls that can be encountered when applying these data. To increase understanding of the nature of land-characteristics data sets, I introduce aspects of source information and data-handling methodology that include the following: ambiguity of land characteristics; temporal considerations and the dynamic nature of the landscape; type of source data versus landscape features of interest; data resolution, scale, and geographic extent; data entry and positional problems; rare landscape features; and interpreter variation. I also include guidance for determining the quality of land-characteristics data sets through metadata or published documentation, visual clues, and independent information. The quality or suitability of the data sets for wildlife applications may be improved with thematic or spatial generalization, avoidance of transitional areas on maps, and merging of multiple data sources. Knowledge of the underlying challenges in compiling such data sets will help wildlife biologists to better assess the strengths and limitations and determine how best to use these data.

Journal of Wildlife Management↗

Digital data sets for map products produced as part of the Black Hills Hydrology Study, western South Dakota

This compact disk contains digital data produced as part of the 1:100,000-scale map products for the Black Hills Hydrology Study conducted in western South Dakota. The digital data include 28 individual Geographic Information System (GIS) data sets: data sets for the hydrogeologic unit map including all mapped hydrogeologic units within the study area (1 data set) and major geologic structure including anticlines and synclines (1 data set); data sets for potentiometric maps including the potentiometric contours for the Inyan Kara, Minnekahta, Minnelusa, Madison, and Deadwood aquifers (5 data sets), wells used as control points for each aquifer (5 data sets), and springs used as control points for the potentiometric contours (1 data set); and data sets for the structure-contour maps including the structure contours for the top of each formation that contains major aquifers (5 data sets), wells and tests holes used as control points for each formation (5 data sets), and surficial deposits (alluvium and terrace deposits) that directly overlie each of the major aquifer outcrops (5 data sets). These data sets were used to produce the maps published by the U.S. Geological Survey.

Open-File Report↗

Global multi-resolution terrain elevation data 2010 (GMTED2010)

In 1996, the U.S. Geological Survey (USGS) developed a global topographic elevation model designated as GTOPO30 at a horizontal resolution of 30 arc-seconds for the entire Earth. Because no single source of topographic information covered the entire land surface, GTOPO30 was derived from eight raster and vector sources that included a substantial amount of U.S. Defense Mapping Agency data. The quality of the elevation data in GTOPO30 varies widely; there are no spatially-referenced metadata, and the major topographic features such as ridgelines and valleys are not well represented. Despite its coarse resolution and limited attributes, GTOPO30 has been widely used for a variety of hydrological, climatological, and geomorphological applications as well as military applications, where a regional, continental, or global scale topographic model is required. These applications have ranged from delineating drainage networks and watersheds to using digital elevation data for the extraction of topographic structure and three-dimensional (3D) visualization exercises (Jenson and Domingue, 1988; Verdin and Greenlee, 1996; Lehner and others, 2008). Many of the fundamental geophysical processes active at the Earth's surface are controlled or strongly influenced by topography, thus the critical need for high-quality terrain data (Gesch, 1994). U.S. Department of Defense requirements for mission planning, geographic registration of remotely sensed imagery, terrain visualization, and map production are similarly dependent on global topographic data. Since the time GTOPO30 was completed, the availability of higher-quality elevation data over large geographic areas has improved markedly. New data sources include global Digital Terrain Elevation Data (DTEDRegistered) from the Shuttle Radar Topography Mission (SRTM), Canadian elevation data, and data from the Ice, Cloud, and land Elevation Satellite (ICESat). Given the widespread use of GTOPO30 and the equivalent 30-arc-second DTEDRegistered level 0, the USGS and the National Geospatial-Intelligence Agency (NGA) have collaborated to produce an enhanced replacement for GTOPO30, the Global Land One-km Base Elevation (GLOBE) model and other comparable 30-arc-second-resolution global models, using the best available data. The new model is called the Global Multi-resolution Terrain Elevation Data 2010, or GMTED2010 for short. This suite of products at three different resolutions (approximately 1,000, 500, and 250 meters) is designed to support many applications directly by providing users with generic products (for example, maximum, minimum, and median elevations) that have been derived directly from the raw input data that would not be available to the general user or would be very costly and time-consuming to produce for individual applications. The source of all the elevation data is captured in metadata for reference purposes. It is also hoped that as better data become available in the future, the GMTED2010 model will be updated.

Open-File Report↗

Defining a data management strategy for USGS Chesapeake Bay studies

The mission of U.S. Geological Survey’s (USGS) Chesapeake Bay studies is to provide integrated science for improved understanding and management of the Chesapeake Bay ecosystem. Collective USGS efforts in the Chesapeake Bay watershed began in the 1980s, and by the mid-1990s the USGS adopted the watershed as one of its national place-based study areas. Great focus and effort by the USGS have been directed toward Chesapeake Bay studies for almost three decades. The USGS plays a key role in using “ecosystem-based adaptive management, which will provide science to improve the efficiency and accountability of Chesapeake Bay Program activities” (Phillips, 2011). Each year USGS Chesapeake Bay studies produce published research, monitoring data, and models addressing aspects of bay restoration such as, but not limited to, fish health, water quality, land-cover change, and habitat loss. The USGS is responsible for collaborating and sharing this information with other Federal agencies and partners as described under the President’s Executive Order 13508—Strategy for Protecting and Restoring the Chesapeake Bay Watershed signed by President Obama in 2009. Historically, the USGS Chesapeake Bay studies have relied on national USGS databases to store only major nationally available sources of data such as streamflow and water-quality data collected through local monitoring programs and projects, leaving a multitude of other important project data out of the data management process. This practice has led to inefficient methods of finding Chesapeake Bay studies data and underutilization of data resources. Data management by definition is “the business functions that develop and execute plans, policies, practices and projects that acquire, control, protect, deliver and enhance the value of data and information.” (Mosley, 2008a). In other words, data management is a way to preserve, integrate, and share data to address the needs of the Chesapeake Bay studies to better manage data resources, work more efficiently with partners, and facilitate holistic watershed science. It is now the goal of the USGS Chesapeake Bay studies to implement an enhanced and all-encompassing approach to data management. This report discusses preliminary efforts to implement a physical data management system for program data that is not replicated nationally through other USGS databases.

Chesapeake Bay↗

Community for Data Integration 2013 Annual Report

The U.S. Geological Survey (USGS) conducts earth science to help address complex issues affecting society and the environment. In 2006, the USGS held the first Scientific Information Management Workshop to bring together staff from across the organization to discuss the data and information management issues affecting the integration and delivery of earth science research and investigate the use of “communities of practice” as mechanisms to share expertise about these issues. Out of this effort emerged the Council for Data Integration, which was conceived as an official organizational function that would help guide data integration activities and formalize communities of practice into working groups. However by 2009, it became apparent that many members of the council had an interest in developing data integration solutions and sharing expertise in a less formal grassroots perspective, thus transforming the “Council” into a “Community” for Data Integration (CDI). Today, the CDI represents a dynamic community of practice focused on advancing science data and information management and integration capabilities across the USGS and the CDI community. The CDI fosters an environment for collaboration and sharing by bringing together expertise from external partners and representatives across USGS who are involved in research, data management, and information technology. Membership is voluntary and open to USGS employees and other individuals and organizations willing to contribute to the community (if interested, contact cdi@usgs.gov). The purpose of the CDI is to advance understanding of Earth systems through enhanced use of data and information including associated tools and techniques provide a forum for people doing work with data integration to come together to share ideas as well as learn new skills and techniques, and grow overall USGS capabilities with data and information by increasing visibility of the work of many people throughout the USGS and the CDI community. To achieve these goals, the CDI operates within four applied areas: monthly forums, annual workshop/webinar series, working groups, and projects. The monthly forums, also known as the Opportunity/Challenge of the Month, provide an open dialogue to share and learn about data integration efforts or to present problems that invite the Community to offer solutions, advice, and support. Since 2010, the CDI has also sponsored annual workshops/webinar series to encourage the exchange of ideas, sharing of activities, presentations of current projects, and networking among members. Stemming from common interests, the working groups are focused on efforts to address data management and technical 2 challenges, including the development of standards and tools, improving interoperability and information infrastructure, and data preservation within USGS and its partners. The growing support for the activities of the working groups led to the CDI’s first formal request for proposals (RFP) process in 2013 to fund projects that produced tangible products. Today the CDI continues to hold an annual RFP that create data management tools and practices, collaboration tools, and training in support of data integration and delivery.

Open-File Report↗

U.S. Geological Survey Community for Data Integration 2017 Workshop Proceedings

Executive Summary The U.S. Geological Survey (USGS) Community for Data Integration (CDI) Workshop was held May 16–19, 2017 at the Denver Federal Center. There were 183 in-person attendees and 35 virtual attendees over four days. The theme of the workshop was “Enabling Integrated Science,” with the purpose of bringing together the community to discuss current topics, shared challenges, and steps forward to advance integrated science at the USGS. The CDI welcomed several keynote speakers, including Bill Werkheiser, USGS Acting Director; Kevin T. Gallagher, USGS Associate Director of the Core Science Systems Mission Area; Bruce Caron, Earth Science Information Partners Community Architect; and Tim Quinn, Chief of the USGS Office of Enterprise Information. Their presentations focused on the importance of collaborative, cross-disciplinary, and open science and the role of the CDI in identifying and supporting new opportunities in these areas for the USGS and its partners. In addition to the stated theme, the workshop agenda was driven by the needs of the CDI, with topics highlighting current resources and technologies that could help attendees in their daily work. Topical sessions were proposed by CDI members and included subjects such as data citation, information technology architecture, legacy data, real-time data, and many more. Plenary speakers from the community talked about USGS activities in data science, elevation and hydrography data integration, advanced scientific computing solutions, cloud computing, data-management training, and data-sharing agreements. Two panels addressed the role of the CDI in enabling integrated science and examples of CDI-supported projects in action. Breakout discussions focused on the workshop theme of “Enabling Integrated Science” and covered five topics: Data and Data Integration, Modeling, Computing Capacity, Science Data Integration, and User Needs and Experience. Sessions on each topic identified actions that could bring the USGS and the broader Earth science community closer to the goal of making integrated science commonplace. The breakouts produced recommendations with the broad themes of improving communication and connections across the USGS, reducing duplication and increasing knowledge transfer, increasing training and testbed opportunities to learn and experiment, and creating community-supported standards to enable better integration and interoperability. The DataBlast poster and live demonstration session showcased 36 projects from around the CDI and included recent CDI-funded projects as well as other USGS and partner initiatives that were related to data and software integration and discovery. Importantly, the CDI workshop provided a forum for scientists, technologists, data and resource managers, program managers, and others to convene face to face to discuss common methods, interests, challenges, and solutions related to scientific data and technologies. As a result of this rare convergence, new connections were made across disciplines, backgrounds, and geographical locations, seeding future activities and collaborations. Sharing of ideas from all attendees was encouraged through the use of a mobile application to collect real-time questions and feedback from the audience The primary outcomes of the workshop are the recommendations from the breakout sessions titled “Roadmap Discussions on Enabling Integrated Science” and from the topical sessions detailed in these proceedings. These sessions, as well as the plenary discussions, identified new areas of collaboration and learning that the CDI will facilitate, such as data science, software development, scientific modeling practices, and user needs and experience. The CDI will build on the results of the workshop to guide its future topics, events, and funding opportunities to support an integrated science capacity for the USGS.

Open-File Report↗

Oyster model inventory: Identifying critical data and modeling approaches to support restoration of oyster reefs in coastal U.S. Gulf of Mexico waters

Executive Summary Along the coast of the U.S. Gulf of Mexico, the eastern oyster ( Crassostrea virginica ) plays important ecological and economic roles. Commercial landings from this region account for more than 50 percent of all U.S. landings; these oyster reefs also provide varied ecosystem services, including nursery habitat for many fish and macroinvertebrate species, shoreline protection, and water-quality maintenance. Declining trends in both total oyster production and functional reef area across this region have spurred investment in restoration of oyster resources, with specific calls for restoration projects to develop a network of reefs and identify broodstock and sanctuary reef restoration sites. Decision making related to restoration and establishment of a network of oyster reefs in the Gulf of Mexico requires information on both the environment and the effects of the environment on the oyster life cycle (including larval movement, survival, oyster recruitment, reproduction, growth, and mortality). Here, we examined the current state of data and model development in this region with the goal of providing an overview of oyster modeling approaches and an inventory of available data and existing oyster models. This report is meant to provide an overview to managers for understanding existing efforts and identify a path forward to most efficiently inform oyster resource management and restoration planning in moving from a single reef management approach to a reef network management approach. Numerous models related to some aspect of the oyster life cycle have been built, calibrated, and validated for various Gulf of Mexico estuaries over the last few decades (over 30 models identified). These models, which could inform site restoration, can be classified into four approaches: (1) oyster Habitat Suitability Index (HSI) models; (2) larval transport models; (3) on-reef oyster models that may include oyster growth, mortality and reproduction, and substrate persistence; and (4) coupled larval transport on-reef metapopulation models that simulate the entire oyster life cycle. The data requirements, model complexity and assumptions, and transferability vary by approach. Specifically, some approaches may offer greater accessibility, flexibility, and transferability spatially or temporally, with minimal data input, but only provide broad information to support site selection. In contrast, other approaches may require significant site-specific data for their construction and validation but may provide more accurate and location-specific data to support site selection for broodstock reefs. Regardless of modeling approach used, data on environmental drivers, such as salinity, water temperature, or water flow impacting oyster metabolism and movement, are required at appropriate spatial and temporal scales. While numerous data collection platforms, environmental models, and research products exist within Gulf of Mexico estuaries to provide important environmental data to use as drivers in the oyster models, significant variability in temporal and spatial coverage of the data, and variation in the availability of future condition models, exists across estuaries. This variation influences the spatial and temporal scales at which oyster models may be developed and impacts the calibration and validation of the oyster models within a given estuary, affecting its potential ability to address specific management or restoration questions. While multiple modeling approaches exist for informing site selection of broodstock or sanctuary oyster reefs, the development, calibration, and validation of a single modeling platform presents the most efficient, transferable, and useful tool for managers across the Gulf of Mexico. The development of a single modeling platform would involve using standardized input variables, governing equations, and assumptions for the modeled oyster processes and outputs, and for standardized calibration and validation procedures that could be applied within each estuary. The differences among estuary applications would require substituting only estuary-specific environmental data, and calibrating and validating the modeling approach with local oyster data. Two modeling approaches likely to be useful include (1) development of a general geospatial HSI modeling framework that could be applied consistently across estuaries and (2) a mechanistic coupled larval transport on-reef metapopulation model requiring only estuarine specific calibration and hydrodynamic models. Both approaches benefit from existing work across multiple Gulf of Mexico estuaries and could provide valuable support for oyster restoration, but may differ in their ability to address specific questions related to oyster restoration. HSI models specifically guide restoration practitioners in determining suitable habitat based on available data. The HSI approach, while currently more widely used and accessible, requires more development of larval suitability and larval input and output components in order to inform reef connectivity. A metapopulation approach considering the full oyster life cycle that simulates both on-reef oyster growth, mortality, reproduction, substrate persistence, and larval transport (ideally with larval growth and mortality) would provide the greatest detail and level of understanding but requires significant up-front investment. The larval oyster model and on-reef oyster model are usually developed independently for systems, although the two approaches can be coupled to represent the entire oyster life cycle in order to characterize and assess a reef metapopulation. This approach may be less accessible and much more data-intensive, however, and it requires some expertise to run and apply to inform oyster resource management. Ultimately, the development of single modeling platforms for each of these approaches would provide flexible tools applicable across all Gulf of Mexico oyster supporting estuaries. By using a single platform for model development, testing, calibrating and validating, and evaluation of modeled future scenarios, oyster restoration scientists and managers would not only be able to examine different scenario outcomes within a single estuary, but could also have comparable modeled results to evaluate potential outcomes, across estuaries and regions, that are not confounded by varying modeled data inputs, governing equations, assumptions, or user judgement.

Alabama, Florida, Louisiana, Mississippi, Texas↗

Compilation and evaluation of data used to identify groundwater sources under the direct influence of surface water in Pennsylvania

A study was conducted to compile and evaluate data used to identify groundwater sources that are under the direct influence of surface water (GUDI) in Pennsylvania. In the early 1990s, the Pennsylvania Department of Environmental Protection (PADEP) implemented the Surface Water Identification Protocol (SWIP) for the identification of GUDI sources. Since the establishment of the SWIP, PADEP has classified more than 500 individual sources across Pennsylvania as GUDI, but Pennsylvania’s complex geology and physiography provide a challenge for a uniform method of GUDI determination. Components used in this study to compile and evaluate data associated with GUDI determination include: (1) a preliminary review of file information for 43 public water-supply wells, (2) quality control and addition of data to PADEP’s database for public water-supply systems to prepare data for analysis, and (3) exploratory evaluation of existing GUDI sources in the database with respect to hydrogeologic and source-construction characteristics that are currently utilized in the assessment methodology. Case files for 43 wells from PADEP’s Northcentral and Southcentral regions were reviewed to: (1) provide a better understanding of how the SWIP was applied in practice, (2) verify and compile missing data, and (3) find additional attributes not previously available that might explain a well’s categorization as GUDI. Review of file information showed that the SWIP outlined in PADEP technical guidance was usually followed, but for some sources, the GUDI determination was more complex and could not be easily summarized. Data compiled for study analyses provided by PADEP include source data derived from public water-supply system case files, a source-information database for public water-supply systems, and Microscopic Particulate Analysis (MPA) results and associated water-quality data for public water-supply system groundwater sources. Data from the Pennsylvania Drinking Water Information System (PADWIS) , which is PADEP’s database for public water-supply systems, were also used for this study. The PADWIS database originally included data for 12,147 groundwater sources (11,812 groundwater sources not under the direct influence of surface water (non-GUDI) wells and 335 GUDI wells). A subset (4,018 wells consisting of 3,842 non-GUDI wells and 175 GUDI wells) of the PADWIS database was created for an analysis and includes only community wells evaluated in accordance with the SWIP. MPA results for 631 community and noncommunity wells were compiled, along with associated water-quality data (alkalinity, chloride, Escherichia coli , fecal coliform, nitrate, pH, sodium, specific conductance, sulfate, total coliform, total dissolved solids, total residue, and turbidity) populated from the PADEP Bureau of Laboratories Sample Information System. Data compiled from sources other than PADEP include spatial data, both naturogenic (for example, average precipitation or distance to closest hydrologic feature) and anthropogenic (for example, percentage of developed or agricultural land cover within a specific vicinity of a public water-supply system well) data representing spatially derived variables. Comparison among wells in the PADWIS dataset subset using the nonparametric Kruskal-Wallis test showed that GUDI wells had significantly older median construction years, shallower depths, and static water levels closer to the land surface than non-GUDI wells and that carbonate aquifers had the highest percentages of wells designated as GUDI (12 percent; 57 wells). Further comparison of wells in the PADWIS database subset using the Spearman’s rho monotonic correlation test illustrated that public water-supply wells designated as GUDI largely occur in unconfined aquifers and have high average yield and shallow static water levels. Assessment of the MPA database subset using the Kruskal-Wallis test showed wells with MPA total risk-factor scores that exceeded zero had older median construction years and shallower casing depths than wells with MPA total risk-factor scores of zero and that carbonate aquifers had the highest percentages of wells with MPA total risk-factor scores exceeding zero (30 percent; 63 wells). Spearman’s rho correlations showed that wells completed in aquifers with depths to major water-bearing zones closer to the land-surface had higher total risk-factor scores resulting from MPA samples. Based on the results of the analyses described in this report, broad conclusions can be drawn regarding site-specific well characteristics as well as anthropogenic and naturogenic factors that could be responsible for a well being designated as GUDI, but the accuracy of these results is dependent on the quality of the data being analyzed. Ultimately, study results serve as an added resource for initial desktop screening of wells to determine if additional site-specific investigation is warranted and underscore the need for field evaluation.

Pennsylvania↗

Data management system for USGS/USEPA urban hydrology studies program

A data management system was developed to store, update, and retrieve data collected in urban stormwater studies jointly conducted by the U.S. Geological Survey and U.S. Environmental Protection Agency in 11 cities in the United States. The data management system is used to retrieve and combine data from USGS data files for use in rainfall, runoff, and water-quality models and for data computations such as storm loads. The system is based on the data management aspect of the Statistical Analysis System (SAS) and was used to create all the data files in the data base. SAS is used for storage and retrieval of basin physiography, land-use, and environmental practices inventory data. Also, storm-event water-quality characteristics are stored in the data base. The advantages of using SAS to create and manage a data base are many with a few being that it is simple, easy to use, contains a comprehensive statistical package, and can be used to modify files very easily. Data base system development has progressed rapidly during the last two decades and the data managment system concepts used in this study reflect the advancement made in computer technology during this era. Urban stormwater data is, however, just one application for which the system can be used. (USGS)

Open-File Report↗

Summary of selected U.S. Geological survey data on domestic well water quality for the Centers for Disease Control's National Environmental Public Health Tracking Program

About 10 to 30 percent of the population in most States uses domestic (private) water supply. In many States, the total number of people served by domestic supplies can be in the millions. The water quality of domestic supplies is inconsistently regulated and generally not well characterized. The U.S. Geological Survey (USGS) has two water-quality data sets in the National Water Information System (NWIS) database that can be used to help define the water quality of domestic-water supplies: (1) data from the National Water-Quality Assessment (NAWQA) Program, and (2) USGS State data. Data from domestic wells from the NAWQA Program were collected to meet one of the Program's objectives, which was to define the water quality of major aquifers in the United States. These domestic wells were located primarily in rural areas. Water-quality conditions in these major aquifers as defined by the NAWQA data can be compared because of the consistency of the NAWQA sampling design, sampling protocols, and water-quality analyses. The NWIS database is a repository of USGS water data collected for a variety of projects; consequently, project objectives and analytical methods vary. This variability can bias statistical summaries of contaminant occurrence and concentrations; nevertheless, these data can be used to define the geographic distribution of contaminants. Maps created using NAWQA and USGS State data in NWIS can show geographic areas where contaminant concentrations may be of potential human-health concern by showing concentrations relative to human-health water-quality benchmarks. On the basis of national summaries of detection frequencies and concentrations relative to U.S. Environmental Protection Agency (USEPA) human-health benchmarks for trace elements, pesticides, and volatile organic compounds, 28 water-quality constituents were identified as contaminants of potential human-health concern. From this list, 11 contaminants were selected for summarization of water-quality data in 16 States (grantee States) that were funded by the Environmental Public Health Tracking (EPHT) Program of the Centers for Disease Control and Prevention (CDC). Only data from domestic-water supplies were used in this summary because samples from these wells are most relevant to human exposure for the targeted population. Using NAWQA data, the concentrations of the 11 contaminants were compared to USEPA human-health benchmarks. Using NAWQA and USGS State data in NWIS, the geographic distribution of the contaminants were mapped for the 16 grantee States. Radon, arsenic, manganese, nitrate, strontium, and uranium had the largest percentages of samples with concentrations greater than their human-health benchmarks. In contrast, organic compounds (pesticides and volatile organic compounds) had the lowest percentages of samples with concentrations greater than human-health benchmarks. Results of data retrievals and spatial analysis were compiled for each of the 16 States and are presented in State summaries for each State. Example summary tables, graphs, and maps based on USGS data for New Jersey are presented to illustrate how USGS water-quality and associated ancillary geospatial data can be used by the CDC to address goals and objectives of the EPHT Program.

Scientific Investigations Report↗

Determination of study reporting limits for pesticide constituent data for the California Groundwater Ambient Monitoring and Assessment Program Priority Basin Project, 2004–2018—Part 1: National Water Quality Schedules 2003, 2032, or 2033, and 2060

The California Groundwater Ambient Monitoring and Assessment Program Priority Basin Project (GAMA-PBP) is a long-term cooperative project designed to assess the quality of groundwater resources used for public and domestic drinking water supplies in the State of California, to monitor and evaluate changes to that quality, to investigate the human and natural factors controlling water quality, and to improve the availability of comprehensive groundwater quality data and information. Between May 18, 2004, and May 3, 2018, the GAMA-PBP collected 3001 groundwater samples for analysis of pesticide constituents by the U.S. Geological Survey (USGS) National Water Quality Laboratory (NWQL)(note that ‘pesticide constituents’ includes parent compounds and degradates). Of these samples, 2994 were analyzed for pesticide constituents on schedules 2003, 2032, or 2033 (65 to 84 constituents), and 840 were analyzed for pesticide constituents on schedule 2060 (58 constituents). The original dataset reported by the NWQL to the USGS National Water Information System (NWIS) database contained a total of 2,688 detections of 78 pesticide constituents and 253,825 non-detections. In this original dataset, 33 percent of the 3,001 samples analyzed had reported detections of one or more pesticide constituents. This report describes the GAMA-PBP data-quality objectives for pesticide data, the procedures used to establish study reporting limits, and use of those reporting limits to censor the data from the NWQL so that the final data published by the GAMA-PBP meet these data-quality objectives. The final GAMA-PBP dataset for samples collected from May 2004 to May 2018, after censoring, had a total of 1,632 detections of 37 pesticide constituents. In the final GAMA-PBP dataset, 25 percent of the 3,001 samples analyzed had detections of one or more pesticide constituents. The presence of pesticides in groundwater is commonly evaluated by calculating detection frequencies. Detection frequencies for pesticides are sensitive to detection limits and method performance for concentrations near those limits; therefore, the two primary data quality issues addressed in the GAMA-PBP data-quality objectives for pesticides are (1) establishing criteria for classifying data from the laboratory as detections or non-detections for the purpose of data reporting by the project and (2) accounting for changes in analytical methods or method performance over time. The GAMA-PBP addresses these issues by developing study reporting limits that are used as the boundary between detections and non-detections for the reporting of GAMA-PBP results. These reporting limits are defined from method detection limits (MDLs) provided by the NWQL, unless examination of results from laboratory set blanks (LSBs) and GAMA-PBP field blanks indicates that a higher concentration censoring limit is warranted. The GAMA-PBP selected the MDL as the primary choice for defining study reporting limits for consistency with U.S. Environmental Protection Agency (EPA) guidelines for reporting detections of pesticides and other organic constituents. A five-step procedure is used to develop study reporting limits and censor the GAMA-PBP dataset accordingly. The effect of the censoring at each step is described to provide information about the relative effect of each step on the overall censoring of the dataset. Steps 1 and 2 can be implemented at the time the data are received, whereas steps 3−5 require information accumulated over an extended period. Step 1: Reject results that were most likely the result of specific contamination instances attributable to unusual field or laboratory conditions during sample collection or processing. Two such instances were identified, leading to rejection of 25 detections, which were assigned a data-quality indicator code of “Q” for “reviewed and rejected” in the NWIS database. Step 2: Use the NWQL MDLs in effect at the time each sample was analyzed as the reporting limit. A total of 506 detections were censored on this basis. Step 3: Use the maximum MDL established by the NWQL during July 2004–August 2018 (MDLmax) as the reporting limit. The rationale for using the MDLmax as the reporting limit is based primarily on the observation that the concentrations of MDLs generally increased over time. A total of 438 detections were censored on this basis. Step 4: Use the LSBs to identify periods of greater potential laboratory contamination bias and define raised reporting limits to be used during those periods. These periods were defined by using a moving average detection frequency approach. For consistency with the NWQL procedures for defining raised reporting limits on the basis of detections in LSBs, the raised reporting limits were defined as equal to three times the highest concentration measured in an LSB during the period. A total of 25 detections in groundwater samples analyzed during periods of increased laboratory contamination bias were censored. Step 5: Use the LSBs and field blanks to identify potential contamination bias from field or laboratory processes outside of the time periods identified in step 4. The NWQL protocols were used to define the MDLs from blanks analyzed outside of the periods identified in step 4. If an MDL defined from blanks was greater than the MDLmax, the MDL defined from blanks was used to censor the data. One constituent had a study reporting limit defined on this basis, and a total of 62 detections in groundwater samples were censored. As of 2019, the USGS NWIS database does not have the capability to store both the original value reported by the NWQL and the final value published by the GAMA-PBP that reflects application of the quality-control censoring described in this report. In the interim, while this capability is developed, the 1,031 results censored in steps 2−5 are blocked from public release in NWIS, and the GAMA-PBP has published the original and final values in a USGS data release accompanying this report. The entire GAMA-PBP final dataset for pesticide constituents on schedules 2003, 2032, or 2033, or on schedule 2060 is publicly available in that USGS data release, through the USGS GAMA-PBP public web portal, and through the California State Water Resources Control Board GAMA public groundwater information system.

California↗

Evaluating hydrologic data products for scientific and management applications related to potential future streamflow conditions in the Upper Mississippi and Illinois Rivers

The hydrology of the Upper Mississippi and Illinois Rivers is a fundamental driver of ecosystem patterns and processes across a large portion of the United States. Quantitative hydrologic data for the main stems of these rivers underlie numerous scientific investigations, statistical models, and decision-making processes for local, State, and Federal agencies involved in the Upper Mississippi River Restoration program. Although historical hydrologic data exist, data representing potential future conditions of the Upper Mississippi and Illinois Rivers lack the resolution necessary to anticipate biotic and abiotic responses to altered hydrology and to determine resilient management actions. A source of future hydrologic scenarios is the readily available LOCA–VIC–mizuRoute hydrologic data products (named for the chain of models the data are produced from—localized constructed analogs, Variable Infiltration Capacity macroscale hydrological model, and the mizuRoute hydrologic routing model—that we shorten further to LVM in this report) that include simulated discharges for historic and future timeframes. The objective of this study is to assess the reliability of the hydrologic data products for their use in Upper Mississippi River Restoration program applications. Key study questions are (1) do the hydrologic data products reproduce characteristics of hydrology necessary to support ecological modeling and restoration decision-making applications within the Upper Mississippi River Restoration program? and (2) are there geographic differences in the reliability of the hydrologic data products? Seven characteristics of river hydrology were selected related to flow magnitude, seasonality, and regime for evaluation. The seven characteristics were calculated using observed and historical simulated hydrologic data at 19 U.S. Geological Survey streamgages throughout the basins of the Upper Mississippi and Illinois Rivers; two streamgages are located on the main stem of the Mississippi River and two streamgages are located on the main stem of the Illinois River. Statistical comparisons between observed and historical simulated characteristics indicated that the hydrologic data products did not reliably represent historical hydrologic conditions in the basin or main stem. The hydrologic data products we evaluated could not reliably capture the overall hydrologic regime or flow magnitudes; the latter is evidenced by substantial underestimates of discharge at most streamgages. Seasonal hydrologic characteristics were captured more reliably than flow magnitude, but overall correspondence was low for most streamgages. A weak latitudinal pattern in seasonal characteristics indicated the hydrologic data products poorly represent streamflow timing in snow-affected regions of the basin. Discrepancies in magnitude, seasonality, and regime indicate the potential for multiple sources of error. Because poor correspondence was present across all 19 streamgages, it was not possible to identify specific drivers of poor performance (that is, drainage area or geography). The modeling chain should be evaluated for biases associated with meteorologic forcing data, as well as hydrologic model formulation and calibration. We conclude that the hydrologic data products we evaluated appear unsuitable for applications tied to habitat and ecosystem restoration and management in the Upper Mississippi and Illinois Rivers. Plans to develop a future hydrology dataset for the Upper Mississippi River Restoration program would benefit from ongoing work to improve global climate model output downscaling methods, to improve hydrologic models, to make use of innovations in machine-learning approaches for projecting hydrology, and other efforts. The framework developed herein to evaluate hydrometeorological outputs generated using global climate models for a specific water resources application is a transferrable approach that could be applied to other data products and river systems.

Illinois, Indiana, Iowa, Minnesota, Missouri, Sout↗