USGS ScienceSearch

SEARCH · USGS Science

Results for “Scientific Data - Nature”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Cloud-native repositories for big scientific data

Scientific data have traditionally been distributed via downloads from data server to local computer. This way of working suffers from limitations as scientific datasets grow toward the petabyte scale. A “cloud-native data repository,” as defined in this article, offers several advantages over traditional data repositories—performance, reliability, cost-effectiveness, collaboration, reproducibility, creativity, downstream impacts, and access and inclusion. These objectives motivate a set of best practices for cloud-native data repositories: analysis-ready data, cloud-optimized (ARCO) formats, and loose coupling with data-proximate computing. The Pangeo Project has developed a prototype implementation of these principles by using open-source scientific Python tools. By providing an ARCO data catalog together with on-demand, scalable distributed computing, Pangeo enables users to process big data at rates exceeding 10 GB/s. Several challenges must be resolved in order to realize cloud computing’s full potential for scientific research, such as organizing funding, training users, and enforcing data privacy requirements.

Computing in Science and Engineering

Tools for discovering and accessing Great Lakes scientific data

The Great Lakes Restoration Initiative (GLRI) is a multidisciplinary and interagency effort focused on the protection and restoration of the Great Lakes (GL) using the best available science and applying lessons learned from previous studies. The U.S. Geological Survey (USGS) contributes to the GLRI effort by providing resource managers with information and tools needed to meet restoration goals. This includes contributing scientific expertise and delivering findings to the GL community through meaningful information products. One of the strengths of the GLRI is its interagency approach; however, this can create challenges when coordinating the large number of restoration activities being performed by GL governments, tribes, academics, nonprofits, and industry. There is a vast array of data being produced by both the USGS and its partners, and it is crucial that scientists, managers, policymakers, and the public can easily locate the biological, geological, geospatial, and water-resources data being generated. The USGS strives to develop data products that are easy to find, easy to understand, and easy to use through Web-accessible tools that allow users to learn about the breadth and scope of GLRI activities being undertaken by the USGS and its partners. By creating tools that enable data to be shared and reused more easily, the USGS can encourage collaboration and assist the GL community in finding, interpreting, and understanding the information created during GLRI science activities.

Great Lakes

Metadata for data rescue and data at risk

Scientific data age, become stale, fall into disuse and run tremendous risks of being forgotten and lost. These problems can be addressed by archiving and managing scientific data over time, and establishing practices that facilitate data discovery and reuse. Metadata documentation is integral to this work and essential for measuring and assessing high priority data preservation cases. The International Council for Science: Committee on Data for Science and Technology (CODATA) has a newly appointed Data-at-Risk Task Group (DARTG), participating in the general arena of rescuing data. The DARTG primary objective is building an inventory of scientific data that are at risk of being lost forever. As part of this effort, the DARTG is testing an approach for documenting endangered datasets. The DARTG is developing a minimal and easy to use set of metadata properties for sufficiently describing endangered data, which will aid global data rescue missions. The DARTG metadata framework supports rapid capture, and easy documentation, across an array of scientific domains. This paper reports on the goals and principles supporting the DARTG metadata schema, and provides a description of the preliminary implementation.

Conference Paper

Proceedings of the first U.S. Geological Survey scientific information management workshop, March 21-23, 2006

In March 2006, the U.S. Geological Survey (USGS) held the first Scientific Information Management (SIM) Workshop in Reston, Virginia. The workshop brought together more than 150 SIM professionals from across the organization to discuss the range and importance of SIM problems, identify common challenges and solutions, and investigate the use and value of “communities of practice” (CoP) as mechanisms to address these issues. The 3-day workshop began with presentations of SIM challenges faced by the Long Term Ecological Research (LTER) network and two USGS programs from geology and hydrology. These presentations were followed by a keynote address and discussion of CoP by Dr. Etienne Wenger, a pioneer and leading expert in CoP, who defined them as "groups of people who share a passion for something that they know how to do and who interact regularly to learn how to do it better." Wenger addressed the roles and characteristics of CoP, how they complement formal organizational structures, and how they can be fostered. Following this motivating overview, five panelists (including Dr. Wenger) with CoP experience in different institutional settings provided their perspectives and lessons learned. The first day closed with an open discussion on the potential intersection of SIM at the USGS with SIM challenges and the potential for CoP. The second session began the process of developing a common vocabulary for both scientific data management and CoP, and a list of eight guiding principles for information management were proposed for discussion and constructive criticism. Following this discussion, 20 live demonstrations and posters of SIM tools developed by various USGS programs and projects were presented. Two community-building sessions were held to explore the next steps in 12 specific areas: Archiving of Scientific Data and Information; Database Networks; Digital Libraries; Emerging Workforce; Field Data for Small Research Projects; Knowledge Capture; Knowledge Organization Systems and Controlled Vocabularies; Large Time Series Data Sets; Metadata; Portals and Frameworks; Preservation of Physical Collections; and Scientific Data from Monitoring Programs. In about two-thirds of these areas, initial steps to forming CoP are now underway. The final afternoon included a panel in which information professionals, managers, program coordinators, and associate directors shared their perspectives on the workshop, on ways in which the USGS could better manage its scientific information, and on the use of CoP as informal mechanisms to complement formal organizational structures. The final session focused on developing the next steps, an action plan, and a communication strategy to ensure continued development.

Scientific Investigations Report

Converting analog interpretive data to digital formats for use in database and GIS applications

There is a growing need by researchers and managers for comprehensive and unified nationwide datasets of scientific data. These datasets must be in a digital format that is easily accessible using database and GIS applications, providing the user with access to a wide variety of current and historical information. Although most data currently being collected by scientists are already in a digital format, there is still a large repository of information in the literature and paper archive. Converting this information into a format accessible by computer applications is typically very difficult and can result in loss of data. However, since scientific data are commonly collected in a repetitious, concise matter (i.e., forms, tables, graphs, etc.), these data can be recovered digitally by using a conversion process that relates the position of an attribute in two-dimensional space to the information that the attribute signifies. For example, if a table contains a certain piece of information in a specific row and column, then the space that the row and column occupies becomes an index of that information. An index key is used to identify the relation between the physical location of the attribute and the information the attribute contains. The conversion process can be achieved rapidly, easily and inexpensively using widely available digitizing and spreadsheet software, and simple programming code. In the geological sciences, sedimentary character is commonly interpreted from geophysical profiles and descriptions of sediment cores. In the field and laboratory, these interpretations were typically transcribed to paper. The information from these paper archives is still relevant and increasingly important to scientists, engineers and managers to understand geologic processes affecting our environment. Direct scanning of this information produces a raster facsimile of the data, which allows it to be linked to the electronic world. But true integration of the content with database and GIS software as point, vector or text information is commonly lost. Sediment core descriptions and interpretation of geophysical profiles are usually portrayed as lines, curves, symbols and text information. They have vertical and horizontal dimensions associated with depth, category, time, or geographic position. These dimensions are displayed in consistent positions, which can be digitized and converted to a digital format, such as a spreadsheet. Once this data is in a digital, tabulated form it can easily be made available to a wide variety of imaging and data manipulation software for compilation and world-wide dissemination.

Open-File Report

On the availability of geoscientific data and scientific collaborators of an in Africa

In addition to the technical papers, the major topic of discussion at the symposium which formed the stepping-off point for this issue of Geoexplorution was the difficult communications between specialists on Africa and the lack of available geological and geophysical data. However, much of this difficulty stems from a lack of advice on where to turn for data and for eollaboration with other interested scientists. This short note attempts to address these problems with some possible solutions.

Geoexploration

Purpose and benefits of U.S. Geological Survey Trusted Digital Repositories

Federal mandates and U.S. Geological Survey (USGS, also known as the Bureau) Fundamental Science Practices (FSP) policies require that publicly funded scientific data, publications, and derivative works be openly accessible to researchers and the public. Open access helps to leverage the public investment by making the acquired data and published information products—collectively referred to as “data assets”—easier to locate, reproduce, and reuse. Open access also provides transparency to the processes used to acquire and analyze the data, thereby helping to ensure the scientific integrity of USGS data and products. The data assets produced by USGS programs, science centers, and projects are preserved digitally in various USGS and non-USGS repositories. To capitalize on the investment expended for data collection, analysis, and interpretation, these systems must remain useful and meaningful. For USGS repositories, the Bureau FSP Advisory Committee has implemented an evaluation process to ensure that the systems being used to preserve these data assets are trustworthy, reliable, and secure and thus provide for data longevity, integrity, and security. A system that is found to meet the reliability and suitability requirements is certified as a USGS Trusted Digital Repository.

Fact Sheet

Bibliography, indices, and data sources of water-related studies, upper Colorado River basin, Colorado and Utah, 1872-1995

As part of the U.S. Geological Survey's National Water-Quality Assessment Program, current water-quality conditions in the Upper Colorado River Basin in Colorado and Utah are being assessed. This report is an initial effort to identify and compile information on water-related studies previously conducted in the basin and consists of a bibliography, coauthor and subject indices, and sources of available water-related data. Computerized literature searches of scientific data bases were carried out to identify past water-related studies in the basin, and government agencies and private organizations were contacted regarding their knowledge or possession of water-related publications and data. Categories of information in the bibliography include: aquatic biology, climate, energy development, geology, land use, limnology, runoff, salinity, surface- and ground-water hydrology, water chemistry, water quality and quantity, and water use and management. The approximately 1,400 indexed references date from 1872 through February 1995 and include books, journal articles, maps, and reports. In many instances, an abstract has been provided for a given reference. Sources of water-related data in the basin are included in a table.

Colorado, Utah

sbtools: A package connecting R to cloud-based data for collaborative online research

The adoption of high-quality tools for collaboration and reproducible research such as R and Github is becoming more common in many research fields. While Github and other version management systems are excellent resources, they were originally designed to handle code and scale poorly to large text-based or binary datasets. A number of scientific data repositories are coming online and are often focused on dataset archival and publication. To handle collaborative workflows using large scientific datasets, there is increasing need to connect cloud-based online data storage to R. In this article, we describe how the new R package sbtools enables direct access to the advanced online data functionality provided by ScienceBase, the U.S. Geological Survey’s online scientific data storage platform.

The R Journal

Stakeholder engagement to guide decision-relevant water data delivery

Water resources management and policy making require access to reliable scientific data. However, water managers may need to overcome various obstacles to accessing data. For example, insufficient technological infrastructures, low data literacy, and data format complexities often inhibit data user access. Thus, it is imperative to include stakeholders in the design of data delivery systems. The United States Geological Survey's Water Resources Mission Area is currently developing Integrated Water Availability Assessments (IWAAs) — multi-extent, stakeholder driven, near real-time water availability census and prediction for human and ecological uses. To provide appropriate user accessibility to data delivery systems developed for IWAAs, a user-centered design process including stakeholder focus groups was used to determine potential water data user needs and preferences. Focus groups identified five types of potential users: Public sector water resources managers, Public sector water resources manager data analysts, Industry and private companies, Tribal Nations, and Nonprofit organizations. Different water data user types depended on diverse spatial and temporal scale data. Public sector water resources managers benefitted most from data synthesized into user-friendly platforms and Public sector water resources data analysts preferred easy access to raw data. These findings can support the development of a water data delivery platform that meets a variety of user needs.

Journal of the American Water Resources Associatio

An interoperability strategy for the next generation of SEEA accounting

The System of Environmental-Economic Accounting (SEEA) is a set of international environmental-economic standards, adopted by the UN Statistical Commission in 2012 (SEEA Central Framework) and 2021 (SEEA Ecosystem Accounting); the latter in particular requires the integration of large and diverse data streams. These include geospatial and other data sources, which have proven challenging for some National Statistical Offices (NSOs) to implement. Although a variety of ecosystem service modelling platforms have been built over the last 15 years to meet various user demands, they often duplicate efforts, rely on data that are siloed, and rarely effectively reuse the knowledge gained from past modelling efforts. By making the data and models that underlie SEEA interoperable, NSOs and the scientific community can advance the accessibility, speed, quality, and transparency of SEEA accounts by making it possible to rapidly integrate and share new scientific data and models. Doing so requires an understanding of the benefits of interoperability, the costs of the status quo, and concrete pathways toward community-endorsed approaches for interoperability. The ARIES Network, which powers the ARIES for SEEA Explorer web application, offers such a path toward interoperability, providing substantial benefits to NSOs and scientific and policy communities.

Report

Assessing decadal-scale coastal change likelihood to define the accuracy and application of scientific information

Defining the accuracy and uncertainties of scientific data products is critical to the usability and trustworthiness of scientific information for environmental management and conservation purposes, such as coastal resource prioritization, design, adaptation, and mitigation. The U.S. Geological Survey has a new decadal-scale coastal change assessment product that synthesizes nearly two dozen coastal datasets. A supervised machine-learning framework is used to combine existing datasets that describe the landscape and the hazards that affect it to determine the coastal change likelihood (CCL) in the coming decade at a resolution of 10 m per pixel for the NE United States from Maine to Virginia. Here, results from a series of statistical tests conducted on source data, the supervised classification, and the CCL outcomes as compared with historical land-cover change are presented. The overall accuracy of the aggregated land-cover dataset that serves as the foundation to which other source datasets are appended is 94%. The supervised learning classification that determines the final CCL output has an overall accuracy of 92%. The CCL predictions of high expected coastal change were consistent with 95% of the coastal and low-elevation landscape change in the last 20 years, as recorded by the Coastal Change Analysis Program land-cover change atlas. Results suggest that CCL provides accurate estimates of coastal landscape change in the next decade that are consistent with recent observed change. Additionally, best practices for applying CCL for planning purposes are outlined, and citing limitations, knowledge gaps, and opportunities for improved accuracy and further investigation are considered.

Journal of Coastal Research

Reproducibility starts at the source: R, Python, and Julia Packages for retrieving USGS hydrologic data

Much of modern science takes place in a computational environment, and, increasingly, that environment is programmed using R, Python, or Julia. Furthermore, most scientific data now live on the cloud, so the first step in many workflows is to query a cloud database and load the response into a computational environment for further analysis. Thus, tools that facilitate programmatic data retrieval represent a critical component in reproducible scientific workflows. Earth science is no different in this regard. To fulfill that basic need, we developed R, Python, and Julia packages providing programmatic access to the U.S. Geological Survey’s National Water Information System database and the multi-agency Water Quality Portal. Together, these packages create a common interface for retrieving hydrologic data in the Jupyter ecosystem, which is widely used in water research, operations, and teaching. Source code, documentation, and tutorials for the packages are available on GitHub. Users can go there to learn, raise issues, or contribute improvements within a single platform, which helps foster better engagement and collaboration between data providers and their users.

Water

U.S. Geological Survey community for data integration: data upload, registry, and access tool

As a leading science and information agency and in fulfillment of its mission to provide reliable scientific information to describe and understand the Earth, the U.S. Geological Survey (USGS) ensures that all scientific data are effectively hosted, adequately described, and appropriately accessible to scientists, collaborators, and the general public. To succeed in this task, the USGS established the Community for Data Integration (CDI) to address data and information management issues affecting the proficiency of earth science research. Through the CDI, the USGS is providing data and metadata management tools, cyber infrastructure, collaboration tools, and training in support of scientists and technology specialists throughout the project life cycle. One of the significant tools recently created to contribute to this mission is the Uploader tool. This tool allows scientists with limited data management resources to address many of the key aspects of the data life cycle: the ability to protect, preserve, publish and share data. By implementing this application inside ScienceBase, scientists also can take advantage of other collaboration capabilities provided by the ScienceBase platform.

Fact Sheet

Exposure pathways and biological receptors: baseline data for the canyon uranium mine, Coconino County, Arizona

Recent restrictions on uranium mining within the Grand Canyon watershed have drawn attention to scientific data gaps in evaluating the possible effects of ore extraction to human populations as well as wildlife communities in the area. Tissue contaminant concentrations, one of the most basic data requirements to determine exposure, are not available for biota from any historical or active uranium mines in the region. The Canyon Uranium Mine is under development, providing a unique opportunity to characterize concentrations of uranium and other trace elements, as well as radiation levels in biota, found in the vicinity of the mine before ore extraction begins. Our study objectives were to identify contaminants of potential concern and critical contaminant exposure pathways for ecological receptors; conduct biological surveys to understand the local food web and refine the list of target species (ecological receptors) for contaminant analysis; and collect target species for contaminant analysis prior to the initiation of active mining. Contaminants of potential concern were identified as arsenic, cadmium, chromium, copper, lead, mercury, nickel, selenium, thallium, uranium, and zinc for chemical toxicity and uranium and associated radionuclides for radiation. The conceptual exposure model identified ingestion, inhalation, absorption, and dietary transfer (bioaccumulation or bioconcentration) as critical contaminant exposure pathways. The biological survey of plants, invertebrates, amphibians, reptiles, birds, and small mammals is the first to document and provide ecological information on .200 species in and around the mine site; this study also provides critical baseline information about the local food web. Most of the species documented at the mine are common to ponderosa pine Pinus ponderosa and pinyon–juniper Pinus–Juniperus spp. forests in northern Arizona and are not considered to have special conservation status by state or federal agencies; exceptions are the locally endemic Tusayan flameflower Phemeranthus validulus, the long-legged bat Myotis volans, and the Arizona bat Myotis occultus. The most common vertebrate species identified at the mine site included the Mexican spadefoot toad Spea multiplicata, plateau fence lizard Sceloporus tristichus, violetgreen swallow Tachycineta thalassina, pygmy nuthatch Sitta pygmaea, purple martin Progne subis, western bluebird Sialia mexicana, deermouse Peromyscus maniculatus, valley pocket gopher Thomomys bottae, cliff chipmunk Tamias dorsalis, black-tailed jackrabbit Lepus californicus, mule deer Odocoileus hemionus, and elk Cervus canadensis. A limited number of the most common species were collected for contaminant analysis to establish baseline contaminant and radiological concentrations prior to ore extraction. These empirical baseline data will help validate contaminant exposure pathways and potential threats from contaminant exposures to ecological receptors. Resource managers will also be able to use these data to determine the extent to which local species are exposed to chemical and radiation contamination once the mine is operational and producing ore. More broadly, these data could inform resource management decisions on mitigating chemical and radiation exposure of biota at high-grade uranium breccia pipes throughout the Grand Canyon watershed.

Arizona

Community for Data Integration 2017 annual report

The Community for Data Integration (CDI) is a group that helps members grow their expertise on all aspects of working with scientific data. The CDI’s activities advance data and information integration capabilities in the U.S. Geological Survey and in the wider Earth and biological sciences. This annual report describes the presentations, activities, collaboration areas, workshop, and other CDI-sponsored events in fiscal year 2017. The report also describes the objectives of the 11 CDI-funded projects in fiscal year 2017. The report shows how the CDI activities fulfill the strategic objective of the U.S. Geological Survey’s Core Science Systems Mission Area to develop a workplace model for interdisciplinary science.

Open-File Report

Highly specialized recreationists contribute the most to the citizen science project eBird

Contributory citizen science projects (hereafter “contributory projects”) are a powerful tool for avian conservation science. Large-scale projects such as eBird have produced data that have advanced science and contributed to many conservation applications. These projects also provide a means to engage the public in scientific data collection. A common challenge across contributory projects like eBird is to maintain participation, as some volunteers contribute just a few times before disengaging. To maximize contributions and manage an effective program that has broad appeal, it is useful to better understand factors that influence contribution rates. For projects capitalizing on recreation activities (e.g., birding), differences in contribution levels might be explained by the recreation specialization framework, which describes how recreationists vary in skill, behavior, and motives. We paired data from a survey of birders across the United States and Canada with data on their eBird contributions (n = 28,926) to test whether those who contributed most are more specialized birders. We assigned participants to 4 contribution groups based on eBird checklist submissions and compared groups’ specialization levels and motivations. More active contribution groups had higher specialization, yet some specialized birders were not active participants. The most distinguishing feature among groups was the behavioral dimension of specialization, with active eBird participants owning specialized equipment and taking frequent trips away from home to bird. Active participants had the strongest achievement motivations for birding (e.g., keeping a life list), whereas all groups had strong appreciation motivations (e.g., enjoying the sights and sounds of birding). Using recreation specialization to characterize eBird participants can help explain why some do not regularly contribute data. Project managers may be able to promote participation, particularly by those who are specialized but not contributing, by appealing to a broader suite of motivations that includes both appreciation and achievement motivations, and thereby increase data for conservation.

Ornithological Applications