USGS Science⌕ Search

SEARCH · USGS Science

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

National Geochemical Database, U.S. Geological Survey RASS (Rock Analysis Storage System) geochemical data for Alaska

This dataset contains geochemical data for Alaska produced by the analytical laboratories of the Geologic Division of the U.S. Geological Survey (USGS). These data represent analyses of stream-sediment, heavy-mineral-concentrate (derived from stream sediment), soil, and organic material samples. Most of the data comes from mineral resource investigations conducted in the Alaska Mineral Resource Assessment Program (AMRAP). However, some of the data were produced in support of other USGS programs. The data were originally entered into the in-house Rock Analysis Storage System (RASS) database. The RASS database, which contains over 580,000 data records, was used by the Geologic Division from the early 1970's through the late 1980's to archive geochemical data. Much of the data have been previously published in paper copy USGS Open-File Reports by the submitter or the analyst but some of the data have never been published. Over the years, USGS scientists recognized several problems with the database. The two primary issues were location coordinates (either incorrect or lacking) and sample media (not precisely identified). This dataset represents a re-processing of the original RASS data to make the data accessible in digital format and more user friendly. This re-processing consisted of checking the information on sample media and location against the original sample submittal forms, the original analytical reports, and published reports. As necessary, fields were added to the original data to more fully describe the sample preparation methods used and sample medium analyzed. The actual analytical data were not checked in great detail, but obvious errors were corrected.

Alaska↗

Surface-water-quality data to support implementation of revised freshwater aluminum water-quality criteria in Massachusetts, 2018–19

The U.S. Geological Survey, in cooperation with the Massachusetts Department of Environmental Protection, performed a study to inform the development of the department’s guidelines for the collection and use of water-chemistry data to support calculation of site-dependent aluminum criteria values. The U.S. Geological Survey collected and analyzed discrete water-quality samples at four wastewater-treatment facilities and seven water-treatment facilities in eastern and central Massachusetts from April 2018 through May 2019. For each of the 11 facilities considered, water-quality samples were collected from treatment-plant effluent and receiving-water bodies. Samples were collected for laboratory analysis of major ions (calcium and magnesium ions are used to calculate total hardness), dissolved organic carbon (DOC), total organic carbon (TOC), and total recoverable aluminum. Field parameters for pH, temperature, and specific conductance were measured in situ concurrently with sample collection. Water-quality conditions differed among monitoring stations. The highest pH values were observed for stations on the Assabet River that receive effluent discharges from wastewater-treatment facilities (the Westborough, Marlborough, Hudson, and Maynard wastewater-treatment facilities). High DOC concentrations (greater than 10 mg/L) were measured in water bodies associated with large areas of riparian wetlands—Lily Pond (Cohasset) and Third Herring Brook (Hanover), and low DOC concentrations (less than 2.5 mg/L) were measured at three water bodies in central Massachusetts—Hocomonco Pond (Westborough), Wyman Pond (Fitchburg), and Monoosnoc Brook (Leominster). Wyman Pond (Fitchburg), Monoosnoc Brook (Leominster), and Lily Pond (Cohasset) also had low pH values and low total hardness concentrations. The monthly discrete pH, DOC, and total hardness data for selected stations on receiving-water bodies were used in the U.S. Environmental Protection Agency Aluminum Criteria Calculator Version 2.0 to estimate site-dependent total recoverable aluminum concentrations that—if not exceeded—would be expected to protect fish, invertebrates, and other aquatic life from adverse effects associated with acute and chronic aluminum exposures. The U.S. Environmental Protection Agency Calculator output provides values for the acute criterion, defined as the criterion maximum concentration (CMC), an estimate of the highest aluminum concentration in surface water to which an aquatic community can be exposed briefly without resulting in an unacceptable effect. This output also provides values for the chronic criterion, defined as the criterion continuous concentration (CCC), an estimate of the highest concentration of aluminum in surface water to which an aquatic community can be exposed indefinitely without resulting in an unacceptable effect. To determine aluminum criteria values typically evaluated for use as protective water-quality criteria, the monthly instantaneous CMC and CCC values were used to calculate the minimum, 5th percentile, and 10th percentile CMC and CCC values for selected monitoring stations. The monthly instantaneous aluminum CMC and CCC values generated using the EPA Calculator varied among stations. Aluminum CMC and CCC values were highest for four ambient (upstream) stations on the Assabet River associated with wastewater-treatment facilities (Westborough, Marlboro, Hudson, and Maynard). Aluminum CMC and CCC values were lower for stations associated with water-treatment facilities, and lowest for selected ambient stations on Lily Pond, Monoosnoc Brook, and Wyman Pond associated with water-treatment facilities in Cohasset, Leominster, and Fitchburg, respectively. For many stations, the highest CMC and CCC instantaneous aluminum criteria values generated using the U.S. Environmental Protection Agency Calculator were for months during the growing season for algae and aquatic macrophytes (April or May through September or October) and the lowest values were for months during the nongrowing season (October or November through March or April), indicating the importance of collecting water-quality data during the nongrowing season. Aluminum CMC and CCC values generated by the U.S. Environmental Protection Agency Calculator are sensitive to variations in the input parameters (pH, DOC, and total hardness). Aluminum solubility is particularly affected by pH. To characterize diel and seasonal variations in pH, multiparameter water-quality monitors recording continuous (15-minute interval) water temperature and pH were installed in the receiving-water body for one station near each facility upstream from the effluent discharge (in rivers) or at a station outside the immediate effect of effluent discharge (in ponds). Continuous water temperature and pH data were collected from April or May 2018 through November or December 2018. Continuous pH data indicated that the pond stations and Assabet River stations had large diel variations in pH during the growing season. Continuous pH data were used together with discrete DOC and total hardness data to evaluate the potential effect of diel variations in pH on calculated site-dependent aluminum criteria values. For the 11 stations, diel variations in pH were determined to correspond to differences in the 10th percentile of CMC values by a median of 160 μg/L, ranging from 0 to 610 μg/L, and differences in the 10th percentile of CCC values by a median of 40 μg/L, ranging from 15 to 210 μg/L. The low monthly instantaneous CMC and CCC values that have the greatest effect on the minimum, 5th percentile, and 10th percentile aluminum values tend to result during the nongrowing season (October or November through March or April) when the range of diel variations in pH is small, thus minimizing the effect of diel variations in pH on the lowest CMC and CCC values. Historical water-quality data on organic carbon in Massachusetts streams were investigated using data retrieved from the USGS National Water Information System database. An assessment of the availability of historical pH, DOC, and hardness data indicated that more data were available for TOC than for DOC. A linear regression equation was developed for the relation between DOC and TOC concentrations to inform the potential use of available data to evaluate water-quality conditions at additional sites across Massachusetts where only pH, hardness, and TOC data are available. DOC and TOC concentrations were well correlated in the 223 samples in which both constituents were analyzed, and the equation had a coefficient of determination ( R 2 ) equal to 0.93.

Massachusetts↗

User engagement to improve coastal data access and delivery

Executive Summary A priority of the U.S. Geological Survey (USGS) Coastal and Marine Hazards and Resources Program focus on coastal change hazards is to provide accessible and actionable science that meets user needs. To understand these needs, 10 virtual Coastal Data Delivery Listening Sessions were completed with 5 coastal data user types that coastal change hazards data are intended to serve: resource managers, consultants, local planners, State planners, and non-USGS researchers. During these listening sessions, participants revealed challenges to coastal data use including being overwhelmed by too many webtools, having a lack of capacity to search for and understand new information, facing difficulties finding data, and not understanding how to apply data. The specific coastal data and information needs described by participants are also detailed in the report and describe data gaps, a need for simpler tools, data needs that differ across spatial and temporal scales, and more outreach on coastal topics and climate change. Participants also suggested leveraging data across study sites and regions to help improve capacity issues and called for more communication and collaboration among and within Federal agencies. The synthesized information from the Coastal Data Delivery Listening Sessions provided in this report can help the USGS and those working on coastal challenges better understand barriers to coastal information use and the exact data requirements of different coastal data users.

Scientific Investigations Report↗

Evaluating oil and gas industry two-dimensional multichannel seismic data for use in near-surface assessment of geologic framework and potential marine minerals resources

Marine seismic reflection data acquired across the Gulf of Mexico during oil and gas exploration are available to the public through an online database archive. The data are archived as two-dimensional multichannel seismic data in two digital formats. The formats include image files in portable document format (PDF), and binary files in industry standard Society for Exploration Geophysicists revision Y (SEG-Y) format. Also included in the database are navigation files and acquisition information associated with the collection of the data. This study examines the data acquired within two geographic areas in the northern Gulf of Mexico. Although the seismic reflection data are acquired for oil and gas exploration many kilometers below the seafloor, this study focuses on the feasibility of using the data for near-surface geologic and seafloor morphologic studies (<100 meters below the seafloor). The report outlines the methodologies used to recover and process the data, including computer processing steps to convert the PDF imagery into SEG-Y format. The report includes two-dimensional profiles of the data to demonstrate the efficacy of the data in near-surface geologic studies. The study found that, for the two areas of interest, the seafloor reflectors in most of the available data are not resolvable. Although the data are readily available and computer processing can adequately image the uppermost reflectors of the seismic profiles, the resolution of the data in most cases are not suitable for near-surface geologic evaluations.

Alabama, Florida, Louisiana, Mississippi, Texas↗

Documentation of a digital spatial data base for hydrologic investigations, Broward County, Florida

Geographic information systems have become an important tool in planning for the protection and development of natural resources, including ground water and surface water. A digital spatial data base consisting of 18 data layers that can be accessed by a geographic information system was developed for Broward County, Florida. Five computer programs, including one that can be used to create documentation files for each data layer and four that can be used to create data layers from data files not already in geographic information system format, were also developed. Four types of data layers have been developed. Data layers for manmade features include major roads, municipal boundaries, the public land-survey section grid, land use, and underground storage tank facilities. The data layer for topographic features consists of surveyed point land-surface elevations. Data layers for hydrologic features include surface-water and rainfall data-collection stations, surface-water bodies, water-control district boundaries, and water-management basins. Data layers for hydrogeologic features include soil associations, transmissivity polygons, hydrogeologic unit depths, and a finite-difference model grid for south-central Broward County. Each data layer is documented as to the extent of the features, number of features, scale, data sources, and a description of the attribute tables where applicable.

Florida↗

Archive of side scan sonar and swath bathymetry data collected during USGS cruise 10CCT01 offshore of Cat Island, Gulf Islands National Seashore, Mississippi, March 2010

In March of 2010, the U.S. Geological Survey (USGS) conducted geophysical surveys east of Cat Island, Mississippi (fig. 1). The efforts were part of the USGS Gulf of Mexico Science Coordination partnership with the U.S. Army Corps of Engineers (USACE) to assist the Mississippi Coastal Improvements Program (MsCIP) and the Northern Gulf of Mexico (NGOM) Ecosystem Change and Hazards Susceptibility Project by mapping the shallow geological stratigraphic framework of the Mississippi Barrier Island Complex. These geophysical surveys will provide the data necessary for scientists to define, interpret, and provide baseline bathymetry and seafloor habitat for this area and to aid scientists in predicting future geomorpholocial changes of the islands with respect to climate change, storm impact, and sea-level rise. Furthermore, these data will provide information for barrier island restoration, particularly in Camille Cut, and provide protection for the historical Fort Massachusetts. For more information refer to http://ngom.usgs.gov/gomsc/mscip/index.html. This report serves as an archive of the processed swath bathymetry and side scan sonar data (SSS). Data products herein include gridded and interpolated surfaces, surface images, and x,y,z data products for both swath bathymetry and side scan sonar imagery. Additional files include trackline maps, navigation files, GIS files, Field Activity Collection System (FACS) logs, and formal FGDC metadata. Scanned images of the handwritten FACS logs and digital FACS logs are also provided as PDF files. Refer to the Acronyms page for expansion of acronyms and abbreviations used in this report or hold the cursor over an acronym for a pop-up explanation. The USGS St. Petersburg Coastal and Marine Science Center assigns a unique identifier to each cruise or field activity. For example, 10CCT01 tells us the data were collected in 2010 for the Coastal Change and Transport (CCT) study and the data were collected during the first field activity for that project in that calendar year. Refer to http://walrus.wr.usgs.gov/infobank/programs/html/definition/activity.html for a detailed description of the method used to assign the field activity ID. Data were collected using a 26-foot (ft) Glacier Bay Catamaran. Side scan sonar and interferometric swath bathymetry data were collected simultaneously along the tracklines. The side scan sonar towfish was towed off the port side just slightly behind the vessel, close to the seafloor. The interferometric swath transducer was sled-mounted on a rail attached between the catamaran hulls. During the survey the sled is secured into position. Navigation was acquired with a CodaOctopus Octopus F190 Precision Attitude and Positioning System and differentially corrected with OmniSTAR. See the digital FACS equipment log for details about the acquisition equipment used. Both raw datasets were stored digitally and processed using CARIS HIPS and SIPS software at the USGS St. Petersburg Coastal and Marine Science Center. For more information on processing refer to the Equipment and Processing page. Post-processing of the swath dataset revealed a motion artifact that is attributed to movement of the pole that the swath transducers are attached to in relation to the boat. The survey took place in the winter months, in which strong winds and rough waves contributed to a reduction in data quality. The rough seas contributed to both the movement of the pole and the very high noise base seen in the raw amplitude data of the side scan sonar. Chirp data were also collected during this survey and are archived separately.

Mississippi↗

A compilation of U.S. Geological Survey pesticide concentration data for water and sediment in the Sacramento–San Joaquin Delta region: 1990–2010

Beginning around 2000, abundance indices of four pelagic fishes (delta smelt, striped bass, longfin smelt, and threadfin shad) within the San Francisco Bay and Sacramento&ndash;San Joaquin Delta began to decline sharply (Sommer and others, 2007). These declines collectively became known as the pelagic organism decline (POD). No single cause has been linked to this decline, and current theories suggest that combinations of multiple stressors are likely to blame. Contaminants (including current-use pesticides) are one potential stressor being investigated for its role in the POD (Anderson, 2007). Pesticide concentration data collected by the U.S. Geological Survey (USGS) at multiple sites in the delta region over the past two decades are critical to understanding the potential effects of current-use pesticides on species of concern as well as the overall health of the delta ecosystem. In April 2010, a compilation of contaminant data for the delta region was published by the State Water Resources Control Board (Johnson and others, 2010). Pesticide occurrence was the major focus of this report, which concluded that &ldquo;there was insufficient high quality data available to make conclusions about the potential role of specific contaminants in the POD.&rdquo; The report cited multiple sources; however, data collected by the USGS were not included in the publication even though these data met all criteria listed for inclusion in the report. What follows is a summary of publicly available USGS data for pesticide concentrations in surface water and sediments within the Sacramento&ndash;San Joaquin Delta region from the years 1990 through 2010. Data were retrieved though the USGS National Water Information System (NWIS) database, a publicly available online-data repository (U.S. Geological Survey, 1998), and from published USGS reports (also available online at http://pubs.er.usgs.gov/). The majority of the data were collected in support of two long term USGS monitoring programs&mdash;National Water Quality Assessment Program (NAWQA; http://water.usgs.gov/ nawqa/) and National Stream Quality Accounting Network (NASQAN; http://water.usgs.gov/nasqan/)&mdash;and through projects associated with the USGS Toxics Substances Hydrology Program (http://toxics.usgs.gov/). In addition, data were collected during multiple research projects that were supported by various federal, state, and local agencies. Although these data have been previously published in some form, it is hoped that by focusing on samples collected within the delta region and presenting these data in a concise format, they will be a valuable resource for scientists, resource managers, and members of the public working to understand the role of pesticides in the POD and their potential effects on the overall health of the delta ecosystem.

California↗

Watershed Data Management (WDM) database for Salt Creek streamflow simulation, DuPage County, Illinois, water years 2005-11

The U.S. Geological Survey (USGS), in cooperation with DuPage County Stormwater Management Division, maintains a USGS database of hourly meteorologic and hydrologic data for use in a near real-time streamflow simulation system, which assists in the management and operation of reservoirs and other flood-control structures in the Salt Creek watershed in DuPage County, Illinois. Most of the precipitation data are collected from a tipping-bucket rain-gage network located in and near DuPage County. The other meteorologic data (wind speed, solar radiation, air temperature, and dewpoint temperature) are collected at Argonne National Laboratory in Argonne, Ill. Potential evapotranspiration is computed from the meteorologic data. The hydrologic data (discharge and stage) are collected at USGS streamflow-gaging stations in DuPage County. These data are stored in a Watershed Data Management (WDM) database. An earlier report describes in detail the WDM database development including the processing of data from January 1, 1997, through September 30, 2004, in SEP04.WDM database. SEP04.WDM is updated with the appended data from October 1, 2004, through September 30, 2011, water years 2005–11 and renamed as SEP11.WDM. This report details the processing of meteorologic and hydrologic data in SEP11.WDM. This report provides a record of snow affected periods and the data used to fill missing-record periods for each precipitation site during water years 2005–11. The meteorologic data filling methods are described in detail in Over and others (2010), and an update is provided in this report.

Illinois↗

Archive of bathymetry data collected at Cape Canaveral, Florida, 2014

Remotely sensed, geographically referenced elevation measurements of the sea floor, acquired by boat- and aircraft-based survey systems, were produced by the U.S. Geological Survey (USGS), St. Petersburg Coastal and Marine Science Center, St. Petersburg, Florida, for the area at Cape Canaveral. The work was conducted as part of a study to describe an updated bathymetric dataset collected in 2014 and compare it to previous data sets. The updated data focus on the bathymetric features and sediment transport pathways that connect the offshore regions to the shoreline and, therefore, are related to the protection of other portions of the coastal environment, such as dunes, that support infrastructure and ecosystems. Cape Canaveral Coastal System (CCCS) is a prominent feature along the Southeast U.S. coastline and is the only large cape south of Cape Fear, North Carolina. Most of the CCCS lies within the Merritt Island National Wildlife Refuge and included within its boundaries are the Cape Canaveral Air Force Station (CCAFS), NASA&rsquo;s Kennedy Space Center (KSC), and a large portion of Canaveral National Seashore. The actual promontory of the modern cape falls within the jurisdictional boundaries of the CCAFS. Hydrographic survey data were collected August 18-20, 2014 ( USGS Field Activity Number 2014-324-FA ). The study covered a 20 kilometer (km) section of shoreline extending from Port Canaveral, Fla., to the northern end of the KSC property, and from the shoreline to about 2.5 km offshore. Data were acquired using both sound navigation and ranging (sonar) and light detection and ranging (lidar) systems. Two jet skis and a 17-foot (ft) outboard motor boat equipped with the USGS SANDS (System for Accurate Nearshore Depth Surveying) hydrographic system collected precision sonar data. The USGS airborne EAARL-B mapping system flown in a twin engine airplane was used to collect lidar data. The missions were synchronized so that there was temporal and spatial overlap between the sonar and lidar operations. Additional data were collected to evaluate water clarity to verify the ability of lidar to receive bathymetric returns. Both systems used differential Global Positioning System GPS and utilized the National Oceanic and Atmospheric Administration/National Geodetic Survey (NOAA/NGS) Continuously Operating Reference Station (CORS) station located at CCAFS was used as the reference station. This data series serves as an archive of processed single-beam sonar and lidar bathymetry data. Graphical Information System (GIS) data products include XYZ point bathymetry data files, a color coded bathymetry map, and interpolated bathymetry grid surface. Additional information includes an error analysis and formal Federal Geographic Data Committee (FGDC) metadata. For more information about similar projects, please visit the Barrier Island Evolution Web site.

Florida↗

Climate, snow, and soil moisture data set for the Tuolumne and Merced river watersheds, California, USA

We present hourly climate data to force land surface process models and assessments over the Merced and Tuolumne watersheds in the Sierra Nevada, California, for the water year 2010–2014 period. Climate data (38 stations) include temperature and humidity (23), precipitation (13), solar radiation (8), and wind speed and direction (8), spanning an elevation range of 333 to 2987 m. Each data set contains raw data as obtained from the source (Level 0), data that are serially continuous with noise and nonphysical points removed (Level 1), and, where possible, data that are gap filled using linear interpolation or regression with a nearby station record (Level 2). All stations chosen for this data set were known or documented to be regularly maintained and components checked and calibrated during the period. Additional time-series data included are available snow water equivalent records from automated stations (8) and manual snow courses (22), as well as distributed snow depth and co-located soil moisture measurements (2–6) from four locations spanning the rain–snow transition zone in the center of the domain. Spatial data layers pertinent to snowpack modeling in this data set are basin polygons and 100 m resolution rasters of elevation, vegetation type, forest canopy cover, tree height, transmissivity, and extinction coefficient. All data are available from online data repositories ( https://doi.org/10.6071/M3FH3D ).

California↗

Are researchers citing their data? A case study from the U.S. Geological Survey

Data citation promotes accessibility and discoverability of data through measures carried out by researchers, publishers, repositories, and the scientific community. This paper examines how a data citation workflow has been implemented by the U.S. Geological Survey (USGS) by evaluating publication and data linkages. Two different methods were used to identify data citations: examining publication structural metadata and examining the full text of the publication. A growing number of USGS researchers are complying with publisher data sharing policies aimed to capture data citation information in a standardized way within associated publications. However, inconsistencies in how data citation information is documented in publications has limited the accessibility and discoverability of the data. This paper demonstrates how organizational evaluations of publication and data linkages can be used to identify obstacles in advancing data citation efforts and improve data citation workflows.

Data Science Journal↗

Water quality data for national-scale aquatic research: The Water Quality Portal

Aquatic systems are critical to food, security, and society. But, water data are collected by hundreds of research groups and organizations, many of which use nonstandard or inconsistent data descriptions and dissemination, and disparities across different types of water observation systems represent a major challenge for freshwater research. To address this issue, the Water Quality Portal (WQP) was developed by the U.S. Environmental Protection Agency, the U.S. Geological Survey, and the National Water Quality Monitoring Council to be a single point of access for water quality data dating back more than a century. The WQP is the largest standardized water quality data set available at the time of this writing, with more than 290 million records from more than 2.7 million sites in groundwater, inland, and coastal waters. The number of data contributors, data consumers, and third-party application developers making use of the WQP is growing rapidly. Here we introduce the WQP, including an overview of data, the standardized data model, and data access and services; and we describe challenges and opportunities associated with using WQP data. We also demonstrate through an example the value of the WQP data by characterizing seasonal variation in lake water clarity for regions of the continental U.S. The code used to access, download, analyze, and display these WQP data as shown in the figures is included as supporting information.

Water Resources Research↗

Estimating ungulate migration corridors from sparse movement data

Many ungulates migrate between distinct summer and winter ranges, and identifying, mapping, and conserving these migration corridors have become a focus of local, regional, and global conservation efforts. Brownian bridge movement models (BBMMs) are commonly used to empirically identify these seasonal migration corridors; however, they require location data sampled at relatively frequent intervals to obtain a robust estimate of an animal’s movement path. Fitting BBMMs to sparse location data violates the assumption of conditional random movement between successive locations, overestimating the area (and width) of a migration corridor when creating individual and population-level occurrence distributions, and precluding the use of low-frequency, or sparse, data in mapping migration corridors. In an effort to expand the utility of BBMMs to include sparse global positioning system (GPS) data, we propose an alternative approach to model migration corridors from sparse GPS data. We demonstrate this method using GPS data collected every 2 hours from four mule deer (Odocoileus hemionus) and four elk (Cervus canadensis) herds within Wyoming and Idaho. First, we used BBMMs to estimate a baseline corridor for the 2-hour data. We then subsampled the 2-hour data to one location every 12 hours (a proxy for sparse data) and fitted BBMMs to the 12-hour data using a fixed motion variance (FMV) value, instead of estimating the Brownian motion variance empirically. A range of FMV values was tested to identify the value that best approximated the baseline migration corridor. FMV values within a species-specific range (mule deer: 400–1,200 m2; elk: 600–1,600 m2) successfully delineated migration corridors similar to the 2-hour baseline corridors; overall, lower values delineated narrower corridors and higher values delineated wider corridors. Optimal FMV values of 800 m2 (mule deer) and 1,000 m2 (elk) decreased the inflation of the 12-hour corridors relative to the 2-hour corridors from traditional BBMMs. This FMV approach thus enables using sparse movement data to approximate realistic migration corridor dimensions, providing an important alternative when movement data are collected infrequently. This approach greatly expands the number of datasets that can be used for migration corridor mapping—a useful tool for management and conservation across the globe.

Idaho, Wyoming↗

Too much and not enough data: Challenges and solutions for generating information in freshwater research and monitoring

Evaluating progress toward achieving freshwater conservation and sustainability goals requires transforming diverse types of data into useful information for scientists, managers, and other interest groups. Despite substantial increases in the volume of freshwater data collected worldwide, many regions and ecosystems still lack sufficient data collection and/or data access. We illustrate how these data challenges result from a diverse set of underlying mechanisms and propose solutions that can be applied by individuals or organizations. We discuss creative approaches to address data scarcity, including the use of community science, remote-sensing, environmental sensors, and legacy datasets. We highlight the importance of coordinated data collection efforts among groups and training programs to improve data access. At the institutional level, we emphasize the power of prioritizing data curation, incentivizing data publication, and promoting research that enhances data coverage and representativeness. Some of these strategies involve technological and analytical approaches, but many necessitate shifting the priorities and incentives of organizations such as academic and government research institutions, monitoring groups, journals, and funding agencies. Our overarching goal is to stimulate discussion to narrow the data disparities hindering the understanding of freshwater processes and their change across spatial scales.

Nahuel Huapi Lake, Lake Tahoe↗

Seabed mapping and characterization of sediment variability using the usSEABED data base

We present a methodology for statistical analysis of randomly located marine sediment point data, and apply it to the US continental shelf portions of usSEABED mean grain size records. The usSEABED database, like many modern, large environmental datasets, is heterogeneous and interdisciplinary. We statistically test the database as a source of mean grain size data, and from it provide a first examination of regional seafloor sediment variability across the entire US continental shelf. Data derived from laboratory analyses ("extracted") and from word-based descriptions ("parsed") are treated separately, and they are compared statistically and deterministically. Data records are selected for spatial analysis by their location within sample regions: polygonal areas defined in ArcGIS chosen by geography, water depth, and data sufficiency. We derive isotropic, binned semivariograms from the data, and invert these for estimates of noise variance, field variance, and decorrelation distance. The highly erratic nature of the semivariograms is a result both of the random locations of the data and of the high level of data uncertainty (noise). This decorrelates the data covariance matrix for the inversion, and largely prevents robust estimation of the fractal dimension. Our comparison of the extracted and parsed mean grain size data demonstrates important differences between the two. In particular, extracted measurements generally produce finer mean grain sizes, lower noise variance, and lower field variance than parsed values. Such relationships can be used to derive a regionally dependent conversion factor between the two. Our analysis of sample regions on the US continental shelf revealed considerable geographic variability in the estimated statistical parameters of field variance and decorrelation distance. Some regional relationships are evident, and overall there is a tendency for field variance to be higher where the average mean grain size is finer grained. Surprisingly, parsed and extracted noise magnitudes correlate with each other, which may indicate that some portion of the data variability that we identify as "noise" is caused by real grain size variability at very short scales. Our analyses demonstrate that by applying a bias-correction proxy, usSEABED data can be used to generate reliable interpolated maps of regional mean grain size and sediment character.

Continental Shelf Research↗

Heed the data gap: Guidelines for using incomplete datasets in annual stream temperature analyses

Stream temperature data are useful for deciphering watershed processes important for aquatic ecosystems. Accurately extracting signal trends from stream temperature is essential for predicting responses of environmental and ecological indicators to change. Missing data periods are common for various reasons, and pose a challenge for scientists using temperature signal analysis to support stream research and ecological management objectives. However, the sensitivity of estimated temperature signal patterns to missing data has not been thoroughly evaluated, despite the potentially large impact on interpretation. In this study, we explored the effects of simulated missing daily data on the characterization of annual water temperature signals measured at headwater sites in the Pacific Northwest and Mid-Atlantic regions of the USA. For each site, we used linear regressions of sine-waves fitted to complete (365-d) and partial (7–357 consecutive missing data points) annual datasets of daily mean water temperature and computed three thermal parameters (mean, phase, and amplitude), which together can indicate thermally and ecologically influential watershed processes (e.g., depth and magnitude of groundwater discharge). Expected values (derived from complete datasets) ranged from 7.0 to 12.6 °C, 205 to 254 d, and 1.9 to 9.5 °C for annual mean, phase, and amplitude, respectively. While annual phase and amplitude could be accurately estimated (i.e., within 95–99% confidence intervals of expected values) with up to approximately two months of consecutively missing data, annual mean temperature required more complete datasets. We found that datasets with less than seven weeks of consecutively missing data enabled estimation of all annual signal parameters with reasonable accuracy (>75% probability of being within the 95–99% confidence intervals of expected values). Imputation of missing data expanded this range to approximately 20 weeks, with the greatest improvements in parameter estimation between 9 and 27 weeks of imputed missing data. However, caution should be exercised when applying this technique. For example, imputation improved the accuracy of parameter estimation for most sites, but accuracy decreased for some sites exhibiting strong groundwater influence. The timing of consecutive missing data points within a year had inconsistent effects on annual thermal parameter estimates among regions, years, and individual parameters. Utilizing sites with more than approximately seven consecutive weeks of missing data or 20 weeks of imputed data increases the probability of mischaracterization of annual stream thermal regimes. Understanding this limitation is vital for identifying the potential of streams to serve as climate refugia for ecological indicator species and effective future management of stream systems.

Ecological Indicators↗

Analysis of simulated advanced spaceborne thermal emission and reflection (ASTER) radiometer data of the Iron Hill, Colorado, study area for mapping lithologies

The advanced spaceborne thermal emission and reflection (ASTER) radiometer was designed to record reflected energy in nine channels with 15 or 30 m resolution, including stereoscopic images, and emitted energy in five channels with 90 m resolution from the NASA Earth Observing System AMI platform. A simulated ASTER data set was produced for the Iron Hill, Colorado, study area by resampling calibrated, registered airborne visible/infrared imaging spectrometer (AVIRIS) data, and thermal infrared multispectral scanner (TIMS) data to the appropriate spatial and spectral parameters. A digital elevation model was obtained to simulate ASTER-derived topographic data. The main lithologic units in the area are granitic rocks and felsite into which a carbonatite stock and associated alkalic igneous rocks were intruded; these rocks are locally covered by Jurassic sandstone, Tertiary rhyolitic tuff, and colluvial deposits. Several methods were evaluated for mapping the main lithologic units, including the unsupervised classification and spectral curve-matching techniques. In the five thermalinfrared (TIR) channels, comparison of the results of linear spectral unmixing and unsupervised classification with published geologic maps showed that the main lithologic units were mapped, but large areas with moderate to dense tree cover were not mapped in the TIR data. Compared to TIMS data, simulated ASTER data permitted slightly less discrimination in the mafic alkalic rock series, and carbonatite was not mapped in the TIMS nor in the simulated ASTER TIR data. In the nine visible and near-infrared channels, unsupervised classification did not yield useful results, but both the spectral linear unmixing and the matched filter techniques produced useful results, including mapping calcitic and dolomitic carbonatite exposures, travertine in hot spring deposits, kaolinite in argillized sandstone and tuff, and muscovite in sericitized granite and felsite, as well as commonly occurring illite/muscovite. However, the distinction made in AVIRIS data between calcite and dolomite was not consistently feasible in the simulated ASTER data. Comparison of the lithologie information produced by spectral analysis of the simulated ASTER data to a photogeologic interpretation of a simulated ASTER color image illustrates the high potential of spectral analysis of ASTER data to geologic interpretation.

Journal of Geophysical Research D: Atmospheres↗

OpenET: Filling a critical data gap in water management for the western United States

The lack of consistent, accurate information on evapotranspiration (ET) and consumptive use of water by irrigated agriculture is one of the most important data gaps for water managers in the western United States (U.S.) and other arid agricultural regions globally. The ability to easily access information on ET is central to improving water budgets across the West, advancing the use of data-driven irrigation management strategies, and expanding incentive-driven conservation programs. Recent advances in remote sensing of ET have led to the development of multiple approaches for field-scale ET mapping that have been used for local and regional water resource management applications by U.S. state and federal agencies. The OpenET project is a community-driven effort that is building upon these advances to develop an operational system for generating and distributing ET data at a field scale using an ensemble of six well-established satellite-based approaches for mapping ET. Key objectives of OpenET include: Increasing access to remotely sensed ET data through a web-based data explorer and data services; supporting the use of ET data for a range of water resource management applications; and development of use cases and training resources for agricultural producers and water resource managers. Here we describe the OpenET framework, including the models used in the ensemble, the satellite, meteorological, and ancillary data inputs to the system, and the OpenET data visualization and access tools. We also summarize an extensive intercomparison and accuracy assessment conducted using ground measurements of ET from 139 flux tower sites instrumented with open path eddy covariance systems. Results calculated for 24 cropland sites from Phase I of the intercomparison and accuracy assessment demonstrate strong agreement between the satellite-driven ET models and the flux tower ET data. For the six models that have been evaluated to date (ALEXI/DisALEXI, eeMETRIC, geeSEBAL, PT-JPL, SIMS, and SSEBop) and the ensemble mean, the weighted average mean absolute error (MAE) values across all sites range from 13.6 to 21.6 mm/month at a monthly timestep, and 0.74 to 1.07 mm/day at a daily timestep. At seasonal time scales, for all but one of the models the weighted mean total ET is within ±8% of both the ensemble mean and the weighted mean total ET calculated from the flux tower data. Overall, the ensemble mean performs as well as any individual model across nearly all accuracy statistics for croplands, though some individual models may perform better for specific sites and regions. We conclude with three brief use cases to illustrate current applications and benefits of increased access to ET data, and discuss key lessons learned from the development of OpenET.

western United States↗