USGS Science⌕ Search

SEARCH · USGS Science

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Evaluation of downscaled, gridded climate data for the conterminous United States

Weather and climate affect many ecological processes, making spatially continuous yet fine-resolution weather data desirable for ecological research and predictions. Numerous downscaled weather data sets exist, but little attempt has been made to evaluate them systematically. Here we address this shortcoming by focusing on four major questions: (1) How accurate are downscaled, gridded climate data sets in terms of temperature and precipitation estimates?, (2) Are there significant regional differences in accuracy among data sets?, (3) How accurate are their mean values compared with extremes?, and (4) Does their accuracy depend on spatial resolution? We compared eight widely used downscaled data sets that provide gridded daily weather data for recent decades across the United States. We found considerable differences among data sets and between downscaled and weather station data. Temperature is represented more accurately than precipitation, and climate averages are more accurate than weather extremes. The data set exhibiting the best agreement with station data varies among ecoregions. Surprisingly, the accuracy of the data sets does not depend on spatial resolution. Although some inherent differences among data sets and weather station data are to be expected, our findings highlight how much different interpolation methods affect downscaled weather data, even for local comparisons with nearby weather stations located inside a grid cell. More broadly, our results highlight the need for careful consideration among different available data sets in terms of which variables they describe best, where they perform best, and their resolution, when selecting a downscaled weather data set for a given ecological application.

Ecological Applications↗

What is the (real) rate of soil health practice adoption? Making sense of three data sources

Conservation stakeholders looking to quantify the impact of their investments to increase soil health practice adoption over time often face challenges in interpreting practice adoption data due to discrepancies in language and results among data sources. Similarly, efforts to estimate environmental outcomes of practice adoption, such as water quality and greenhouse gas emissions, can vary depending on different practice adoption input data. To help make sense of different adoption data sources, we compared county-level adoption data for winter cover crops (WCC), no-till (NT), and reduced tillage (RT) in three areas of the United States with contrasting climates and production systems: central Illinois (CIL), southern Illinois (SIL), and western New York (WNY). We analyzed data available during 2015 through 2022 from the Operational Tillage Information System (OpTIS, remote sensing), US Census of Agriculture (AgCensus, a farmer survey), and, specifically in Illinois, the Illinois Soil Conservation Transect Survey (Transect, a roadside survey). The magnitude of differences between the datasets depended on the practice and geographic location. For example, OpTIS and AgCensus tillage data were much more similar in Illinois (average difference of less than 4 percentage points) compared to New York (average differences of 20 percentage points). Similarly, there was less variability and smaller differences between OpTIS and AgCensus WCC data in Illinois compared to WNY. AgCensus tended to report lower WCC adoption for Illinois and greater adoption in WNY compared to OpTIS. All data sources agreed that the rate of change in tillage practices is slow (mainly –1% to 1%) and that adoption of WCC is low (assuming linear growth, it could take nearly a century to reach 50% WCC adoption in CIL). Differences among the datasets were attributed to definitional inconsistencies for RT and NT and how WCC data were acquired. For example, the AgCensus asks if a WCC was planted, whereas OpTIS and Transect evaluate the presence of a standing WCC. Data sources also reflect different time periods (calendar years or crop years) and types of cropland assessed (corn [ Zea mays L.], soybean [ Glycine max {L.} Merr.], or all cropland). We propose two recommendations to improve interpretation and consistency: (1) a working group to harmonize definitions and protocols and develop educational materials for data users, and (2) a research effort that integrates different adoption data types and produces publicly available adoption data at HUC-10 and county scales. Such activities could help improve data access and utility for evidence-based conservation decision-making and enhance the accuracy of environmental models that rely on adoption data as input.

Journal of Soil and Water Conservation↗

Challenges and opportunities for data integration to improve estimation of migratory connectivity

Understanding migratory connectivity, or the linkage of populations between seasons, is critical for effective conservation and management of migratory wildlife. A growing number of tools are available for understanding where migratory individuals and populations occur throughout the annual cycle. Integration of the diverse measures of migratory movements can help elucidate migratory connectivity patterns with methodology that accounts for differences in sampling design, directionality, effort, precision and bias inherent to each data type. The R package MigConnectivity was developed to estimate population-specific connectivity and the range-wide strength of those connections. New functions allow users to integrate intrinsic markers, tracking and long-distance reencounter data, collected from the same or different individuals, to estimate population-specific transition probabilities (estTransition) and the range-wide strength of those transition probabilities (estStrength). We used simulation and real-world case studies to explore the challenges and limitations of data integration based on data from three migratory bird species, Painted Bunting ( Passerina ciris ), Yellow Warbler ( Setophaga petechia ) and Bald Eagle ( Haliaeetus leucocephalus ), two of which had bidirectional data. We found data integration is useful for quantifying migratory connectivity, as single data sources are less likely to be available across the species range. Furthermore, accurate strength estimates can be obtained from either breeding-to-nonbreeding or nonbreeding-to-breeding data. For bidirectional data, integration can lead to more accurate estimates when data are available from all regions in at least one season. The ability to conduct combined analyses that account for the unique limitations and biases of each data type is a promising possibility for overcoming the challenge of range-wide coverage that has been hard to achieve using single data types. The best-case scenario for data integration is to have data from all regions, especially if the question is range-wide or data are bidirectional. Multiple data types on animal movements are becoming increasingly available and integration of these growing datasets will lead to a better understanding of the full annual cycle of migratory animals.

Methods in Ecology and Evolution↗

Digital data grids for the magnetic anomaly map of North America

The digital magnetic anomaly database and map for the North American continent is the result of a joint effort by the Geological Survey of Canada (GSC), U. S. Geological Survey (USGS), and Consejo de Recursos Minerales of Mexico (CRM). This integrated, readily accessible, modern digital database of magnetic anomaly data is a powerful tool for further evaluation of the structure, geologic processes, and tectonic evolution of the continent and may also be used to help resolve societal and scientific issues that span national boundaries. The North American magnetic anomaly map derived from the digital database provides a comprehensive magnetic view of continental-scale trends not available in individual data sets, helps link widely separated areas of outcrop, and unifies disparate geologic studies. This open-file report presents three unique, gridded data sets used to make the magnetic anomaly map of North America. Subsets of these three grids that span only the United States were also created, giving a total of six grids. Details on the data processing and compilation procedures used to produce the grids are described in the booklet that accompanies the North American magnetic anomaly map. All three grids have 1-km spacing and are projected to the DNAG projection (spherical transverse mercator, central meridian of 100 o W, base latitude of 0o, scale factor of 0.926 and Earth radius of 6,371,204 m.) More details are given in the metadata files that accompany the gridded data files. These grids are presented in Geosoft binary grid format, with two files describing each of the six grids (suffixes .grd and .gi). This format can be easily converted to numerous other formats using the free conversion software offered by this company at http://www.geosoft.com/. The first grids (NAmag_origmrg.grd and USmag_origmrg.grd) show the magnetic field at 305 m. above terrain. For the second grids (NAmag_hp500.grd and USmag_hp500.grd) we removed long-wavelength anomalies (500 km and greater) from the first grid. This grid was used for the published map. Although the North American merged grid represents a significant upgrade to older compilations, the existing patchwork of surveys is inherently unable to accurately represent anomalies with long (greater than roughly 150 km) wavelengths, particularly in the US and Canada (U.S. Magnetic-Anomaly Data Set Task Group, 1994). The lack of information about long wavelength anomalies is primarily related to datum shifts between merged surveys, caused by data acquisition at widely different times and by differences in merging procedures. Therefore, we removed anomalies with wavelengths greater than 500 km from the merged grid to reduce the effects caused by the spurious long wavelengths but still maintain the continuity of anomalies. The correction was accomplished by transforming the merged grid to the frequency domain, filtering the transformed data with a long-wavelength cutoff at 500 km, and subtracting the long-wavelength data grid from the merged grid. In addition to the 500-km high pass filter, an equivalent source method, based on long-wavelength characterization using satellite data (CHAMP satellite anomalies, Maus and others, 2002), was also used to correct for spurious shifts in the original magnetic anomaly grid (Ravat and others, 2002). These results are presented in the third grids (NAmag_CM.grd and USmag_CM.grd), in which the wavelengths longer than 500 km have been replaced by downward-continued satellite data. The steps used to create the third long-wavelength-corrected grid are: 0. The North American 1-km merged grid was decimated to 5 km. 1. This 5-km grid was converted to a 0.05 degree grid and was low-pass filtered using a Gaussian filter with a 500-km cutoff, then decimated to 1 degree. 2. A joint inversion of this 1-degree low-pass aeromagnetic grid and satellite data, with the aeromagnetic data weighted very low, was used to produce a stabilized downward continuation of the satellite data. 3. The inverted data were interpolated to 0.05 degrees and again low-pass filtered using the same Gaussian 500-km filter to remove short-wavelength artifacts. 4. The low-pass grid from step 1 was subtracted from the original 0.05-degree aeromagnetic grid to create a 500-km high-pass aeromagnetic grid. This grid was added to the low-pass inverted grid from step 3 to get a corrected 0.05-degree aeromagnetic grid. 5. The corrected 0.05-degree aeromagnetic grid was projected to the DNAG projection and regridded to 5 km. This was subtracted from the decimated 5-km aeromagnetic grid to generate a 5-km correction grid. A matched filter was used to remove short-wavelength artifacts resulting from the projection and regridding process. 6. The resulting 5-km correction grid was regridded to the original 1-km grid and subtracted from the original 1-km aeromagnetic grid to generate the final 1-km corrected aeromagnetic grid. The six grids described in this report are available for download. Two metadata files, one for the North American grids and one for the United States grids, are also included with the gridded data.

Open-File Report↗

Wave data processing toolbox manual

Researchers routinely deploy oceanographic equipment in estuaries, coastal nearshore environments, and shelf settings. These deployments usually include tripod-mounted instruments to measure a suite of physical parameters such as currents, waves, and pressure. Instruments such as the RD Instruments Acoustic Doppler Current Profiler (ADCP(tm)), the Sontek Argonaut, and the Nortek Aquadopp(tm) Profiler (AP) can measure these parameters. The data from these instruments must be processed using proprietary software unique to each instrument to convert measurements to real physical values. These processed files are then available for dissemination and scientific evaluation. For example, the proprietary processing program used to process data from the RD Instruments ADCP for wave information is called WavesMon. Depending on the length of the deployment, WavesMon will typically produce thousands of processed data files. These files are difficult to archive and further analysis of the data becomes cumbersome. More imperative is that these files alone do not include sufficient information pertinent to that deployment (metadata), which could hinder future scientific interpretation. This open-file report describes a toolbox developed to compile, archive, and disseminate the processed wave measurement data from an RD Instruments ADCP, a Sontek Argonaut, or a Nortek AP. This toolbox will be referred to as the Wave Data Processing Toolbox. The Wave Data Processing Toolbox congregates the processed files output from the proprietary software into two NetCDF files: one file contains the statistics of the burst data and the other file contains the raw burst data (additional details described below). One important advantage of this toolbox is that it converts the data into NetCDF format. Data in NetCDF format is easy to disseminate, is portable to any computer platform, and is viewable with public-domain freely-available software. Another important advantage is that a metadata structure is embedded with the data to document pertinent information regarding the deployment and the parameters used to process the data. Using this format ensures that the relevant information about how the data was collected and converted to physical units is maintained with the actual data. EPIC-standard variable names have been utilized where appropriate. These standards, developed by the NOAA Pacific Marine Environmental Laboratory (PMEL) (http://www.pmel.noaa.gov/epic/), provide a universal vernacular allowing researchers to share data without translation.

Open-File Report↗

Using the U.S. Geological Survey National Water Quality Laboratory LT-MDL to Evaluate and Analyze Data

A long-term method detection level (LT-MDL) and laboratory reporting level (LRL) are used by the U.S. Geological Survey?s National Water Quality Laboratory (NWQL) when reporting results from most chemical analyses of water samples. Changing to this method provided data users with additional information about their data and often resulted in more reported values in the low concentration range. Before this method was implemented, many of these values would have been censored. The use of the LT-MDL and LRL presents some challenges for the data user. Interpreting data in the low concentration range increases the need for adequate quality assurance because even small contamination or recovery problems can be relatively large compared to concentrations near the LT-MDL and LRL. In addition, the definition of the LT-MDL, as well as the inclusion of low values, can result in complex data sets with multiple censoring levels and reported values that are less than a censoring level. Improper interpretation or statistical manipulation of low-range results in these data sets can result in bias and incorrect conclusions. This document is designed to help data users use and interpret data reported with the LTMDL/ LRL method. The calculation and application of the LT-MDL and LRL are described. This document shows how to extract statistical information from the LT-MDL and LRL and how to use that information in USGS investigations, such as assessing the quality of field data, interpreting field data, and planning data collection for new projects. A set of 19 detailed examples are included in this document to help data users think about their data and properly interpret lowrange data without introducing bias. Although this document is not meant to be a comprehensive resource of statistical methods, several useful methods of analyzing censored data are demonstrated, including Regression on Order Statistics and Kaplan-Meier Estimation. These two statistical methods handle complex censored data sets without resorting to substitution, thereby avoiding a common source of bias and inaccuracy.

Open-File Report↗

Community for Data Integration 2014 annual report

The U.S. Geological Survey (USGS) researches Earth science to help address complex issues affecting society and the environment. In 2006, the USGS held the first Scientific Information Management Workshop to bring together staff from across the organization to discuss the data and information management issues affecting the integration and delivery of Earth science research and investigate the use of “communities of practice” as mechanisms to share expertise about these issues. Out of this effort emerged the Council for Data Integration, which was conceived as an official organizational function that would help guide data integration activities and formalize communities of practice into working groups; however, by 2009 it became evident that many members of the Council for Data Integration had an interest in developing data integration solutions and sharing expertise in a less formal, grassroots manner, which transformed the Council into a Community for Data Integration (CDI). As of 2014, the CDI represents a dynamic community of practice focused on advancing science data and information management and integration capabilities across the USGS and the CDI community. The CDI fosters an environment for collaboration and sharing by bringing together expertise from external partners and representatives across the USGS who are involved in research, data management, and information technology. Membership is voluntary and open to USGS employees and other individuals and organizations willing to contribute to the community (if interested, contact cdi@usgs.gov). The purpose of the CDI is to do the following: • advance understanding of Earth systems through enhanced use of data and information including associated tools and techniques, • provide a forum for people doing work with data integration to come together to share ideas and learn new skills and techniques, and • grow overall USGS capabilities with data and information by increasing visibility of the work of many people throughout the USGS and the CDI community. To achieve these goals, the CDI operates within four applied areas: monthly forums, annual workshop/webinar series, working groups, and projects. The monthly forums, also known as the Opportunity/Challenge of the Month, provide an open dialogue to share and learn about data integration efforts or to present problems that invite the community to offer solutions, advice, and support. Since 2010, the CDI has also sponsored annual workshops/webinar series to encourage the exchange of ideas, sharing of activities, presentations of current projects, and networking among members. Stemming from common interests, the working groups are focused on efforts to address data management and technical challenges including the development of standards and tools, improving interoperability and information infrastructure, and data preservation within USGS and its partners. The growing support for the activities of the working groups led to the CDI’s first formal request for proposals (RFP) process in 2013 to fund projects that produced tangible products. As of 2014, the CDI continues to hold an annual RFP that creates data management tools and practices, collaboration tools, and training in support of data integration and delivery.

Open-File Report↗

A computerized data base of nitrate concentrations in Indiana ground water

As part of a cooperative study with the Indiana Department of Environmental Management, the U.S. Geological Survey compiled a computerized data base of nitrate concentrations in Indiana ground water. The data included nitrate determinations from more than 29 studies by five Federal and State agencies during June 1973 through August 1991. The National Water Information System software of the U.S. Geological Survey was used to store the data at the U.S. Geological Survey office in Indianapolis, Indiana. Electronic data sets were converted to a standard format of well data, sample data, and analytical data. Data were screened by several error-checking procedures before they were retained in the data base; they were examined for potential duplicates of well location and name. The data base of nitrate concentrations in Indiana ground water contains records of 5,525 samples collected from 4,448 wells in 88 of 92 counties during 1973-91. Those wells included 3,832 drinking-water wells; 536 monitoring wells, 38 livestock-supply wells; and 42 irrigation wells. Nitrate concentrations greater than minimum reporting limits of 0.0 to 0.5 milligrams per liter (mg/L) were determined in 2,453 samples (44 percent of the total). Nitrate in ground water at concentrations greater than 3 mg/L have been considered to be the result of human activities. Nitrate concentrations ranged from 0.005 to 380 mg/L with a median nitrate concentration of 0.3 mg/L. Nitrate concentrations were greater than or equal to 3 mg/L in 704 samples (13 percent of the total). Nitrate concentrations were greater than or equal to the U.S. Environmental Protection Agency Maximum Contaminant Level of 10 mg/L in 188 samples (3.4 percent of the total). Of the 3,832 drinking- water wells in the data base, 147 had at least one sample in which a nitrate concentration was greater than the Maximum Contaminant Level. The percentage of samples with nitrate concentrations greater than or equal to 3 mg/L and greater than or equal to 10 mg/L generally increased during the period 1973 through 1991. The nitrate data base was compiled from numerous data sets that were readily accessible in electronic format. The uses of these data may be limited because they were neither comprehensive nor of a single statistical design. Nonetheless, the nitrate data can be used in several ways: (1) to identify geographic areas with and without nitrate data; (2) to evaluate assumptions, models, and maps of ground-water-contamination potential; and (3) to investigate the relation between environmental factors, land-use types, and the occurrence of nitrate.

Indiana↗

Global cropland-extent product at 30-m resolution (GCEP30) derived from Landsat satellite time-series data for the year 2015 using multiple machine-learning algorithms on Google Earth Engine cloud

Executive Summary Global food and water security analysis and management require precise and accurate global cropland-extent maps. Existing maps have limitations, in that they are (1) mapped using coarse-resolution remote-sensing data, resulting in the lack of precise mapping location of croplands and their accuracies; (2) derived by collecting and collating national statistical data that are often subjective, leading to substantial uncertainties in cropland-area estimates, as well as their locations; and (3) extracted from one or more classes of a land use–land cover product in which cropland classes are not the focus of mapping, leading to their mixing with other classes and creating significant errors of omission and commission. These limitations can be overcome by producing high-resolution cropland-extent maps using satellite-sensor data, such as Landsat 30-m resolution or higher. The most fundamental cropland product is the high-resolution cropland-extent map because all higher level cropland products, such as crop-watering method (that is, whether crops are irrigated or rainfed), crop types, cropping intensities, cropland fallows, crop productivity, and crop-water productivity, are dependent on a precise and accurate cropland-extent product. Given these realities, the overarching goal of this study was to produce a Landsat satellite-derived global cropland-extent product at 30-m resolution. The work, which involved a paradigm shift in how global cropland-extent maps are produced, involved the following five key steps: (1) petabyte-scale computing that involved multiyear, 8- to 16-day, time-series Landsat 30-m resolution data for the global land surface; (2) composition of analysis-ready data (ARD) cubes; (3) creation of a large global-reference data hub for machine learning; (4) use of multiple machine-learning algorithms (MLAs) by writing software and computing in the cloud; and (5) Google Earth Engine (GEE) cloud computing. The five key steps involved nine distinct phases. First, the world was segmented into 74 agroecological zones (AEZs). Second, Landsat 8- to 16-day data were used to time-composite 10-band (blue, green, red, near-infrared, short-wave infrared band 1, short-wave infrared band 2, thermal infrared, enhanced vegetation index, normalized difference water index, and normalized difference vegetation index) Landsat 30-m resolution data cubes for every 2- to 4-month time period during 3- to 4-year periods (stated as nominal-year 2015 or, simply, 2015), along with two additional 30-m resolution bands (Shuttle Radar Topography Mission elevation, and slope) in each of the 74 AEZs. Third, more than 100,000 reference-training data samples were collected using ground data (some of which were collected using a mobile application), as well as submeter- to 5-m-resolution, very high-resolution imagery sourced from other reliable sources. Fourth, reference-training data were used to create a knowledge base for separating cropland from noncropland. Fifth, MLAs such as the pixel-based supervised random forest and support-vector machines were written on the GEE using Python and JavaScript. Sixth, object-based recursive hierarchical segmentation algorithm was used, in addition to MLAs, to overcome uncertainties. Seventh, MLAs used the knowledge base to classify and separate cropland from noncropland. Eighth, accuracy assessment was conducted by generating error matrices for each of the 74 AEZs using 19,171 independent validation-data samples. Ninth, cropland areas were computed for all countries of the world and compared with United Nation’s (UN’s) Food and Agricultural Organization (FAO) and other national statistics. The outcome was a Landsat-derived global cropland-extent product at 30-m resolution (GCEP30), which has an overall accuracy of 91.7 percent. For the cropland class, producer’s accuracy was 83.4 percent, and user’s accuracy was 78.3 percent. GCEP30 calculated (using direct pixel count) the global net-cropland area (GNCA) for the year 2015 as 1.873 billion hectares (~12.6 percent of the Earth’s terrestrial area). The continental cropland distribution as a percentage of GNCA was Asia, 33 percent; Europe, 25.5 percent; Africa, 16.7 percent; North America, 14.4 percent; South America, 8.1 percent; and Australia and Oceania, 2.4 percent. The worldwide cropland areas in GCEP30 for 2015 were higher by 236 to 299 million hectares (Mha) compared to national statistics reported elsewhere for the same year (for example, in Food and Agriculture Organization’s corporate statistical database [FAOSTAT] and in the monthly irrigated and rainfed crop areas [MIRCA] database). The global cropland area reported for 2015 increased by 344 Mha (22.5 percent), compared to the year 2000. During the same period (2000–2015), the world’s population increased by 20 percent. Whereas some of these areal increases are real increases in cropland areas, others are due to the types of data, methods, and approaches used. Using the highest known resolution (compared to previous coarse-resolution global products) enabled this study to capture fragmented croplands. Coarse-resolution data compute areas on the basis of subpixels, which, for a large proportion of certain land use–land cover classes, will show only a certain percentage of the total pixel area as actual area. Subpixel areas can lead to substantial uncertainties in area computation, as determining the exact fraction of cropland areas within a coarse-resolution pixel is resource intensive and subject to errors. Other innovations in GCEP30 include reference-data hubs, machine learning, and cloud computing. Cropland areas in 214 countries, territories, departments, and regions were calculated for the year 2015 using GCEP30, on the basis of UN’s global administrative unit layers (GAUL) boundaries. The 10 leading countries in terms of cropland area (as a percentage of the GNCA) were India (9.6 percent), United States (8.95 percent), China (8.82 percent), Russia (8.32 percent), Brazil (3.42 percent), Ukraine (2.32 percent), Canada (2.29 percent), Argentina (2.05 percent), Indonesia (2 percent), and Nigeria (1.91 percent). Together, these 10 countries occupy 50 percent of the global cropland, and they have 52 percent of the global population. Their combined cropland area increased by 2 percent between 2000 and 2015, compared to the substantial increase in population of 517 million (15.5 percent). Together, India, United States, China, and Russia encompass 36 percent of the total area. In the United States and Canada, from 2000 to 2015, cropland decreased by about 2 percent, whereas their populations increased by 14 and 13 percent, respectively. The additional food requirements in these 10 countries, which are caused by increased populations, as well as increasing nutritional demands, are met by production increases in existing cropland or through virtual food trade, or both. More than 18 countries, territories, departments, or regions had 60 percent or more of their geographic area as cropland: Republic of Moldova, San Marino, and Hungary had more than 80 percent of the country’s area as cropland; Denmark, Ukraine, Ireland, and Bangladesh, 70 to 80 percent; and Uruguay, Netherlands, United Kingdom, Spain, Lithuania, Poland, Gaza Strip, Czechia, Italy, India, and Azerbaijan, 60 to 70 percent. Europe and South Asia can be considered agricultural capitals of the world, on the basis of their percentages of geographic area as cropland. United States, China, and Russia, which all have high cropland areas, are ranked second, third, and fourth in the world; India is ranked first. However, the amount of cropland as a percentage of the country’s geographic area is relatively very low for United States (18.3 percent), China (17.7 percent), and Russia (9.5 percent), whereas it is 60.5 percent for India. Most African and South American countries, territories, departments, or regions have less than 15 percent of their geographic area as cropland. China and India together house 36 percent of the world’s population; however, between 2000 and 2015, the amount of China’s cropland area fell by 18.9 percent, owing to urban expansion and the abandonment of farmlands caused by demographic changes (that is, the movement of population from villages to cities). In contrast, China’s population grew by 10 percent. The amount of India’s cropland increased by 8.5 percent, whereas its population grew by 20 percent. This study showed that, out of the 10 leading cropland countries, Ukraine, Nigeria, Russia, and Indonesia showed an 18 to 31 percent increase in cropland areas, on the basis of GCEP30 by the year 2015, compared to 2000. Nigeria’s cropland area increased by 25 percent, and its population increased by 31 percent in the same period. In these countries, food security is maintained by cropland expansion, productivity increases, and virtual food trade. Nevertheless, this trend of increasing net-cropland area and productivity will likely become difficult to maintain, owing to diminishing arable lands and plateauing of 50 years of continual yield increases, requiring policymakers to explore novel and data-supported approaches to solving future food security issues. The GCEP30 product, which can be browsed at full resolution at www.croplands.org , has been released for public download and use through U.S. Geological Survey (USGS)–National Aeronautics and Space Administration (NASA) Land Processes Distributed Active Archive Center (see https://lpdaac.usgs.gov/news/release-of-gfsad-30-meter-cropland-extent-products/ ).

Professional Paper↗

Using inferential sensors for quality control of Everglades Depth Estimation Network water-level data

The Everglades Depth Estimation Network (EDEN), with over 240 real-time gaging stations, provides hydrologic data for freshwater and tidal areas of the Everglades. These data are used to generate daily water-level and water-depth maps of the Everglades that are used to assess biotic responses to hydrologic change resulting from the U.S. Army Corps of Engineers Comprehensive Everglades Restoration Plan. The generation of EDEN daily water-level and water-depth maps is dependent on high quality real-time data from water-level stations. Real-time data are automatically checked for outliers by assigning minimum and maximum thresholds for each station. Small errors in the real-time data, such as gradual drift of malfunctioning pressure transducers, are more difficult to immediately identify with visual inspection of time-series plots and may only be identified during on-site inspections of the stations. Correcting these small errors in the data often is time consuming and water-level data may not be finalized for several months. To provide daily water-level and water-depth maps on a near real-time basis, EDEN needed an automated process to identify errors in water-level data and to provide estimates for missing or erroneous water-level data. The Automated Data Assurance and Management (ADAM) software uses inferential sensor technology often used in industrial applications. Rather than installing a redundant sensor to measure a process, such as an additional water-level station, inferential sensors, or virtual sensors, were developed for each station that make accurate estimates of the process measured by the hard sensor (water-level gaging station). The inferential sensors in the ADAM software are empirical models that use inputs from one or more proximal stations. The advantage of ADAM is that it provides a redundant signal to the sensor in the field without the environmental threats associated with field conditions at stations (flood or hurricane, for example). In the event that a station does malfunction, ADAM provides an accurate estimate for the period of missing data. The ADAM software also is used in the quality assurance and quality control of the data. The virtual signals are compared to the real-time data, and if the difference between the two signals exceeds a certain tolerance, corrective action to the data and (or) the gaging station can be taken. The ADAM software is automated so that, each morning, the real-time EDEN data are compared to the inferential sensor signals and digital reports highlighting potential erroneous real-time data are generated for appropriate support personnel. The development and application of inferential sensors is easily transferable to other real-time hydrologic monitoring networks.

Florida↗

The search for reliable aqueous solubility (Sw) and octanol-water partition coefficient (Kow) data for hydrophobic organic compounds; DDT and DDE as a case study

The accurate determination of an organic contaminant’s physico-chemical properties is essential for predicting its environmental impact and fate. Approximately 700 publications (1944–2001) were reviewed and all known aqueous solubilities (S w ) and octanol-water partition coefficients (K ow ) for the organochlorine pesticide, DDT, and its persistent metabolite, DDE were compiled and examined. Two problems are evident with the available database: 1) egregious errors in reporting data and references, and 2) poor data quality and/or inadequate documentation of procedures. The published literature (particularly the collative literature such as compilation articles and handbooks) is characterized by a preponderance of unnecessary data duplication. Numerous data and citation errors are also present in the literature. The percentage of original S w and K ow data in compilations has decreased with time, and in the most recent publications (1994–97) it composes only 6–26 percent of the reported data. The variability of original DDT/DDE S w and K ow data spans 2–4 orders of magnitude, and there is little indication that the uncertainty in these properties has declined over the last 5 decades. A criteria-based evaluation of DDT/DDE S w and K ow data sources shows that 95–100 percent of the database literature is of poor or unevaluatable quality. The accuracy and reliability of the vast majority of the data are unknown due to inadequate documentation of the methods of determination used by the authors. [For example, estimates of precision have been reported for only 20 percent of experimental S w data and 10 percent of experimental K ow data.] Computational methods for estimating these parameters have been increasingly substituted for direct or indirect experimental determination despite the fact that the data used for model development and validation may be of unknown reliability. Because of the prevalence of errors, the lack of methodological documentation, and unsatisfactory data quality, the reliability of the DDT/ DDE S w and K ow database is questionable. The nature and extent of the errors documented in this study are probably indicative of a more general problem in the literature of hydrophobic organic compounds. Under these circumstances, estimation of critical environmental parameters on the basis of S w and K ow (for example, bioconcentration factors, equilibrium partition coefficients) is inadvisable because it will likely lead to incorrect environmental risk assessments. The current state of the database indicates that much greater efforts are needed to: 1) halt the proliferation of erroneous data and references, 2) initiate a coordinated program to develop improved methods of property determination, 3) establish and maintain consistent reporting requirements for physico-chemical property data, and 4) create a mechanism for archiving reliable data for widespread use in the scientific/regulatory community.

Water-Resources Investigations Report↗

Analysis of ground-water-quality data of the Upper Colorado River basin, water years 1972-92

As part of the U.S. Geological Survey's National Water-Quality Assessment program, an analysis of the existing ground-water-quality data in the Upper Colorado River Basin study unit is necessary to provide information on the historic water-quality conditions. Analysis of the historical data provides information on the availability or lack of data and water-quality issues. The information gathered from the historical data will be used in the design of ground-water-quality studies in the basin. This report includes an analysis of the ground-water data (well and spring data) available for the Upper Colorado River Basin study unit from water years 1972 to 1992 for major cations and anions, metals and selected trace elements, and nutrients. The data used in the analysis of the ground-water quality in the Upper Colorado River Basin study unit were predominantly from the U.S. Geological Survey National Water Information System and the Colorado Department of Public Health and Environment data bases. A total of 212 sites representing alluvial aquifers and 187 sites representing bedrock aquifers were used in the analysis. The available data were not ideal for conducting a comprehensive basinwide water-quality assessment because of lack of sufficient geographical coverage. Evaluation of the ground-water data in the Upper Colorado River Basin study unit was based on the regional environmental setting, which describes the natural and human factors that can affect the water quality. In this report, the ground-water-quality information is evaluated on the basis of aquifers or potential aquifers (alluvial, Green River Formation, Mesaverde Group, Mancos Shale, Dakota Sandstone, Morrison Formation, Entrada Sandstone, Leadville Limestone, and Precambrian) and land-use classifications for alluvial aquifers. Most of the ground-water-quality data in the study unit were for major cations and anions and dissolved-solids concentrations. The aquifer with the highest median concentrations of major ions was the Mancos Shale. The U.S. Environmental Protection Agency secondary maximum contaminant level of 500 milligrams per liter for dissolved solids in drinking water was exceeded in about 75 percent of the samples from the Mancos Shale aquifer. The guideline by the Food and Agriculture Organization of the United States for irrigation water of 2,000 milligrams per liter was also exceeded by the median concentration from the Mancos Shale aquifer. For sulfate, the U.S. Environmental Protection Agency proposed maximum contaminant level of 500 milligrams per liter for drinking water was exceeded by the median concentration for the Mancos Shale aquifer. A total of 66 percent of the sites in the Mancos Shale aquifer exceeded the proposed maximum contaminant level. Metal and selected trace-element data were available for some sites, but most of these data also were below the detection limit. The median concentrations for iron for the selected aquifers and land-use classifications were below the U.S. Environmental Protection Agency secondary maximum contaminant level of 300 micrograms per liter in drinking water. Median concentration of manganese for the Mancos Shale exceeded the U.S. Environmental Protection Agency secondary maximum contaminant level of 50 micrograms per liter in drinking water. The highest selenium concentrations were in the alluvial aquifer and were associated with rangeland. However, about 22 percent of the selenium values from the Mancos Shale exceeded the U.S. Environmental Protection Agency maximum contaminant level of 50 micrograms per liter in drinking water. Few nutrient data were available for the study unit. The only nutrient species presented in this report were nitrate-plus-nitrite as nitrogen and orthophosphate. Median concentrations for nitrate-plus-nitrite as nitrogen were below the U.S. Environmental Protection Agency maximum contaminant level of 10 milligrams per liter in drinking water except for 0.02 percent of the sites in the alluvial aquifer and 0.03 percent of the sites in the Mancos Shale. Concentrations of orthophosphate did not vary significantly among aquifers or land-use classifications. Historic water-quality data from wells and springs helped to characterize the regional distribution of ground-water quality information in the Upper Colorado River Basin study unit. The historical ground-water data summarized in this report will be used in the design of a ground-water-quality network. Because ground-water-quality issues in the study unit are related to high dissolved solids, sulfate, selenium, and nutrients, this report discusses some of the important findings related to these issues.

Colorado↗

On the documentation, independence, and stability of widely used seismological data products

Earthquake scientists have traditionally relied on relatively small data sets recorded on small numbers of instruments. With advances in both instrumentation and computational resources, the big-data era, including an established norm of open data-sharing, allows seismologists to explore important issues using data volumes that would have been unimaginable in earlier decades. Alongside with these developments, the community has moved towards routine production of interpreted data products such as seismic moment tensor catalogs that have provided an additional boon to earthquake science. As these products have become increasingly familiar and useful, it is important to bear in mind that they are not data, but rather interpreted data products. As such, they differ from data in ways that can be important, but not always appreciated. Important - and sometimes surprising - issues can arise if methodology is not fully described, data from multiple sources are included, or data products are not versioned (time-stamped). The line between data and data products is sometimes blurred, leading to an underappreciation of issues that affect data products. This note illustrates examples from two widely used data products: moment tensor catalogs and Did You Feel It? (DYFI) macroseismic intensity values. These examples show that increasing a data product’s documentation, independence, and stability can make it even more useful. To ensure the reproducibility of studies using data products, time-stamped products should be preserved, for example as electronic supplements to published papers, or, ideally, a more permanent repository.

Frontiers in Earth Science↗

Monitoring landscape dynamics in central U.S. grasslands with harmonized Landsat-8 and Sentinel-2 time series data

Remotely monitoring changes in central U.S. grasslands is challenging because these landscapes tend to respond quickly to disturbances and changes in weather. Such dynamic responses influence nutrient cycling, greenhouse gas contributions, habitat availability for wildlife, and other ecosystem processes and services. Traditionally, coarse-resolution satellite data acquired at daily intervals have been used for monitoring. Recently, the harmonized Landsat-8 and Sentinel-2 (HLS) data increased the temporal frequency of the data. Here we investigated if the increased data frequency provided adequate observations to characterize highly dynamic grassland processes. We evaluated HLS data available for 2016 to (1) determine if data from Sentinel-2 contributed to an improvement in characterizing landscape processes over Landsat-8 data alone, and (2) quantify how observation frequency impacted results. Specifically, we investigated into estimating annual vegetation phenology, detecting burn scars from fire, and modeling within-season wetland hydroperiod and growth of aquatic vegetation. We observed increased sensitivity to the start of the growing season (SOST) with the HLS data. Our estimates of the grassland SOST compared well with ground estimates collected at a phenological camera site. We used the Continuous Change Detection and Classification (CCDC) algorithm to assess if the HLS data improved our detection of burn scars following grassland fires and found that detection was considerably influenced by the seasonal timing of the fires. The grassland burned in early spring recovered too quickly to be detected as change events by CCDC; instead, the spectral characteristics following these fires were incorporated as part of the ongoing time-series models. In contrast, the spectral effects from late-season fires were detected both by Landsat-8 data and HLS data. For wetland-rich areas, we used a modified version of the CCDC algorithm to track within-season dynamics of water and aquatic vegetation. The addition of Sentinel-2 data provided the potential to build full time series models to better distinguish different wetland types, suggesting that the temporal density of data was sufficient for within-season characterization of wetland dynamics. Although the different data frequency, in both the spatial and temporal dimensions, could cause inconsistent model estimation or sensitivity sometimes; overall, the temporal frequency of the HLS data improved our ability to track within-season grassland dynamics and improved results for areas prone to cloud contamination. The results suggest a greater frequency of observations, such as from harmonizing data across all comparable Landsat and Sentinel sensors, is still needed. For our study areas, at least a 3-day revisit interval during the early growing season (weeks 14–17) is required to provide a >50% probability of obtaining weekly clear observations.

Minnesota↗

Considerations for blending data from various sensors

A project is being proposed at the EROS Data Center to blend the information from sensors aboard various satellites. The problems of, and considerations for, blending data from several satellite-borne sensors are discussed. System descriptions of the sensors aboard the HCMM, TIROS-N, GOES-D, Landsat 3, Landsat D, Seasat, SPOT, Stereosat, and NOSS satellites, and the quantity, quality, image dimensions, and availability of these data are summaries to define attributes of a multi-sensor satellite data base. Unique configurations of equipment, storage, media, and specialized hardware to meet the data system requirement are described as well as archival media and improved sensors that will be on-line within the next 5 years. Definitions and rigor required for blending various sensor data are given. Problems of merging data from the same sensor (intrasensor comparison) and from different sensors (intersensor comparison), the characteristics and advantages of cross-calibration of data, and integration of data into a product matrix field are addressed. Data processing considerations as affected by formation, resolution, and problems of merging large data sets, and organization of data bases for blending data are presented. Examples utilizing GOES and Landsat data are presented to demonstrate techniques of data blending, and recommendations for future implementation of a set of standard scenes and their characteristics necessary for optimal data blending are discussed.

Sixth Annual Pecora Symposium and Exposition↗

Alaska Geochemical Database, Version 2.0 (AGDB2)–Including “best value” data compilations for rock, sediment, soil, mineral, and concentrate sample mediaI

The Alaska Geochemical Database Version 2.0 (AGDB2) contains new geochemical data compilations in which each geologic material sample has one “best value” determination for each analyzed species, greatly improving speed and efficiency of use. Like the Alaska Geochemical Database (AGDB, http://pubs.usgs.gov/ds/637/) before it, the AGDB2 was created and designed to compile and integrate geochemical data from Alaska in order to facilitate geologic mapping, petrologic studies, mineral resource assessments, definition of geochemical baseline values and statistics, environmental impact assessments, and studies in medical geology. This relational database, created from the Alaska Geochemical Database (AGDB) that was released in 2011, serves as a data archive in support of present and future Alaskan geologic and geochemical projects, and contains data tables in several different formats describing historical and new quantitative and qualitative geochemical analyses. The analytical results were determined by 85 laboratory and field analytical methods on 264,095 rock, sediment, soil, mineral and heavy-mineral concentrate samples. Most samples were collected by U.S. Geological Survey personnel and analyzed in U.S. Geological Survey laboratories or, under contracts, in commercial analytical laboratories. These data represent analyses of samples collected as part of various U.S. Geological Survey programs and projects from 1962 through 2009. In addition, mineralogical data from 18,138 nonmagnetic heavy-mineral concentrate samples are included in this database. The AGDB2 includes historical geochemical data originally archived in the U.S. Geological Survey Rock Analysis Storage System (RASS) database, used from the mid-1960s through the late 1980s and the U.S. Geological Survey PLUTO database used from the mid-1970s through the mid-1990s. All of these data are currently maintained in the National Geochemical Database (NGDB). Retrievals from the NGDB were used to generate most of the AGDB data set. These data were checked for accuracy regarding sample location, sample media type, and analytical methods used. This arduous process of reviewing, verifying and, where necessary, editing all U.S. Geological Survey geochemical data resulted in a significantly improved Alaska geochemical dataset. USGS data that were not previously in the NGDB because the data predate the earliest U.S. Geological Survey geochemical databases, or were once excluded for programmatic reasons, are included here in the AGDB2 and will be added to the NGDB. The AGDB2 data provided here are the most accurate and complete to date, and should be useful for a wide variety of geochemical studies. The AGDB2 data provided in the linked database may be updated or changed periodically.

Alaska↗

Water resources data, New York, water year 1996; Volume 1. Eastern New York; excluding Long Island

Introduction Water-resources data for the 1996 water year for New York consist of records of stage, discharge, and water quality of streams; stage, contents, and water quality of lakes and reservoirs; ground-water levels; and precipitation quality. This volume contains records for water discharge at 122 gaging stations; stage only at 7 gaging stations; stage and contents at 4 gaging stations, and 18 other lakes and reservoirs; water quality at 28 gaging stations and 1 precipitation-quality station; and water levels at 3 observation wells. Also included are data for 33 crest-stage partial-record stations. Additional water data were collected at various sites not involved in the systematic data-collection program, and are published as miscellaneous measurements and analyses in this volume. These data together with the data in Volumes 2 and 3 represent that part of the National Water Data System operated by the U.S. Geological Survey in cooperation with State, Municipal, and Federal agencies in New York. Records of discharge and stage of streams, and contents and stage of lakes and reservoirs, were first published in a series of U.S. Geological Survey water-supply papers entitled, “Surface Water Supply of the United States.” Through September 30, 1960, these water-supply papers were in an annual series and then in a 5-year series for 1961-65 and 1966-70. Records of water quality, water temperatures, and suspended sediment were published from 1941 to 1970 in an annual series of water-supply papers entitled “Quality of Surface Waters of the United States.” Records of ground-water levels were published from 1935 to 1974 in a series of water-supply papers entitled “Ground-Water Levels in the United States.” Water-supply papers may be consulted in the libraries of the principal cities and universities in the United States or may be purchased from the U.S. Geological Survey, Branch of Distribution, 604 South Pickett Street, Alexandria, VA 22304. Since the 1961 water year, streamflow data and since the 1964 water year, water-quality data have been released by the Geological Survey in annual reports on a State-boundary basis. These reports provided rapid release of water data in each state shortly after the end of the water year. Through 1970 the data were also released in the water-supply paper series mentioned above. Streamflow and water-quality data beginning with the 1971 water year, and ground-water data beginning with the 1975 water year are published only in reports on a State-boundary basis. Beginning with the 1975 water year, these Survey reports carry an identification number consisting of the two-letter State abbreviation, the last two digits of the water year, and the volume number. For example, this volume is identified as “U.S. Geological Survey Water-Data Report NY-96-1.” Water-data reports are for sale in paper copy or in microfiche by the National Technical Information Service, U.S. Department of Commerce, Springfield, VA 22161. Additional information, including current prices for ordering specific reports, may be obtained from the District Office at the address given on the back of the title page or by telephone (518) 285-5600.

New York↗

Estimating the drivers of species distributions with opportunistic data using mediation analysis

Ecological occupancy modeling has historically relied on high-quality, low-quantity designed-survey data for estimation and prediction. In recent years, there has been a large increase in the amount of high-quantity, unknown-quality opportunistic data. This has motivated research on how best to combine these two data sources in order to optimize inference. Existing methods can be infeasible for large datasets or require opportunistic data to be located where designed-survey data exist. These methods map species occupancies, motivating a need to properly evaluate covariate effects (e.g., land cover proportion) on their distributions. We describe a spatial estimation method for supplementarily including additional opportunistic data using mediation analysis concepts. The opportunistic data mediate the effect of the covariate on the designed-survey data response, decomposing it into a direct and indirect effect. A component of the indirect effect can then be quickly estimated via regressing the mediator on the covariate, while the other components are estimated through a spatial occupancy model. The regression step allows for use of large quantities of opportunistic data that can be collected in locations with no designed-survey data available. Simulation results suggest that the mediated method produces an improvement in relative MSE when the data are of reasonable quality. However, when the simulated opportunistic data are poorly correlated with the true spatial process, the standard, unmediated method is still preferable. A spatiotemporal extension of the method is also developed for analyzing the effect of deciduous forest land cover on red-eyed vireo distribution in the southeastern United States and find that including the opportunistic data do not lead to a substantial improvement. Opportunistic data quality remains an important consideration when employing this method, as with other data integration methods.

eastern United States↗