USGS ScienceSearch

SEARCH · USGS Science

Results for “Data Series”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

664 records · Page 7Linked to original sources

From critical minerals to food security, the benefits of data collaboration

The volume of data in the public geoscience sphere is rapidly and continually expanding. At Geoscience Australia (GA) we saw an over 500% increase in data points within our relational databases between 2018 and 2024, over the life of the Exploring for the Future (EFTF) program. With the Resourcing Australia’s Prosperity initiative, a continued increase in data quantity will be seen for the next 10 to 35 years. At the same time, a broadening audience for geoscience data is increasing the desire to enhance the diversity of delivery streams. This ranges from data-dense highly technical outputs for geoscience specialists to curated interpretive products for people who are non-geoscientists. Development of these curated outputs has contributed to our awareness of the need for data to be collected and compiled in a way that ensures its reuse, with a focus on quality metadata and data provenance.

Conference Paper

A model uncertainty quantification protocol for evaluating the value of observation data

The history-matching approach to parameter estimation with models enables a powerful offshoot analysis of data worth—using the uncertainty of a model forecast as a metric for the worth of data. Adding observation data will either have no impact on forecast uncertainty or will reduce it. Removing existing data will either have no impact on forecast uncertainty or will increase it. The history-matching framework makes it possible to perform this quantitative analysis leveraging the connections among observations, model parameters, and model forecasts. We show this behavior on a specific groundwater flow model of the Mississippi Alluvial Plain and show where the analysis can be informative for considering the potential design of an observation network based on existing or potential observations.

Scientific Investigations Report

Overcoming the data limitations in landslide susceptibility modelling

Data-driven models widely used for assessing landslide susceptibility are severely limited by the landslide and environmental data needed to create them. They rely on inventories of past landslide locations, which are difficult to collect and often nonrepresentative. Furthermore, susceptibility maps are most needed in regions without the means to assemble an inventory. To overcome these challenges, we develop a method for assessing shallow landslide susceptibility based on a probabilistic morphometric analysis of the landscape’s topography, rather than the characteristics of landslides. The model assumes that hillslopes with higher relief and gradient compared to the surrounding landscape are more prone to landslides. We demonstrate the superior performance of this approach over contrasting data-driven models across the northwestern United States. As our morphometric model only requires elevation data, it overcomes the major limitations of data-driven models and facilitates the creation of effective susceptibility models in areas where it was previously unfeasible.

Oregon, Washington

Effect of mineral deposit data on predictions from the three-part approach to quantitative mineral resource assessment—A study of 16 previous U.S. Geological Survey assessments

The three-part approach to quantitative mineral resource assessment requires information about the properties of undiscovered mineral deposits in an assessment area. These properties are unknown, so the properties of discovered mineral deposits of the same mineral deposit type are used instead. In the three-part approach, these discovered mineral deposits come from around the world, and their properties constitute the pooled data for that mineral deposit type. Alternatively, these discovered mineral deposits could come from the assessment area, and their properties constitute the tract data for that mineral deposit type. Tract data may be more representative of the undiscovered mineral deposits in the assessment area than the pooled data. The goal of this study was to determine whether resource predictions using pooled data are equivalent to resource predictions using tract data. To this end, 16 previous U.S Geological Survey assessments were studied. For each assessment, resources were predicted for one undiscovered mineral deposit in the assessment area. One set of predictions used pooled data, and another used tract data. The two sets of predictions were compared with an equivalence test, using the six assessment statistics that are commonly reported for mineral resource assessments. Practical equivalence is the condition that two corresponding assessment statistics are within a factor of 1.5 of one another. For each of 2 assessments, all 6 assessment statistics were practically equivalent. For both assessments, the assessment statistics from the pooled data, relative to the corresponding assessment statistics from the tract data, ranged from 1.30 times smaller to 1.03 times larger. For each of 14 assessments, 1 or more of the 6 assessment statistics were not practically equivalent. The assessment statistics from the pooled data, relative to the corresponding assessment statistics from the tract data, ranged from 26.6 times smaller to 5.53 times larger. The use of pooled data has been a standard procedure in the three-part approach since at least 1986. The 16 assessments in this study are not a representative sample of those prior assessments that used pooled data. So, it is inappropriate to use the study results to infer whether pooled data affected the resource predictions for those prior assessments.

Scientific Investigations Report

Detecting earthquakes in noisy real-time GNSS data with deep learning for improved PGD magnitude estimation

To disseminate accurate and useful warnings, earthquake early warning (EEW) systems must quickly determine the size and location of an earthquake to estimate expected shaking. Traditional seismic‐based algorithms tend to underestimate the true magnitudes of large earthquakes, a phenomenon known as magnitude saturation. This limitation motivated the recent inclusion of Global Navigation Satellite Systems (GNSS) data into the U.S. Geological Survey’s ShakeAlert EEW system with the Geodetic First Approximation of Size and Time (GFAST) algorithm because GNSS data do not saturate with large ground motions. However, the noise levels of GNSS data are very high compared with traditional seismic data, which obscures P ‐wave arrivals and can result in less accurate magnitude estimations if displacement amplitudes are low, such as for lower magnitude earthquakes or large source–station distances. In this study, we develop a deep‐learning model that detects earthquakes in GNSS data and use the Ridgecrest, California, earthquake sequence as a case study to demonstrate how the model could act as a filter to reduce the amount of low‐quality data that enters an algorithm like GFAST. To preserve our limited real earthquake data for model inference, we generated a training dataset composed of >700,000 synthetic displacement waveforms. We combined the synthetic waveforms with real‐time GNSS noise to produce realistically noisy training waveforms and then tested our model on additional synthetic data and performed inference using the real data that were held back. We discuss the performance of our trained model on both the unseen synthetic data and real inference data. Our model can be used to selectively filter only high‐quality data where an earthquake signal is observed for input into an algorithm like GFAST (outperforming a simple signal‐to‐noise ratio–based filter) to reduce the error in GFAST’s real‐time earthquake magnitude estimations.

California

Relationship of basin structure and bedrock lithology to faulting in the 2019 Ridgecrest earthquake region, California, from gravity and aeromagnetic data

We investigate patterns of cumulative offsets on the faults that ruptured in 2019 and along the Garlock Fault in the Ridgecrest region, California using recently published gravity and aeromagnetic data. We also examine the relationship of basin structure and bedrock structure to the 2019 M7.1 Ridgecrest earthquake ruptures (Fig. 1A), which were primarily along a dextral northwest-striking fault system, and along a sinistral northeast-striking fault, which ruptured hours earlier with a M6.4 event.

California

Toward a new framework to evaluate process-based model configurations and quantify data worth prior to calibration

Model criticism, discrimination, and selection methods often rely on calibrated model outputs. Because calibration can be computationally expensive, model criticism can first be undertaken by assessing model outputs obtained from limited prior parameter ensembles. However, such prior-based methods are often heuristic and do not formalize the notion of balancing model consistency with data and model complexity (i.e., model adequacy). We present a new framework to discriminate among candidate models prior to calibration that formalizes prior-to-calibration model adequacy into a metric to implicitly balance prior model output data coverage with model complexity represented by prior output (co)variance. The prior model adequacy metric “Mahalanobis distance deviation” quantifies the deviation of (a) the set of squared Mahalanobis distances of data from a prior model output distribution from (b) the set of squared Mahalanobis distances of data from their own distribution. A new data worth metric “discernment value” is also presented which quantifies the value of data for screening less-adequate models prior to calibration. Discernment value is calculated from the change in variance of a weighted average of prior model outputs from all candidate models due to less-adequate model outputs receiving lower weight. The framework is demonstrated using a one-dimensional groundwater flow model with eight possible configurations. A synthetic data network is used to test the framework. Results show the framework identifies the candidate models most similar to the true model used to create the synthetic data. Discernment values show variation in the value of different data types and locations for screening less-adequate models.

Water Resources Research

Distinguishing natural sources from anthropogenic events in seismic data

As seismic data are increasingly used to investigate a diverse range of subsurface phenomena beyond regular fast-rupturing earthquakes (Peng and Gomberg, 2010; Beroza and Ide, 2011), it is important to acknowledge that human-generated ground vibrations may be mistaken for naturally generated subsurface processes (Larose et al., 2015; Li et al., 2018). Correct discrimination of natural processes from anthropogenic noise is especially pressing given the trend in seismic detection research toward automated algorithms and machine learning methods (Yoon et al., 2015; Kong et al., 2019;Mousavi and Beroza, 2022) and the growth in seismic data collection in new environments such as urban and industry settings (e.g., Díaz et al.,2017).

Seismological Research Letters

The Sedimentary Geochemistry and Paleoenvironments Project Phase 2 data release: An open data resource for the study of Earth's environmental history

Geochemical data from sedimentary rocks are the primary source of information regarding Earth's surface evolution through time, including its air and water envelopes and interactions with life and deep Earth processes. The Sedimentary Geochemistry and Paleoenvironments Project (SGP) is a scientific consortium centered around open data and community-driven development of cyberinfrastructure tools and resources for sedimentary geochemistry and Earth history. Here we describe the SGP Phase 2 data release, which focused on incorporating Paleoproterozoic and Mesoproterozoic (2500–1000 million years ago) data and better accommodating carbonate data. This data release was built through the involvement of >200 researchers worldwide in academia, government, and industry, and provides the largest available public data resource for our user community in the academic fields of geochemistry, sedimentology, tectonics, paleontology, Earth history, and paleoclimate, as well as the petroleum and minerals industries. The dataset now encompasses 126,006 samples and 4,132,371 geochemical analyses. In addition to direct entry by SGP Team Members, we have ingested and incorporated datasets from the Geoscience Australia OZCHEM database, the Alberta Geological Survey, and the Deep-Time Marine Sedimentary Element Database (DM-SED) compilation. This paper details sampling in the Phase 2 dataset with respect to age, geography, lithology, and other geological characteristics, documents access via our search website and API, discusses possible issues and/or biases in the dataset that could impact analyses, describes plans for governance and stewardship of data from Indigenous lands, and serves as the citable reference paper for the data release.

Chemical Geology

Assessment of density pattern retention of generalized data for 1:100,000-scale United States topographic maps

Cartographic generalization reduces the complexity of geographic data to produce legible, smaller-scale displays that retain essential information and logical geographic patterns. Generalization is a vital process in topographic map production. An important challenge in this process is managing and evaluating consistency across scale in the density and spatial distribution of map features such as buildings, roads, streams, water bodies, and elevation contours. Density patterns in these features reflect underlying physiographic conditions, which include factors such as bedrock geology, tectonics, climate, and landforms. Assessments of an acceptable level of change in feature density patterns are critical to ensuring the readability, usability, and accuracy of generalized maps and data. Preserving realistic density patterns across mapping scales also supports sustainable development goals in cartography, by helping to prioritize and communicate the relative reliability of geospatial data at specific scales.

Conference Paper

Bayesian mapping of regionally grouped, sparse, univariate earth science data

Some earth science data are naturally grouped by region, and it is often desirable to map these data by region. However, if there are only a few samples within each region, then the map should be smoothed in an appropriate way to mitigate the problems that arise from having only a few samples. A smoothing algorithm based on a Bayesian hierarchical model is developed and presented in this report. This algorithm has several features that make it especially suitable for mapping earth science data: it can account for measurements that are censored, it can process multiple datasets with different measurement errors and different censoring thresholds, and it can calculate the uncertainty in any statistic that is mapped. The algorithm is demonstrated by mapping gold concentrations that are measured in streambed sediments in the Taylor Mountains quadrangle in southwestern Alaska.

Alaska

New developments at the Center for Engineering Strong-Motion Data (CESMD)

The Center for Engineering Strong-Motion Data (CESMD), an internationally utilized joint center of the U.S. Geological Survey (USGS) and the California Geological Survey (CGS), provides a single access point for earthquake strong-motion records and station metadata from the CGS California Strong-Motion Instrumentation Program (CSMIP), the USGS National Strong-Motion Project (NSMP), the USGS Advanced National Seismic System, and other affiliates. The CESMD has been continuously improving its webtools to facilitate the access of strong-motion data and metadata for use in post-earthquake response and for scientific and engineering research applications. The Center provides raw and processed strong-motion data via the Engineering Data Center (EDC) and the Virtual Data Center (VDC) web portals. This paper focuses on the strong-motion products provided by the EDC where more than 48,000 records with peak ground accelerations greater than 0.1% g from over 2400 earthquakes are currently hosted. and on the ongoing efforts to develop data access tools and applications. The new developments and ongoing efforts in the EDC include: 1) enhancements to the CESMD webservices to facilitate access to station metadata, earthquake information, and strong motion records 2) new features to the interactive map interface, improving the visualization and access to earthquake, station, and record information, 3) efforts to develop a new web application tool for data format conversion from a number of data formats, 4) efforts to unify varying waveform data formats into a consistent format, 5) ongoing efforts to compile seismic station site geology, measured or inferred Vs30 values, shear-wave profiles, NEHRP site class, and available structural instrument deployment schematics, and 6) a special studies pages for research topic-specific ground motion datasets that offer uniform processing of records from a variety of sources.

Conference Paper

U.S. Geological Survey geomagnetic variometer data: Capitalizing on seismic infrastructure

The U.S. Geological Survey’s Geomagnetism Program is collaborating with the Earthquake Hazards Program and Global Seismographic Network Program to densify magnetic field observations. This collaboration focuses on the installation of magnetometers, or magnetic variometers, at existing seismic stations. Along with improving the density of space weather observations for hazard monitoring, these data can be used to correct colocated magnetic field induced noise in seismic data. Such corrections are especially useful during time periods of large magnetic storms where the magnetic field‐induced instrument noise can be of similar amplitude to earthquake ground‐motion records.

contiguous United States

Constraining mean landslide occurrence rates for non-temporal landslide inventories using high-resolution elevation data

Constraining landslide occurrence rates can help to generate landslide hazard models that predict the spatial and temporal occurrence of landslides. However, most landslide inventories do not include any temporal data due to the difficulties of dating landslide deposits. Here we introduce a method for estimating the mean landslide occurrence rate of deep-seated rotational and translational slides derived solely from high-resolution (≤3 m) elevation data and globally available estimates of the diffusion coefficient for sediment flux. The method applies a linear diffusion model to the roughest landslide deposits until they reach a representative non-landslide roughness distribution. This estimates the time for a landslide deposit to be unrecognizable in high-resolution digital elevation data, which we term the mean lifetime of the landslide. Using the mean lifetime and number of landslides within an area of interest, we can estimate the mean occurrence rate of landslides over that domain. We validate this approach using a comprehensive temporal inventory of landslides in western Oregon created using age-roughness curves that are calibrated with high-resolution elevation data and radiocarbon data. We find good agreement between our diffusion method and the existing age-roughness-derived estimates, producing mean lifetimes of 4500 and 5200 years (4% difference), respectively. Hazard maps produced using the two methodologies generally agree, with the maximum differences in landslide probability reaching 0.1. Due to the relative abundance of high-resolution elevation data compared with age-dated landslides, our method could help constrain landslide occurrence rates in areas previously considered unfeasible.

Oregon

Updating regional‐scale geospatial liquefaction models with locally available geotechnical data

We present a method to update the geospatial liquefaction model used by the U.S. Geological Survey’s near‐real‐time ground failure product with subsurface geotechnical data. The geospatial model estimates liquefaction probability from peak ground velocity (via ShakeMap) and geospatial susceptibility proxies. In many regions, additional information relevant to constraining liquefaction likelihood is also available, including surface geology maps and subsurface geotechnical measurements. There is currently no mechanism to use these data in the ground failure product liquefaction model, even though these data could provide more precise constraints on spatial variations in the lithologic character of the soil (surface geology) and direct measurements of the subsurface mechanical properties that affect liquefaction occurrence and severity (geotechnical measurements). In this study, we develop a method to integrate these data with the geospatial model and assess how these data can improve regional‐scale predictions. We develop a Bayesian updating framework and apply it to the 1989 magnitude 6.9 Loma Prieta, California, earthquake, for which mapped observations are available to evaluate performance. We constrain the Bayesian framework with 373 Northern California cone penetration tests and liquefaction susceptibility classes based on the mapped surface geology. This Bayesian model incorporates geotechnical information into the geospatial model and more accurately predicts liquefaction occurrences than the geospatial model, while sacrificing less accuracy in terms of predicting the absence of liquefaction than the geotechnical model. In future applications, this approach could be adapted to update other geospatial models using locally available subsurface data.

California

Selected water-quality data from the Cedar River and Cedar Rapids well fields, Cedar Rapids, Iowa, 2017–22

The Cedar River alluvial aquifer is the source of drinking water in Cedar Rapids, Iowa. Production wells are completed in the alluvial aquifer approximately 40 to 80 feet below land surface. The City of Cedar Rapids and the U.S. Geological Survey have studied the groundwater-flow system and water quality of the aquifer in the vicinity of Cedar Rapids since 1992. Results of these studies documented hydrologic conditions, water quality, and geochemistry of the alluvial aquifer and interactions with the Cedar River. Water-quality samples were collected for studies involving well field monitoring, trends, source-water protection, groundwater geochemistry, surface-water–groundwater interaction, and pesticides in groundwater and surface water. Water quality was analyzed for dissolved major ions (boron, bromide, calcium, chloride, fluoride, iron, magnesium, manganese, potassium, silica, sodium, sulfate, and total dissolved solids), dissolved nutrients (ammonia as nitrogen, ammonia plus organic nitrogen as nitrogen, nitrite plus nitrate as nitrogen, nitrite as nitrogen, orthophosphate as phosphorus, and phosphorus), dissolved organic carbon, and selected pesticides. Physical characteristics (alkalinity, dissolved oxygen, pH, specific conductance, and water temperature) were measured on site and recorded for each water sample collected. This report presents the results of routine water-quality data-collection activities from October 2017 through September 2022. Methods of data collection, quality assurance, water-quality analyses, and statistical procedures are presented. Data include the results of water-quality analyses from quarterly sampling from monitoring wells, production wells, two water treatment plants, and the Cedar River at Blairs Ferry Road at Palo, Iowa, streamgage (U.S. Geological Survey station number 05464420), as well as monthly nutrient sampling from the Cedar River and Morgan Creek near Covington, Iowa, streamgage (U.S. Geological Survey station number 05464475).

Iowa

Assessing the state of hydrologic science in the Upper Klamath Basin—A comprehensive review of data, tools, and models

Water demand in the Upper Klamath Basin (UKB) from various stakeholders and ecological needs often outstrips available supply, leading to persistent management challenges. This study reviews the state of hydrologic science within the UKB as of 2025—specifically, the tools, data, and models available for assessing five key components of the water system: (1) surface water; (2) precipitation; (3) evapotranspiration; (4) groundwater; and (5) water use. The UKB water supply is critical for Native American communities, regional agriculture, and federally listed fishes and faces challenges from competing needs, climate variability, and operational/regulatory requirements. We assess existing datasets, regional and national models, and historical studies to understand the available resources and identify gaps that may hinder integrated water assessments and management. Our findings indicate areas where improvements in data collection and model precision could improve the accuracy of water-availability forecasts and support water-management practices. This review can inform near-term forecasting, assist in optimizing water-resource data collection and management strategies, and support regional water-availability assessments of the basin.

Oregon

Data gap analysis for estimation of agricultural return flows in the Upper Gunnison River Basin, Colorado

The Gunnison River and many tributaries in the Upper Gunnison River Basin provide water to irrigate agricultural crops. The application of irrigation water can recharge some aquifers locally by water percolating below the root zone and eventually flowing back to the stream or river through the subsurface. Diverting surface water for irrigation reduces streamflow during the irrigation season but can provide temporary storage of water and supplement streamflow after the snowmelt runoff season. Understanding the timing and quantity of agricultural return flows could help resource managers make informed decisions and adapt to potential changes in water management and availability that could affect irrigation practices. In 2024, the U.S. Geological Survey, in cooperation with the Upper Gunnison River Water Conservancy District, began a study to characterize agricultural return flows in the Upper Gunnison River Basin by using endmember mixing analysis and developing a groundwater model. Both approaches require data from multiple sources, but data gaps exist in the East River study reach and other reaches of interest (Ohio Creek, Tomichi Creek, and Cochetopa Creek). The East River Basin, which is the initial focus of the study, has fewer data gaps than the other basins. Data gaps could be addressed by installing additional surface water and groundwater monitoring sites, making regular streamflow measurements on tributaries, and completing tests to characterize local aquifer properties.

Colorado