USGS Science⌕ Search

SEARCH · USGS Science

Results for “Spatial Statistics”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Lithology-derived structure classification from the joint interpretation of magnetotelluric and seismic models

Magnetotelluric and seismic methods provide complementary information about the resistivity and velocity structure of the subsurface on similar scales and resolutions. No global relation, however, exists between these parameters, and correlations are often valid for only a limited target area. Independently derived inverse models from these methods can be combined using a classification approach to map geologic structure. The method employed is based solely on the statistical correlation of physical properties in a joint parameter space and is independent of theoretical or empirical relations linking electrical and seismic parameters. Regions of high correlation (classes) between resistivity and velocity can in turn be mapped back and re-examined in depth section. The spatial distribution of these classes, and the boundaries between them, provide structural information not evident in the individual models. This method is applied to a 10 km long profile crossing the Dead Sea Transform in Jordan. Several prominent classes are identified with specific lithologies in accordance with local geology. An abrupt change in lithology across the fault, together with vertical uplift of the basement suggest the fault is sub-vertical within the upper crust. ?? 2007 The Authors Journal compilation ?? 2007 RAS.

Geophysical Journal International↗

Northwest Forest Plan — The first 25 years (1994–2018): Watershed condition status and trends

This report describes status and trends in watershed condition across the Northwest Forest Plan (NWFP) area over the first 25 years since its inception in 1994. The program charged with this task is the Aquatic and Riparian Effectiveness Monitoring Program (AREMP), which has assembled information from field data collection, spatial datasets, and a host of landscape models to evaluate the status and trends in aquatic resources in streams and watersheds. Field data included hydrologic measurements (stream wetted widths and temperatures), geomorphic responses (instream wood and sediment), and biological responses (macroinvertebrates and aquatic organism passage). Novel statistical models were used to estimate trends in these measured responses. A suite of complementary modeled results was also employed to describe hydrometeorological drivers (e.g., drought indices and stream discharge), forest cover (upslope and riparian vegetation), and geomorphic conditions (e.g., road-related estimates of chronic and shallow landslide sediment delivery risk). Collectively, information on these responses allowed us to rigorously evaluate instream responses and hypothesize watershed drivers of those responses across the NWFP area and over time. The majority of responses we observed indicated widespread and incremental improvements from active management of forests, forest roads, and road-stream crossings as envisioned by the aquatic conservation strategy of the NWFP. Additionally, many of the responses we observed were consistent with those expected under the influences of changing climates in the Pacific Northwest. Ultimately, the long-term, broad-scale information provided by AREMP is a critical foundation for evaluating the effectiveness of federal land management and the effects of changing climates on water resources that sustain the Pacific Northwest’s human and natural landscapes.

California, Oregon, Washington↗

Large scale wildlife monitoring studies: Statistical methods for design and analysis

Techniques for estimation of absolute abundance of wildlife populations have received a lot of attention in recent years. The statistical research has been focused on intensive small-scale studies. Recently, however, wildlife biologists have desired to study populations of animals at very large scales for monitoring purposes. Population indices are widely used in these extensive monitoring programs because they are inexpensive compared to estimates of absolute abundance. A crucial underlying assumption is that the population index ( C ) is directly proportional to the population density ( D ). The proportionality constant, β , is simply the probability of 'detection' for animals in the survey. As spatial and temporal comparisons of indices are crucial, it is necessary to also assume that the probability of detection is constant over space and time. Biologists intuitively recognize this when they design rigid protocols for the studies where the indices are collected. Unfortunately, however in many field studios the assumption is clearly invalid. We believe that the estimation of detection probability should be built into the monitoring design through a double sampling approach. A large sample of points provides an abundance index, and a smaller sub-sample of the same points is used to estimate detection probability. There is an important need for statistical research on the design and analysis of these complex studies. Some basic concepts based on actual avian, amphibian, and fish monitoring studies are presented in this article.

Environmetrics↗

Well predictive performance of play-wide and Subarea Random Forest models for Bakken productivity

In recent years, geologists and petroleum engineers have struggled to clearly identify the mechanisms that drive productivity in horizontal, hydraulically-fractured oil wells producing from the middle member of the Bakken formation. This paper fills a gap in the literature by showing how this play’s heterogeneity affects factors that drive well productivity. It is important because understanding the relative strength of productivity drivers and how predictors vary spatially facilitates best-practices for well site selection and well completion design. The paper describes an application of the Random Forest (RF) machine learning technique to identify these mechanisms and to evaluate their importance across 9 subareas of the North Dakota portion of the Bakken play. The study examined productivity of 7311 wells initiating production from 2010 through 2017. Well productivity varied considerably across the 9 subareas within the play, so it was not surprising that the dominant predictors, the initial 180-day water cut and the 30-day initial gas production, vary spatially to mirror local conditions that strongly affect well productivity. The relative importance of well completion predictor variables, that is, the numbers of fractures stages per well, volume of injected proppant per stage, volume of injected fluids per stage, and lateral length, varied considerably across the subareas. Statistical permutation tests are presented that generally confirm the importance rankings. Subarea Random Forest models explained from 50 percent to 82 percent of the variation in productivity test samples while the play-wide model explained 73 percent of the test sample well productivity. Weakness in the predictive ability of the Random Forest models are traced to the limited variability in the training data. Implications of the empirical findings regarding the Bakken play for operators and for research and government institutions are discussed in the concluding section.

Montana, North Dakota, South Dakota↗

Novel approach for computing photosynthetically active radiation for productivity modeling using remotely sensed images in the Great Plains, United States

Gross primary production (GPP) is a key indicator of ecosystem performance, and helps in many decision-making processes related to environment. We used the Eddy covariancelight use efficiency (EC-LUE) model for estimating GPP in the Great Plains, United States in order to evaluate the performance of this model. We developed a novel algorithm for computing the photosynthetically active radiation (PAR) based on net radiation. A strong correlation ( R 2 =0.94, N =24) was found between daily PAR and Landsat-based mid-day instantaneous net radiation. Though the Moderate Resolution Spectroradiometer (MODIS) based instantaneous net radiation was in better agreement ( R 2 =0.98, N =24) with the daily measured PAR, there was no statistical significant difference between Landsat based PAR and MODIS based PAR. The EC-LUE model validation also confirms the need to consider biological attributes (C 3 versus C 4 plants) for potential light use efficiency. A universal potential light use efficiency is unable to capture the spatial variation of GPP. It is necessary to use C 3 versus C 4 based land use/land cover map for using EC-LUE model for estimating spatiotemporal distribution of GPP.

Journal of Applied Remote Sensing↗

Carnivore hotspots in Peninsular Malaysia and their landscape attributes

Mammalian carnivores play a vital role in ecosystem functioning. However, they are prone to extinction because of low population densities and growth rates, and high levels of persecution or exploitation. In tropical biodiversity hotspots such as Peninsular Malaysia, rapid conversion of natural habitats threatens the persistence of this vulnerable group of animals. Here, we carried out the first comprehensive literature review on 31 carnivore species reported to occur in Peninsular Malaysia and updated their probable distribution. We georeferenced 375 observations of 28 species of carnivore from 89 unique geographic locations using records spanning 1948 to 2014. Using the Getis-Ord Gi*statistic and weighted survey records by IUCN Red List status, we identified hotspots of species that were of conservation concern and built regression models to identify environmental and anthropogenic landscape factors associated with Getis-Ord Gi* z scores. Our analyses identified two carnivore hotspots that were spatially concordant with two of the peninsula’s largest and most contiguous forest complexes, associated with Taman Negara National Park and Royal Belum State Park. A cold spot overlapped with the southwestern region of the Peninsula, reflecting the disappearance of carnivores with higher conservation rankings from increasingly fragmented natural habitats. Getis-Ord Gi* z scores were negatively associated with elevation, and positively associated with the proportion of natural land cover and distance from the capital city. Malaysia contains some of the world’s most diverse carnivore assemblages, but recent rates of forest loss are some of the highest in the world. Reducing poaching and maintaining large, contiguous tracts of lowland forests will be crucial, not only for the persistence of threatened carnivores, but for many mammalian species in general.

PLoS ONE↗

A topology of mineralization and its meaning for prospecting

Epigenetic mineral deposits are universal members of an orderly spatial-temporal arrangement of igneous rocks, endomorphic rocks, and hydrothermally altered rocks. The association and sequence of these rocks is invariant whereas the metric relations and configurations of the properties of these rocks are unlimited in variety. This characterization satisfies the doctrines of topology. Metric relations are statistical, and their modes are among the better guides to optimal areas for exploration. Metric configurations are graphically irregular and unpredictable mathematical surfaces like mountain topography. Each mineral edifice must be mapped to locate its mineral deposits. All measurements and observations are only positive or neutral for the occurrence of a mineral deposit. Effective prospecting is based on an increasing density of positive data with proximity to the mineral deposit. This means sampling for maximal numbers of positive data, pragmatically the highest ore-element assays at each site, by selecting rock showing maximal development of lode attributes.

Open-File Report↗

Mechanisms of aquatic species invasions across the South Atlantic Landscape Conservation Cooperative region

Invasive species are a global issue, and the southeastern United States is not immune to the problems they present. Therefore, various analyses using modeling and exploratory statistics were performed on the U.S. Geological Survey Nonindigenous Aquatic Species (NAS) Database with the primary objective of determining the most appropriate use of presence-only data as related to invasive species in the South Atlantic Landscape Conservation Cooperative (SALCC) region. A hierarchical model approach showed that a relatively small amount of high-quality data from planned surveys can be used to leverage the information in presence-only observations, having a broad spatial coverage and high biases of observer detection and in site selection. Because a variety of sampling protocols can be used in planned surveys, this approach to the analysis of presence-only data is widely applicable. An important part of the management of natural landscapes is the preservation of designated protected areas. When the hydrologic connection was considered in this analysis, the number of potential invaders that could spread to each protected area within the SALCC region was greatly increased, with a mean exceeding 30 species and the maximum reaching 57 species. Nearly all protected areas are hydrologically connected to at least 20 nonindigenous aquatic species. To examine possible factors which may contribute to nonindigenous aquatic species richness in the SALCC region, a set of exploratory statistics was employed. The best statistical model that included a combination of three anthropogenic variables (densities of housing, roads, and reservoirs) and two environmental variables (elevation range and longitude) explained approximately 62 percent of the variation in introduced species richness. Highest nonindigenous aquatic species richness occurred in the more upland, mountainous regions, where elevation range favored reservoirs and attracted urban centers. Lastly, patterns seen in a diffusion model may reflect less about the diffusion process of the organism and more about the opportunistic nature of the data collection process. These results of the model are considered exploratory in nature.

Alabama, Florida, Georgia, North Carolina, South C↗

Seasonal and interannual variability in the taxonomic composition and production dynamics of phytoplankton assemblages in Crater Lake, Oregon

Taxonomic composition and production dynamics of phytoplankton assemblages in Crater Lake, Oregon, were examined during time periods between 1984 and 2000. The objectives of the study were (1) to investigate spatial and temporal patterns in species composition, chlorophyll concentration, and primary productivity relative to seasonal patterns of water circulation; (2) to explore relationships between water column chemistry and the taxonomic composition of the phytoplankton; and (3) to determine effects of primary and secondary consumers on the phytoplankton assemblage. An analysis of 690 samples obtained on 50 sampling dates from 14 depths in the water column found a total of 163 phytoplankton taxa, 134 of which were identified to genus and 101 were identified to the species or variety level of classification. Dominant species by density or biovolume included Nitzschia gracilis, Stephanodiscus hantzschii, Ankistrodesmus spiralis, Mougeotia parvula, Dinobryon sertularia, Tribonema affine, Aphanocapsa delicatissima, Synechocystis sp., Gymnodinium inversum , and Peridinium inconspicuum . When the lake was thermally stratified in late summer, some of these species exhibited a stratified vertical distribution in the water column. A cluster analysis of these data also revealed a vertical stratification of the flora from the middle of the summer through the early fall. Multivariate test statistics indicated that there was a significant relationship between the species composition of the phytoplankton and a corresponding set of chemical variables measured for samples from the water column. In this case, concentrations of total phosphorus, ammonia, total Kjeldahl nitrogen, and alkalinity were associated with interannual changes in the flora; whereas pH and concentrations of dissolved oxygen, orthophosphate, nitrate, and silicon were more closely related to spatial variation and thermal stratification. The maximum chlorophyll concentration when the lake was thermally stratified in August and September was usually between depths of 100 m and 120 m. In comparison, the depth of maximum primary production ranged from 60 m to 80 m at this time of year. Regression analysis detected a weak negative relationship between chlorophyll concentration and Secchi disk depth, a measure of lake transparency. However, interannual changes in chlorophyll concentration and the species composition of the phytoplankton could not be explained by the removal of the septic field near Rim Village or by patterns of upwelling from the deep lake. An alternative trophic hypothesis proposes that the productivity of Crater Lake is controlled primarily by long-term patterns of climatic change that regulate the supply of allochthonous nutrients.

Crater Lake↗

Predicting aquatic habitat connectivity across watershed boundaries: Implications for interbasin spread of nonindigenous aquatic species.

Understanding habitat connectivity is critical for managing nonindigenous aquatic species (NAS) spread. Dams and watershed boundaries can be impassable to NAS during typical conditions but may become temporarily passable during flooding. The goal of our project was to develop an approach for identifying locations of aquatic connectivity at a fine spatial scale along watershed boundaries using readily available data. To develop this approach, we focused on the potential for range expansion of invasive fish in the United States via possible cross-boundary habitat connections. First, we developed an index using metrics of elevation, watershed size, and geology at regular points along a watershed boundary to stratify points by likelihood of connectivity during high precipitation (>20 mm of precipitation in a 3-day period). We then used a subset of points across a gradient of connectivity likelihoods to gather Landsat-derived observed surface water data and developed a statistical model to predict surface water presence from landscape characteristics. We applied the model throughout the entire watershed boundary to identify locations of hydrologic connectivity during high-water events. The presence of surface water on watershed boundaries was predicted by the interactions between watershed boundary point elevation relative to the minimum adjacent HUC-12 elevations and watershed boundary point elevation relative to neighboring point elevations (marginal R 2 = 0.94). Our approach can be used to identify potential areas of surface water connectivity between watersheds quickly and easily at a fine spatial scale using readily available, remotely sensed data that can inform conservation and management actions across disciplines.

North Dakota, South Dakota↗

A topology of mineralization and its meaning for prospecting

Epigenetic mineral deposits are universal members of an orderly spatial and temporal arrangement of igneous rocks, endomorphic rocks, and hydrothermally altered rocks. The association and sequence of these rocks is invariant whereas the metric relations and configurations of the properties of these rocks are unlimited in variety. This characterization satisfies the doctrines of topology. Metric relations are statistical, and their modes are among the better guides to optimal areas for exploration. Metric configurations are graphically irregular and unpredictable mathematical surfaces like mountain topography. Each mineral edifice must be mapped to locate its mineral deposits. All measurements and observations are only positive or neutral for the occurrence of a mineral deposit. Effective prospecting is based on an increasing density of positive data with proximity to the mineral deposit. This means sampling for maximal numbers of positive data, pragmatically the highest ore-element assays at each site, by selecting rock showing maximal development of lode attributes.

Book chapter↗

Detecting influential observations in nonlinear regression modeling of groundwater flow

Nonlinear regression is used to estimate optimal parameter values in models of groundwater flow to ensure that differences between predicted and observed heads and flows do not result from nonoptimal parameter values. Parameter estimates can be affected, however, by observations that disproportionately influence the regression, such as outliers that exert undue leverage on the objective function. Certain statistics developed for linear regression can be used to detect influential observations in nonlinear regression if the models are approximately linear. This paper discusses the application of Cook's D , which measures the effect of omitting a single observation on a set of estimated parameter values, and the statistical parameter DFBETAS, which quantifies the influence of an observation on each parameter. The influence statistics were used to (1) identify the influential observations in the calibration of a three-dimensional, groundwater flow model of a fractured-rock aquifer through nonlinear regression, and (2) quantify the effect of omitting influential observations on the set of estimated parameter values. Comparison of the spatial distribution of Cook's D with plots of model sensitivity shows that influential observations correspond to areas where the model heads are most sensitive to certain parameters, and where predicted groundwater flow rates are largest. Five of the six discharge observations were identified as influential, indicating that reliable measurements of groundwater flow rates are valuable data in model calibration. DFBETAS are computed and examined for an alternative model of the aquifer system to identify a parameterization error in the model design that resulted in overestimation of the effect of anisotropy on horizontal hydraulic conductivity.

Water Resources Research↗

A computer program (MODFLOWP) for estimating parameters of a transient, three-dimensional ground-water flow model using nonlinear regression

This report documents a new version of the U.S. Geological Survey modular, three-dimensional, finite-difference, ground-water flow model (MODFLOW) which, with the new Parameter-Estimation Package that also is documented in this report, can be used to estimate parameters by nonlinear regression. The new version of MODFLOW is called MODFLOWP (pronounced MOD-FLOW*P), and functions nearly identically to MODFLOW when the ParameterEstimation Package is not used. Parameters are estimated by minimizing a weighted least-squares objective function by the modified Gauss-Newton method or by a conjugate-direction method. Parameters used to calculate the following MODFLOW model inputs can be estimated: Transmissivity and storage coefficient of confined layers; hydraulic conductivity and specific yield of unconfined layers; vertical leakance; vertical anisotropy (used to calculate vertical leakance); horizontal anisotropy; hydraulic conductance of the River, Streamflow-Routing, General-Head Boundary, and Drain Packages; areal recharge rates; maximum evapotranspiration; pumpage rates; and the hydraulic head at constant-head boundaries. Any spatial variation in parameters can be defined by the user. Data used to estimate parameters can include existing independent estimates of parameter values, observed hydraulic heads or temporal changes in hydraulic heads, and observed gains and losses along head-dependent boundaries (such as streams). Model output includes statistics for analyzing the parameter estimates and the model; these statistics can be used to quantify the reliability of the resulting model, to suggest changes in model construction, and to compare results of models constructed in different ways.

Open-File Report↗

Defining surfaces for skewed, highly variable data

Skewness of environmental data is often caused by more than simply a handful of outliers in an otherwise normal distribution. Statistical procedures for such datasets must be sufficiently robust to deal with distributions that are strongly non-normal, containing both a large proportion of outliers and a skewed main body of data. In the field of water quality, skewness is commonly associated with large variation over short distances. Spatial analysis of such data generally requires either considerable effort at modeling or the use of robust procedures not strongly affected by skewness and local variability. Using a skewed dataset of 675 nitrate measurements in ground water, commonly used methods for defining a surface (least-squares regression and kriging) are compared to a more robust method (loess). Three choices are critical in defining a surface: (i) is the surface to be a central mean or median surface? (ii) is either a well-fitting transformation or a robust and scale-independent measure of center used? (iii) does local spatial autocorrelation assist in or detract from addressing objectives? Published in 2002 by John Wiley & Sons, Ltd.

Environmetrics↗

Modelling gully-erosion susceptibility in a semi-arid region, Iran: Investigation of applicability of certainty factor and maximum entropy models

Gully erosion susceptibility mapping is a fundamental tool for land-use planning aimed at mitigating land degradation. However, the capabilities of some state-of-the-art data-mining models for developing accurate maps of gully erosion susceptibility have not yet been fully investigated. This study assessed and compared the performance of two different types of data-mining models for accurately mapping gully erosion susceptibility at a regional scale in Chavar, Ilam, Iran. The two methods evaluated were: Certainty Factor (CF), a bivariate statistical model; and Maximum Entropy (ME), an advanced machine learning model. Several geographic and environmental factors that can contribute to gully erosion were considered as predictor variables of gully erosion susceptibility. Based on an existing differential GPS survey inventory of gully erosion, a total of 63 eroded gullies were spatially randomly split in a 70:30 ratio for use in model calibration and validation, respectively. Accuracy assessments completed with the receiver operating characteristic curve method showed that the ME-based regional gully susceptibility map has an area under the curve (AUC) value of 88.6% whereas the CF-based map has an AUC of 81.8%. According to jackknife tests that were used to investigate the relative importance of predictor variables, aspect, distance to river, lithology and land use are the most influential factors for the spatial distribution of gully erosion susceptibility in this region of Iran. The gully erosion susceptibility maps produced in this study could be useful tools for land managers and engineers tasked with road development, urbanization and other future development.

Science of the Total Environment↗

Estimation of aquifer scale proportion using equal area grids: assessment of regional scale groundwater quality

The proportion of an aquifer with constituent concentrations above a specified threshold (high concentrations) is taken as a nondimensional measure of regional scale water quality. If computed on the basis of area, it can be referred to as the aquifer scale proportion. A spatially unbiased estimate of aquifer scale proportion and a confidence interval for that estimate are obtained through the use of equal area grids and the binomial distribution. Traditionally, the confidence interval for a binomial proportion is computed using either the standard interval or the exact interval. Research from the statistics literature has shown that the standard interval should not be used and that the exact interval is overly conservative. On the basis of coverage probability and interval width, the Jeffreys interval is preferred. If more than one sample per cell is available, cell declustering is used to estimate the aquifer scale proportion, and Kish's design effect may be useful for estimating an effective number of samples. The binomial distribution is also used to quantify the adequacy of a grid with a given number of cells for identifying a small target, defined as a constituent that is present at high concentrations in a small proportion of the aquifer. Case studies illustrate a consistency between approaches that use one well per grid cell and many wells per cell. The methods presented in this paper provide a quantitative basis for designing a sampling program and for utilizing existing data.

Water Resources Research↗

A multi-scale soil moisture monitoring strategy for California: Design and validation

A multi‐scale soil moisture monitoring strategy for California was designed to inform water resource management. The proposed workflow classifies soil moisture response units (SMRUs) using publicly available datasets that represent soil, vegetation, climate, and hydrology variables, which control soil water storage. The SMRUs were classified, using principal component analysis and unsupervised K‐means clustering within a geographic information system, and validated, using summary statistics derived from measured soil moisture time series. Validation stations, located in the Sierra Nevada, include transect of sites that cross the rain‐to‐snow transition and a cluster of sites located at similar elevations in a snow‐dominated watershed. The SMRUs capture unique responses to varying climate conditions characterized by statistical measures of central tendency, dispersion, and extremes. A topographic position index and landform classification is the final step in the workflow to guide the optimal placement of soil moisture sensors at the local‐scale. The proposed workflow is highly flexible and can be implemented over a range of spatial scales and input datasets can be customized. Our approach captures a range of soil moisture responses to climate across California and can be used to design and optimize soil moisture monitoring strategies to support runoff forecasts for water supply management or to assess landscape conditions for forest and rangeland management.

California↗

Normal streamflows and water levels continue—Summary of hydrologic conditions in Georgia, 2014

The U.S. Geological Survey (USGS) South Atlantic Water Science Center (SAWSC) Georgia office, in cooperation with local, State, and other Federal agencies, maintains a long-term hydrologic monitoring network of more than 350 real-time, continuous-record, streamflow-gaging stations (streamgages). The network includes 14 real-time lake-level monitoring stations, 72 real-time surface-water-quality monitors, and several water-quality sampling programs. Additionally, the SAWSC Georgia office operates more than 204 groundwater monitoring wells, 39 of which are real-time. The wide-ranging coverage of streamflow, reservoir, and groundwater monitoring sites allows for a comprehensive view of hydrologic conditions across the State. One of the many benefits this monitoring network provides is a spatially distributed overview of the hydrologic conditions of creeks, rivers, reservoirs, and aquifers in Georgia. Streamflow and groundwater data are verified throughout the year by USGS hydrographers and made available to water-resource managers, recreationists, and Federal, State, and local agencies. Hydrologic conditions are determined by comparing the statistical analyses of data collected during the current water year to historical data. Changing hydrologic conditions underscore the need for accurate, timely data to allow informed decisions about the management and conservation of Georgia’s water resources for agricultural, recreational, ecological, and water-supply needs and in protecting life and property.

Georgia↗