USGS ScienceSearch

USGS · 70015394

Response of selected binomial coefficients to varying degrees of matrix sparseness and to matrices with known data interrelationships

Abstract

Numerous departures from ideal relationships are revealed by Monte Carlo simulations of widely accepted binomial coefficients. For example, simulations incorporating varying levels of matrix sparseness (presence of zeros indicating lack of data) and computation of expected values reveal that not only are all common coefficients influenced by zero data, but also that some coefficients do not discriminate between sparse or dense matrices (few zero data). Such coefficients computationally merge mutually shared and mutually absent information and do not exploit all the information incorporated within the standard 2 ?? 2 contingency table; therefore, the commonly used formulae for such coefficients are more complicated than the actual range of values produced. Other coefficients do differentiate between mutual presences and absences; however, a number of these coefficients do not demonstrate a linear relationship to matrix sparseness. Finally, simulations using nonrandom matrices with known degrees of row-by-row similarities signify that several coefficients either do not display a reasonable range of values or are nonlinear with respect to known relationships within the data. Analyses with nonrandom matrices yield clues as to the utility of certain coefficients for specific applications. For example, coefficients such as Jaccard, Dice, and Baroni-Urbani and Buser are useful if correction of sparseness is desired, whereas the Russell-Rao coefficient is useful when sparseness correction is not desired. ?? 1989 International Association for Mathematical Geology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A.W. Archer, C.G. Maples. 1989. Response of selected binomial coefficients to varying degrees of matrix sparseness and to matrices with known data interrelationships. https://doi.org/10.1007/bf00893319

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Declustering of clustered preferential sampling for histogram and semivariogram inference

Measurements of attributes obtained more as a consequence of business ventures than sampling design frequently result in samplings that are preferential both in location and value, typically in the form of clusters along the pay. Preferential sampling requires preprocessing for the purpose of properly inferring characteristics of the parent population, such as the cumulative distribution and the semivariogram. Consideration of the distance to the nearest neighbor allows preparation of resampled sets that produce comparable results to those from previously proposed methods. Clustered sampling of size 140, taken from an exhaustive sampling, is employed to illustrate this approach. ?? International Association for Mathematical Geology 2007.

Mathematical Geology

Comparison of two probability distributions used to model sizes of undiscovered oil and gas accumulations: Does the tail wag the assessment?

Undiscovered oil and gas assessments are commonly reported as aggregate estimates of hydrocarbon volumes. Potential commercial value and discovery costs are, however, determined by accumulation size, so engineers, economists, decision makers, and sometimes policy analysts are most interested in projected discovery sizes. The lognormal and Pareto distributions have been used to model exploration target sizes. This note contrasts the outcomes of applying these alternative distributions to the play level assessments of the U.S. Geological Survey's 1995 National Oil and Gas Assessment. Using the same numbers of undiscovered accumulations and the same minimum, medium, and maximum size estimates, substitution of the shifted truncated lognormal distribution for the shifted truncated Pareto distribution reduced assessed undiscovered oil by 16% and gas by 15%. Nearly all of the volume differences resulted because the lognormal had fewer larger fields relative to the Pareto. The lognormal also resulted in a smaller number of small fields relative to the Pareto. For the Permian Basin case study presented here, reserve addition costs were 20% higher with the lognormal size assumption. ?? 2002 International Association for Mathematical Geology.

Mathematical Geology

Uncertainty estimation for resource assessment-an application to coal

The U.S. Geological Survey is conducting a national assessment of coal resources. As part of that assessment, a geostatistical procedure has been developed to estimate the uncertainty of coal resources for the historical categories of geological assurance: measured, indicated, inferred, and hypothetical coal. Data consist of spatially clustered coal thickness measurements from coal beds and/or zones that cover, in some cases, several thousand square kilometers. Our procedure involved trend removal, an examination of spatial correlation, computation of a sample semivariogram, and fitting a semivariogram model. This model provided standard deviations for the uncertainty estimates. The number of sample points (drill holes) in each historical category also was estimated. Measurement error in the thickness of the coal bed/zone was obtained from the fitted model or supplied exogenously. From this information approximate estimates of uncertainty on the historical categories were computed. We illustrate the methodology using drill hole data from the Harmon coal bed located in southwestern North Dakota. The methodology will be applied to approximately 50 coal data sets.

Mathematical Geology