USGS ScienceSearch

USGS · 70197439

Small values in big data: The continuing need for appropriate metadata

Abstract

Compiling data from disparate sources to address pressing ecological issues is increasingly common. Many ecological datasets contain left-censored data – observations below an analytical detection limit. Studies from single and typically small datasets show that common approaches for handling censored data — e.g., deletion or substituting fixed values — result in systematic biases. However, no studies have explored the degree to which the documentation and presence of censored data influence outcomes from large, multi-sourced datasets. We describe left-censored data in a lake water quality database assembled from 74 sources and illustrate the challenges of dealing with small values in big data, including detection limits that are absent, range widely, and show trends over time. We show that substitutions of censored data can also bias analyses using ‘big data’ datasets, that censored data can be effectively handled with modern quantitative approaches, but that such approaches rely on accurate metadata that describe treatment of censored data from each source.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Craig A. Stow, Katherine E. Webster, Tyler Wagner, Noah R. Lottig, Patricia A. Soranno, YoonKyung Cha. 2018. Small values in big data: The continuing need for appropriate metadata. https://doi.org/10.1016/j.ecoinf.2018.03.002

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

A Bayesian hierarchical modeling approach for species diversity in ecology

Species diversity is the foundation of many ecological disciplines. This metric is often approximated using species richness and evenness, even though actual richness likely exceeds observations due to imperfect sampling methods. Estimating the “true” species richness, which includes identifying the number of missing species, has intrigued ecologists for decades. We adopted a parametric model that appeared in Fisher et al. (1943), which models the numbers of individuals from different species as random samples from a negative binomial distribution, and developed a Bayesian computational approach to directly estimate the distribution model parameters. The model parameters represent species abundance and evenness, and can be used to derive species richness. We evaluated our parametric approach using (1) a simulation study and (2) three historical data sets. Furthermore, we illustrated the hierarchical modeling approach to combine data from multiple parallel studies using a biannual fishery survey data set. Our parametric model formulation is computationally efficient, and the hierarchical structure facilitates embedding diversity estimation into broader application, such as assessing spatial and temporal trends in species diversity associated with environmental stressors. Additionally, because the two parameters of the negative binomial distribution model represent species abundance and evenness of a community, this parametric approach facilitates a deeper understanding of the ecological systems under study. The negative binomial distribution model works with a wide range of species frequency distribution types. As a result, our emphasis on a parametric model can help us characterize the structure of an ecosystem and provide a greater depth of ecologically meaningful information.

Ecological Informatics

Hierarchical mixture models and high-resolution monitoring data can inform siting and operational strategies to mitigate bat fatalities at wind turbines

Bats provide critical ecosystem services, but bat fatalities due to wind energy development may imperil some bat populations. Statistical models are used to estimate the total fatalities that occur based on carcasses observed during monitoring surveys. Current models often estimate fatalities aggregated across species, time, and/or turbines, but fall short of reliably informing siting and operational collision mitigation strategies that account for species-specific fatality patterns on a fine spatiotemporal scale. We developed a hierarchical mixture model for estimating species-specific covariate effects and total fatalities per species at each turbine on weekly intervals. We applied the model to a high-resolution dataset of bat carcasses found during turbine searches across nineteen wind facilities in Iowa over two years. Our model explains species-specific variation in bat fatalities at individual wind turbines according to turbine proximity to bat habitat, turbine design specifications, seasonal trends, and weather conditions such as nightly air temperature, air pressure, and wind speed. Turbines located on the edge of wind facilities had higher fatalities, and proximity to roosting and foraging habitat accounted for variation in species-specific fatality estimates. These insights into turbine placement effects can inform siting strategies. We also discovered species-specific relationships with average nightly wind speed and air temperature, among other weather conditions, that could inform operational mitigation strategies such as smart curtailment. Our model can transform observations of carcasses found during turbine searches across multiple facilities, years, and variable search efforts into estimates of total fatalities per species associated with species-specific spatial, temporal, and environmental covariate effects.

Ecological Informatics

Two-stage approach to automatic detection with machine learning for improved surveillance of the invasive Cuban treefrog

The Cuban treefrog ( Osteopilus septentrionalis ), as an invasive species in the southern United States, presents a need for effective surveillance. Automated detection expedites processing of audio data for large-scale surveillance and monitoring programs. However, current available methods commonly used for anuran species have not been sufficient to detect Cuban treefrogs. Here, we present results from a two-stage method for automated detection that employs both cross-correlation template matching and secondary supervised learning classifiers. In the first stage, audio data are screened for initial detections using template matching, in which the detections contain both true and false positives. In the second stage, the false positives are screened out using classifier algorithms. We used this method to process 139,985 audio recordings, consisting of 596,046 total minutes, collected at 13 locations in Louisiana and Florida from 2014 to 2022. From the stage 1 template matching, we detected 83,191 Cuban treefrog signals across recordings. The stage 2 machine learning model was able to identify stage 1 false positive detections with a testing accuracy of 98.46% and a testing false positive rate of 1.116%. After pruning false positive detections, a total of 20,271 individual Cuban treefrog detections remained, distributed mainly across 3 sites in an area with known presence. Locations with presumed absence had an easily verifiable number of false positive detections ( n = 109 across all other sites). The two-stage methodology utilizing both template matching and machine learning algorithms can be integrated into wildlife surveillance or monitoring programs for species with distinctive, conserved calls as an effective way to achieve sensitive species detection with a low incidence of false positives.

Florida, Louisiana