USGS ScienceSearch

USGS · 70255284

Greater than the sum of its parts: Computationally flexible Bayesian hierarchical modeling

Abstract

We propose a multistage method for making inference at all levels of a Bayesian hierarchical model (BHM) using natural data partitions to increase efficiency by allowing computations to take place in parallel form using software that is most appropriate for each data partition. The full hierarchical model is then approximated by the product of independent normal distributions for the data component of the model. In the second stage, the Bayesian maximum a posteriori (MAP) estimator is found by maximizing the approximated posterior density with respect to the parameters. If the parameters of the model can be represented as normally distributed random effects, then the second-stage optimization is equivalent to fitting a multivariate normal linear mixed model. We consider a third stage that updates the estimates of distinct parameters for each data partition based on the results of the second stage. The method is demonstrated with two ecological data sets and models, a generalized linear mixed effects model (GLMM) and an integrated population model (IPM). The multistage results were compared to estimates from models fit in single stages to the entire data set. In both cases, multistage results were very similar to a full MCMC analysis. Supplementary materials accompanying this paper appear online.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Devin S. Johnson, Brian M. Brost, Mevin Hooten. 2022-03-09. Greater than the sum of its parts: Computationally flexible Bayesian hierarchical modeling. https://doi.org/10.1007/s13253-021-00485-9

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Estimating submersed aquatic vegetation biomass using a hierarchical Bayesian model of ordinal measurements and structural zeros

This study explores the prediction of a continuous target variable that is primarily measured as an ordinal outcome. Interest in such prediction may occur given data from double- or two-phase sampling designs, where an auxiliary variable is measured at all sampling units and a target and predictor variable, X , at a subset of those units. This study focuses on the case where X is nonnegative, is conditional on a zero process, and is sampled using a cluster sampling design. We model the ordinal outcome using cumulative logistic regression and the conditional regressor X , following log transformation, in a Bayesian setting. Estimation of model parameters and prediction of missing X is generally accurate and precise, particularly with sufficiently large Pr(X > 0) and associated variances. This study is motivated by a need to accurately estimate the biomass of submersed aquatic vegetation on the Upper Mississippi River, where two types of measurements for biomass are available: inexpensive, ordinal biomass scores and expensive, continuous diver-harvested biomass data. Here, X represents plant biomass and Pr(X > 0) is the usual species site occupancy measure. When fit using biomass data from the submerged aquatic vegetation species Vallisneria americana Michx, our model estimates parameters within the parameter space region where the model performed acceptably using synthetic data. We expect this model to appeal to investigators with ordered outcomes, cluster designs, and one or more continuous, positive predictors.

Journal of Agricultural, Biological and Environmen

Nonstationary demographic state-space models using unreplicated counts for species undergoing environmental stressors

A fundamental task in ecological statistics is to estimate abundance and growth rate distributions from wildlife monitoring data to inform conservation management. Modeling time series of wildlife populations presents a number of challenges from both statistical and ecological perspectives, including discreteness; lack of replication; nonstationarity; and observation, demographic, and other phenomenological processes. Nonstationary dynamics are often exhibited by populations undergoing environmental stressors. Models must account for these characteristics to produce reliable estimates of abundance and trends, yet estimation can be challenging with unreplicated data. We propose nonstationary demographic state-space models using unreplicated counts for populations undergoing environmental stressors. A reduced growth rate model matches the complexity of the unreplicated count data, and a fecundity bound on growth rate distributions allows the separation of processes affecting growth rates like environmental stressors from those affecting abundance external to growth rates like migration. NDSSMs allow for the embedding of nonstationary model components, and we explore the use of changepoints, volatility clustering, and migration processes. We apply the proposed nonstationary models in case studies of herons affected by predator/competitor reestablishment and three bat species affected by a fungal pathogen causing white-nose syndrome. Nonstationary models outperform stationary models and generalized linear mixed effects models according to model scoring and visual inspection of predictions, and provide estimates more consistent with published values. Incorporating migration improves model fit universally, even with approximate one-way immigration, most likely because populations are extirpated, recolonized, and increase multiple-fold over the upper bound set by species fecundity. In addition, estimates of the timing and severity of the environmental stressor differed for models with migration. Including nonstationary and demographic components in a fecundity-bounded growth rate model improves inference and benefits interpretability of hyperparameters. In turn, this adjusts uncertainties in predictions of abundance and growth rates over time, providing the ingredients needed for informed conservation analysis and for directing future monitoring of at-risk species.

Journal of Agricultural, Biological and Environmen

Bayesian approaches to proxy uncertainty quantification in paleoecology: A mathematical justification and practical integration

Paleoenvironmental data are essential for reconstructing environmental conditions in the distant past, and these reconstructions strongly depend on proxies and age–depth models. Proxies are indirect measurements that substitute for variables that cannot be directly measured, such as past precipitation. Conversely, an age–depth model is a tool that correlates the observed proxy with a specific moment in time. Bayesian age–depth modelling has proved to be a powerful method for estimating sediment ages and their associated uncertainties. However, there remains considerable potential for further integration into proxy analysis. In this paper, we explore a mathematical justification and a computational approach that integrates uncertainty at the age–depth level and propagates it to the proxy scale in the form of a posterior predictive distribution. This method mitigates potential biases and errors by removing the need to assign a single age to a given proxy measurement. It allows for quantifying the likelihood that proxy data values correspond to modelled ages, thus enabling the quantification of uncertainty in both the temporal and proxy value domains. The use of Bayesian statistics in proxy analysis represents a relatively recent advancement. We aim to mathematically justify incorporating the Markov chain Monte Carlo output from age–depth models into proxy analysis and to present a novel methodology for constructing environmental reconstructions using this approach.

Journal of Agricultural, Biological and Environmen