USGS ScienceSearch

Geology topics

Ephraim M. Hanks

Publications and source records attributed to Ephraim M. Hanks.

At least 19 recordsLinked to original sources

A flexible movement model for partially migrating species

We propose a flexible model for a partially migrating species, which we demonstrate using yearly paths for golden eagles ( Aquila chrysaetos ). Our model relies on a smoothly time-varying potential surface defined by a number of attractors. We compare our proposed approach using varying coefficients to a latent-state model, which we define differently for migrating, dispersing, and local individuals. While latent-state models are more common in the existing animal movement literature, varying coefficient models have various benefits including the ability to fit a wide range of movement strategies without the need for major model adjustments. We compare simulations from the models for three individuals to illustrate the ability of our model to better describe movement behavior for specific movement strategies. We also demonstrate the flexibility of our model by fitting several individuals whose movement behavior is less stereotypical.

Spatial Statistics

Ecological prediction at macroscales using big data: Does sampling design matter?

Although ecosystems respond to global change at regional to continental scales (i.e., macroscales), model predictions of ecosystem responses often rely on data from targeted monitoring of a small proportion of sampled ecosystems within a particular geographic area. In this study, we examined how the sampling strategy used to collect data for such models influences predictive performance. We subsampled a large and spatially-extensive dataset to investigate how macroscale sampling strategy affects prediction of ecosystem characteristics in 6,784 lakes across a 1.8 million km2 area. We estimated model predictive performance for different subsets of the dataset to mimic three common sampling strategies for collecting observations of ecosystem characteristics: random sampling design, stratified random sampling design, and targeted sampling. We found that sampling strategy influenced model predictive performance such that (1) stratified random sampling designs did not improve predictive performance compared to simple random sampling designs and (2) although one of the scenarios that mimicked targeted (non-random) sampling had the poorest performing predictive models, the other targeted sampling scenarios resulted in models with similar predictive performance to that of the random sampling scenarios. Our results suggest that although potential biases in datasets from some forms of targeted sampling may limit predictive performance, compiling existing spatially-extensive datasets can result in models with good predictive performance that may inform a wide range of science questions and policy goals related to global change.

Ecological Applications

A novel quantitative framework for riverscape genetics

Riverscape genetics, which applies concepts in landscape genetics to riverine ecosystems, lack appropriate quantitative methods that address the spatial autocorrelation structure of linear stream networks and account for bidirectional geneflow. To address these challenges, we present a general framework for the design and analysis of riverscape genetic studies. Our framework starts with the estimation of pairwise genetic distance at sample sites and the development of a spatially structured ecological network (SSEN) on which riverscape covariates are measured. We then introduce the novel bidirectional geneflow in riverscapes (BGR) model that uses principles of isolation-by-resistance to quantify the effects of environmental covariates on genetic connectivity, with spatial covariance defined using simultaneous autoregressive models on the SSEN and the generalized Wishart distribution to model pairwise distance matrices arising through a random walk model of geneflow. We highlight the utility of this framework in an analysis of riverscape genetics for brook trout ( Salvelinus fontinalis ) in north central Pennsylvania, USA. Using the fixation index ( F ST ) as the measure of genetic distance, we estimated the effects of 12 riverscape covariates on geneflow by evaluating the relative support of eight competing BGR models. We then compared the performance of the top-ranked BGR model to results obtained from comparable analyses using multiple regression on distance matrices (MRM) and the program STRUCTURE. We found that the BGR model had more power to detect covariate effects, particularly for variables that were only partial barriers to geneflow and/or uncommon in the riverscape, making it more informative for assessing patterns of population connectivity and identifying threats to species conservation. This case study highlights the utility of our modeling framework over other quantitative methods in riverscape genetics, particularly the ability to rigorously test hypotheses about factors that influence geneflow and probabilistically estimate the effect of riverscape covariates, including stream flow direction. This framework is flexible across taxa and riverine networks, is easily executable, and provides intuitive results that can be used to investigate the likely outcomes of current and future management scenarios.

Ecological Applications

Increasing accuracy of lake nutrient predictions in thousands of lakes by leveraging water clarity data

Aquatic scientists require robust, accurate information about nutrient concentrations and indicators of algal biomass in unsampled lakes in order to understand and predict the effects of global climate and land-use change. Historically, lake and landscape characteristics have been used as predictor variables in regression models to generate nutrient predictions, but often with significant uncertainty. An alternative approach to improve predictions is to leverage the observed relationship between water clarity and nutrients, which is possible because water clarity is more commonly measured than lake nutrients. We used a joint-nutrient model that conditioned predictions of total phosphorus, nitrogen, and chlorophyll a on observed water clarity. Our results demonstrated substantial reductions (8–27%; median = 23%) in prediction error when conditioning on water clarity. These models will provide new opportunities for predicting nutrient concentrations of unsampled lakes across broad spatial scales with reduced uncertainty.

Limnology and Oceanography Letters

Identifying and characterizing extrapolation in multivariate response data

Faced with limitations in data availability, funding, and time constraints, ecologists are often tasked with making predictions beyond the range of their data. In ecological studies, it is not always obvious when and where extrapolation occurs because of the multivariate nature of the data. Previous work on identifying extrapolation has focused on univariate response data, but these methods are not directly applicable to multivariate response data, which are common in ecological investigations. In this paper, we extend previous work that identified extrapolation by applying the predictive variance from the univariate setting to the multivariate case. We propose using the trace or determinant of the predictive variance matrix to obtain a scalar value measure that, when paired with a selected cutoff value, allows for delineation between prediction and extrapolation. We illustrate our approach through an analysis of jointly modeled lake nutrients and indicators of algal biomass and water clarity in over 7000 inland lakes from across the Northeast and Mid-west US. In addition, we outline novel exploratory approaches for identifying regions of covariate space where extrapolation is more likely to occur using classification and regression trees. The use of our Multivariate Predictive Variance (MVPV) measures and multiple cutoff values when exploring the validity of predictions made from multivariate statistical models can help guide ecological inferences.

PLoS ONE

Simultaneous autoregressive (SAR) model

Simultaneous autoregressive (SAR) models are useful for accommodating various forms of dependence among data that have discrete support in a space of interest. These models are often specified hierarchically as mixed-effects regression models with first-moment structure controlled by a conventional linear regression term and second-moment structure induced by correlated random effects. In their general form, SAR models resemble conditional autoregressive (CAR) models, and can be made equivalent but are often parameterized differently. Importantly, SAR models can be specified by simultaneously regressing a discrete spatial process on itself. Thus, they allow one to construct statistical models for processes with directional graphical properties that pertain to data generating mechanisms. Most commonly SAR models have been used to account for structure among data with areal spatial support in applications involving ecology, epidemiology, sociology, and environmental science.

Book chapter

Confronting models with data: The challenges of estimating disease spillover

For pathogens known to transmit across host species, strategic investment in disease control requires knowledge about where and when spillover transmission is likely. One approach to estimating spillover is to directly correlate observed spillover events with covariates. An alternative is to mechanistically combine information on host density, distribution, and pathogen prevalence to predict where and when spillover events are expected to occur. We use several case studies at the wildlife-livestock disease interface to highlight the challenges, and potential solutions, to estimating spatio-temporal variation in spillover risk. Datasets on multiple host species often do not align in space, time or resolution, and may have no estimates of observation error. Linking these datasets requires they be related to a common spatial and temporal resolution and appropriately propagating errors in predictions can be difficult. Hierarchical models are one potential solution, but for fine-resolution predictions at broad spatial scales many models become computationally challenging. Despite these limitations, the confrontation of mechanistic predictions with observed events is an important avenue for developing a better understanding of pathogen spillover. Systems where data have been collected at all levels in the spillover process are rare, or non-existent, and require investment and sustained effort across disciplines.

Philosophical Transactions of the Royal Society B:

Extreme value-based methods for modeling elk yearly movements

Species range shifts and the spread of diseases are both likely to be driven by extreme movements, but are difficult to statistically model due to their rarity. We propose a statistical approach for characterizing movement kernels that incorporate landscape covariates as well as the potential for heavy-tailed distributions. We used a spliced distribution for distance travelled paired with a resource selection function to model movements biased toward preferred habitats. As an example, we used data from 704 annual elk movements around the Greater Yellowstone Ecosystem from 2001 to 2015. Yearly elk movements were both heavy-tailed and biased away from high elevations during the winter months. We then used a simulation to illustrate how these habitat effects may alter the rate of disease spread using our estimated movement kernel relative to a more traditional approach that does not include landscape covariates. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme

Time-varying predatory behavior is primary predictor of fine-scale movement of wildland-urban cougars

Background While many species have suffered from the detrimental impacts of increasing human population growth, some species, such as cougars ( Puma concolor ), have been observed using human-modified landscapes. However, human-modified habitat can be a source of both increased risk and increased food availability, particularly for large carnivores. Assessing preferential use of the landscape is important for managing wildlife and can be particularly useful in transitional habitats, such as at the wildland-urban interface. Preferential use is often evaluated using resource selection functions (RSFs), which are focused on quantifying habitat preference using either a temporally static framework or researcher-defined temporal delineations. Many applications of RSFs do not incorporate time-varying landscape availability or temporally-varying behavior, which may mask conflict and avoidance behavior. Methods Contemporary approaches to incorporate landscape availability into the assessment of habitat selection include spatio-temporal point process models, step selection functions, and continuous-time Markov chain (CTMC) models; in contrast with the other methods, the CTMC model allows for explicit inference on animal movement in continuous-time. We used a hierarchical version of the CTMC framework to model speed and directionality of fine-scale movement by a population of cougars inhabiting the Front Range of Colorado, U.S.A., an area exhibiting rapid population growth and increased recreational use, as a function of individual variation and time-varying responses to landscape covariates. Results We found evidence for individual- and daily temporal-variability in cougar response to landscape characteristics. Distance to nearest kill site emerged as the most important driver of movement at a population-level. We also detected seasonal differences in average response to elevation, heat loading, and distance to roads. Motility was also a function of amount of development, with cougars moving faster in developed areas than in undeveloped areas. Conclusions The time-varying framework allowed us to detect temporal variability that would be masked in a generalized linear model, and improved the within-sample predictive ability of the model. The high degree of individual variation suggests that, if agencies want to minimize human-wildlife conflict management options should be varied and flexible. However, due to the effect of recursive behavior on cougar movement, likely related to the location and timing of potential kill-sites, kill-site identification tools may be useful for identifying areas of potential conflict.

Colorado

Examining speed versus selection in connectivity models using elk migration as an example

Context Landscape resistance is vital to connectivity modeling and frequently derived from resource selection functions (RSFs). RSFs estimate relative probability of use and tend to focus on understanding habitat preferences during slow, routine animal movements (e.g., foraging). Dispersal and migration, however, can produce rarer, faster movements, in which case models of movement speed rather than resource selection may be more realistic for identifying habitats that facilitate connectivity. Objective To compare two connectivity modeling approaches applied to resistance estimated from models of movement rate and resource selection. Methods Using movement data from migrating elk, we evaluated continuous time Markov chain (CTMC) and movement-based RSF models (i.e., step selection functions [SSFs]). We applied circuit theory and shortest random path (SRP) algorithms to CTMC, SSF and null (i.e., flat) resistance surfaces to predict corridors between elk seasonal ranges. We evaluated prediction accuracy by comparing model predictions to empirical elk movements. Results All connectivity models predicted elk movements well, but models applied to CTMC resistance were more accurate than models applied to SSF and null resistance. Circuit theory models were more accurate on average than SRP models. Conclusions CTMC can be more realistic than SSFs for estimating resistance for fast movements, though SSFs may demonstrate some predictive ability when animals also move slowly through corridors (e.g., stopover use during migration). High null model accuracy suggests seasonal range data may also be critical for predicting direct migration routes. For animals that migrate or disperse across large landscapes, we recommend incorporating CTMC into the connectivity modeling toolkit.

Landscape Ecology

Spatial autoregressive models for statistical inference from ecological data

Ecological data often exhibit spatial pattern, which can be modeled as autocorrelation. Conditional autoregressive (CAR) and simultaneous autoregressive (SAR) models are network‐based models (also known as graphical models) specifically designed to model spatially autocorrelated data based on neighborhood relationships. We identify and discuss six different types of practical ecological inference using CAR and SAR models, including: (1) model selection, (2) spatial regression, (3) estimation of autocorrelation, (4) estimation of other connectivity parameters, (5) spatial prediction, and (6) spatial smoothing. We compare CAR and SAR models, showing their development and connection to partial correlations. Special cases, such as the intrinsic autoregressive model (IAR), are described. Conditional autoregressive and SAR models depend on weight matrices, whose practical development uses neighborhood definition and row‐standardization. Weight matrices can also include ecological covariates and connectivity structures, which we emphasize, but have been rarely used. Trends in harbor seals ( Phoca vitulina ) in southeastern Alaska from 463 polygons, some with missing data, are used to illustrate the six inference types. We develop a variety of weight matrices and CAR and SAR spatial regression models are fit using maximum likelihood and Bayesian methods. Profile likelihood graphs illustrate inference for covariance parameters. The same data set is used for both prediction and smoothing, and the relative merits of each are discussed. We show the nonstationary variances and correlations of a CAR model and demonstrate the effect of row‐standardization. We include several take‐home messages for CAR and SAR models, including (1) choosing between CAR and IAR models, (2) modeling ecological effects in the covariance matrix, (3) the appeal of spatial smoothing, and (4) how to handle isolated neighbors. We highlight several reasons why ecologists will want to make use of autoregressive models, both directly and in hierarchical models, and not only in explicit spatial settings, but also for more general connectivity models.

Ecological Monographs

The Bayesian group lasso for confounded spatial data

Generalized linear mixed models for spatial processes are widely used in applied statistics. In many applications of the spatial generalized linear mixed model (SGLMM), the goal is to obtain inference about regression coefficients while achieving optimal predictive ability. When implementing the SGLMM, multicollinearity among covariates and the spatial random effects can make computation challenging and influence inference. We present a Bayesian group lasso prior with a single tuning parameter that can be chosen to optimize predictive ability of the SGLMM and jointly regularize the regression coefficients and spatial random effect. We implement the group lasso SGLMM using efficient Markov chain Monte Carlo (MCMC) algorithms and demonstrate how multicollinearity among covariates and the spatial random effect can be monitored as a derived quantity. To test our method, we compared several parameterizations of the SGLMM using simulated data and two examples from plant ecology and disease ecology. In all examples, problematic levels multicollinearity occurred and influenced sampling efficiency and inference. We found that the group lasso prior resulted in roughly twice the effective sample size for MCMC samples of regression coefficients and can have higher and less variable predictive accuracy based on out-of-sample data when compared to the standard SGLMM.

Journal of Agricultural, Biological, and Environme

A dynamic spatio-temporal model for spatial data

Analyzing spatial data often requires modeling dependencies created by a dynamic spatio-temporal data generating process. In many applications, a generalized linear mixed model (GLMM) is used with a random effect to account for spatial dependence and to provide optimal spatial predictions. Location-specific covariates are often included as fixed effects in a GLMM and may be collinear with the spatial random effect, which can negatively affect inference. We propose a dynamic approach to account for spatial dependence that incorporates scientific knowledge of the spatio-temporal data generating process. Our approach relies on a dynamic spatio-temporal model that explicitly incorporates location-specific covariates. We illustrate our approach with a spatially varying ecological diffusion model implemented using a computationally efficient homogenization technique. We apply our model to understand individual-level and location-specific risk factors associated with chronic wasting disease in white-tailed deer from Wisconsin, USA and estimate the location the disease was first introduced. We compare our approach to several existing methods that are commonly used in spatial statistics. Our spatio-temporal approach resulted in a higher predictive accuracy when compared to methods based on optimal spatial prediction, obviated confounding among the spatially indexed covariates and the spatial random effect, and provided additional information that will be important for containing disease outbreaks.

Wisconsin

Hierarchical animal movement models for population-level inference

New methods for modeling animal movement based on telemetry data are developed regularly. With advances in telemetry capabilities, animal movement models are becoming increasingly sophisticated. Despite a need for population-level inference, animal movement models are still predominantly developed for individual-level inference. Most efforts to upscale the inference to the population level are either post hoc or complicated enough that only the developer can implement the model. Hierarchical Bayesian models provide an ideal platform for the development of population-level animal movement models but can be challenging to fit due to computational limitations or extensive tuning required. We propose a two-stage procedure for fitting hierarchical animal movement models to telemetry data. The two-stage approach is statistically rigorous and allows one to fit individual-level movement models separately, then resample them using a secondary MCMC algorithm. The primary advantages of the two-stage approach are that the first stage is easily parallelizable and the second stage is completely unsupervised, allowing for an automated fitting procedure in many cases. We demonstrate the two-stage procedure with two applications of animal movement models. The first application involves a spatial point process approach to modeling telemetry data, and the second involves a more complicated continuous-time discrete-space animal movement model. We fit these models to simulated data and real telemetry data arising from a population of monitored Canada lynx in Colorado, USA.

Environmetrics

Latent spatial models and sampling design for landscape genetics

We propose a spatially-explicit approach for modeling genetic variation across space and illustrate how this approach can be used to optimize spatial prediction and sampling design for landscape genetic data. We propose a multinomial data model for categorical microsatellite allele data commonly used in landscape genetic studies, and introduce a latent spatial random effect to allow for spatial correlation between genetic observations. We illustrate how modern dimension reduction approaches to spatial statistics can allow for efficient computation in landscape genetic statistical models covering large spatial domains. We apply our approach to propose a retrospective spatial sampling design for greater sage-grouse ( Centrocercus urophasianus ) population genetics in the western United States.

Annals of Applied Statistics

Animal movement constraints improve resource selection inference in the presence of telemetry error

Multiple factors complicate the analysis of animal telemetry location data. Recent advancements address issues such as temporal autocorrelation and telemetry measurement error, but additional challenges remain. Difficulties introduced by complicated error structures or barriers to animal movement can weaken inference. We propose an approach for obtaining resource selection inference from animal location data that accounts for complicated error structures, movement constraints, and temporally autocorrelated observations. We specify a model for telemetry data observed with error conditional on unobserved true locations that reflects prior knowledge about constraints in the animal movement process. The observed telemetry data are modeled using a flexible distribution that accommodates extreme errors and complicated error structures. Although constraints to movement are often viewed as a nuisance, we use constraints to simultaneously estimate and account for telemetry error. We apply the model to simulated data, showing that it outperforms common ad hoc approaches used when confronted with measurement error and movement constraints. We then apply our framework to an Argos satellite telemetry data set on harbor seals ( Phoca vitulina ) in the Gulf of Alaska, a species that is constrained to move within the marine environment and adjacent coastlines.

Ecology

Continuous-time discrete-space models for animal movement

The processes influencing animal movement and resource selection are complex and varied. Past efforts to model behavioral changes over time used Bayesian statistical models with variable parameter space, such as reversible-jump Markov chain Monte Carlo approaches, which are computationally demanding and inaccessible to many practitioners. We present a continuous-time discrete-space (CTDS) model of animal movement that can be fit using standard generalized linear modeling (GLM) methods. This CTDS approach allows for the joint modeling of location-based as well as directional drivers of movement. Changing behavior over time is modeled using a varying-coefficient framework which maintains the computational simplicity of a GLM approach, and variable selection is accomplished using a group lasso penalty. We apply our approach to a study of two mountain lions ( Puma concolor ) in Colorado, USA.

Annals of Applied Statistics

Reconciling resource utilization and resource selection functions

Summary: 1. Analyses based on utilization distributions (UDs) have been ubiquitous in animal space use studies, largely because they are computationally straightforward and relatively easy to employ. Conventional applications of resource utilization functions (RUFs) suggest that estimates of UDs can be used as response variables in a regression involving spatial covariates of interest. 2. It has been claimed that contemporary implementations of RUFs can yield inference about resource selection, although to our knowledge, an explicit connection has not been described. 3. We explore the relationships between RUFs and resource selection functions from a hueristic and simulation perspective. We investigate several sources of potential bias in the estimation of resource selection coefficients using RUFs (e.g. the spatial covariance modelling that is often used in RUF analyses). 4. Our findings illustrate that RUFs can, in fact, serve as approximations to RSFs and are capable of providing inference about resource selection, but only with some modification and under specific circumstances. 5. Using real telemetry data as an example, we provide guidance on which methods for estimating resource selection may be more appropriate and in which situations. In general, if telemetry data are assumed to arise as a point process, then RSF methods may be preferable to RUFs; however, modified RUFs may provide less biased parameter estimates when the data are subject to location error.

Journal of Animal Ecology