Estimation of wildlife population size and spatial dynamics is central to ecology and conservation. Spatial capture-recapture (SCR) models estimate abundance, using detections at sensors such as camera traps, by linking detection probability to the distance between detectors and latent individual activity centres. However, standard SCR formulations assume detections are conditionally independent over time given activity centres without explicitly modelling movement between detection events. This assumption can be problematic when individuals exhibit movement-driven dependence in detections, potentially leading to biased inference on population size. We address this important issue by developing a continuous-time framework that integrates movement into spatial capture-recapture. Individual movement is modelled as a continuous-time Markov chain over a discretised landscape, and detections arise as state-dependent Poisson events. This yields a Markov-modulated marked Poisson process representation, in which detections provide information about an individual's latent location at the time of observation and allow likelihood-based inference in continuous time. We show through simulation studies that ignoring movement-driven dependence can lead to positively biased estimates of population size, whereas the proposed model recovers unbiased estimates and provides additional inference on space use. An application to camera-trap data of American martens illustrates how the framework yields new insights into movement and density. These results demonstrate that explicitly modelling movement is critical for reliable inference in spatial capture-recapture studies.
Survivorship (or selection) bias arises within statistical analyses where the observed data are subject to some underlying selection process prior to entry into the sampled data. For example, within capture-recapture studies, a primary selection mechanism is the survival until initial capture time. The common Cormack-Jolly-Seber model conditions on the first time an individual is observed, leading to potential survivorship bias. However, while the issue of survivorship bias has been well studied in many fields, there has been little exploration within the capture-recapture framework. In particular, we focus on individual (continuous) random effect Cormack-Jolly-Seber models, where it is assumed that individuals have different survival probabilities, specified to be from some common underlying distribution. We discuss the implications of the survivorship bias within the data collection process, and describe a novel modeling approach that accounts for the survivorship bias within an ecologically sensible manner. Using simulated data, we demonstrate the significant impact of ignoring the survivorship bias present in the data. We fit the corrected model to a guillemot data set and demonstrate that even with relatively mild selection bias, the individual heterogeneity variability is substantially underestimated when ignoring this survivorship bias.
Prevalence estimates for opioid use disorder (OUD) are essential for guiding investments in prevention and treatment. Previous estimates have been produced using modeling approaches including capture-recapture (CRC). Empirical evaluations of CRC are needed to assess population size impacts of methodological practices utilized. We linked administrative data sources including New York State (NYS) adults with OUD by indication between January 1-December 31, 2020 and applied Bayesian CRC methods to evaluate how methodological practices impact OUD prevalence: (1) using sensitive versus specific case definition characteristics to identify OUD, (2) including versus excluding inferred referrals from emergency medical services to healthcare settings, and (3) including versus excluding overdose mortality as a source list. We estimated OUD prevalence overall and by demographic characteristics and treatment coverage with medications for OUD (MOUD). Across methodological scenarios, OUD prevalence estimates ranged from 4.6%-9.9%. The largest estimate resulted from a model in which the more sensitive OUD case definition was applied, while the smallest resulted from a model using the more specific case definition. An estimated 4.2% of NYS adults had OUD in 2020, with 20% receiving MOUD. Findings demonstrate that CRC methodological decisions can meaningfully influence OUD prevalence estimates, with recommended methodological practices provided.
We consider the challenge of estimating the model parameters and latent states of general state-space models within a Bayesian framework. We extend the commonly applied particle Gibbs framework by proposing an efficient particle generation scheme for the latent states. The approach efficiently samples particles using an approximate hidden Markov model (HMM) representation of the general state-space model via a deterministic grid on the state space. We refer to the approach as the grid particle Gibbs with ancestor sampling algorithm. We discuss several computational and practical aspects of the algorithm in detail and highlight further computational adjustments that improve the efficiency of the algorithm. The efficiency of the approach is investigated via challenging regime-switching models, including a post-COVID tourism demand model, and we demonstrate substantial computational gains compared to previous particle Gibbs with ancestor sampling methods.
The DECOVID database contains harmonized pseudonymized electronic health record (EHR) data on all adult (>= 18 years old) patients presenting to two large, digitally mature centers in the United Kingdom between 1 January 2020 and 28 February 2021, with follow-up until at least 28 March 2021. The database was originally developed to support the COVID-19 response but is now available via the PIONEER data hub for researchers to explore a wide range of research questions, including exploratory analyses, risk factor assessment, prediction modeling, and comparative effectiveness studies. Raw data were extracted from local EHRs and transformed into a standardized form (Observational Health Data Sciences and Informatics-Common Data Model version 5.3.1). The database includes 165,420 patients across 256,804 hospital presentations. For these patients, highly granular data are available, including patient demographics, longitudinal vital signs, physiology, treatments, laboratory findings, clinical diagnoses, and outcomes. There are 10,030 patients with COVID-19, of whom 1472 died in hospital.
Bayesian inference for statistical models with a hierarchical structure is often characterized by specification of priors for parameters at different levels of the hierarchy. When higher level parameters are functions of the lower level parameters, specifying a prior on the lower level parameters leads to induced priors on the higher level parameters. However, what are deemed uninformative priors for lower level parameters can induce strikingly non-vague priors for higher level parameters. Depending on the sample size and specific model parameterization, these priors can then have unintended effects on the posterior distribution of the higher level parameters. Here we focus on Bayesian inference for the Bernoulli distribution parameter θ which is modeled as a function of covariates via a logistic regression, where the coefficients are the lower level parameters for which priors are specified. A specific area of interest and application is the modeling of survival probabilities in capture-recapture data and occupancy and detection probabilities in presence-absence data. In particular we propose alternative priors for the coefficients that yield specific induced priors for θ. We address three induced prior cases. The simplest is when the induced prior for θ is Uniform(0,1). The second case is when the induced prior for θ is an arbitrary Beta(α, β) distribution. The third case is one where the intercept in the logistic model is to be treated distinct from the partial slope coefficients; e.g., E[θ] equals a specified value on (0,1) when all covariates equal 0. Simulation studies were carried out to evaluate performance of these priors and the methods were applied to a real presence/absence data set and occupancy modelling.
Obtaining abundance and density estimates is a particularly important aspect within wildlife conservation and management. To monitor wildlife populations, the use of motion-sensor camera traps is becoming increasing popular due to its non-invasive nature. However, animal identification is not always feasible in practice due to poor quality images and/or individuals not having uniquely identifiable physical characteristics. Spatially explicit models for unmarked individuals permit the estimation of animal density when individuals cannot be uniquely identified. Due to the structure of these models, a Bayesian super-population (data augmentation) approach is often used to fit the models to data, which involves specifying some reasonably large upper limit for the population. However, this approach presents substantial computational challenges for larger populations, as demonstrated by the motivating dataset relating to barking deer ( Muntiacus muntjak ) collected in Ujung Kulon National Park, Indonesia (with a population size in the low thousands). We develop a new and computationally efficient Bayesian algorithm for fitting the models to data that does not require specifying an upper population limit a priori . We apply the new algorithm to the large barking deer dataset, where the standard super-population approach is computationally expensive, and demonstrate a substantial improvement in computational efficiency.Supplementary material to this paper is provided online.
Obtaining reliable and precise estimates of wildlife species abundance and distribution is essential for the conservation and management of animal populations and natural reserves. Spatial capture-recapture (SCR) models provide estimates of population size and spatial density from data collected from remote sensors such as camera traps. Such data contain spatial correlation between observations of the same individual, which SCR models partly account for through a latent individual-specific activity centre, a location near which the individual is more likely detected. However, SCR models assume that the observations of an individual are independent over time and space, conditional on its activity centre, so that observed sightings at a given time and location do not influence the probability of being seen at future times and/or locations. This assumption is ecologically unrealistic given the smooth movement of animals over space through time. We propose a new continuous-time modelling framework that incorporates both an individual's (latent) activity centre and its (known) previous location and time of detection. By formulating the detections of an individual as an inhomogeneous temporal Poisson process, we develop a model drawing inspiration from the Ornstein-Uhlenbeck process, which is commonly used to model animal movement. Applying our model to a camera-trap survey of American martens, we observe a substantial improvement in model fit and notable differences in the estimated spatial distribution of activity centres. A simulation study shows that standard SCR models can produce substantially biased population estimates when spatio-temporal dependence is ignored, while the memory-based model remains robust. These findings highlight the importance of accounting for memory of previous detections in SCR models to improve ecological interpretation and inference.
Understanding animal movement and behaviour can aid spatial planning and inform conservation management. However, it is difficult to directly observe behaviours in remote and hostile terrain such as the marine environment. Different underlying states can be identified from telemetry data using hidden Markov models (HMMs). The inferred states are subsequently associated with different behaviours, using ecological knowledge of the species. However, the inferred behaviours are not typically validated due to difficulty obtaining 'ground truth' behavioural information. We investigate the accuracy of inferred behaviours by considering a unique data set provided by Joint Nature Conservation Committee. The data consist of simultaneous proxy movement tracks of the boat (defined as visual tracks as birds are followed by eye) and seabird behaviour obtained by observers on the boat. We demonstrate that visual tracking data is suitable for our study. Accuracy of HMMs ranging from 71% to 87% during chick-rearing and 54% to 70% during incubation was generally insensitive to model choice, even when AIC values varied substantially across different models. Finally, we show that for foraging, a state of primary interest for conservation purposes, identified missed foraging bouts lasted for only a few seconds. We conclude that HMMs fitted to tracking data have the potential to accurately identify important conservation-relevant behaviours, demonstrated by a comparison in which visual tracking data provide a 'gold standard' of manually classified behaviours to validate against. Confidence in using HMMs for behavioural inference should increase as a result of these findings, but future work is needed to assess the generalisability of the results, and we recommend that, wherever feasible, validation data be collected alongside GPS tracking data to validate model performance. This work has important implications for animal conservation, where the size and location of protected areas are often informed by behaviours identified using HMMs fitted to movement data.
We consider the challenges that arise when fitting ecological individual heterogeneity models to "large" data sets. In particular, we focus on dividual heterogeneity present in ecological populations within the context of capture-recapture data, although the approach is more widely applicable to more general latent variable models. Within such models the associated likelihood is expressible only as an analytically intractable integral. Common techniques for fitting such models to data include, for example, the use of numerical approximations for the integral or a Bayesian data augmentation approach. However, as the size of the data set increases (i.e., the number of individuals increases), these computational tools may become computationally infeasible. We present an efficient Bayesian model-fitting approach, whereby we initially sample from the posterior distribution of a smaller subsample of the data, before correcting this sample to obtain estimates of the posterior distribution of the full data set using an importance sampling approach. We consider several practical issues, including the subsampling mechanism, computational efficiencies (including the ability to parallelise the algorithm) and combining subsampling estimates using multiple subsampled data sets. We initially demonstrate the feasibility (and accuracy) of the approach via simulated data before considering a challenging real data set of approximately 30,000 guillemots and, using the proposed algorithm, obtain posterior estimates of the model parameters in substantially reduced computational time, compared to the standard Bayesian model-fitting approach.
We constructed annual abundance of a migratory baleen whale at an oceanic stopover site to elucidate temporal changes in Bermuda, an area with increasing anthropogenic activity. The annual abundance of North Atlantic humpback whales visiting Bermuda between 2011 and 2020 was estimated using photo-identification capture-recapture data for 1,204 whales, collected between December 2009 and May 2020. Owing to a sparse data set, we combined a Cormack-Jolly-Seber (CJS) model, fit through maximum likelihood estimation, with a Horvitz-Thompson estimator to calculate abundance and used stratified bootstrap resampling to derive 95% confidence intervals (CI). We accounted for temporal heterogeneity in detection and sighting rates via a catch-effort model and, guided by goodness-of-fit testing, considered models that accounted for transience. A model incorporating modified sighting effort and time-varying transience was selected using (corrected) Akaike’s Information Criterion (AICc). The survival probability of non-transient animals was 0.97 (95% CI 0.91-0.98), which is comparable with other studies. The rate of transience increased gradually from 2011 to 2018, before a large drop in 2019. Abundance varied from 786 individuals (95% CI 593-964) in 2016 to 1,434 (95% CI 924-1,908) in 2020, with a non-significant linear increase across the period and interannual fluctuations. These abundance estimates confirm the importance of Bermuda for migrating North Atlantic humpback whales and should encourage a review of cetacean conservation measures in Bermudian waters, including area-based management tools. Moreover, in line with the time series presented here, regional abundance estimates should be updated across the North Atlantic to facilitate population monitoring over the entire migratory range.
State‐space models are an increasingly common and important tool in the quantitative ecologists’ armoury, particularly for the analysis of time‐series data. This is due to both their flexibility and intuitive structure, describing the different individual processes of a complex system, thus simplifying the model specification step. State‐space models are composed of two processes (a) the system (or state) process that describes the dynamics of the true underlying state of the system over time; and (b) the observation process that links the observed data with the current true state of the system at that time. Specification of the general model structure consists of considering each distinct ecological process within the system and observation processes, which are then automatically combined within the state‐space structure. There is typically a trade‐off between the complexity of the model and the associated model‐fitting process. Simpler model specifications permit the application of simpler model‐fitting tools; whereas more complex model specifications, with nonlinear dynamics and/or non‐Gaussian stochasticity often require more sophisticated model‐fitting algorithms to be applied. We provide a brief overview of general state‐space models before focusing on the different model‐fitting tools available. In particular for different general state‐space model structures we discuss established model‐fitting tools that are available. We also offer practical guidance for choosing a specific fitting procedure.
State-space models (SSMs) are commonly used to model time series data where the observations depend on an unobserved latent process. However, inference on the model parameters of an SSM can be challenging, especially when the likelihood of the data given the parameters is not available in closed-form. One approach is to jointly sample the latent states and model parameters via Markov chain Monte Carlo (MCMC) and/or sequential Monte Carlo approximation. These methods can be inefficient, mixing poorly when there are many highly correlated latent states or parameters, or when there is a high rate of sample impoverishment in the sequential Monte Carlo approximations. We propose a novel block proposal distribution for Metropolis-within-Gibbs sampling on the joint latent state and parameter space. The proposal distribution is informed by a deterministic hidden Markov model (HMM), defined such that the usual theoretical guarantees of MCMC algorithms apply. We discuss how the HMMs are constructed, the generality of the approach arising from the tuning parameters, and how these tuning parameters can be chosen efficiently in practice. We demonstrate that the proposed algorithm using HMM approximations provides an efficient alternative method for fitting state-space models, even for those that exhibit near-chaotic behavior.
1Department of Mathematics and Statistics, Lancaster University, Lancaster, UK; 2School of Mathematics and Maxwell Institute for Mathematical Sciences, University of Edinburgh, Edinburgh, UK; 3Geography, Earth & Environmental Sciences, University of Birmingham, Birmingham, UK; 4Biodiversity, Ecology & Conservation Group, International Institute for Applied Systems Analysis, Vienna, Austria; 5Department of Biosciences, Swansea University, Swansea, UK and 6Centre for Biomathematics, Swansea University, Swansea, UK
The statistical analysis of group studies in neuroscience is particularly challenging due to the complex spatio-temporal nature of the data, its multiple levels and the inter-individual variability in brain responses. In this respect, traditional ANOVA-based studies and linear mixed effects models typically provide only limited exploration of the dynamic of the group brain activity and variability of the individual responses potentially leading to overly simplistic conclusions and/or missing more intricate patterns. In this study we propose a novel Bayesian model-based clustering method for functional data to simultaneously assess group effects and individual deviations over the most important temporal features in the data. To this aim, we develop an innovative multilevel partition prior to model the functional scores of a functional Principal Components decomposition of neuroscientific recordings; this approach returns a thorough exploration of group differences and individual deviations without compromising on the spatio-temporal nature of the data. By means of a simulation study we demonstrate that the proposed model returns correct classification in different clustering scenarios under low and high noise levels in the data. Finally we consider a case study using Electroencephalogram data recorded during an object recognition task where our approach provides new insights into the underlying brain mechanisms generating the data and their variability.
We propose a novel efficient model-fitting algorithm for state space models. State space models are an intuitive and flexible class of models, frequently used due to the combination of their natural separation of the different mechanisms acting on the system of interest: the latent underlying system process; and the observation process. This flexibility, however, often comes at the price of more complicated model-fitting algorithms due to the associated analytically intractable likelihood. For the general case a Bayesian data augmentation approach is often employed, where the true unknown states are treated as auxiliary variables and imputed within the MCMC algorithm. However, standard “vanilla” MCMC algorithms may perform very poorly due to high correlation between the imputed states and/or parameters, often leading to model-specific bespoke algorithms being developed that are nontransferable to alternative models. The proposed method addresses the inefficiencies of traditional approaches by combining data augmentation with numerical integration in a Bayesian hybrid approach. This approach permits the use of standard “vanilla” updating algorithms that perform considerably better than the traditional approach in terms of improved mixing and lower autocorrelation, and has the potential to be incorporated into bespoke model-specific algorithms. To demonstrate the ideas, we apply our semi-complete data augmentation algorithm to different application areas and models, leading to distinct implementation schemes and improved mixing and demonstrating improved mixing of the model parameters. Supplementary materials for this article are available online.
We consider the challenges that arise when fitting complex ecological models to “large” data sets. In particular, we focus on random effect models which are commonly used to describe individual heterogeneity, often present in ecological populations under study. In general, these models lead to a likelihood that is expressible only as an analytically intractable integral. Common techniques for fitting such models to data include, for example, the use of numerical approximations for the integral, or a Bayesian data augmentation approach. However, as the size of the data set increases (i.e. the number of