Estimating species abundance under imperfect detection is a key challenge in biodiversity conservation. The N-mixture model, widely recognized for its ability to distinguish between abundance and individual detection probability without marking individuals, is constrained by its stringent closure assumption, which leads to biased estimates when violated in real-world settings. To address this limitation, we propose an extended framework based on a development of the mixed Gamma-Poisson model, incorporating a community parameter that represents the proportion of individuals consistently present throughout the survey period. This flexible framework generalizes both the zero-inflated type occupancy model and the standard N-mixture model as special cases, corresponding to community parameter values of 0 and 1, respectively. The model's effectiveness is validated through simulations and applications to real-world datasets, specifically with 5 species from the North American Breeding Bird Survey and 46 species from the Swiss Breeding Bird Survey, demonstrating its improved accuracy and adaptability in settings where strict closure may not hold.
Local biodiversity monitoring is important to assess the effects of global change, but also to evaluate the performance of landscape and wildlife protection, since large-scale assessments may buffer local fluctuations, rare species tend to be underrepresented, and management actions are usually implemented on local scales. We estimated population trends of 58 bird species using open-population N-mixture models based on count data in two localities in southeastern Spain, which have been collected according to a citizen science monitoring program (SACRE, Monitoring Common Breeding Birds in Spain) over 21 and 15 years, respectively. We performed different abundance models for each species and study area, accounting for imperfect detection of individuals in replicated counts. After selecting the best models for each species and study area, empirical Bayes methods were used for estimating abundances, which allowed us to calculate population growth rates ( λ ) and finally population trends. We also compared the two local population trends and related them with national and European trends, and species functional traits (phenological status, dietary, and habitat specialization characteristics). Our results showed increasing trends for most species, but a weak correlation between populations of the same species from both study areas. In general, local population trends were consistent with the trends observed at national and continental scales, although contrasting patterns exist for several species, mainly with increasing local trends and decreasing Spanish and European trends. Moreover, we found no evidence of a relationship between population trends and species traits. We conclude that using open-population N-mixture models is an appropriate method to estimate population trends, and that citizen science-based monitoring schemes can be a source of data for such analyses. This modeling approach can help managers to assess the effectiveness of their actions at the local level in the context of global change.
Point counts (PCs) are widely used in biodiversity surveys, but despite numerous advantages, simple PCs suffer from several problems: detectability, and therefore abundance, is unknown; systematic spatiotemporal variation in detectability produces biased inferences, and unknown survey area prevents formal density estimation and scaling-up to the landscape level. We introduce integrated distance sampling (IDS) models that combine distance sampling (DS) with simple PC or detection/nondetection (DND) data and capitalize on the strengths and mitigate the weaknesses of each data type. Key to IDS models is the view of simple PC and DND data as aggregations of latent DS surveys that observe the same underlying density process. This enables estimation of separate detection functions, along with distinct covariate effects, for all data types. Additional information from repeat or time-removal surveys, or variable survey duration, enables separate estimation of the availability and perceptibility components of detectability. IDS models reconcile spatial and temporal mismatches among data sets and solve the above-mentioned problems of simple PC and DND data. To fit IDS models, we provide JAGS code and the new IDS() function in the R package unmarked. Extant citizen-science data generally lack adjustments for detection biases, but IDS models address this shortcoming, thus greatly extending the utility and reach of these data. In addition, they enable formal density estimation in hybrid designs, which efficiently combine distance sampling with distance-free, point-based PC or DND surveys. We believe that IDS models have considerable scope in ecology, management, and monitoring.
Species interactions shape biodiversity patterns, community assemblage, and the dynamics of wildlife populations. Ecological theory posits that the strength of interspecific interactions is fundamentally underpinned by the population sizes of the involved species. Nonetheless, prevalent approaches for modeling species interactions predominantly center around occupancy states. Here, we use simulations to illuminate the inadequacies of modeling species interactions solely as a function of occupancy, as is common practice in ecology. We demonstrate erroneous inference into species interactions due to error in parameter estimates when considering species occupancy alone. To address this critical issue, we propose, develop, and demonstrate an abundance-mediated interaction framework designed explicitly for modeling species interactions involving two or more species from detection/non-detection data. We present Markov chain Monte Carlo (MCMC) samplers tailored for diverse ecological scenarios, including intraguild predation, disease- or predator-mediated competition, and trophic cascades. Illustrating the practical implications of our approach, we compare inference from modeling the interactions in a three-species network involving coyotes (Canis latrans), fishers (Pekania pennanti), and American marten (Martes americana) in North America as a function of occupancy states and as a function of abundance. When modeling interactions as a function of abundance rather than occupancy, we uncover previously unidentified interactions. Our study emphasizes that accounting for abundance-mediated interactions rather than simple co-occurrence patterns can fundamentally alter our comprehension of system dynamics. Through an empirical case study and comprehensive simulations, we demonstrate the importance of accounting for abundance when modeling species interactions, and we present a statistical framework equipped with MCMC samplers to achieve this paradigm shift in ecological research.
The U.S. Fish and Wildlife Service monitors species-specific waterfowl (ducks, seaducks, geese, and brant) harvest through two hunter surveys, one that estimates the total harvest for each waterfowl group, and a second that estimates the species composition of each waterfowl group. Point estimates for species-specific harvest can be computed by multiplying the estimated total harvest by the estimated proportion of the total harvest of each species. However, to date, no uncertainty estimates have been available. Here, we combine these two data sources to provide species-specific harvest estimates at the state and flyway level while characterizing the uncertainty via Bayesian estimation. We take a similar approach to Smith et al. (2022), providing both estimates that treat yearly data as independent and estimates that share information across years via a random walk process. We then discuss the advantages and disadvantages of each approach. ### Competing Interest Statement The authors have declared no competing interest.
Understanding how detection probability varies over time, space, or in response to measurable covariates is important to inform the monitoring and assessment of many species. A standard model to understand detectability – the availability/perception model – admits that detection probability is the composite of two components: availability and perception. Availability is largely affected by environmental and behavioral factors, whereas perception is primarily affected by attributes of individual observers and survey protocols, and thus can potentially be partially controlled by survey design. We designed and implemented a field study to understand the perception component of detection for Eastern box turtles ( Terrapene carolina carolina ) using visual encounter surveys. We obtained and deployed museum specimens of Eastern box turtle shells and subjected them to visual search surveys by observers in realistic field situations. Overall, about 50% of the box turtle shells were detected by observers including, 41.5% in the ‘partially visible’ state and 63% in the ‘fully visible’ state. There were significant differences among observers, which may be due to observer-specific variation in search technique -- the observers varied in how well they achieved the protocol guidance.
Ivory poaching continues to threaten African elephants. We (1) used criminology theory and literature evidence to generate hypotheses about factors that may drive, facilitate or motivate poaching, (2) identified datasets representing these factors, and (3) tested those factors with strong hypotheses and sufficient data quality for empirical associations with poaching. We advance on previous analyses of correlates of elephant poaching by using additional poaching data and leveraging new datasets for previously untested explanatory variables. Using data on 10 286 illegally killed elephants detected at 64 sites in 30 African countries (2002–2020), we found strong evidence to support the hypotheses that the illegal killing of elephants is associated with poor national governance, low law enforcement capacity, low household wealth and health, and global elephant ivory prices. Forest elephant populations suffered higher rates of illegal killing than savannah elephants. We found only weak evidence that armed conflicts may increase the illegal killing of elephants, and no evidence for effects of site accessibility, vegetation density, elephant population density, precipitation or site area. Results suggest that addressing wider systemic challenges of human development, corruption and consumer demand would help reduce poaching, corroborating broader work highlighting these more ultimate drivers of the global illegal wildlife trade.
Classifying species into risk categories is a ubiquitous process in conservation decision-making affecting regulatory procedures, conservation actions, and guiding resource allocation at global, national, and regional scales. However, monitoring programs often do not provide data required for accurate species classification decisions. Misclassification can lead to otherwise preventable species extinctions, undue regulatory burden, poor allocation of limited conservation resources, and can undermine species conservation legislation. We developed a framework that evaluates monitoring designs based on the ability to correctly inform a species classification decision, where minimizing the risk of misclassification is the central objective. We further evaluated monitoring designs by calculating the expected value of information and explored the relationship between statistical power to detect trends and misclassification. Our measure of misclassification risk, which can be tailored to the decision context, clarified the costs of over- and under-protection. High power to detect trends often corresponded to accurate species classification decisions. However, in several scenarios power to detect trends was low but the ability to correctly inform the classification decision was high. The value of information generally increased with monitoring intensity and quantified the tradeoffs between spatial and temporal replication. Our framework allows managers to assess monitoring program performance with direct implications for conservation decision-making. Our framework affords practitioners an opportunity to evaluate the effectiveness of monitoring programs a priori focusing on improving conservation decisions. We demonstrate that prioritizing monitoring to minimize misclassification errors can improve monitoring efficiency and conservation decision-making with considerable practical applications and benefits for species conservation.
Species distribution models (SDMs) are widely applied to understand the processes governing spatial and temporal variation in species abundance and distribution but often do not account for measurement errors such as false negatives and false positives. We describe unmarked, a package for the freely available and open-source R software that provides a complete workflow for modelling species distribution and abundance while explicitly accounting for measurement errors. Here we focus on recent advances in unmarked functionality to support multi-species, multi-state, and multi-season data, as well as support for fitting models with random effects. For illustration, we present an analysis of Acadian Flycatcher Empidonax virescens abundance on Roanoke River National Wildlife Refuge, North Carolina, USA, over 18 years. We found that Acadian Flycatcher abundance was initially greater in hardwood plantation habitat relative to bottomland hardwood forest along river levees but that abundance declined over time in both habitats. We plan for unmarked development to keep pace with advances in hierarchical modelling in ecology, including better handling of continuous-time data from camera trap and automated recording units and integrated models for multiple data streams.
In heterogeneous landscapes, resource selection constitutes a crucial link between landscape and population-level processes such as density. We conducted a non-invasive genetic study of white-tailed deer in southern Finland in 2016 and 2017 using fecal DNA samples to understand factors influencing white-tailed deer density and space use in late summer prior to the hunting season. We estimated deer density as a function of landcover types using a spatial capture-recapture (SCR) model with individual identities established using microsatellite markers. The study revealed second-order habitat selection with highest deer densities in fields and mixed forest, and third-order habitat selection (detection probability) for transitional woodlands (clear-cuts) and closeness to fields. Including landscape heterogeneity improved model fit and increased inferred total density compared with models assuming a homogenous landscape. Our findings underline the importance of including habitat covariates when estimating density and exemplifies that resource selection can be studied using non-invasive methods.
N-mixture models were born in 2004 of the necessity to model animal population size from point counts with imperfect detection of individuals, where capture-recapture methods are infeasible. Initially developed for applications where population size was assumed constant, N-mixture models were extended in 2011 to include population dynamics, allowing application to populations whose size fluctuates during the study. A further extension in 2014 accommodates populations with multiple "states" such as age class or sex. More recent extensions model spatial movement of animals among habitat patches or the spatial spread of infectious disease in a human population. The core idea underlying this class of models is a hierarchical structure, where the observation model is defined conditional on the model for true abundance. This hierarchy allows researchers to incorporate information about observation and abundance processes, while permitting distinct inferences about elements affecting detection and those affecting abundance. Another benefit of the hierarchical approach is the ability to accommodate many existing sampling protocols such as removal sampling and distance sampling. One drawback to N-mixture models is that since they estimate both abundance and detection from replicated but unmarked counts, model parameters may not be clearly identifiable. A second drawback is that when observed counts are large, calculating the N-mixture likelihood is computationally infeasible. This difficulty motivated an approximate likelihood based on the normal approximation to the binomial. The normal approximation provides a diagnostic of parameter estimability based on the closed-form expression of the Fisher information matrix for a multivariate normal likelihood.This article is categorized under:Data: Types and Structure > Image and Spatial DataStatistical Learning and Exploratory Methods of the Data Sciences > Modeling Methods
Aerial count surveys of wildlife populations are a prominent monitoring method for many wildlife species. Traditionally, these surveys utilize human observers to detect, count, and classify observations to species. However, given recent technological advances, many research groups are exploring the combined use of remote sensing and deep learning methods to replace human observers in order to improve data quality and reproducibility, reduce disturbance to wildlife, and increase aircrew safety. Given that deep learning detection and classification are not perfect and that statistical inference from ecological models is generally very sensitive to misclassification, we require study designs and statistical models to accommodate these observation errors. As part of an ongoing effort by the U.S. Fish and Wildlife Service, Bureau of Ocean Energy Management, and U.S. Geological Survey to survey marine birds and other marine wildlife using digital aerial imagery and deep learning object detection and classification, we developed a general hierarchical model for estimating species-specific abundance that accommodates object-level errors in classification. We consider hierarchical deep learning classification at multiple taxonomic levels subject to misclassification, hierarchically-structured human validation data subject to partial and erroneous misclassification, and an image censoring process leading to preferential sampling. We demonstrate that this model can estimate species-specific abundance and habitat relationships without bias when the assumptions are met, and we discuss the plausibility of these assumptions in practice for this study and others like it. Finally, we use this model to demonstrate the relevance of the features of the ecological systems under study to the classification task itself. In models that couple the ecological and classification processes into a single, hierarchical model, the true classes are treated as latent variables to be estimated and are informed by both the classification probability parameters and the ecological parameters that determine the expected frequencies of each class at the level the data are being modeled (e.g., site or site by occasion). We show that ignoring the expected frequencies of each class (when they are imbalanced) can cause correction for misclassification to produce biased parameter estimates, but coupling the ecological and classification models allows for the variability in relative class frequency across space and time due to ecological and sampling conditions to be accommodated with spatial or temporal covariates. As a result, bias is removed, classification is more accurate, and uncertainty is propagated between the ecological and classification models. We therefore argue that ability of deep learning classifiers, and classifiers more generally, to produce reliable ecological inference depends, in part, on the ecological system under study.
Meeting food/wood demands with increasing human population and per-capita consumption is a pressing conservation issue, and is often framed as a choice between land sparing and land sharing. Although most empirical studies comparing the efficacy of land sparing and sharing supported land sparing, land sharing may be more efficient if its performance is tested by rigorous experimental design and habitat structures providing crucial resources for various species-keystone structures-are clearly involved. We launched a manipulative experiment to retain naturally regenerated broad-leaved trees when harvesting conifer plantations in central Hokkaido, northern Japan. We surveyed birds in harvested treatments, unharvested plantation controls, and natural forest references 1-year before the harvest and for three consecutive postharvest years. We developed a hierarchical community model separating abundance and space use (territorial proportion overlapping treatment plots) subject to imperfect detection to assess population consequences of retention harvesting. Application of the model to our data showed that retaining some broad-leaved trees increased the total abundance of forest birds over the harvest rotation cycle. Specifically, a preharvest survey showed that the amount of broad-leaved trees increased forest bird abundance in a concave manner (i.e., in the form of diminishing returns). After harvesting, a small amount of retained broad-leaved trees mitigated negative harvesting impacts on abundance, although retention harvesting reduced the space use. Nevertheless, positive retention effects on the postharvest bird density as the product of abundance and space use exhibited a concave form. Thus, small profit reductions were shown to yield large increases in forest bird abundance. The difference in bird abundance between clearcutting and low amounts of broad-leaved tree retention increased slightly from the first to second postharvesting years. We conclude that retaining a small amount of broad-leaved trees may be a cost-effective on-site conservation approach for the management of conifer plantations. The retention of 20-30 broad-leaved trees per ha may be sufficient to maintain higher forest bird abundance than clearcutting over the rotation cycle. Retention approaches can be incorporated into management systems using certification schemes and best management practices. Developing an awareness of the roles and values of naturally regenerated trees is needed to diversify plantations.
Having an accurate estimate of population size and density is imperative to the conservation of chelonian species and a central objective of many monitoring programs. Capture-recapture and related methods are widely used to obtain information about population size of chelonians. However, classical capture-recapture methods have strict spatial sampling requirements and do not account for lack of geographic closure caused by movement of individuals in and out of the surveyed landscape. Newly developed spatial capture-recapture (SCR) models address these limitations by specification of explicit models for spatial sampling as well as the spatial distribution of individuals in the population. Spatial capture-recapture models have not yet been applied to the study of chelonian populations. Here we demonstrate their application to a population of box turtles in Maryland that has been studied for 75 yr. Results support dramatic declines in population size of box turtles since the 1940s.
Animal movement is a fundamental ecological process affecting the survival and reproduction of individuals, the structure of populations, and the dynamics of communities. Methods to quantify animal movement and spatiotemporal abundances, however, are generally separate and therefore omit linkages between individual-level and population-level processes. We describe an integrated spatial capture-recapture (SCR) movement model to jointly estimate (1) the number and distribution of individuals in a defined spatial region and (2) movement of those individuals through time. We applied our model to a study of polar bears (Ursus maritimus) in a 28,125 km(2) survey area of the eastern Chukchi Sea, USA in 2015 that incorporated capture-recapture and telemetry data. In simulation studies, the model provided unbiased estimates of movement, abundance, and detection parameters using a bivariate normal random walk and correlated random walk movement process. Our case study provided detailed evidence of directional movement persistence for both male and female bears, where individuals regularly traversed areas larger than the survey area during the 36-day study period. Scaling from individual- to population-level inferences, we found that densities varied from <0.75 bears/625 km(2) grid cell/day in nearshore cells to 1.6-2.5 bears/grid cell/day for cells surrounded by sea ice. Daily abundance estimates ranged from 53 to 69 bears, with no trend across days. The cumulative number of unique bears that used the survey area increased through time due to movements into and out of the area, resulting in an estimated 171 individuals using the survey area during the study (95% credible interval 124-250). Abundance estimates were similar to a previous multiyear integrated population model using capture-recapture and telemetry data (2008-2016; Regehr et al., Scientific Reports 8:16780, 2018). Overall, the SCR-movement model successfully quantified both individual- and population-level space use, including the effects of landscape characteristics on movement, abundance, and detection, while linking the movement and abundance processes to directly estimate density within a prescribed spatial region and temporal period. Integrated SCR-movement models provide a generalizable approach to incorporate greater movement realism into population dynamics and link movement to emergent properties including spatiotemporal densities and abundances.
Hidden Markov models (HMMs) are broadly applicable hierarchical models that derive their utility from separating state processes from observation processes yielding the data. Multistate models such as mark–recapture and dynamic multistate occupancy models are HMMs frequently used in ecology. In their early formulations, states, such as pathogen infection status, were assumed to be perfectly observed without ambiguity. However, state uncertainty is a pervasive feature of many ecological studies, and multievent models were developed to explicitly account for it. We developed a novel extended multievent mark–recapture model that incorporates state uncertainty at multiple levels of detection. Using a disease‐structured example, both false negative and false positive state assignment errors are modelled at two levels of state assignment—the pathogen sampling process and the diagnostic process that samples are subjected to. We additionally describe methods to jointly model infection intensity to integrate heterogeneity in ecological parameters, such as mortality and infection dynamics, and the pathogen detection processes. We provide code to simulate and analyse datasets with various underlying ecological processes and fit our model to a mark–recapture dataset of Mixophyes fleayi (Fleay's barred frog) infected with the amphibian chytrid fungus ( Batrachochytrium dendrobatidis , Bd ). In our case study, we found evidence for various state assignment errors: the sampling protocol performed poorly in detecting Bd , pathogen detection was highly dependent on infection intensity and false positives were non‐negligible. Incorporating state uncertainty yielded significantly higher estimates of infection prevalence and 4–5 times lower rates of infection state transitions compared to those obtained from a traditional multistate model. Our results highlight that incorporating state assignment errors improves inference on the ecological process, especially when sensitivity and specificity of the state assignment processes are low. The general model structure can be applied to other HMMs, providing a foundation for modelling state uncertainty in related models. For disease‐structured multistate models, we recommend conducting robust design surveys and collecting samples during each capture event to facilitate incorporating pathogen detection errors.
Poaching is a global driver of wildlife population decline, including inside protected areas (PAs). Reducing poaching requires an understanding of its cryptic drivers and accurately quantifying poaching scales and intensity. There is little quantification of how poaching is affected by law enforcement intensity (e.g., ranger stations) versus economic factors (e.g., unemployment), while simultaneously accounting for imperfect detection. Using extensive data of poaching events (i.e., seizures) and censuses of nine ungulate species across the PAs and unprotected lands of Iran from 2010 to 2018, we developed a single-visit hierarchical (N-mixture) model to accurately estimate annual poaching of Iranian ungulates and to differentiate between social and ecological effects on annual poaching intensity. We found that poaching detectability increased with numbers of ranger stations. A recent surge in poaching (2013–2018) coincides with rising unemployment rate. We estimated that 19,727 ungulates (95% confidence interval 11,178–36,195) were poached across the country during 2010–2018. Poaching intensity was positively related to unemployment rate, road density, and ungulate abundance. Our simulations demonstrated that the Poisson and Negative binomial N-mixture models had adequate performance when the conditions of Sólymos et al. (2012) were satisfied, in particular, when at least one covariate is unique to both the detection and abundance parts of the model. Overall, we suggest that single-visit models offer unique insights into understanding the link between poaching intensity, economic conditions, and law enforcement in large-scale landscapes while accounting for imperfect detection of poaching events.
Large‐scale, long‐term biodiversity monitoring is essential to conservation, land management and identifying threats to biodiversity. However, multispecies surveys are prone to various types of observation error, including false‐positive/false‐negative detection and misclassification, where a species is thought to have been encountered but not correctly identified. Previous methods assume an imperfect classifier produces species‐level classifications, but in practice, particularly with human observers, we may end up with extraspecific classifications including ‘unknown’, morphospecies designations and taxonomic identifications coarser than species. Disregarding these types of species misclassification in biodiversity monitoring datasets can bias estimates of ecologically important quantities such as demographic rates, occurrence and species richness. Here we present a joint classification‐occupancy model that accounts for species non‐detection and misclassification. Our framework accommodates extinction and colonization dynamics, allows for additional uncertain ‘morphospecies’ designations and makes use of individual specimens with known species identities in a semi‐supervised setting. We compare the performance of our model to a classification‐only model that discards information about occupancy and encounter rate. We illustrate our model with an empirical case study of the carabid beetle (Carabidae) community at the National Ecological Observatory Network Niwot Ridge Mountain Research Station, near Boulder, CO, USA. We also use simulations to evaluate model performance through validation metrics where varying fractions of the data are confirmed. The model supported imperfect classifier accuracy and favoured certain true species classifications strongly for some morphospecies. The model outperformed (e.g. precision) the reduced model that discarded occupancy information, and these differences were most pronounced for abundant species. Spatial and temporal dynamics from modelled occupancy and encounter rates may inform species misclassification probability, but this idea has not yet been tested. Our statistical framework explores this opportunity, and can be applied to datasets with imperfect species detection and classification, limited verification data and non‐species classifications.