Abstract Capture–recapture models provide a statistical framework for estimating demographic parameters from incomplete observation data, where not all individuals in a population are detected during sampling. Assessing the fit of such models is crucial for reliable inference. However, capture–recapture models often involve non‐standard probability distributions, making standard goodness‐of‐fit (GOF) tools difficult to apply and frequently requiring bespoke programming. To address this challenge, we adopt a partial likelihood approach that conditions on the initial capture of individuals and explicitly models the subsequent recapture process. This formulation recasts the problem as a binomial generalized linear model for post‐entry capture events, thereby enabling the use of the widely adopted R‐package DHARMa for simulation‐based GOF assessment of the recapture component in closed population capture–recapture models. We provide practical guidance for implementing the partial likelihood approach to assess GOF in capture–recapture models. The diagnostic capabilities of the DHARMa package are illustrated using both simulated and real ecological datasets. Our results show that DHARMa, within the partial likelihood framework, effectively detects misspecification in the recapture process, including heterogeneity, temporal variation, and Markovian behavioural (trap‐response) effects in (re)capture probabilities. Furthermore, diagnostic power increases with the amount of recapture information, indicating that studies with either extensive sampling effort or high capture probabilities are particularly well suited to this framework.
Zero-inflated count data models are widely used in various fields such as ecology, epidemiology, and transportation, where count data with a large proportion of zeros is prevalent. Despite their widespread use, their theoretical properties have not been extensively studied. This study aims to investigate the impact of ignoring heterogeneity in event count intensity and susceptibility probability on zero-inflated count data analysis within the zero-inflated Poisson framework. To address this issue, we propose a novel conditional likelihood approach that uses positive count data only to estimate event count intensity parameters and develop a consistent estimator for estimating the average susceptibility probability. Our approach is compared with the maximum likelihood approach, and we demonstrate our findings through a comprehensive simulation study and real data analysis. The results can also be extended to zero-inflated binomial and geometric models with similar conclusions. These findings contribute to the understanding of the theoretical properties of zero-inflated count data models and provide a practical approach to handling heterogeneity in such models.
While log-Gaussian Cox process regression models are useful tools for modeling point patterns, they can be technically difficult to fit and require users to learn/adopt bespoke software. We show that, for suitably formatted data, we can actually fit these models using generalized additive model software, via a simple line of code, demonstrated on R by the popular mgcv package. We are able to do this because a common and computationally efficient way to fit a log-Gaussian Cox process model is to use a basis function expansion to approximate the Gaussian random field, as is provided by a generic bivariate smoother over geographic space. We further show that if basis functions are parameterized appropriately then we can estimate parameters in the spatial covariance function for the latent random field using a generalized additive model. We use simulation to show that this approach leads to model fits of comparable quality to state-of-the-art software, often more quickly. But we see the main advance from this work as lowering the technology barrier to spatial statistics for applied researchers, many of whom are already familiar with generalized additive model software.
In regression modelling, measurement error models are often needed to correct for uncertainty arising from measurements of covariates/predictor variables. The literature on measurement error (or errors-in-variables) modelling is plentiful, however, general algorithms and software for maximum likelihood estimation of models with measurement error are not as readily available, in a form that they can be used by applied researchers without relatively advanced statistical expertise. In this study, we develop a novel algorithm for measurement error modelling, which could in principle take any regression model fitted by maximum likelihood, or penalised likelihood, and extend it to account for uncertainty in covariates. This is achieved by exploiting an interesting property of the Monte Carlo Expectation-Maximization (MCEM) algorithm, namely that it can be expressed as an iteratively reweighted maximisation of complete data likelihoods (formed by imputing the missing values). Thus we can take any regression model for which we have an algorithm for (penalised) likelihood estimation when covariates are error-free, nest it within our proposed iteratively reweighted MCEM algorithm, and thus account for uncertainty in covariates. The approach is demonstrated on examples involving generalized linear models, point process models, generalized additive models and capture-recapture models. Because the proposed method uses maximum (penalised) likelihood, it inherits advantageous optimality and inferential properties, as illustrated by simulation. We also study the model robustness of some violations in predictor distributional assumptions. Software is provided as the refitME package on R, whose key function behaves like a refit() function, taking a fitted regression model object and re-fitting with a pre-specified amount of measurement error.
Population size estimation is an important research field in biological sciences. In practice, covariates are often measured upon capture on individuals sampled from the population. However, some biological measurements, such as body weight, may vary over time within a subject's capture history. This can be treated as a population size estimation problem in the presence of covariate measurement error. We show that if the unobserved true covariate and measurement error are both normally distributed, then a naïve estimator without taking into account measurement error will under-estimate the population size. We then develop new methods to correct for the effect of measurement errors. In particular, we present a conditional score and a nonparametric corrected score approach that are both consistent for population size estimation. Importantly, the proposed approaches do not require the distribution assumption on the true covariates; furthermore, the latter does not require normality assumptions on the measurement errors. This is highly relevant in biological applications, as the distribution of covariates is often non-normal or unknown. We investigate finite sample performance of the new estimators via extensive simulated studies. The methods are applied to real data from a capture–recapture study. Supplementary materials accompanying this paper appear on-line.
Site occupancy models are routinely used to estimate the probability of species presence from either abundance or presence–absence data collected across sites with repeated sampling occasions. In the last two decades, a broad class of occupancy models have been developed, but little attention has been given in examining the effects of heterogeneity in parameter estimation. This study focuses on occupancy models where heterogeneity is present in detection intensity and the presence probability. We show that the presence probability will be underestimated if detection heterogeneity is ignored. On the other hand, the behaviour is different if heterogeneity in the presence probability is ignored; notably, an estimate of the average presence probability may be unbiased or overor under-estimated depending on the relationship between detection *Corresponding author Email: wenhan@nchu.edu.tw 1 ar X iv :2 20 4. 00 12 6v 1 [ st at .M E ] 3 1 M ar 2 02 2 and presence probabilities. In addition, when heterogeneity in the detection intensity is related to covariates, we propose a conditional likelihood approach to estimate the detection intensity parameters. This alternative method shares an optimal estimating function property and it ensures robustness against model specification on the presence probability. We then propose a consistent estimator for the average presence probability, provided that the detection intensity component model is correctly specified. We illustrate the bias effects and estimator performance in simulation studies and real data analysis.
Negative binomial modelling is one of the most commonly used statistical tools for analysing count data in ecology and biodiversity research. This is not surprising given the prevalence of overdispersion (i.e., evidence that the variance is greater than the mean) in many biological and ecological studies. Indeed, overdispersion is often indicative of some form of biological aggregation process (e.g., when species or communities cluster in groups). If overdispersion is ignored, the precision of model parameters can be severely overestimated and can result in misleading statistical inference. In this article, we offer some insight as to why the negative binomial distribution is becoming, and arguably should become, the default starting distribution (as opposed to assuming Poisson counts) for analysing count data in ecology and biodiversity research. We begin with an overview of traditional uses of negative binomial modelling, before examining several modern applications and opportunities in modern ecology/biodiversity where negative binomial modelling is playing a critical role, from generalisations based on exploiting its Poisson-gamma mixture formulation in species distribution models and occurrence data analysis, to estimating animal abundance in negative binomial N-mixture models, and biodiversity measures via rank abundance distributions. Comparisons to other common models for handling overdispersion on real data are provided. We also address the important issue of software, and conclude with a discussion of future directions for analysing ecological and biological data with negative binomial models. In summary, we hope this overview will stimulate the use of negative binomial modelling as a starting point for the analysis of count data in ecology and biodiversity studies.
Site occupancy models are routinely used to estimate the probability of species presence from either abundance or presence-absence data collected across sites with repeated sampling occasions. In the last two decades, a broad class of occupancy models has been developed, but little attention has been given to examining the effects of heterogeneity in parameter estimation. This study focuses on occupancy models where heterogeneity is present in detection intensity and the presence probability. We show that the presence probability will be underestimated if detection heterogeneity is ignored. On the other hand, the behavior is different if heterogeneity in the presence probability is ignored; notably, an estimate of the average presence probability may be unbiased or over- or under-estimated depending on the relationship between detection and presence probabilities. In addition, when heterogeneity in the detection intensity is related to covariates, we propose a conditional likelihood approach to estimate the detection intensity parameters. This alternative method shares an optimal estimating function property and it ensures robustness against model specification on the presence probability. We then propose a consistent estimator for the average presence probability, provided that the detection intensity component model is correctly specified. We illustrate the bias effects and estimator performance in simulation studies and real data analysis.
Spatial or temporal clustering commonly arises in various biological and ecological applications, for example, species or communities may cluster in groups. In this paper, we develop a new clustered occurrence data model where presence-absence data are modeled under a multivariate negative binomial framework. We account for spatial or temporal clustering by introducing a community parameter in the model that controls the strength of dependence between observations thereby enhancing the estimation of the mean and dispersion parameters. We provide conditions to show the existence of maximum likelihood estimates when cluster sizes are homogeneous and equal to 2 or 3 and consider a composite likelihood approach that allows for additional robustness and flexibility in fitting for clustered occurrence data. The proposed method is evaluated in a simulation study and demonstrated using forest plot data from the Center for Tropical Forest Science. Finally, we present several examples using multiple visit occupancy data to illustrate the difference between the proposed model and those of N-mixture models.
Multiple imputation and maximum likelihood estimation (via the expectation‐maximization algorithm) are two well‐known methods readily used for analyzing data with missing values. While these two methods are often considered as being distinct from one another, multiple imputation (when using improper imputation) is actually equivalent to a stochastic expectation‐maximization approximation to the likelihood. In this article, we exploit this key result to show that familiar likelihood‐based approaches to model selection, such as Akaike's information criterion (AIC) and the Bayesian information criterion (BIC), can be used to choose the imputation model that best fits the observed data. Poor choice of imputation model is known to bias inference, and while sensitivity analysis has often been used to explore the implications of different imputation models, we show that the data can be used to choose an appropriate imputation model via conventional model selection tools. We show that BIC can be consistent for selecting the correct imputation model in the presence of missing data. We verify these results empirically through simulation studies, and demonstrate their practicality on two classical missing data examples. An interesting result we saw in simulations was that not only can parameter estimates be biased by misspecifying the imputation model, but also by overfitting the imputation model. This emphasizes the importance of using model selection not just to choose the appropriate type of imputation model, but also to decide on the appropriate level of imputation model complexity.
In many applications involving regression analysis, explanatory variables (or covariates) may be imprecisely measured or may contain missing values. Although there exists a vast literature on measurement error modeling to account for errors-in-variables, and on missing data methodology to handle missingness, very few methods have been developed to simultaneously address both. In this paper, we consider likelihood-based multiple imputation to handle missing data, and combine this with two well-known functional measurement error methods: simulation-extrapolation and corrected score. This unified approach has several appealing characteristics: the model fitting procedure is easy to understand and off-the-shelf software can be incorporated into the modeling framework; no calibration data or a validation subset is required in the model fitting procedure; and the missing data component of the proposed approach is likelihood-based which allows standard likelihood machinery. We demonstrate our methods on simulated datasets and apply them to daily ozone pollution measurements in Los Angeles where observed covariates consist of missing data and imprecise measurements. We conclude that the proposed methods substantially reduce bias and mean squared errors in regression coefficients, in comparison to methods that ignore either measurement error or missingness in covariates.
Conducting complete surveys on flora and fauna species within a sampling unit (or quadrat) of interest can be costly, particularly if there are several species in high abundance. A commonly used approach, which aims to reduce time and costs, consists of occurrence data reflecting the status of occupancy of a species– e.g., rather than counting every individual, the survey is stopped as soon as one individual has been observed. Although this approach is cheaper to conduct than a complete survey, some statistical efficiency in model estimators is lost. In this study, we consider occurrence data as a special case of right-censored count data where the collecting process stops until some set threshold on the number of observed individuals is reached. We then propose a new class of regression estimation models for right-censored count data that incorporate information from detection times (or catch effort) collected during sampling. First, we show that incorporating ancillary information in the form of detection times can greatly improve statistical efficiency over, say, right-censored Poisson or negative binomial models. Furthermore, the proposed models retain the same cost-effectiveness as censored-type models. We also consider zero-truncated and zero-inflated models for a variety of count data types. These models can be extended to a more general class of mixed Poisson models. We investigate model performance on simulated data and give two examples consisting of plant abundance data and bat acoustics data. Supplementary materials accompanying this paper appear online.
Zero-truncated data arises in various disciplines where counts are observed but the zero count category cannot be observed during sampling. Maximum likelihood estimation can be used to model these data; however, due to its nonstandard form it cannot be easily implemented using well-known software packages, and additional programming is often required. Motivated by the Rao-Blackwell theorem, we develop a weighted partial likelihood approach to estimate model parameters for zero-truncated binomial and Poisson data. The resulting estimating function is equivalent to a weighted score function for standard count data models, and allows for applying readily available software. We evaluate the efficiency for this new approach and show that it performs almost as well as maximum likelihood estimation. The weighted partial likelihood approach is then extended to regression modelling and variable selection. We examine the performance of the proposed methods through simulation and present two case studies using real data.
In capture–recapture experiments, covariates collected on individuals, such as body weight and length, are often measured imprecisely or are missing at random. Furthermore, the number of recorded covariate measurements collected on each observed individual is usually equal to or less than the individual’s capture frequency. Correcting for multiple error-prone covariate is seldom seen in capture–recapture models and even fewer research have considered cases where individual’s have no measurements at all. In this paper, we develop an unbiased estimating equation using the conditional score within the capture–recapture framework. We then extend this approach to simultaneously account for both measurement error and missing data using two well-known missing data methods: (1) inverse probability weighting; and (2) multiple imputation. These new methods are shown to yield consistent and asymptotically normal estimators, ∗Corresponding author Email: wenhan@nchu.edu.tw 1 Statistica Sinica: Newly accepted Paper (accepted version subject to English editing)
ABSTRACT Multivariate adaptive regression splines (MARS) is a popular nonparametric regression tool often used for prediction and for uncovering important data patterns between the response and predictor variables. The standard MARS algorithm assumes responses are normally distributed and independent, but in this article we relax both of these assumptions by extending MARS to generalized estimating equations. We refer to this MARS-for-GEEs algorithm as “MARGE.” Our algorithm makes use of fast forward selection techniques, such that in the univariate case, MARGE has similar computation speed to a standard MARS implementation. Through simulation we show that the proposed algorithm has improved predictive performance than the original MARS algorithm when using correlated and/or nonnormal response data. MARGE is also competitive with alternatives in the literature, especially for problems with multiple interacting predictors. We apply MARGE to various ecological examples with different data types. Supplementary material for this article is available online.
A presence–absence map consists of indicators of the occurrence or nonoccurrence of a given species in each cell over a grid, without counting the number of individuals in a cell once it is known it is occupied. They are commonly used to estimate the distribution of a species, but our interest is in using these data to estimate the abundance of the species. In practice, certain types of species (in particular flora types) may be spatially clustered. For example, some plant communities will naturally group together according to similar environmental characteristics within a given area. To estimate abundance, we develop an approach based on clustered negative binomial models with unknown cluster sizes. Our approach uses working clusters of cells to construct an estimator which we show is consistent. We also introduce a new concept called super-clustering used to estimate components of the standard errors and interval estimators. A simulation study is conducted to examine the performance of the estimators and they are applied to real data.