In ecology and evolution, meta-analysis is an important tool to synthesise findings across separate studies and identify sources of heterogeneity. However, ecological and evolutionary data often exhibit complex dependence structures, such as shared sources of variation within studies, phylogenetic relationships and hierarchical sampling designs. Recent statistical advancements offer approaches for handling such complexities in dependence, yet these methods remain under-utilised or unfamiliar to ecologists and evolutionary biologists. We conducted extensive simulations to evaluate modelling approaches for handling dependence in effect sizes and sampling errors in ecological and evolutionary meta-analyses. We assessed the performance of multilevel models, incorporating an assumed sampling error variance-covariance (VCV) matrix (which account for within-study correlation), cluster robust variance estimation (CRVE) methods and their combination across different true within-study correlations. Finally, we showcased the applications of these models in two case studies of published meta-analyses. Multilevel models produced unbiased regression coefficient estimates, and when a sampling VCV matrix was used, it provided accurate random effect variance components estimates within and among studies. However, the latter had no impact on regression coefficient estimates if the model was misspecified. In simulations involving phylogenetic multilevel meta-analysis, models using CRVE methods generated narrower confidence intervals and lower coverage rates than the nominal expectations. The case study results showed the importance of considering a sampling error VCV matrix to improve the model fit. Our results provide clear modelling recommendations for ecologists and evolutionary biologists conducting meta-analyses. To improve the precision of variance component estimates, we recommend constructing a VCV matrix that accounts for dependencies in sampling errors within studies. Although CRVE methods provide robust inference under certain conditions, we caution against their use with crossed random effects, such as phylogenetic multilevel meta-analyses, as CRVE methods currently do not account for multi-way clustering and may inflate Type I error rates. Finally, we recommend using multilevel meta-analytic models to account for heterogeneity at all relevant hierarchical levels and to follow guidance on inference methods to ensure accurate coverage of the overall mean.
Muscle volume must increase substantially during childhood growth to generate the power required to propel the growing body. One unresolved but fundamental question about childhood muscle growth is whether muscles grow at equal rates; that is, if muscles grow in synchrony with each other. In this study, we used magnetic resonance imaging (MRI) and advances in artificial intelligence methods (deep learning) for medical image segmentation to investigate whether human lower leg muscles grow in synchrony. Muscle volumes were measured in 10 lower leg muscles in 208 typically developing children (eight infants aged less than 3 months and 200 children aged 5 to 15 years). We tested the hypothesis that human lower leg muscles grow synchronously by investigating whether the volume of individual lower leg muscles, expressed as a proportion of total lower leg muscle volume, remains constant with age. There were substantial age-related changes in the relative volume of most muscles in both boys and girls (p < 0.001). This was most evident between birth and five years of age but was still evident after five years. The medial gastrocnemius and soleus muscles, the largest muscles in infancy, grew faster than other muscles in the first five years. The findings demonstrate that muscles in the human lower leg grow asynchronously. This finding may assist early detection of atypical growth and allow targeted muscle-specific interventions to improve the quality of life, particularly for children with neuromotor conditions such as cerebral palsy.
Abstract Integrated distribution models (IDMs) predict where species might occur using data from multiple sources, a technique thought to be especially useful when data from any individual source are scarce. Recent advances allow us to fit such models with latent terms to account for dependence within and between data sources, but they are computationally challenging to fit. We propose a fast new methodology for fitting integrated distribution models using presence/absence and presence‐only data, via a spatial random effects approach combined with automatic differentiation. We have written an R package (called scampr) for straightforward implementation of our approach. We use simulation to demonstrate that our approach has comparable performance to INLA—a common framework for fitting IDMs—but with computation times up to an order of magnitude faster. We also use simulation to look at when IDMs can be expected to outperform models fitted to a single data source, and find that the amount of benefit gained from using an IDM is a function of the relative amount of additional information available from incorporating a second data source into the model. We apply our method to predict 29 plant species in NSW, Australia, and find particular benefit in predictive performance when data from a single source are scarce and when compared to models for presence‐only data. Our faster methods of fitting IDMs make it feasible to more deeply explore the model space (e.g. comparing different ways to model latent terms), and in future work, to consider extensions to more complex models, for example the multi‐species setting.
Abstract Repeat measurement surveys of tree size are used in forests to estimate growth behaviour, biomass and population dynamics. Although size is measured with error and individuals vary in their growth trajectories, current size‐based growth modelling approaches do not usually or fully account for both of these features, and therefore under‐utilise available data. We present a new method that leverages the auto‐correlation structure of repeat surveys into a hierarchical Bayesian longitudinal growth model. This new structure allows users to correct for measurement error and capture individual‐level variation in growth trajectories and parameters. To demonstrate the new method we applied it to a sample of tropical tree survey data from long‐term monitoring sites at Barro Colorado Island. We were able to reduce estimated error in size and growth, and extract individual‐ and population‐level growth parameter estimates. We used simulation to evaluate the ability of the new method to improve estimates of growth rate and size, and estimate individual and species‐level parameters. Our method substantially improved the root mean squared error (RMSE) for growth by an average of 61% compared to existing approaches using pairwise differences and reduced RMSE in estimated size RMSE compared to ‘observed’ values in simulated data. Better numerical integration methods (Runge–Kutta fourth order in comparison to Euler and midpoint) provided better estimates of parameters, but did not improve the estimation of size and growth. The choice of a positive growth function eliminated all negative increments without data exclusion. Overall, this study shows how we can gain new and improved insights on growth, using repeat forest surveys. Our new method offers improved biomass dynamics estimation through reduced error in sizes over time, coupled with novel information about within‐species variation in growth behaviour that is inaccessible with species average models, such as individual parameters for the growth function which allows for relationships between parameters to be considered for the first time.
While log-Gaussian Cox process regression models are useful tools for modeling point patterns, they can be technically difficult to fit and require users to learn/adopt bespoke software. We show that, for suitably formatted data, we can actually fit these models using generalized additive model software, via a simple line of code, demonstrated on R by the popular mgcv package. We are able to do this because a common and computationally efficient way to fit a log-Gaussian Cox process model is to use a basis function expansion to approximate the Gaussian random field, as is provided by a generic bivariate smoother over geographic space. We further show that if basis functions are parameterized appropriately then we can estimate parameters in the spatial covariance function for the latent random field using a generalized additive model. We use simulation to show that this approach leads to model fits of comparable quality to state-of-the-art software, often more quickly. But we see the main advance from this work as lowering the technology barrier to spatial statistics for applied researchers, many of whom are already familiar with generalized additive model software.
Abstract Simulation studies are essential tools to assess statistical methods. Functioning as controlled experiments, simulations generate data from known underlying processes. However, unclear or incomplete reporting of simulation studies can impact their interpretability and reproducibility, potentially leading to the misuse of statistical methods. While Morris et al. (2019, Stat Med, 38, p. 2074) recently provided guidance on the planning and conduct of simulation studies for statistical method evaluation, there is currently no comprehensive set of reporting guidelines in ecology and evolutionary biology. Here, we propose 11 reporting items for statistical simulation studies extending on Morris and colleagues' guidance. These items span across three stages: planning, coding and analysis. We also clarify the terminology related to statistical components and the broad purposes of statistical simulation studies. To highlight our proposed reporting items with current practices, we surveyed 100 articles in ecology and evolution journals that included a simulation study evaluating a statistical method. Our survey found room for improvement in more transparent reporting to ensure clear evaluation of statistical methods. Most notably, only a small proportion of articles reported a Monte Carlo uncertainty (17%; 17 out of 98), and 32% (32 out of 100) articles did not provide code. Beyond the proposed reporting items, we discuss the benefits of open science tools to enhance the reproducibility of simulation studies. Specifically, we propose the registration of statistical simulation studies to enhance planning, reporting and collaboration. We aim to instigate discussions to improve the reporting of simulation studies for statistical method research. The reporting items we propose, along with open science tools, serve as a template for developing standards and guidelines for simulation studies evaluating statistical methods. These reporting guidelines will help enhance reproducibility and indirectly encourage more consideration in the design and conduct of these simulation studies.
In regression modelling, measurement error models are often needed to correct for uncertainty arising from measurements of covariates/predictor variables. The literature on measurement error (or errors-in-variables) modelling is plentiful, however, general algorithms and software for maximum likelihood estimation of models with measurement error are not as readily available, in a form that they can be used by applied researchers without relatively advanced statistical expertise. In this study, we develop a novel algorithm for measurement error modelling, which could in principle take any regression model fitted by maximum likelihood, or penalised likelihood, and extend it to account for uncertainty in covariates. This is achieved by exploiting an interesting property of the Monte Carlo Expectation-Maximization (MCEM) algorithm, namely that it can be expressed as an iteratively reweighted maximisation of complete data likelihoods (formed by imputing the missing values). Thus we can take any regression model for which we have an algorithm for (penalised) likelihood estimation when covariates are error-free, nest it within our proposed iteratively reweighted MCEM algorithm, and thus account for uncertainty in covariates. The approach is demonstrated on examples involving generalized linear models, point process models, generalized additive models and capture-recapture models. Because the proposed method uses maximum (penalised) likelihood, it inherits advantageous optimality and inferential properties, as illustrated by simulation. We also study the model robustness of some violations in predictor distributional assumptions. Software is provided as the refitME package on R, whose key function behaves like a refit() function, taking a fitted regression model object and re-fitting with a pre-specified amount of measurement error.
The accurate extraction of species-abundance information from DNA-based data (metabarcoding, metagenomics) could contribute usefully to diet analysis and food-web reconstruction, the inference of species interactions, the modelling of population dynamics and species distributions, the biomonitoring of environmental state and change, and the inference of false positives and negatives. However, multiple sources of bias and noise in sampling and processing combine to inject error into DNA-based data sets. To understand how to extract abundance information, it is useful to distinguish two concepts. (i) Within-sample across-species quantification describes relative species abundances in one sample. (ii) Across-sample within-species quantification describes how the abundance of each individual species varies from sample to sample, such as over a time series, an environmental gradient or different experimental treatments. First, we review the literature on methods to recover across-species abundance information (by removing what we call "species pipeline biases") and within-species abundance information (by removing what we call "pipeline noise"). We argue that many ecological questions can be answered with just within-species quantification, and we therefore demonstrate how to use a "DNA spike-in" to correct for pipeline noise and recover within-species abundance information. We also introduce a model-based estimator that can be used on data sets without a physical spike-in to approximate and correct for pipeline noise.
Log-Gaussian Cox processes (LGCPs) offer a framework for regression-style modeling of point patterns that can accommodate spatial latent effects. These latent effects can be used to account for missing predictors or other sources of clustering that could not be explained by a Poisson process. Fitting LGCP models can be challenging because the marginal likelihood does not have a closed form and it involves a high dimensional integral to account for the latent Gaussian field. We propose a novel methodology for fitting LGCP models that addresses these challenges using a combination of variational approximation and reduced rank interpolation. Additionally, we implement automatic differentiation to obtain exact gradient information, for computationally efficient optimization and to consider integral approximation using the Laplace method. We demonstrate the method’s performance through both simulations and a real data application, with promising results in terms of computational speed and accuracy compared to that of existing approaches. Supplementary material for this article is available online.
Residual plots are often used to interrogate regression model assumptions, but interpreting them requires an understanding of how much sampling variation to expect when assumptions are satisfied. In this article, we propose constructing global envelopes around data (or around trends fitted to data) on residual plots, exploiting recent advances that enable construction of global envelopes around functions by simulation. While the proposed tools are primarily intended as a graphical aid, they can be interpreted as formal tests of model assumptions, which enables the study of their properties via simulation experiments. We considered three model scenarios-fitting a linear model, generalized linear model or generalized linear mixed model-and explored the power of global simulation envelope tests constructed around data on quantile-quantile plots, or around trend lines on residual versus fits plots or scale-location plots. Global envelope tests compared favorably to commonly used tests of assumptions at detecting violations of distributional and linearity assumptions. Freely available R software (ecostats::plotenvelope) enables application of these tools to any fitted model that has methods for the simulate, residuals and predict functions.
Sample size estimation through power analysis is a fundamental tool in planning an ecological study, yet there are currently no well‐established procedures for when multivariate abundances are to be collected. A power analysis procedure would need to address three challenges: designing a parsimonious simulation model that captures key community data properties; measuring effect size in a realistic yet interpretable fashion; and ensuring computational feasibility when simulation is used both for power estimation and significance testing. Here, we propose a power analysis procedure that addresses these three challenges by: using for simulation a Gaussian copula model with factor analytical structure, fitted to pilot data; assuming a common effect size across all taxa, but applied in different directions according to expert opinion (to “increaser”, “decreaser” or “no effect” taxa); using a critical value approach to estimate power, which reduces computation time by a factor of 500 (if we would otherwise use 999 resamples to estimate each p ‐value) with minor loss of accuracy. The procedure is demonstrated on pilot data from fish assemblages in a restoration study, where it was found that the planned study design would only be capable of detecting relatively large effects (change in abundance by a factor of 1.7 or more). The methods outlined in this paper are available in accompanying R software (the ecopower package), which allows researchers with pilot data to answer a wide range of design questions to assist them in planning their studies.
1. We introduce community-level basis function models (CBFMs) as an approach for spatiotemporal joint distribution modelling. CBFMs can be viewed as related to spatiotemporal latent variable models, where the latent variables are replaced by a set of pre-specified spatiotemporal basis functions which are common across species. 2. In a CBFM, the coefficients that link the basis functions to each species are treated as random slopes. As such, the CBFM can be formulated to have a similar structure to a generalised additive model. This allows us to adapt existing techniques to fit CBFMs efficiently. 3. CBFMs can be used for a variety of reasons, such as inferring patterns of habitat use in space and time, understanding how residual covariation between species varies spatially and/or temporally, and spatiotemporal predictions of species-and community-level quantities. 4. A simulation study and an application to data from a bottom trawl survey conducted across the U. S. Northeast shelf show that CBFMs can achieve similar and sometimes better predictive performance compared to existing approaches for spatiotemporal joint species distribution modelling, while being computationally more scalable.
Diffusion tensor magnetic resonance imaging (DT-MRI) was used to investigate the three-dimensional architecture and diffusion properties of medial gastrocnemius muscles in living human infants aged 2-3 months. Mean muscle volume, physiological cross-sectional area and fascicle length in infants were 1.8%, 3.8% and 47.2% of values previously obtained in 8 adult muscles. Radial diffusivity in infant muscle was half that in adult muscle, presumably because infant muscle fibres have much smaller transverse dimensions.
New technologies for acquiring biological information such as eDNA, acoustic or optical sensors, make it possible to generate spatial community observations at unprecedented scales. The potential of these novel community data to standardize community observations at high spatial, temporal, and taxonomic resolution and at large spatial scale ('many rows and many columns') has been widely discussed, but so far, there has been little integration of these data with ecological models and theory. Here, we review these developments and highlight emerging solutions, focusing on statistical methods for analyzing novel community data, in particular joint species distribution models; the new ecological questions that can be answered with these data; and the potential implications of these developments for policy and conservation.
The life span of leaves increases with their mass per unit area (LMA). It is unclear why. Here, we show that this empirical generalization (the foundation of the worldwide leaf economics spectrum) is a consequence of natural selection, maximizing average net carbon gain over the leaf life cycle. Analyzing two large leaf trait datasets, we show that evergreen and deciduous species with diverse construction costs (assumed proportional to LMA) are selected by light, temperature, and growing-season length in different, but predictable, ways. We quantitatively explain the observed divergent latitudinal trends in evergreen and deciduous LMA and show how local distributions of LMA arise by selection under different environmental conditions acting on the species pool. These results illustrate how optimality principles can underpin a new theory for plant geography and terrestrial carbon dynamics.
Unmeasured or latent variables are often the cause of correlations between multivariate measurements, which are studied in a variety of fields such as psychology, ecology, and medicine. For Gaussian measurements, there are classical tools such as factor analysis or principal component analysis with a well-established theory and fast algorithms. Generalized Linear Latent Variable models (GLLVMs) generalize such factor models to non-Gaussian responses. However, current algorithms for estimating model parameters in GLLVMs require intensive computation and do not scale to large datasets with thousands of observational units or responses. In this article, we propose a new approach for fitting GLLVMs to high-dimensional datasets, based on approximating the model using penalized quasi-likelihood and then using a Newton method and Fisher scoring to learn the model parameters. Computationally, our method is noticeably faster and more stable, enabling GLLVM fits to much larger matrices than previously possible. We apply our method on a dataset of 48,000 observational units with over 2,000 observed species in each unit and find that most of the variability can be explained with a handful of factors. We publish an easy-to-use implementation of our proposed fitting algorithm.
This book on ecology introduces modern tools for data analysis, particularly with extensions to multivariate analyses.
Visualising data is a key step in data analysis, allowing researchers to find patterns, and assess and communicate the results of statistical modelling. In ecology, visualisation is often challenging when there are many variables (often for different species or other taxonomic groups) and they are not normally distributed (often counts or presence–absence data). Ordination is a common and powerful way to overcome this hurdle by reducing data from many response variables to just two or three, to be easily plotted. Ordination is traditionally done using dissimilarity‐based methods, most commonly non‐metric multidimensional scaling (nMDS). In the last decade, however, model‐based methods for unconstrained ordination have gained popularity. These are primarily based on latent variable models, with latent variables estimating the underlying, unobserved ecological gradients. Despite some major benefits, a drawback of model‐based ordination methods is their speed, as they typically take much longer to return a result than dissimilarity‐based methods, especially for large sample sizes. We introduce copula ordination, a new, scalable model‐based approach to unconstrained ordination. This method has all the desirable properties of model‐based ordination methods, with the added advantage that it is computationally far more efficient. In particular, simulations show copula ordination is an order of magnitude faster than current model‐based methods, and can even be faster than nMDS for large sample sizes, while being able to produce similar ordination plots and trends as these methods.