Forecasting ecosystem changes due to disturbances or conservation interventions is essential to improve ecosystem management and anticipate unintended consequences of conservation decisions. Mathematical models allow practitioners to understand the potential effects and unintended consequences via simulation. However, calibrating these models is often challenging due to a paucity of appropriate ecological data. Ensemble ecosystem modelling (EEM) is a quantitative method used to parameterize models from theoretical ecosystem features rather than data. Two approaches have been considered to find parameter values satisfying those features: a standard accept–reject algorithm, appropriate for small ecosystem networks, and a sequential Monte Carlo (SMC) algorithm that is more computationally efficient for larger ecosystem networks. In practice, using SMC for EEM generation requires advanced statistical and mathematical knowledge, as well as strong programming skills, which might limit its uptake. In addition, current EEM approaches have been developed for only one model structure (generalised Lotka–Volterra). To facilitate the usage of EEM methods, we introduce EEMtoolbox, an R package for calibrating quantitative ecosystem models. Our package allows the generation of parameter sets satisfying ecosystem features by using either the standard accept–reject algorithm or the novel SMC procedure. Our package extends the existing EEM methodology, originally developed for the generalised Lotka–Volterra model, to two additional model structures (the multispecies Gompertz and the Bimler–Baker model) and additionally allows users to define their own model structures. We demonstrate the usage of EEMtoolbox by simulating changes in species abundance immediately after the release of the sihek ( Todiramphus cinnamominus , extinct‐in‐the‐wild species) on Palmyra Atoll in the Pacific Ocean. With its simple interface, our package facilitates straightforward generation of EEM parameter sets, thus unlocking advanced statistical methods supporting conservation decisions using ecosystem network models.
Mathematical models connect theory with the real world through data, enabling us to interpret, understand, and predict complex phenomena. However, scientific knowledge often extends beyond what can be empirically measured, offering valuable insights into complex and uncertain systems. Here, we introduce a statistical framework for calibrating mathematical models using non-empirical information. Through examples in ecology, biology, and medicine, we demonstrate how expert knowledge, scientific theory, and qualitative observations can meaningfully constrain models. In each case, these non-empirical insights guide models toward more realistic dynamics and more informed predictions than empirical data alone could achieve. Now, our understanding of the systems represented by mathematical models is not limited by the data that can be obtained; they instead sit at the edge of scientific understanding.
Woodstock and Harris raise concerns about generalisations in our paper 'Calibrated Ecosystem Models Cannot Predict the Consequences of Conservation Management Decisions'. They implied that our conclusions are invalidated by the exclusion of density dependent compensatory feedback and the existence of automated calibration approaches with overfitting penalties-claims that overlook the broader structural issues we identified. Here, we explore their key criticism in more detail, clarify our position and explain why we continue to stand by our findings.
The effectiveness of natural hazard warnings relies on transforming the available data into actionable knowledge for the public. Real-time warning communication and emergency response therefore need to be evaluated from a data-driven perspective. However, gaps exist between established data science best practices and their application in supporting natural hazard warnings. This perspective reviews existing data-driven approaches, highlighting limitations in hazard and impact warnings. Four main themes emerge for enhancing warning communication and supporting decision-making: (1) identifying data-barriers to effective warnings, (2) applying best-practice principles in visualizing warnings, (3) utilizing novel data for more localized forecasts and warnings, and (4) improving data-driven decision-making using uncertainty. These themes are illustrated using the extensive flooding in Australia in 2022 as a case study. This perspective reveals opportunities for improving the efficacy of natural hazard warnings using data science, and the collaborative potential between the data science and natural hazards communities.
In ecological and environmental contexts, management actions must sometimes be chosen urgently. Value of information (VoI) analysis provides a quantitative toolkit for projecting the improved management outcomes expected after making additional measurements. However, traditional VoI analysis reports metrics as expected values (i.e.risk-neutral). This can be problematic because expected values hide uncertainties in projections. The true value of a measurement will only be known after the measurement’s outcome is known, leaving large uncertainty in the measurement’s value before it is performed. As a result, the expected value metrics produced in traditional VoI analysis may not align with the priorities of a risk-averse decision-maker who wants to avoid low-value measurement outcomes. In the present work, we introduce four new VoI metrics that can address a decision-maker’s risk-aversion to different measurement outcomes. We demonstrate the benefits of the new metrics with two ecological case studies for which traditional VoI analysis has been previously applied. In the first case study concerning a test for disease presence at a potential frog translocation site, traditional VoI analysis predicts the test yields an additional expected gain of approximately 10 frogs. However, our new VoI metrics also highlight a 40% risk that the test is valueless; this knowledge may deter a risk-averse decision-maker from doing the test. In the second case study concerning the design of a trial release prior to a large-scale turtle reintroduction, traditional and new VoI metrics have consistent predictions of which design to choose. However, whilst the best trial release design will increase expected turtle survival in the wild by only 3%, the new VoI metrics find that this trial design has a 94% probability of improving the design of the large-scale turtle reintroduction. Using the new metrics, we also demonstrate a clear mathematical link between the often-separated environmental decision-making disciplines of VoI and optimal design of experiments. This mathematical link has the potential to catalyse future collaborations between ecologists and statisticians to work together to quantitatively address environmental decision-making questions of fundamental importance. Overall, the introduced VoI metrics complement existing metrics to provide decision-makers with a comprehensive view of the value of, and risks associated with, a proposed monitoring or measurement activity. This is critical for improved environmental outcomes when decisions must be urgently made.
Quantitative population modelling is an invaluable tool for identifying the cascading effects of conservation on an ecosystem. When population data from monitoring programs is not available, deterministic ecosystem models have often been calibrated using the theoretical assumption that ecosystems have a stable, coexisting equilibrium. However, a growing body of literature suggests these theoretical assumptions are inappropriate for conservation contexts. Here, we develop an alternative for data-free population modelling that relies on expert-elicited knowledge of species populations. Our new Bayesian algorithm systematically removes model parameters that lead to impossible predictions, as defined by experts, without incurring excessive computational costs. We demonstrate our framework on an ordinary differential equation model by limiting predicted population sizes and their ability to change rapidly, utilising readily available knowledge from field observations and experts rather than relying on theoretical ecosystem properties. Our results show that using only coexistence and stability requirements can lead to unrealistic population dynamics, which can be avoided by switching to expert-derived information. We demonstrate how this change can dramatically impact population predictions, expected responses to management, conservation decision-making, and long-term ecosystem behaviour. Without data, we argue that field observations and expert knowledge are more trustworthy for representing ecosystems observed in nature, improving the precision and confidence in predictions.
The potential effects of conservation actions on threatened species can be predicted using ensemble ecosystem models by forecasting populations with and without intervention. These model ensembles commonly assume stable coexistence of species in the absence of available data. However, existing ensemble-generation methods become computationally inefficient as the size of the ecosystem network increases, preventing larger networks from being studied. We present a novel sequential Monte Carlo sampling approach for ensemble generation that is orders of magnitude faster than existing approaches. We demonstrate that the methods produce equivalent parameter inferences, model predictions, and tightly constrained parameter combinations using a novel sensitivity analysis method. For one case study, we demonstrate a speed-up from 108 days to 6 hours, while maintaining equivalent ensembles. Additionally, we demonstrate how to identify the parameter combinations that strongly drive feasibility and stability, drawing ecological insight from the ensembles. Now, for the first time, larger and more realistic networks can be practically simulated and analysed.
: Conservation planning requires an analysis of the potential risks of action, including the risk that the conservation is ineffective or causes unintended consequences. In this respect, quantitative ecosystem models are a valuable decision-making tool for predicting ecosystem responses to actions. Systems of differential equations, such as the generalised Lotka-Volterra equations, can be combined with information about species interactions via food webs to model species abundances. However, parameterising these models is challenging – particularly for data-limited ecosystems. Ensemble ecosystem modelling (EEM) is an increasingly popular method that uses expected dynamic systems constraints to parameterise ecosystem models without time-series abundance data (Baker et al., 2017). This method works by randomly sampling the parameter space to identify an ensemble of models with a stable equilibrium that has coexistence (feasibility) for all species in the ecosystem defined by the inputted food web. But this process is computationally intensive because there can be a very low probability of randomly sampling ecosystem models which meet the constraints (Peterson and Bode, 2021), limiting EEM’s usefulness to highly simplified, low-dimensional ecosystem networks. To overcome this computational burden, we connected EEM to approximate Bayesian computation techniques, allowing us to take advantage of established efficient parameterisation algorithms in the field. We develop a new algorithm called SMC-EEM, adapted from Drovandi and Pettitt’s (2011) sequential Monte Carlo approximate Bayesian inference algorithm – an efficient parameterisation process for low probability constraints, given intractable likelihoods (Sisson et al., 2007). Through simulation studies, we show that our SMC-EEM algorithm is several orders of magnitude faster than the standard-EEM method for high-dimensional ecosystem models. We show that the methods produce equivalent results though a combination of standard sample comparison methods,
The effectiveness and adequacy of natural hazard warnings hinges on the availability of data and its transformation into actionable knowledge for the public. Real-time warning communication and emergency response therefore need to be evaluated from a data science perspective. However, there are currently gaps between established data science best practices and their application in supporting natural hazard warnings. This Perspective reviews existing data-driven approaches that underpin real-time warning communication and emergency response, highlighting limitations in hazard and impact forecasts. Four main themes for enhancing warnings are emphasised: (i) applying best-practice principles in visualising hazard forecasts, (ii) data opportunities for more effective impact forecasts, (iii) utilising data for more localised forecasts, and (iv) improving data-driven decision-making using uncertainty. Motivating examples are provided from the extensive flooding experienced in Australia in 2022. This Perspective shows the capacity for improving the efficacy of natural hazard warnings using data science, and the collaborative potential between the data science and natural hazards communities.
It can be difficult to identify ways to reduce the complexity of large models whilst maintaining predictive power, particularly where there are hidden parameter interdependencies. Here, we demonstrate that the analysis of model sloppiness can be a new invaluable tool for strategically simplifying complex models. Such an analysis identifies parameter combinations which strongly and/or weakly inform model behaviours, yet the approach has not previously been used to inform model reduction. Using a case study on a coral calcification model calibrated to experimental data, we show how the analysis of model sloppiness can strategically inform model simplifications which maintain predictive power. Additionally, when comparing various approaches to analysing sloppiness, we find that Bayesian methods can be advantageous when unambiguous identification of the best-fit model parameters is a challenge for standard optimisation procedures.
This work introduces a comprehensive approach to assess the sensitivity of model outputs to changes in parameter values, constrained by the combination of prior beliefs and data. This approach identifies stiff parameter combinations strongly affecting the quality of the model-data fit while simultaneously revealing which of these key parameter combinations are informed primarily by the data or are also substantively influenced by the priors. We focus on the very common context in complex systems where the amount and quality of data are low compared to the number of model parameters to be collectively estimated, and showcase the benefits of this technique for applications in biochemistry, ecology, and cardiac electrophysiology. We also show how stiff parameter combinations, once identified, uncover controlling mechanisms underlying the system being modeled and inform which of the model parameters need to be prioritized in future experiments for improved parameter inference from collective model-data fitting.