Ecosystem maps support a vast array of applications in conservation, land management and policy. The capacity of an ecosystem map to support these applications is determined by its ability to accurately represent ecosystem distributions, which is heavily influenced by the model used to produce them. Here, we evaluated the influence of key modelling decisions made whilst developing a new and comprehensive ecosystem map using a recently developed ecosystem typology for the remote Tiwi Islands, Australia. We collated a reference set of training points from diverse datasets and employed a pixel-based, random forest model to classify and predict ecosystem distributions. We tested decisions at three stages of the model formulation. First, we tested the number of classes by aggregating ecosystem types (finest scale, n = 11) into functional groups (n = 10) and biomes (coarsest, n = 8) according to the Global Ecosystem Typology. Second, we compared data acquired from the Sentinel-2 satellite using the MSI sensor and Landsat-9 with the OLI-2 sensor. Finally, we tested covariates from satellite image bands only or satellite imagery combined with additional covariates describing other ecological characteristics. We evaluated these decisions using a range of model performance metrics, including overall, by-class and spatially explicit estimates. Our study found that using covariates additional to those from satellite images improved all evaluation metrics for all model decisions. Acquisitions from Landsat-9 tended to improve model performance over Sentinel-2, although the effect was variable. Developing maps at the biome scale (coarsest resolution) slightly improved overall performance but hinders applications that need to differentiate between ecosystem types. Including additional relevant covariates or considering alternative satellites are better options for improving map performance than simplifying the classes. Producing spatially explicit evaluation of ecosystem maps is a rapid and achievable method to communicate limitations and support users to make informed decisions.
Effective ecosystem conservation for biodiversity and human well-being relies on accurate information. Consistent approaches to classifying, describing, and assessing ecosystems can improve understanding of ecological processes, threats, and management. We explored how the International Union for Conservation of Nature (IUCN) Global Ecosystem Typology-a global classification framework based on ecosystem function-could support the development of a classification of ecosystems for the Tiwi Islands, Australia, by incorporating scientific information and Indigenous Tiwi knowledge to facilitate environmental management and conservation. We synthesized ecosystem information from previous research, field data, reports, and Tiwi knowledge authorities to develop a classification, descriptions, and a map of 14 terrestrial ecosystem types. These ecosystem types were defined and described based on ecological processes and were broader yet largely congruent with existing vegetation classifications. Including functional properties accounted for variation in the vegetation physiognomy exhibited by dynamic and disturbance-prone ecosystems, such as savannas. Because we considered Tiwi knowledge authorities and the IUCN Global Ecosystem Typology, our inventory included ecosystem types that were typically omitted from previous classifications, which should allow for more comprehensive assessments and management. Relating the new ecosystem typology to the IUCN Global Ecosystem Typology will facilitate comparisons among similar ecosystems, regarding, for example, effective threat abatement options. Describing the biota and processes opens new avenues for monitoring. More collaborative work is needed to explore how Western scientific ecosystem inventories operate alongside and in connection with management of Tiwi murrakupuni enacted by Tiwi people. Given the ongoing loss of biodiversity, ecosystem management must draw on information across domains, scales, and knowledge systems. We demonstrated an approach to this task and provided socioecologically relevant ecosystem information.
Understanding patterns of species occupancy across landscapes and throughout time is a long-standing objective of ecological research that has inspired the development of numerous quantitative modelling approaches. However, estimating occupancy can be a challenge, particularly when contending with issues like imperfect detection and shifting distributions. Dynamic occupancy models (DOMs) offer a framework for occupancy estimation that explicitly accounts for observation error while capturing the mechanisms driving occupancy dynamics by estimating colonisation and local extinction processes. In light of increasing interest in more process-explicit models for understanding species occurrence, here we examine how DOMs have been applied to field ecological data in the two decades since their introduction. Following a general introduction to the model, we present the results of a systematic review exploring where and how DOMs have been applied. We interrogate how authors have built their models, with particular emphasis on how covariates are incorporated to describe variation in occupancy dynamics. Our findings indicate that DOMs are a flexible tool readily applied to diverse study systems and data types, with their usage expanding in recent years as more studies apply them to make spatial and temporal predictions of species occupancy. DOMs are also amenable to extension, further broadening their utility. However, model complexity in DOMs tends to be low; most studies consider relatively few covariates and these are typically represented as simple linear relationships. Approaches to covariate selection also vary considerably, and there remains little research 1on how these choices may influence model performance. Furthermore, only a fraction of articles report evaluating DOMs and little guidance exists on how to approach this task. These uncertainties in the modelling process should be key priorities for future research on DOMs given their increasing use in applied ecological research.
Species distribution models are widely used to identify potential and high-quality habitat of endangered species to inform conservation decisions. However, their usefulness is constrained by the amount and quality of biodiversity data and the approaches for dealing with data deficiencies. Presence-only data, used in presence/background modelling methods, are widely available but are often affected by sampling bias. Presence/absence modelling methods are less affected by biases, but data are less common. We modelled the distribution of a widely distributed, endangered species from Australia - the greater glider - and tested how predictions were influenced by data treatment and modelling framework. We collated available species data and fitted generalized linear models and boosted regression trees using presence/absence data, as well as using an augmented dataset that included additional presences alongside absences inferred from survey data. We also fitted presence/background models, adopting three common strategies for bias correction. We compared model performance quantitatively through evaluation metrics calculated internally and on held out data, and qualitatively by identifying areas of agreement of spatial predictions. We found that presence/background models with bias correction performed better than not corrected, though evaluation metrics did not favour a single strategy. Presence/absence models outperformed presence/background models in comparable metrics and delivered different spatial predictions. Importantly, differences in spatial predictions between models had the potential to substantially alter decisions about where to protect high-quality habitat. The approach to inferring absences proved useful, as models fitted with these outperformed all other models. Dealing with sampling bias requires additional time and data management strategies, but we found that the time invested allowed improvement of models and more reliable predictions. Our results suggest that ancillary occurrence data and careful data handling can improve both presence/background and presence/absence models.
Effective ecosystem management for biodiversity and human well-being relies on accurate information. Consistent approaches to classifying, describing, and assessing ecosystems can improve the understanding of the ecological processes, threats, and management. We explored how the Global Ecosystem Typology – a global classification framework based on ecosystem function – could support the development of a local ecosystem inventory for the Tiwi Islands, Australia, to facilitate management by the Indigenous Tiwi peoples and government agencies by incorporating Tiwi knowledge and scientific information. We synthesized ecosystem information from previous research, field data, and reports, together with input from Tiwi knowledge authorities, to develop a classification, descriptions, and map for 14 terrestrial ecosystem types. These ecosystem types were defined and described by ecological processes and were broader, yet largely congruent, with previous classifications. Including functional properties accounted for variation in the vegetation physiognomy exhibited by dynamic and disturbance-prone ecosystems, such as savannas. By bringing together Tiwi knowledge authorities’ input, regional information and the Global Ecosystem Typology, we included in our inventory ecosystem types that were typically omitted from previous classifications. Describing the biota within each ecosystem type ensured local relevance and opened new avenues for monitoring, while the Global Ecosystem Typology facilitated comparisons to similar ecosystems in terms of effective threat and management options. Many of the ecosystem types aligned with terms in Tiwi language, significantly enhancing the ways in which global frameworks can support ecological management suitable for Tiwi Country (murrakapuni). Beyond this, more collaborative work is needed to explore how the ecosystem inventories and global ecosystem management approaches may operate alongside, and in connection with, the ways of managing Tiwi murrakapuni currently enacted by Tiwi people. With the current and ongoing loss of biodiversity, managing ecosystems must span interdisciplinary knowledges and bridge local and global understandings for the shared goal of conservation.
AimTo assess whether flexible species distribution models that perform well at nearby testing locations still perform strongly when evaluated on spatially separated testing data. LocationAustralian Wet Tropics (AWT), Ontario, Canada (CAN), north-east New South Wales, Australia (NSW), New Zealand (NZ), five countries of South America (SA), and Switzerland (SWI). Time periodMost species data were collected between 1950 and 2000. Major taxa studiedBirds, mammals, plants and reptiles. MethodsWe compared 10 species distribution modelling methods with varying flexibility in terms of the allowed complexity of their fitted functions [boosted regression trees (BRT), generalized additive model (GAM), multivariate adaptive regression splines (MARS), maximum entropy (MaxEnt), support vector machine (SVM), variants of generalized linear model (GLM) and random forest (RF), and an Ensemble model]. We used established practices for model selection to avoid overfitting, including parameter tuning in learning methods. Models were trained on presence-background data for 171 species and tested on presence-absence data. Training and testing data were separated using both random and spatial partitioning, the latter based on 75-km blocks. We calculated the average performance and mean rank of the methods (focussing on the area under the receiver operating characteristic and precision-recall gain curves, and correlation) and assessed the statistical significance of the differences between them. Results The ranking of methods did not change when evaluated on spatially separated testing data. Methods with the strongest predictive performance were nonparametric methods known to be flexible. An ensemble formed by averaging predictions of five pre-selected modelling methods was the best model in both random and spatial partitioning, followed by MaxEnt and a variant of random forest. Main conclusionsWhilst some modellers expect methods limited to simple smooth functions to predict better spatially separated data, we found no evidence of that using blocks of 75 km. We conclude that flexible models that are tuned well enough to avoid overfitting are effective at predicting to spatially distinct areas.
Predicting novel ranges of non-native species is a critical component to understanding the biosecurity threat posed by pests and diseases on economic, environmental and social assets. Species distribution models (SDMs) are often employed to predict the potential ranges of exotic pests and diseases in novel environments and geographic space. To date, researchers have focused on model complexity, data available for model fitting, the size of the geographic area to be considered and how the choice of model impacts results. These investigations are coupled with considerable examination of how model evaluation methods and test scores are influenced by these choices. An area that remains under-discussed is how to account for uncertainty in predictor selection while also selecting variables that increase a model’s ability to predict to novel environments (model transferability). Here we propose a novel method to finesse this problem by using multiple simple (bivariate) models to search for the candidate sets of predictor variables that are likely to produce transferable models. Once identified, each set is then used to construct 2-dimensional niche envelopes of pest presence/absence. This process ultimately results in a number of possible models that can be used to predict pest potential distributions, however, rather than relying on a single model, we ensemble these models in an attempt to account for predictor uncertainty. We apply this method to both virtual species and real species data, and find that it generally performs well against conventional approaches for statistically fitting numerous variables in a single model. While our methods only consider simple ecological relationships of species to environmental predictors, they allow for increased model transferability because they reduce the likelihood of over-fitting and collinearity issues. Simple models are also likely to be more conservative (over-predict potential distributions) relative to complex models containing many covariates – making them more appropriate for risk-averse applications such as biosecurity. The approach we have explored transforms a model selection problem, for which there is no true correct answer amongst the typically distal covariates on offer, to one of model uncertainty. We argue that increased model transferability at the expense of model interpretation is perhaps more important for effective rapid predictions and management of non-native species and biological invasions.
Species distribution modeling (SDM) is widely used in ecology and conservation. Currently, the most available data for SDM are species presence-only records (available through digital databases). There have been many studies comparing the performance of alternative algorithms for modeling presence-only data. Among these, a 2006 paper from Elith and colleagues has been particularly influential in the field, partly because they used several novel methods (at the time) on a global data set that included independent presence-absence records for model evaluation. Since its publication, some of the algorithms have been further developed and new ones have emerged. In this paper, we explore patterns in predictive performance across methods, by reanalyzing the same data set (225 species from six different regions) using updated modeling knowledge and practices. We apply well-established methods such as generalized additive models and MaxEnt, alongside others that have received attention more recently, including regularized regressions, point-process weighted regressions, random forests, XGBoost, support vector machines, and the ensemble modeling framework biomod. All the methods we use include background samples (a sample of environments in the landscape) for model fitting. We explore impacts of using weights on the presence and background points in model fitting. We introduce new ways of evaluating models fitted to these data, using the area under the precision-recall gain curve, and focusing on the rank of results. We find that the way models are fitted matters. The top method was an ensemble of tuned individual models. In contrast, ensembles built using the biomod framework with default parameters performed no better than single moderate performing models. Similarly, the second top performing method was a random forest parameterized to deal with many background samples (contrasted to relatively few presence records), which substantially outperformed other random forest implementations. We find that, in general, nonparametric techniques with the capability of controlling for model complexity outperformed traditional regression methods, with MaxEnt and boosted regression trees still among the top performing models. All the data and code with working examples are provided to make this study fully reproducible.
Extreme weather can have significant impacts on plant species demography; however, most studies have focused on responses to a single or small number of extreme events. Long-term patterns in climate extremes, and how they have shaped contemporary distributions, have rarely been considered or tested. BIOCLIM variables that are commonly used in correlative species distribution modelling studies cannot be used to quantify climate extremes, as they are generated using long-term averages and therefore do not describe year-to-year, temporal variability. We evaluated the response of 37 plant species to base climate (long-term means, equivalent to BIOCLIM variables), variability (standard deviations) and extremes of varying return intervals (defined using quantiles) based on historical observations. These variables were generated using fine-grain (approx. 250 m), time-series temperature and precipitation data for the hottest, coldest and driest months over 39 years. Extremes provided significant additive improvements in model performance compared to base climate alone and were more consistent than variability across all species. Models that included extremes frequently showed notably different mapped predictions relative to those using base climate alone, despite often small differences in statistical performance as measured as a summary across sites. These differences in spatial patterns were most pronounced at the predicted range margins, and reflect the influence of coastal proximity, continentality, topography and orographic barriers on climate extremes. Species occupying hotter and drier locations that are exposed to severe maximum temperature extremes were associated with better predictive performance when modelled using extremes. Understanding how plant species have historically responded to climate extremes may provide valuable insights into our understanding of contemporary distributions and help to make more accurate predictions under a changing climate.
Species distribution models (SDMs) are increasingly used in conservation and land-use planning as inputs to describe biodiversity patterns. These models can be built in different ways, and decisions about data preparation, selection of predictor variables, model fitting, and evaluation all alter the resulting predictions. Commonly, the true distribution of species is unknown and independent data to verify which SDM variant to choose are lacking. Such model uncertainty is of concern to planners. We analyzed how 11 routine decisions about model complexity, predictors, bias treatment, and setting thresholds for predicted values altered conservation priority patterns across 25 species. Models were created with MaxEnt and run through Zonation to determine the priority rank of sites. Although all SDM variants performed well (area under the curve >0.7), they produced spatially different predictions for species and different conservation priority solutions. Priorities were most strongly altered by decisions to not address bias or to apply binary thresholds to predicted values; on average 40% and 35%, respectively, of all grid cells received an opposite priority ranking. Forcing high model complexity altered conservation solutions less than forcing simplicity (14% and 24% of cells with opposite rank values, respectively). Use of fewer species records to build models or choosing alternative bias treatments had intermediate effects (25% and 23%, respectively). Depending on modeling choices, priority areas overlapped as little as 10-20% with the baseline solution, affecting top and bottom priorities differently. Our results demonstrate the extent of model-based uncertainty and quantify the relative impacts of SDM building decisions. When it is uncertain what the best SDM approach and conservation plan is, solving uncertainty or considering alterative options is most important for those decisions that change plans the most.
Open-access occurrence data are useful for studying spatial patterns of fungi, but often have quality issues. These include errors in taxonomy and geo-coordinates, and incomplete coverage across areas and taxonomic groups. We identify 15 quality issues that can lead to incorrect biogeographic inference, and develop a reproducible pipeline that flags and removes problematic entries. This pipeline tests accuracy of geographic records and names. Then, if information on non-native status is unavailable or unreliable, it detects non-native species via a predictive model. Finally, it identifies spatial and environmental outliers and removes them when biologically improbable. We test the pipeline by cleaning data for Australian fungi, with 251,642 records retained after cleaning the initial 1,034,601 records. Exploratory analysis showed that the cleaned data is useful for analyses such as biogeographic regionalisation, but recording gaps and lack of saturation in collection effort also caution that more surveys are needed to improve collection completeness.
Predictions of species' current and future ranges are needed to effectively manage species under environmental change. Species ranges are typically estimated using correlative species distribution models (SDMs), which have been criticized for their static nature. In contrast, dynamic occupancy models (DOMs) explicitily describe temporal changes in species’ occupancy via colonization and local extinction probabilities, estimated from time series of occurrence data. Yet, tests of whether these models improve predictive accuracy under current or future conditions are rare. Using a long‐term data set on 69 Swiss birds, we tested whether DOMs improve the predictions of distribution changes over time compared to SDMs. We evaluated the accuracy of spatial predictions and their ability to detect population trends. We also explored how predictions differed when we accounted for imperfect detection and parameterized models using calibration data sets of different time series lengths. All model types had high spatial predictive performance when assessed across all sites (mean AUC > 0.8), with flexible machine learning SDM algorithms outperforming parametric static and DOMs. However, none of the models performed well at identifying sites where range changes are likely to occur. In terms of estimating population trends, DOMs performed best, particularly for species with strong population changes and when fit with sufficient data, while static SDMs performed very poorly. Overall, our study highlights the importance of considering what aspects of performance matter most when selecting a modelling method for a particular application and the need for further research to improve model utility. While DOMs show promise for capturing range dynamics and inferring population trends when fitted with sufficient data, computational constraints on variable selection and model fitting can lead to reduced spatial accuracy of predictions, an area warranting more attention.
The random forest (RF) algorithm is an ensemble of classification or regression trees and is widely used, including for species distribution modelling (SDM). Many researchers use implementations of RF in the R programming language with default parameters to analyse species presence‐only data together with ‘background' samples. However, there is good evidence that RF with default parameters does not perform well for such ‘presence‐background' modelling. This is often attributed to the disparity between the number of presence and background samples, also known as 'class imbalance', and several solutions have been proposed. Here, we first set the context: the background sample should be large enough to represent all environments in the region. We then aim to understand the drivers of poor performance of RF when models are fitted to presence‐only species data alongside background samples. We show that 'class overlap' (where both classes occur in the same environment) is an important driver of poor performance, alongside class imbalance. Class overlap can even degrade performance for presence–absence data. We explain, test and evaluate suggested solutions. Using simulated and real presence‐background data, we compare performance of default RF with other weighting and sampling approaches. Our results demonstrate clear evidence of improvement in the performance of RFs when techniques that explicitly manage imbalance are used. We show that these either limit or enforce tree depth. Without compromising the environmental representativeness of the sampled background, we identify approaches to fitting RF that ameliorate the effects of imbalance and overlap and allow excellent predictive performance. Understanding the problems of RF in presence‐background modelling allows new insights into how best to fit models, and should guide future efforts to best deal with such data.
Predictive performance is important to many applications of species distribution models (SDMs). The SDM ‘ensemble’ approach, which combines predictions across different modelling methods, is believed to improve predictive performance, and is used in many recent SDM studies. Here, we aim to compare the predictive performance of ensemble species distribution models to that of individual models, using a large presence–absence dataset of eucalypt tree species. To test model performance, we divided our dataset into calibration and evaluation folds using two spatial blocking strategies (checkerboard‐pattern and latitudinal slicing). We calibrated and cross‐validated all models within the calibration folds, using both repeated random division of data (a common approach) and spatial blocking. Ensembles were built using the software package ‘biomod2’, with standard (‘untuned’) settings. Boosted regression tree (BRT) models were also fitted to the same data, tuned according to published procedures. We then used evaluation folds to compare ensembles against both their component untuned individual models, and against the BRTs. We used area under the receiver‐operating characteristic curve (AUC) and log‐likelihood for assessing model performance. In all our tests, ensemble models performed well, but not consistently better than their component untuned individual models or tuned BRTs across all tests. Moreover, choosing untuned individual models with best cross‐validation performance also yielded good external performance, with blocked cross‐validation proving better suited for this choice, in this study, than repeated random cross‐validation. The latitudinal slice test was only possible for four species; this showed some individual models, and particularly the tuned one, performing better than ensembles. This study shows no particular benefit to using ensembles over individual tuned models. It also suggests that further robust testing of performance is required for situations where models are used to predict to distant places or environments.
Species distribution models (SDMs) are widely used to predict and study distributions of species. Many different modeling methods and associated algorithms are used and continue to emerge. It is important to understand how different approaches perform, particularly when applied to species occurrence records that were not gathered in structured surveys (e.g. opportunistic records). This need motivated a large-scale, collaborative effort, published in 2006, that aimed to create objective comparisons of algorithm performance. As a benchmark, and to facilitate future comparisons of approaches, here we publish that dataset: point location records for 226 anonymized species from six regions of the world, with accompanying predictor variables in raster (grid) and point formats. A particularly interesting characteristic of this dataset is that independent presence-absence survey data are available for evaluation alongside the presence-only species occurrence data intended for modeling. The dataset is available on Open Science Framework and as an R package and can be used as a benchmark for modeling approaches and for testing new ways to evaluate the accuracy of SDMs.
Species distribution models (SDMs) constitute the most common class of models across ecology, evolution and conservation. The advent of ready-to-use software packages and increasing availability of digital geoinformation have considerably assisted the application of SDMs in the past decade, greatly enabling their broader use for informing conservation and management, and for quantifying impacts from global change. However, models must be fit for purpose, with all important aspects of their development and applications properly considered. Despite the widespread use of SDMs, standardisation and documentation of modelling protocols remain limited, which makes it hard to assess whether development steps are appropriate for end use. To address these issues, we propose a standard protocol for reporting SDMs, with an emphasis on describing how a study's objective is achieved through a series of modeling decisions. We call this the ODMAP (Overview, Data, Model, Assessment and Prediction) protocol, as its components reflect the main steps involved in building SDMs and other empirically-based biodiversity models. The ODMAP protocol serves two main purposes. First, it provides a checklist for authors, detailing key steps for model building and analyses, and thus represents a quick guide and generic workflow for modern SDMs. Second, it introduces a structured format for documenting and communicating the models, ensuring transparency and reproducibility, facilitating peer review and expert evaluation of model quality, as well as meta-analyses. We detail all elements of ODMAP, and explain how it can be used for different model objectives and applications, and how it complements efforts to store associated metadata and define modelling standards. We illustrate its utility by revisiting nine previously published case studies, and provide an interactive web-based application to facilitate its use. We plan to advance ODMAP by encouraging its further refinement and adoption by the scientific community.