Optimal design facilitates intelligent data collection. In this paper, we introduce a fully Bayesian design approach for spatial processes with complex covariance structures, like those typically exhibited in natural ecosystems. Coordinate exchange algorithms are commonly used to find optimal design points. However, collecting data at specific points is often infeasible in practice. Currently, there is no provision to allow for flexibility in the choice of design. Accordingly, we also propose an approach to find Bayesian sampling windows, rather than points, via Gaussian process emulation to identify regions of high design efficiency across a multi-dimensional space. These developments are motivated by two ecological case studies: monitoring water temperature in a river network system in the northwestern United States and monitoring submerged coral reefs off the north-west coast of Australia.
The SSN2 R package provides tools for spatial statistical modeling, parameter estimation, and prediction on stream (river) networks. SSN2 is the successor to the SSN R package (Ver Hoef, Peterson, Clifford, & Shah, 2014), which was archived alongside broader changes in the R-spatial ecosystem (Nowosad, 2023) that included 1) the retirement of rgdal (Bivand, Keitt, & Rowlingson, 2021), rgeos (Bivand & Rundel, 2020), and maptools (Bivand & Lewin-Koh, 2021) and 2) the lack of active development of sp (Bivand, Pebesma, & Gómez-Rubio, 2013). SSN2 maintains compatibility with the input data file structures used by the SSN R package but leverages modern R-spatial tools like sf (Pebesma, 2018). SSN2 also provides many useful features that were not available in the SSN R package, including new modeling and helper functions, enhanced fitting algorithms, and simplified syntax consistent with other R generic functions.
Crowdsourcing methods facilitate the production of scientific information by non-experts. This form of citizen science (CS) is becoming a key source of complementary data in many fields to inform data-driven decisions and study challenging problems. However, concerns about the validity of these data often constrain their utility. In this paper, we focus on the use of citizen science data in addressing complex challenges in environmental conservation. We consider this issue from three perspectives. First, we present a literature scan of papers that have employed Bayesian models with citizen science in ecology. Second, we compare several popular majority vote algorithms and introduce a Bayesian item response model that estimates and accounts for participants' abilities after adjusting for the difficulty of the images they have classified. The model also enables participants to be clustered into groups based on ability. Third, we apply the model in a case study involving the classification of corals from underwater images from the Great Barrier Reef, Australia. We show that the model achieved superior results in general and, for difficult tasks, a weighted consensus method that uses only groups of experts and experienced participants produced better performance measures. Moreover, we found that participants learn as they have more classification opportunities, which substantially increases their abilities over time. Overall, the paper demonstrates the feasibility of CS for answering complex and challenging ecological questions when these data are appropriately analysed. This serves as motivation for future work to increase the efficacy and trustworthiness of this emerging source of data.
The use of in-situ digital sensors for water quality monitoring is becoming increasingly common worldwide. While these sensors provide near real-time data for science, the data are prone to technical anomalies that can undermine the trustworthiness of the data and the accuracy of statistical inferences, particularly in spatial and temporal analyses. Here we propose a framework for detecting anomalies in sensor data recorded in stream networks, which takes advantage of spatial and temporal autocorrelation to improve detection rates. The proposed framework involves the implementation of effective data imputation to handle missing data, alignment of time-series to address temporal disparities, and the identification of water quality events. We explore the effectiveness of a suite of state-of-the-art statistical methods including posterior predictive distributions, finite mixtures, and Hidden Markov Models (HMM). We showcase the practical implementation of automated anomaly detection in near-real time by employing a Bayesian recursive approach. This demonstration is conducted through a comprehensive simulation study and a practical application to a substantive case study situated in the Herbert River, located in Queensland, Australia, which flows into the Great Barrier Reef. We found that methods such as posterior predictive distributions and HMM produce the best performance in detecting multiple types of anomalies. Utilizing data from multiple sensors deployed relatively near one another enhances the ability to distinguish between water quality events and technical anomalies, thereby significantly improving the accuracy of anomaly detection. Thus, uncertainty and biases in water quality reporting, interpretation, and modelling are reduced, and the effectiveness of subsequent management actions improved.
Adaptive design methods can be used to make changes to survey designs in ecosystem monitoring to ensure that informative data are collected in an ongoing, cost-effective, and flexible manner. Such methods are of particular benefit in environmental monitoring as such monitoring is often very costly and in many cases consists of only a few sampling sites from which inference about a larger geographical region is needed. In addition, ecological processes are continuously changing, and monitoring programs must account for both known and unknown drivers, so making changes to data collection plans over time may be needed based on the current state and understanding of the process of interest. Through considering a Long-term Monitoring Program of Australia’s Great Barrier Reef, this paper aims to develop adaptive design approaches to efficiently monitor coral health through the consideration of a statistical model that accounts for both spatial variability and time-varying disturbance patterns. In particular, to develop this model, we considered time-varying disturbance data that have been reproduced at a fine spatial resolution for uniform representation over the study region. By adopting our proposed approach, we show that adaptive designs are able to significantly reduce survey effort while still remaining effective in, for example, quantifying the effects of different environmental disturbances.
ObjectiveSea Lamprey Petromyzon marinus provide important ecological services within their native range, such as nutrient cycling, and can also act as a prey source for other species. Adult Sea Lamprey must access freshwater rivers to spawn, and because of this they are susceptible to changes in river connectivity. Human-made structures, such as dams, can exclude them from usable habitat. Sea Lamprey dam passage has not been extensively studied in Maine, despite Maine being within the native range of this species. The goals of this study were to evaluate upstream passage efficiency at the Milford Dam on the Penobscot River, Maine, and to provide comprehensive information about adult Sea Lamprey passage at five other dams throughout the Penobscot River watershed.MethodsIn 2020-2021 we captured and tagged 150 Sea Lamprey at the Milford Dam, the lowest dam in the Penobscot River, Maine, and displaced them downstream to assess passage efficiency at this dam and five upstream dams. In 2020, 50 Sea Lamprey were released on the east shore of the river downstream of Milford Dam; in 2021, the east shore release was repeated with an additional 50 fish and another 50 fish were released on the west shore.ResultBetween 70-82% of Sea Lamprey were observed passing Milford Dam again after mean delay times of 9-11 days. The release location did not affect dam passage success or the amount of time that was required to locate and use the passage structures. Sea Lampreys from both release groups were equally likely to approach the entrance to the fishway upon returning to Milford Dam, despite the fishway being located against the eastern shore of the river. However, high flows shortly after release may have resulted in higher attraction to the fishway in 2020. Passage success at dams upstream of Milford was highly variable. All Sea Lamprey were able to successfully navigate past West Enfield Dam (100% passage, n = 63), whereas Brownsmill Dam apparently acted as a complete barrier to further migration (0% passage, n = 7). Fish from all years and release groups together had a median upstream migration distance of 38.8 km after fish had passed Milford Dam, and a maximum observed upstream travel distance of approximately 100 km, indicating that most tagged Sea Lamprey ended their migration in the vicinity of a dam.ConclusionThe results of this study indicate that Sea Lamprey have high passage efficiency at the Milford Dam and highlight areas within the Penobscot River basin-such as the Brownsmill Dam-where passage facilities are currently inadequate for Sea Lamprey. Sea Lamprey are an ecologically important species in their native range. Although they contribute to nutrient cycling and serve as prey for other species, little is known about how damming has affected them. We studied migratory movements of adult Sea Lamprey in the Penobscot River, Maine, a heavily dammed coastal river system.Impact statement
Spatio-temporal models are widely used in many research areas from ecology to epidemiology. However, most covariance functions describe spatial relationships based on Euclidean distance only. In this paper, we introduce the R package SSNbayes for fitting Bayesian spatio-temporal models and making predictions on branching stream networks. SSNbayes provides a linear regression framework with multiple options for incorporating spatial and temporal autocorrelation. Spatial dependence is captured using stream distance and flow connectivity while temporal autocorrelation is modelled using vector autoregression approaches. SSNbayes provides the functionality to make predictions across the whole network, compute exceedance probabilities and other probabilistic estimates such as the proportion of suitable habitat. We illustrate the functionality of the package using a stream temperature dataset collected in Idaho, USA.
Real-time monitoring using in-situ sensors is becoming a common approach for measuring water-quality within watersheds. High-frequency measurements produce big datasets that present opportunities to conduct new analyses for improved understanding of water-quality dynamics and more effective management of rivers and streams. Of primary importance is enhancing knowledge of the relationships between nitrate, one of the most reactive forms of inorganic nitrogen in the aquatic environment, and other water-quality variables. We analysed high-frequency water-quality data from in-situ sensors deployed in three sites from different watersheds and climate zones within the National Ecological Observatory Network, USA. We used generalised additive mixed models to explain the nonlinear relationships at each site between nitrate concentration and conductivity, turbidity, dissolved oxygen, water temperature, and elevation. Temporal auto-correlation was modelled with an auto-regressive-moving-average (ARIMA) model and we examined the relative importance of the explanatory variables. Total deviance explained by the models was high for all sites (99%). Although variable importance and the smooth regression parameters differed among sites, the models explaining the most variation in nitrate contained the same explanatory variables. This study demonstrates that building a model for nitrate using the same set of explanatory water-quality variables is achievable, even for sites with vastly different environmental and climatic characteristics. Applying such models will assist managers to select cost-effective water-quality variables to monitor when the goals are to gain a spatial and temporal in-depth understanding of nitrate dynamics and adapt management plans accordingly.
We consider four main goals when fitting spatial linear models: 1) estimating covariance parameters, 2) estimating fixed effects, 3) kriging (making point predictions), and 4) block-kriging (predicting the average value over a region). Each of these goals can present different challenges when analyzing large spatial data sets. Current research uses a variety of methods, including spatial basis functions (reduced rank), covariance tapering, etc, to achieve these goals. However, spatial indexing, which is very similar to composite likelihood, offers some advantages. We develop a simple framework for all four goals listed above by using indexing to create a block covariance structure and nearest-neighbor predictions while maintaining a coherent linear model. We show exact inference for fixed effects under this block covariance construction. Spatial indexing is very fast, and simulations are used to validate methods and compare to another popular method. We study various sample designs for indexing and our simulations showed that indexing leading to spatially compact partitions are best over a range of sample sizes, autocorrelation values, and generating processes. Partitions can be kept small, on the order of 50 samples per partition. We use nearest-neighbors for kriging and block kriging, finding that 50 nearest-neighbors is sufficient. In all cases, confidence intervals for fixed effects, and prediction intervals for (block) kriging, have appropriate coverage. Some advantages of spatial indexing are that it is available for any valid covariance matrix, can take advantage of parallel computing, and easily extends to non-Euclidean topologies, such as stream networks. We use stream networks to show how spatial indexing can achieve all four goals, listed above, for very large data sets, in a matter of minutes, rather than days, for an example data set.
Atlantic salmon ( Salmo salar) return to rivers in spring for an energetically costly upstream migration for spawning. These fish are often delayed in the lower river below dams, subjecting them to warmer waters than occur in upstream sections of river, that may increase metabolic costs. We sought to quantify the energetic cost of dam-mediated delays in migrating adults in the Penobscot and Kennebec rivers, ME. We radio-tagged fish at the lower most dams, released them downstream (18 and 14 km), and tracked their movements back upstream. We used a Distell Fish Fatmeter as a noninvasive measurement of full-body energy at tagging and then again after re-ascending the fish-way at the dams. We found that adults ( n = 99) experienced average delays of 16–23 days at dams, losing 11%–22% of initial fat reserves. Using linear regressions, we showed thermal experience as a strong predictor of fat loss. Delay time was also a contributing factor. Extensive delays at dams expose migrating Atlantic salmon to warmer temperatures and increase the depletion rate of energy reserves required for spawning and post-spawn survival.
Objectives Clinical supervision is essential for ensuring effective service delivery. International imperatives to demonstrate professional competence has increased attention on the role of supervision in enhancing client outcomes. Although supervisor competency tools are recognised as important components in effective supervision, there remains a shortage of tools that are evidenced-based, applicable across workforces and freely accessible. Design An expert multidisciplinary group developed the Generic Supervision Assessment Tool (GSAT) to assess supervisor competencies across a range of professions. Initially the GSAT consisted of 32 items responded to by either a supervisor (GSAT-SR) or supervisee (GSAT-SE). The current study, using surveys, employed a cross-sectional design to test the reliability and construct validity of the GSAT. Methods The study consisted of two phases and included 12 professional groups across Australasia. In 2018, exploratory factor analysis (EFA) was undertaken with survey data from 479 supervisors and 447 supervisees. In 2019 survey data from 182 supervisors and 186 supervisees were used to conduct confirmatory factor analysis (CFA). The results were used to refine and validate the GSAT. Results The final GSAT-SR has four factors with 26 competency items. The final GSAT-SE has two factors with 21 competency items. The EFA and CFA confirmed that the GSAT-SR and the GSAT-SE are psychometrically valid tools that supervisors and supervisees can utilise to assess competencies. Conclusion As a non-discipline specific supervision tool, the GSAT is a validated, freely available tool for benchmarking the competencies of clinical supervisors across professions, potentially optimising supervisory evaluation processes and strengthening supervision effectiveness. Practitioner points Supervisor competency tools are recognised as important components of safe and effective supervision provision yet there is a dearth of valid, reliable and effective measures. The Generic Supervision Assessment Tool (GSAT-SR and GSAT-SE) are unique psychometrically valid, and reliable measures of supervisor competence. The GSAT-SR and the GSAT-SE can enhance translation of evidence-based supervision competency skills into regular practice. Validated with a broad cross section of professionals in diverse practice settings the GSAT provides a comprehensive conceptualization of supervisor competence.
Spatio-temporal models are widely used in many research areas including ecology. The recent proliferation of the use of in-situ sensors in streams and rivers supports space-time water quality modelling and monitoring in near real-time. A new family of spatio-temporal models is introduced. These models incorporate spatial dependence using stream distance while temporal autocorrelation is captured using vector autoregression approaches. Several variations of these novel models are proposed using a Bayesian framework. The results show that our proposed models perform well using spatio-temporal data collected from real stream networks, particularly in terms of out-of-sample RMSPE. This is illustrated considering a case study of water temperature data in the northwestern United States.
Many research domains use data elicited from 'citizen scientists' when a direct measure of a process is expensive or infeasible. However, participants may report incorrect estimates or classifications due to their lack of skill. We demonstrate how Bayesian hierarchical models can be used to learn about latent variables of interest, while accounting for the participants' abilities. The model is described in the context of an ecological application that involves crowdsourced classifications of georeferenced coral-reef images from the Great Barrier Reef, Australia. The latent variable of interest is the proportion of coral cover, which is a common indicator of coral reef health. The participants' abilities are expressed in terms of sensitivity and specificity of a correctly classified set of points on the images. The model also incorporates a spatial component, which allows prediction of the latent variable in locations that have not been surveyed. We show that the model outperforms traditional weighted-regression approaches used to account for uncertainty in citizen science data. Our approach produces more accurate regression coefficients and provides a better characterisation of the latent process of interest. This new method is implemented in the probabilistic programming language Stan and can be applied to a wide number of problems that rely on uncertain citizen science data.
Virtual reality (VR) technology is an emerging tool that is supporting the connection between conservation research and public engagement with environmental issues. The use of VR in ecology consists of interviewing diverse groups of people while they are immersed within a virtual ecosystem to produce better information than more traditional surveys. However, at present, the relatively high level of expertise in specific programming languages and disjoint pathways required to run VR experiments hinder their wider application in ecology and other sciences. We present R2VR, a package for implementing and performing VR experiments in R with the aim of easing the learning curve for applied scientists including ecologists. The package provides functions for rendering VR scenes on web browsers with A-Frame that can be viewed by multiple users on smartphones, laptops, and VR headsets. It also provides instructions on how to retrieve answers from an online database in R. Three published ecological case studies are used to illustrate the R2VR workflow, and show how to run a VR experiments and collect the resulting datasets. By tapping into the popularity of R among ecologists, the R2VR package creates new opportunities to address the complex challenges associated with conservation, improve scientific knowledge, and promote new ways to share better understanding of environmental issues. The package could also be used in other fields outside of ecology.
Many research domains use data elicited from ‘citizen scientists’ when a direct measure of a process is expensive or infeasible. However, participants may report incorrect estimates or classifications due to their lack of skill. We demonstrate how Bayesian hierarchical models can be used to learn about latent variables of interest, while accounting for the participants’ abilities. The model is described in the context of an ecological application that involves crowdsourced classifications of georeferenced coral-reef images from the Great Barrier Reef, Australia. The latent variable of interest is the proportion of coral cover, which is a common indicator of coral reef health. The participants’ abilities are expressed in terms of sensitivity and specificity of a correctly classified set of points on the images. The model also incorporates a spatial component, which allows prediction of the latent variable in locations that have not been surveyed. We show that the model outperforms traditional weighted-regression approaches used to account for uncertainty in citizen science data. Our approach produces more accurate regression coefficients and provides a better characterisation of the latent process of interest. This new method is implemented in the probabilistic programming language Stan and can be applied to a wide number of problems that rely on uncertain citizen science data.
In situ sensors that collect high-frequency data are used increasingly to monitor aquatic environments. These sensors are prone to technical errors, resulting in unrecorded observations and/or anomalous values that are subsequently removed and create gaps in time series data. We present a framework based on generalized additive and auto-regressive models to recover these missing data. To mimic sporadically missing (i) single observations and (ii) periods of contiguous observations, we randomly removed (i) point data and (ii) day- and week-long sequences of data from a two-year time series of nitrate concentration data collected from Arikaree River, USA, where synoptically collected water temperature, turbidity, conductance, elevation, and dissolved oxygen data were available. In 72% of cases with missing point data, predicted values were within the sensor precision interval of the original value, although predictive ability declined when sequences of missing data occurred. Precision also depended on the availability of other water quality covariates. When covariates were available, even a sudden, event-based peak in nitrate concentration was reconstructed well. By providing a promising method for accurate prediction of missing data, the utility and confidence in summary statistics and statistical trends will increase, thereby assisting the effective monitoring and management of fresh waters and other at-risk ecosystems.