Abstract Pacific harbor seals ( Phoca vitulina richardii ) were surveyed throughout their range along Alaska coasts over a 28-year period. Twelve management stocks, delineated on the basis of evidence for demographic independence, were monitored for abundance and trends. We employed a two-stage Bayesian hierarchical analysis that integrated aerial survey counts with satellite-linked bio-logger haul-out timelines to account for the proportion of seals in the water and not visible to survey observers and cameras. Overall, harbor seals were abundant (201,122 seals in 2023) but, during the span of our study the population (stock) trends varied. The first half of our study was characterised primarily by growth, with a peak abundance of approximately 225,000 seals in 2015. Since then trends have been variable, with some stocks (e.g., Bristol Bay) showing continued growth while others (e.g. Prince William Sound, South Kodiak) declined. The recent regional declines coincided with observed climate anomalies (e.g., marine heatwaves) and the rapid retreat of tidewater glaciers. Our results provide a basis for managing this species in Alaska, which is protected under the Marine Mammal Protection Act; a vital nutritional and cultural resource for Alaska Native communities; and an important sentinel of change in the Northeast Pacific marine ecosystem.
In ecology and related sciences, missing data are common and occur in a variety of different contexts. When missing data are not handled properly, subsequent statistical estimates tend to be biased, inefficient, and lack proper confidence interval coverage. Missing data are often grouped into three categories: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). We review each category and compare their benefits and drawbacks. We review several approaches to handling missing data including complete case analysis, imputation, inverse probability weighting, and data augmentation. We clarify what types of variables should accompany imputation methods and how those variables are influenced by the analysis methods. Additionally, we discuss missing data that lack a formal basis for measurement and hence are fundamentally different from MCAR, MAR, and MNAR missing data. Throughout, we introduce concepts and numeric examples using both simulated data and data from the United States Environmental Protection Agency's 2016 National Wetland Condition Assessment. We conclude by providing five considerations for ecologists and other scientists handling missing data.
Aim Fundamental to species conservation efforts is the development of accurate distribution models. Doing so is challenging for many stream organisms, where limited funding may necessitate the compilation of incidental observations from multiple sources which lack an overall sampling design and are often spatially clustered. We demonstrate the application of specialised spatial-stream-network models (SSNMs), which incorporate autocorrelation among observations and have the potential to improve species distribution models for many organisms. Location Rocky Mountains in west-central North America. Methods We compiled a comprehensive presence-absence dataset for Idaho giant salamander (IGS; Dicamptodon aterrimus) from previous studies, natural resource agencies, museum collections and new surveys, and linked these data to geospatial habitat covariates. The dataset was modelled using a suite of candidate SSNMs, and results were compared to those from non-spatial generalised linear models (GLMs). The top-ranked models were used to predict range-wide IGS occurrence probabilities for scenarios that represented historical baselines and futures associated with two model covariates (water temperature and riparian tree canopy density) that are changing with environmental trends in the study area. Results The classification accuracy of salamander observations was higher with SSNMs than GLMs (90.8% vs. 63.2%) and the spatial models identified fewer significant habitat relationships, which simplified model interpretation. Baseline range estimates from the models were similar (13,090-14,114 stream km) and both predicted small range expansions (2.0%-24.8%) with future warming because many streams were sub-optimally cold for IGS. However, these expansions were partially offset in scenarios which included decreases in riparian canopy density. Main Conclusions SSNMs significantly improved species distribution models on stream networks by incorporating spatial autocorrelation and provide an inexpensive means of developing new information from many existing datasets. This incentivises aggregation of datasets, which could be further leveraged to create efficient monitoring and inventory programs using the spatially explicit outputs from SSNMs.
In the spirit of so-called "sightability models" for estimating population abundance, we developed a Bayesian hierarchical model that combines survey counts for animals (or plants) and a separate data set for detection to account for individuals that were missed during surveys. Our case study consisted of harbor seal (Phoca vitulina richardii) aerial survey counts from 1996 to 2023 for the Prince William Sound (PWS) stock in Alaska and haul-out data from bio-logged individuals. Detection (i.e., haul-out probability) was modeled using logistic regression with temporally autocorrelated latent random effects. The probability of detection informed binomial count models, where true abundances were temporally autocorrelated Poisson models, leading to a logistic-binomial-Poisson hierarchical model. To speed computations, we coupled a two-stage sampling with first-order autoregressive (AR1) and random walk models for autocorrelation. We found time-of-year and time-from-low-tide to be the most important predictors for detection, and our population abundance analysis showed a significant decline (1996-2001), followed by an increase (2001-2015), and then another decline (2015-2023) for the PWS stock. Our approach can be used for other organisms and surveys that have separate detection and count data sets, such as those commonly used in sightability models, as part of long-term population monitoring programs.
Ice-associated seals rely on sea ice for a variety of activities, including pupping, breeding, molting, and resting. In the Arctic, many of these activities occur in spring (April through June) as sea ice begins to melt and retreat northward. Rapid acceleration of climate change in Arctic ecosystems is therefore of concern as the quantity and quality of suitable habitat is forecast to decrease. Robust estimates of seal population abundance are needed to properly monitor the impacts of these changes over time. Aerial surveys of seals on ice are an efficient method for counting seals but must be paired with estimates of the proportion of seals out of the water to derive population abundance. In this paper, we use hourly percent-dry data from satellite-linked bio-loggers deployed between 2005 and 2021 to quantify the proportion of seals hauled out on ice. This information is needed to accurately estimate abundance from aerial survey counts of ice-associated seals (i.e., to correct for the proportion of animals that are in the water, and so are not counted, while surveys are conducted). In addition to providing essential data for survey ‘availability’ calculations, our analysis also provides insights into the seasonal timing and environmental factors affecting haul-out behavior by ice-associated seals. We specifically focused on bearded (Erignathus barbatus), ribbon (Histriophoca fasciata), and spotted seals (Phoca largha) in the Bering and Chukchi seas. Because ringed seals (Phoca (pusa) hispida) can be out of the water but hidden from view in snow lairs analysis of their ‘availability’ to surveys requires special consideration; therefore, they were not included in this analysis. Using generalized linear mixed pseudo-models to properly account for temporal autocorrelation, we fit models with covariates of interest (e.g., day-of-year, solar hour, age and sex class, wind speed, barometric pressure, temperature, precipitation) to examine their ability to explain variation in haul-out probability. We found evidence for strong diel and within-season patterns in haul-out behavior, as well as strong weather effects (particularly wind and temperature). In general, seals were more likely to haul out on ice in the middle of the day and when wind speed was low and temperatures were higher. Haul-out probability increased through March and April, peaking in May and early June before declining again. The timing and frequency of haul-out events also varied based on species and age-sex class. For ribbon and spotted seals, models with year effects were highly supported, indicating that the timing and magnitude of haul-out behavior varied among years. However, we did not find broad evidence that haul-out timing was linked to annual sea-ice extent. Our analysis emphasizes the importance of accounting for seasonal and temporal variation in haul-out behavior, as well as associated environmental covariates, when interpreting the number of seals counted in aerial surveys.
The SSN2 R package provides tools for spatial statistical modeling, parameter estimation, and prediction on stream (river) networks. SSN2 is the successor to the SSN R package (Ver Hoef, Peterson, Clifford, & Shah, 2014), which was archived alongside broader changes in the R-spatial ecosystem (Nowosad, 2023) that included 1) the retirement of rgdal (Bivand, Keitt, & Rowlingson, 2021), rgeos (Bivand & Rundel, 2020), and maptools (Bivand & Lewin-Koh, 2021) and 2) the lack of active development of sp (Bivand, Pebesma, & Gómez-Rubio, 2013). SSN2 maintains compatibility with the input data file structures used by the SSN R package but leverages modern R-spatial tools like sf (Pebesma, 2018). SSN2 also provides many useful features that were not available in the SSN R package, including new modeling and helper functions, enhanced fitting algorithms, and simplified syntax consistent with other R generic functions.
Many applied researchers equate spatial statistics with prediction or mapping, but this book naturally extends linear models, which includes regression and ANOVA as pillars of applied statistics, to achieve a more comprehensive treatment of the analysis of spatially autocorrelated data. Spatial Linear Models for Environmental Data, aimed at students and professionals with a master's level training in statistics, presents a unique, applied, and thorough treatment of spatial linear models within a statistics framework. Two subfields, one called geostatistics and the other called areal or lattice models, are extensively covered. Zimmerman and Ver Hoef present topics clearly, using many examples and simulation studies to illustrate ideas. By mimicking their examples and R code, readers will be able to fit spatial linear models to their data and draw proper scientific conclusions. Topics covered include: Exploratory methods for spatial data including outlier detection, (semi)variograms, Moran's I, and Geary's c. Ordinary and generalized least squares regression methods and their application to spatial data. Suitable parametric models for the mean and covariance structure of geostatistical and areal data. Model-fitting, including inference methods for explanatory variables and likelihood-based methods for covariance parameters. Practical use of spatial linear models including prediction (kriging), spatial sampling, and spatial design of experiments for solving real world problems. All concepts are introduced in a natural order and illustrated throughout the book using four datasets. All analyses, tables, and figures are completely reproducible using open-source R code provided at a GitHub site. Exercises are given at the end of each chapter, with full solutions provided on an instructor's FTP site supplied by the publisher.
Conductivity is an important indicator of the health of aquatic ecosystems. We model large amounts of lake conductivity data collected as part of the United States Environmental Protection Agency’s National Lakes Assessment using spatial indexing, a flexible and efficient approach to fitting spatial statistical models to big data sets. Spatial indexing is capable of accommodating various spatial covariance structures as well as features like random effects, geometric anisotropy, partition factors, and non-Euclidean topologies. We use spatial indexing to compare lake conductivity models and show that calcium oxide rock content, crop production, human development, precipitation, and temperature are strongly related to lake conductivity. We use this model to predict lake conductivity at hundreds of thousands of lakes distributed throughout the contiguous United States. We find that lake conductivity models fit using spatial indexing are nearly identical to lake conductivity models fit using traditional methods but are nearly 50 times faster (sample size 3,311). Spatial indexing is readily available in the spmodel R package.
The use of in-situ digital sensors for water quality monitoring is becoming increasingly common worldwide. While these sensors provide near real-time data for science, the data are prone to technical anomalies that can undermine the trustworthiness of the data and the accuracy of statistical inferences, particularly in spatial and temporal analyses. Here we propose a framework for detecting anomalies in sensor data recorded in stream networks, which takes advantage of spatial and temporal autocorrelation to improve detection rates. The proposed framework involves the implementation of effective data imputation to handle missing data, alignment of time-series to address temporal disparities, and the identification of water quality events. We explore the effectiveness of a suite of state-of-the-art statistical methods including posterior predictive distributions, finite mixtures, and Hidden Markov Models (HMM). We showcase the practical implementation of automated anomaly detection in near-real time by employing a Bayesian recursive approach. This demonstration is conducted through a comprehensive simulation study and a practical application to a substantive case study situated in the Herbert River, located in Queensland, Australia, which flows into the Great Barrier Reef. We found that methods such as posterior predictive distributions and HMM produce the best performance in detecting multiple types of anomalies. Utilizing data from multiple sensors deployed relatively near one another enhances the ability to distinguish between water quality events and technical anomalies, thereby significantly improving the accuracy of anomaly detection. Thus, uncertainty and biases in water quality reporting, interpretation, and modelling are reduced, and the effectiveness of subsequent management actions improved.