Is there spatial pattern? Spatial pattern implies that observations from units closer to each other are more similar than those recorded in units farther away. Do we want to smooth the data? Perhaps to adjust for low population sizes (or sample sizes) in certain units? How much do we want to smooth? Inference for new areal units? Is prediction meaningful here? If we modify the areal units to new units (e.g. from zip codes to county values), what can we say about the new counts we expect for the latter given those for the former? This is the Modifiable Areal Unit Problem (MAUP) or Misalignment.
OVERVIEW OF SPATIAL DATA PROBLEMS Introduction to Spatial Data and Models Fundamentals of Cartography Exercises BASICS OF POINT-REFERENCED DATA MODELS Elements of Point-Referenced Modeling Spatial Process Models Exploratory Approaches for Point-Referenced Data Classical Spatial Prediction Computer Tutorials Exercises BASICS OF AREAL DATA MODELS Exploratory Approaches for Areal Data Brook's Lemma and Markov Random Fields Conditionally Autoregressive (CAR) Models Simultaneous Autoregressive (SAR) Models Computer Tutorials Exercises BASICS OF BAYESIAN INFERENCE Introduction to Hierarchical Modeling and Bayes Theorem Bayesian Inference Bayesian Computation Computer Tutorials Exercises HIERARCHICAL MODELING FOR UNIVARIATE SPATIAL DATA Stationary Spatial Process Models Generalized Linear Spatial Process Modeling Nonstationary Spatial Process Models Areal Data Models General Linear Areal Data Modeling Exercises SPATIAL MISALIGNMENT Point-Level Modeling Nested Block-Level Modeling Nonnested Block-Level Modeling Misaligned Regression Modeling Exercises MULTIVARIATE SPATIAL MODELING Separable Models Coregionalization Models Other Constructive Approaches Multivariate Models for Areal Data Exercises SPATIOTEMPORAL MODELING General Modeling Formulation Point-Level Modeling with Continuous Time Nonseparable Spatio-Temporal Models Dynamic Spatio-Temporal Models Block-Level Modeling Exercises SPATIAL SURVIVAL MODELS Parametric Models Semiparametric Models Spatio-Temporal Models Multivariate Models Spatial Cure Rate Models Exercises SPECIAL TOPICS IN SPATIAL PROCESS MODELING Process Smoothness Revisited Spatially Varying Coefficient Models Spatial CDFs APPENDICES Matrix Theory and Spatial Computing Methods Answers to Selected Exercises REFERENCES AUTHOR INDEX SUBJECT INDEX Short TOC
Disease mapping attempts to explain observed health event counts across areal units, typically using Markov random field models. These models rely on spatial priors to account for variation in raw relative risk or rate estimates. Spatial priors introduce some degree of smoothing, wherein, for any particular unit, empirical risk or incidence estimates are either adjusted towards a suitable mean or incorporate neighbor-based smoothing. While model explanation may be the primary focus, the literature lacks a comparison of the amount of smoothing introduced by different spatial priors. Additionally, there has been no investigation into how varying the parameters of these priors influences the resulting smoothing. This study examines seven commonly used spatial priors through both simulations and real data analyses. Using areal maps of peninsular Spain and England, we analyze smoothing effects using two datasets with associated populations at risk. We propose empirical metrics to quantify the smoothing achieved by each model and theoretical metrics to calibrate the expected extent of smoothing as a function of model parameters. We employ areal maps in order to quantitatively characterize the extent of smoothing within and across the models as well as to link the theoretical metrics to the empirical metrics.
North Atlantic right whales are an endangered species. Their entire population is estimated to be approximately 372 individuals, and they are subject to major anthropogenic threats. They feed on zooplankton species whose distribution shifts in a dynamic and warming oceanic environment. Because right whales in turn follow their shifting food resource, it is necessary to jointly study the distribution of whales and their prey. The innovative joint species distribution modelling (JSDM) contribution here is different from anything in the large JDSM literature, reflecting the processes and data we have to work with. Specifically, our JSDM supplies a geostatistical model for the expected amount of zooplankton collected at a site. We require a point pattern model for the intensity of right whale abundance. The two process models are linked through a latent conditional-marginal specification. Furthermore, each species has two data sources informing its respective distribution, necessitating a novel data fusion approach. The result is a complex multi-level model. Through simulation, we demonstrate that our joint specification effectively identifies model unknowns and improves the estimation of species distributions compared to modelling them separately. We then apply our model to real data from Cape Cod Bay, Massachusetts, USA.
We will use forest inventory data from the U.S. Department of Agriculture Forest Service, Bartlett Experimental Forest (BEF), Bartlett, NH. This dataset holds 1991 and 2002 forest inventory data for 437 plots. Variables include species specific basal area and total tree biomass; inventory plot coordinates; slope; elevation; and tasseled cap brightness (TC1), greenness (TC2), and wetness (TC3) components from spring, summer, and fall 2002 Landsat images. We use these data to demonstrate some basics of spatial data manipulation, visualization, and univariate spatial regression analysis. The regression model will be used to make prediction of biomass for every image pixel across the BEF. We begin by removing non-forest inventory plots, converting biomass measurements from kilograms per hectare to the log of metric tons per hectare, and taking a look at plot locations across the forest.
A frequent challenge encountered in real-world applications is data having a high proportion of zeros. Focusing on ecological abundance data, much attention has been given to zero-inflated count data. Models for non-negative continuous abundance data with an excess of zeros are rarely discussed. Work presented here considers the creation of a point mass at zero through a left-censoring approach or through a hurdle approach. We incorporate both mechanisms to capture the analogue of zero-inflation for count data. Additionally, primary attention has been given to univariate zero-inflated modeling (e.g., single species), whereas data often arise jointly (e.g., a collection of species). With multivariate abundance data, a key issue is to capture dependence among the species at a site, both in terms of positive abundance as well as absence. Therefore, our contribution is a model for multivariate zero-inflated continuous data that are non-negative. Working in a Bayesian framework, we discuss the issue of separating the two sources of zeros and offer model comparison metrics for multivariate zero-inflated data. In an application, we model the total biomass for five tree species obtained from plots established in the Forest Inventory Analysis database in the Northeast region of the United States.
Sound is assumed to be the primary modality of communication among marine mammal species. Analyzing acoustic recordings helps to understand the function of the acoustic signals as well as the possible impact of anthropogenic noise on acoustic behavior. Motivated by a dataset from a network of hydrophones in Cape Cod Bay, Massachusetts, utilizing automatically detected calls in recordings, we study the communication process of the endangered North Atlantic right whale. For right whales an "up-call" is known as a contact call, and ensuing counter-calling between individuals is presumed to facilitate group cohesion. We present novel spatiotemporal excitement modeling consisting of a background process and a counter-call process. The background process intensity incorporates the influences of diel patterns and ambient noise on occurrence. The counter-call intensity captures potential excitement, that calling elicits calling behavior. Call incidence is found to be clustered in space and time; a call seems to excite more calls nearer to it in time and space. We find evidence that whales make more calls during twilight hours, respond to other whales nearby, and are likely to remain quiet in the presence of increased ambient noise.
Record-breaking temperature events are now frequently in the news, proffered as evidence of climate change, and often bring significant economic and human impacts. Our previous work undertook the first substantial spatial modelling investigation of temperature record-breaking across years for any given day within the year, employing a dataset consisting of over 60 years of daily maximum temperatures across peninsular Spain. That dataset also supplies daily minimum temperatures (which, in fact, are now available through 2023). Here, the dataset is converted into a daily pair of binary events, indicators, for that day, of whether a yearly record was broken for the daily maximum temperature and/or for the daily minimum temperature. Joint modelling addresses several inference issues: (i) defining/modelling record-breaking with bivariate time series of yearly indicators, (ii) strength of relationship between record-breaking events, (iii) prediction of joint, conditional and marginal record-breaking, (iv) persistence in record-breaking across days, and (v) spatial interpolation across peninsular Spain. We substantially expand our previous work to enable investigation of these issues. We observe strong correlation between both processes but a growing trend of climate change that is well differentiated between them both spatially and temporally as well as different strengths of persistence and spatial dependence.
Marine mammals are increasingly vulnerable to human disturbance and climate change. Their diving behavior leads to limited visual access during data collection, making studying the abundance and distribution of marine mammals challenging. In theory, using data from more than one observation modality should lead to better informed predictions of abundance and distribution. With focus on North Atlantic right whales, we consider the fusion of two data sources to inform about their abundance and distribution. The first source is aerial distance sampling which provides the spatial locations of whales detected in the region. The second source is passive acoustic monitoring (PAM), returning calls received at hydrophones placed on the ocean floor. Due to limited time on the surface and detection limitations arising from sampling effort, aerial distance sampling only provides a partial realization of locations. With PAM, we never observe numbers or locations of individuals. To address these challenges, we develop a novel thinned point pattern data fusion. Our approach leads to improved inference regarding abundance and distribution of North Atlantic right whales throughout Cape Cod Bay, Massachusetts in the US. We demonstrate performance gains of our approach compared to that from a single source through both simulation and real data.
Abstract Turnover, or change in the composition of species over space and time, is one of the primary ways to define beta diversity. Inferring what factors impact beta diversity is not only important for understanding biodiversity processes but also for conservation planning. At present, a popular approach to understanding the drivers of compositional turnover is through generalized dissimilarity modelling (GDM). We argue that the current GDM approach suffers several limitations and provide an alternative modelling approach that remedies these issues. We propose using generative spatial random effects models implemented in a Bayesian framework. We offer hierarchical specifications to yield full regression and spatial predictive inference, both with associated full uncertainties. The approach is illustrated by examining dissimilarity in three datasets: tree survey data from Panama's Barro Colorado Island (BCI), plant occurrence data from southwest Australia and plant abundance surveys from the Greater Cape Floristic Region (GCFR) of South Africa. We select a best model using out‐of‐sample predictive performance. We find that the form of the best model differs across the three datasets, but our models provide performance ranging from comparable to significant improvement over GDMs. Within the GCFR, the spatial random effects play a more important role in the modelling than all the environmental variables. We have proposed a model that provides several improvements to the current GDM framework. This includes advantages such as a flexible spatially varying mean function, spatial random effects that capture dependence unaccounted for by explanatory variables, and spatially heterogeneous variance structure. All these features are offered in a model that can adequately handle a large incidence of total dissimilarity through ‘one‐inflation’, as would be expected from highly biodiverse areas with steep turnover gradients.
Record-breaking temperature events are now very frequently in the news, viewed as evidence of climate change. With this as motivation, we undertake the first substantial spatial modeling investigation of temperature record-breaking across years for any given day within the year. We work with a dataset consisting of over sixty years (1960-2021) of daily maximum temperatures across peninsular Spain. Formal statistical analysis of record-breaking events is an area that has received attention primarily within the probability community, dominated by results for the stationary record-breaking setting with some additional work addressing trends. Such effort is inadequate for analyzing actual record-breaking data. Effective analysis requires rich modeling of the indicator events which define record-breaking sequences. Resulting from novel and detailed exploratory data analysis, we propose hierarchical conditional models for the indicator events. After suitable model selection, we discover explicit trend behavior, necessary autoregression, significance of distance to the coast, useful interactions, helpful spatial random effects, and very strong daily random effects. Illustratively, the model estimates that global warming trends have increased the number of records expected in the past decade almost two-fold, 1.93 (1.89,1.98), but also estimates highly differentiated climate warming rates in space and by season.
Ecological modelling often involves addressing challenges such as dependence in responses, e.g., spatial and/or temporal correlation, heterogeneity of variance, and hierarchical structures inherent in ecological processes and data. A constant challenge is the inadequacy of the data to well address the questions of interest. What is observable may not be sufficiently informative. What has been observed may not have been well designed. Carefully conceived modelling can take us to improved inference compared with adopting standard inference tools like basic regression and analysis of variance. Hierarchical modelling techniques provide a powerful framework for capturing these complexities by explicitly modelling the multi-level structure of ecological systems.In this paper we focus on good modelling practice in the hierarchical Bayesian framework. We discuss good modelling practice in the form of 10 steps, elaborated with a running example, to aid quantitative ecologists interested in employing Bayesian modelling. Two of the authors are statisticians who view themselves primarily as stochastic modellers. The other two are quantitative ecologists who view such stochastic tools as essential for advancing ecological knowledge. Hence, implicitly, we pay careful attention to good modelling practice. Further, we argue the benefits that accrue to working in a Bayesian framework as the paper is developed.It is worth noting that the steps we propose, though presented in the context of ecological modelling, are appropriate for effective model building across most fields of application.
Cuvier's beaked whales (Ziphius cavirostris) are the deepest diving marine mammal, consistently diving to depths exceeding 1,000m for durations longer than an hour, making them difficult animals to study. They are important to study because they are sensitive to disturbances from naval sonar. Satellite-linked telemetry devices provide up to 14-day long records of dive behavior. However, the time series of depths is discretized to coarse bins due to bandwidth limitations. We analyze telemetry data from beaked whales that were exposed to moderate levels of sonar within controlled exposure experiments (CEEs) to study behavioral responses to sound exposure. We model the data as a hidden Markov model (HMM) over the time series of discrete depth bins, introducing partially observed movement types and recent diving activity covariates to model marginal non-stationarity. Movement types provide more flexible modeling for CEEs than partially observed dive stages, which are more commonly used in dive behavior HMMs. We estimate the proposed model within a hierarchical Bayesian framework, using HMM methods to compute marginalized likelihoods and posterior predictive distributions. We assess behavioral response by comparing observed post-exposure behavior to usual unexposed behavior via the posterior predictive distribution. The model quantifies patterns in baseline diving behavior and finds evidence that beaked whales deviate in response to sound. We find evidence that (i) beaked whales initially shorten the time they spend between deep dives, which may have physiological effects and (ii) subsequently avoid deep dives, which can result in lost foraging opportunities.
The investigation of leaf-level traits in response to varying environmental conditions has immense importance for understanding plant ecology. Remote sensing technology enables measurement of the reflectance of plants to make inferences about underlying traits along environmental gradients. While much focus has been placed on understanding how reflectance and traits are related at the leaf-level, the challenge of modelling the dependence of this relationship while accounting for environmental gradients has limited this line of inquiry. Here, we take up the problem of jointly modeling traits and reflectance given environment. Our objective is to assess not only response to environmental regressors but also dependence between trait levels and the reflectance spectrum in the context of this regression. We jointly model the response vector of traits with reflectance, which is a function of wavelength. To conduct this investigation, we employ a dataset from a global biodiversity hotspot, the Greater Cape Floristic Region in South Africa.
There is continuing interest in the investigation of change in temperature over space and time. For this analysis, we offer statistical tools to illuminate changes temporally, at desired temporal resolution, and spatially, using data generated from suitable space–time models. The proposed tools can be used with the output from any suitable model fitted to any set of spatially referenced time series data. The tools to assess space and time changes include spatial surfaces of probabilities and spatial extents for events defined by exceeding a threshold. The spatial surfaces capture the spatial variation in the probability or risk of an exceedance event, while the spatial extents capture the expected proportion of incidence of an event for a region of interest. This approach is used analyse the changes in daily maximum temperature in an inland Mediterranean region (NE of Spain) in the period 1956–2015. The area is very heterogeneous in orography and climate, including the central Ebro valley and part of the Pyrenees. We use a collection of daily temperature series obtained from simulation under a Bayesian daily temperature model fitted to 18 stations in that area. The results for the summer period show that, although there is an increasing risk in all the events used to quantify the effects of climate change, it is not spatially homogeneous, with the largest increase arising in the centre of the Ebro valley and the Eastern Pyrenees area. The risk of an increase in the average daily maximum temperature from 1966–1975 to 2006–2015 higher than 1°C is higher than 0.5 over all of the region, and close to 1 in the previous areas. The extent of daily maximum temperature higher than the reference mean has increased 3.5% per decade. The mean of the extent indicates that 95% of the area under study has suffered a positive increment of the average temperature, and almost 70% an increment higher than 1°C.
There is continuing interest in the investigation of change in temperature over space and time. For this analysis, we offer statistical tools to illuminate changes temporally, at desired temporal resolution, and spatially, using data generated from suitable space-time models. The proposed tools can be used with the output from any suitable model fitted to any set of spatially referenced time series data. The tools to assess space and time changes include spatial surfaces of probabilities and spatial extents for events defined by exceeding a threshold. The spatial surfaces capture the spatial variation in the probability or risk of an exceedance event, while the spatial extents capture the expected proportion of incidence of an event for a region of interest. This approach is used analyse the changes in daily maximum temperature in an inland Mediterranean region (NE of Spain) in the period 1956-2015. The area is very heterogeneous in orography and climate, including the central Ebro valley and part of the Pyrenees. We use a collection of daily temperature series obtained from simulation under a Bayesian daily temperature model fitted to 18 stations in that area. The results for the summer period show that, although there is an increasing risk in all the events used to quantify the effects of climate change, it is not spatially homogeneous, with the largest increase arising in the centre of the Ebro valley and the Eastern Pyrenees area. The risk of an increase in the average daily maximum temperature from 1966-1975 to 2006-2015 higher than 1degree celsius is higher than 0.5 over all of the region, and close to 1 in the previous areas. The extent of daily maximum temperature higher than the reference mean has increased 3.5% per decade. The mean of the extent indicates that 95% of the area under study has suffered a positive increment of the average temperature, and almost 70% an increment higher than 1degree celsius.
Regression is the most widely used modeling tool in statistics. Quantile regression offers a strategy for enhancing the regression picture beyond customary mean regression. With time-series data, we move to quantile autoregression and, finally, with spatially referenced time series, we move to spacetime quantile regression. Here, we are concerned with the spatiotemporal evolution of daily maximum temperature, particularly with regard to extreme heat. Our motivating data set is 60 years of daily summer maximum temperature data over Aragon in Spain. Hence, we work with time on two scales- days within summer season across years-collected at geocoded station locations. For a specified quantile, we fit a very flexible, mixed-effects autoregressive model, introducing four spatial processes. We work with asymmetric Laplace errors to take advantage of the available conditional Gaussian representation for these distributions. Further, while the autoregressive model yields conditional quantiles, we demonstrate how to extract marginal quantiles with the asymmetric Laplace specification. Thus, we are able to interpolate quantiles for any days within years across our study region.