Spatial prediction of exposure to air pollution in a large city such as Santiago de Chile is a challenging problem because of the lack of a dense air‐quality monitoring network. Statistical spatiotemporal models exploit the space–time correlation in the pollution data and other relevant meteorological and land‐use information to generate accurate predictions in both space and time. In this paper, we develop a Bayesian modeling method to accurately predict hourly PM 2.5 concentrations in a 1‐km high‐resolution grid covering the city. The modeling method combines a spatiotemporal land‐use regression model for PM 2.5 and a Bayesian calibration model for the input meteorological variables used in the land‐use regression model. Using a 3‐month winter‐time pollution data set, the output of sample validation results obtained in this paper shows a substantial increase in accuracy due to the incorporation of the linear calibration model. The proposed Bayesian modeling method is then used to provide short‐term spatiotemporal predictions of PM 2.5 concentrations on a fine (1 km 2 ) spatial grid covering the city. Along with the paper, we publish the R code used and the output of sample predictions for future scientific use.
Recently, there has been a surge of interest in Bayesian space-time modeling of daily maximum eight-hour average ozone concentration levels. Hierarchical models based on well known time series modeling methods such as the dynamic linear models (DLM) and the auto-regressive (AR) models are often used in the literature. The DLM, developed as a result of the popularity of Kalman filtering methods, provide a dynamical state-space system that is thought to evolve from a pair of state and observation equations. The AR models, on the other hand, cast in a Bayesian hierarchical setting, have recently been developed through a pair of models where a measurement error model is formulated at the top level and an AR model for the true ozone concentration levels is postulated at the next level. Each of the modeling scenarios is set in an appropriate multivariate setting to model the spatial dependence. This paper compares these two methods in hierarchical Bayesian settings. A simplified skeletal version of the DLM taken from Dou et al. (2010) [5] is compared theoretically with a matching hierarchical AR model. The comparisons reveal many important differences in the induced space-time correlation structures. Further comparisons of the variances of the predictive distributions by conditioning on different sets of data for each model show superior performances of the AR models under certain conditions. These theoretical investigations are followed up by a simulation study and a real data example implemented using Markov chain Monte Carlo (MCMC) methods for modeling daily maximum eight-hour average ozone concentration levels observed in the state of New York in the months of July and August, 2006. The hierarchical AR model is chosen using all the model choice criteria considered in this example. (C) 2011 Elsevier B.V. All rights reserved.
Summary The problem motivating the paper is the determination of sample size in clinical trials under normal likelihoods and at the substantive testing stage of a financial audit where normality is not an appropriate assumption. A combination of analytical and simulation-based techniques within the Bayesian framework is proposed. The framework accommodates two different prior distributions: one is the general purpose fitting prior distribution that is used in Bayesian analysis and the other is the expert subjective prior distribution, the sampling prior which is believed to generate the parameter values which in turn generate the data. We obtain many theoretical results and one key result is that typical non-informative prior distributions lead to very small sample sizes. In contrast, a very informative prior distribution may either lead to a very small or a very large sample size depending on the location of the centre of the prior distribution and the hypothesized value of the parameter. The methods that are developed are quite general and can be applied to other sample size determination problems. Some numerical illustrations which bring out many other aspects of the optimum sample size are given.
Summary Short-term forecasts of air pollution levels in big cities are now reported in news-papers and other media outlets. Studies indicate that even short-term exposure to high levels of an air pollutant called atmospheric particulate matter can lead to long-term health effects. Data are typically observed at fixed monitoring stations throughout a study region of interest at different time points. Statistical spatiotemporal models are appropriate for modelling these data. We consider short-term forecasting of these spatiotemporal processes by using a Bayesian kriged Kalman filtering model. The spatial prediction surface of the model is built by using the well-known method of kriging for optimum spatial prediction and the temporal effects are analysed by using the models underlying the Kalman filtering method. The full Bayesian model is implemented by using Markov chain Monte Carlo techniques which enable us to obtain the optimal Bayesian forecasts in time and space. A new cross-validation method based on the Mahalanobis distance between the forecasts and observed data is also developed to assess the forecasting performance of the model implemented.
The authors develop a new class of distributions by introducing skewness in multivariate elliptically symmetric distributions. The class, which is obtained by using transformation and conditioning, contains many standard families including the multivariate skew-normal and t distributions. The authors obtain analytical forms of the densities and study distributional properties. They give practical applications in Bayesian regression models and results on the existence of the posterior distributions and moments under improper priors for the regression coefficients. They illustrate their methods using practical examples.
The authors propose a procedure for determining the unknown number of components in mixtures by generalizing a Bayesian testing method proposed by Mengersen & Robert (1996). The testing criterion they propose involves a Kullback-Leibler distance, which may be weighted or not. They give explicit formulas for the weighted distance for a number of mixture distributions and propose a stepwise testing procedure to select the minimum number of components adequate for the data. Their procedure, which is implemented using the BUGS software, exploits a fast collapsing approach which accelerates the search for the minimum number of components by avoiding full refitting at each step. The performance of their method is compared, using both distances, to the Bayes factor approach.
A new method of construction of Markov chains with a given stationary distribution isproposed. The method is based on constructing an auxiliary chain with some other stationarydistribution and picking elements of this auxiliary chain a suitable number of times.The proposed method is easy to implement and analyse; it may be more efficient thanother related Markov chain Monte Carlo techniques. The main attractive feature of theassociated Markov chain is that it regenerates whenever it accepts a new proposed point.This makes the algorithm easy to adapt and tune for practical problems. A theoreticalstudy and numerical comparisons with some other available Markov chain Monte Carlotechniques are presented.
Item response models are essential tools for analyzing results from many educational and psychological tests. Such models are used to quantify the probability of correct response as a function of unobserved examinee ability and other parameters explaining the difficulty and the discriminatory power of the questions in the test. Some of these models also incorporate a threshold parameter for the probability of the correct response to account for the effect of guessing the correct answer in multiple choice type tests. In this article we consider fitting of such models using the Gibbs sampler. A data augmentation method to analyze a normal-ogive model incorporating a threshold guessing parameter is introduced and compared with a Metropolis-Hastings sampling method. The proposed method is an order of magnitude more efficient than the existing method. Another objective of this paper is to develop Bayesian model choice techniques for model discrimination. A predictive approach based on a variant of the Bayes factor is used and compared with another decision theoretic method which minimizes an expected loss function on the predictive space. A classical model choice technique based on a modified likelihood ratio test statistic is shown as one component of the second criterion. As a consequence the Bayesian methods proposed in this paper are contrasted with the classical approach based on the likelihood ratio test. Several examples are given to illustrate the methods.
This article aims to provide a method for approximately predetermining convergence properties of the Gibbs sampler. This is to be done by first finding an approximate rate of convergence for a normal approximation of the target distribution. The rates of convergence for different implementation strategies of the Gibbs sampler are compared to find the best one. In general, the limiting convergence properties of the Gibbs sampler on a sequence of target distributions (approaching a limit) are not the same as the convergence properties of the Gibbs sampler on the limiting target distribution. Theoretical results are given in this article to justify that under conditions, the convergence properties of the Gibbs sampler can be approximated as well. A number of practical examples are given for illustration.
SUMMARY For many years, archaeologists have postulated that the numbers of various artefact types found within excavated features should give insight about their relative dates of deposition even when stratigraphic information is not present. A typical data set used in such studies can be reported as a cross-classification table (often called an abundance matrix or, equivalently, a contingency table) of excavated features against artefact types. Each entry of the table represents the number of a particular artefact type found in a particular archaeological feature. Methodologies for attempting to identify temporal sequences on the basis of such data are commonly referred to as seriation techniques. Several different procedures for seriation including both parametric and non-parametric statistics have been used in an attempt to reconstruct relative chronological orders on the basis of such contingency tables. We develop some possible model-based approaches that might be used to aid in relative, archaeological chronology building. We use the recently developed Markov chain Monte Carlo method based on Langevin diffusions to fit some of the models proposed. Predictive Bayesian model choice techniques are then employed to ascertain which of the models that we develop are most plausible. We analyse two data sets taken from the literature on archaeological seriation.
Summary In this paper many convergence issues concerning the implementation of the Gibbs sampler are investigated. Exact computable rates of convergence for Gaussian target distributions are obtained. Different random and non-random updating strategies and blocking combinations are compared using the rates. The effect of dimensionality and correlation structure on the convergence rates are studied. Some examples are considered to demonstrate the results. For a Gaussian image analysis problem several updating strategies are described and compared. For problems in Bayesian linear models several possible parameterizations are analysed in terms of their convergence rates characterizing the optimal choice.
The generality and easy programmability of modern sampling-based methods for maximisation of likelihoods and summarisation of posterior distributions have led to a tremendous increase in the complexity and dimensionality of the statistical models used in practice. However, these methods can often be extremely slow to converge, due to high correlations between, or weak identifiability of, certain model parameters. We present simple hierarchical centring reparametrisations that often give improved convergence for a broad class of normal linear mixed models. In particular, we study the two-stage hierarchical normal linear model, the Laird-Ware model for longitudinal data, and a general structure for hierarchically nested linear models. Using analytical arguments, simulation studies, and an example involving clinical markers of acquired immune deficiency syndrome (AIDS), We indicate when reparametrisation is likely to provide substantial gains in efficiency.
The scan test for clustering in time is based on the maximum number of events in an interval (window) of width w as the window moves across the entire time frame. Power estimates of the scan statistic are simulated for a variety of epidemiologically motivated situations. Two cluster configurations are used: a rectangular pulse, and a triangular pulse designed to emulate environmental contamination. For a rectangular pulse, the relative risk R of disease in the cluster region is R-fold as high as it is for the background region. The power is strongly influenced by the sample size, the relative risk, and the width or duration of the cluster region, whereas the effect of the cluster configuration is small. Using a 5 per cent significance level, a relative risk of 4, a standardized cluster duration of 0.10, a relative window width of 1.5, and a (non-random) sample size of 50, the simulated power is approximately 80 per cent, indicating that the minimum sample size in the cluster region for adequate power is in the 12-32 range for values of the parameters used in this study.
We revisit the bounded maximal risk point estimation problem as well as the fixed-width confidence interval estimation problem for the largest mean amongk(≥2) independent normal populations having unknown means and unknown but equal variance. In the point estimation setup, we devise appropriate two-stage and modified two-stage methodologies so that the associatedmaximal risk can bebounded from aboveexactly by a preassigned positive number. Kuo and Mukhopadhyay (1990), however, emphasized only the asymptotics in this context. We have also introduced, in both point and interval estimation problems,accelerated sequential methodologies thereby saving sampling operations tremendously over the purely sequential schemes considered in Kuo and Mukhopadhyay (1990), but enjoying at the same time asymptotic second-order characteristics, fairly similar to those of the purely sequential ones.