This chapter covers a mix of topics including splines, functional data, extreme values and density estimation. Being associated with continuous time monitoring processes, functional data are usually smooth curves or surfaces and are often treated as realizations of underlying random functions. The basic philosophy of functional data is to consider the observed curves as single entities, rather than only as a sequence of individual observations. Comprehensive surveys of statistical techniques for analyzing functional data can be found in Ramsay and Silverman and Ferraty and Vieu. Goldsmith study Diffusion Tensor Imaging (DTI) metrics of multiple sclerosis (MS) patients over multiple clinical visits. The data consist of 100 subjects, aged between 21 and 70 years at first visit. The number of visits per subject ranges from 2 to 8, and a total of 340 visits were recorded. Most statistical modeling is concerned with the mean …
SummaryWe define residuals for point process models fitted to spatial point pattern data, and we propose diagnostic plots based on them. The residuals apply to any point process model that has a conditional intensity; the model may exhibit spatial heterogeneity, interpoint interaction and dependence on spatial covariates. Some existing ad hoc methods for model checking (quadrat counts, scan statistic, kernel smoothed intensity and Berman's diagnostic) are recovered as special cases. Diagnostic tools are developed systematically, by using an analogy between our spatial residuals and the usual residuals for (non-spatial) generalized linear models. The conditional intensity λ plays the role of the mean response. This makes it possible to adapt existing knowledge about model validation for generalized linear models to the spatial point process context, giving recommendations for diagnostic plots. A plot of smoothed residuals against spatial location, or against a spatial covariate, is effective in diagnosing spatial trend or co-variate effects. Q–Q-plots of the residuals are effective in diagnosing interpoint interaction.
This paper describes a method for estimating the risk from a disease over a set of contiguous geographical regions, when data on a potentially important covariate, such as race, are not available. Conditions under which the extra margin can be recovered are suggested. An application to prostate cancer mortality among the non-white population in the counties of the U.S.A. is discussed.
This paper combines existing models for longitudinal and spatial data in a hierarchical Bayesian framework, with particular emphasis on the role of time- and space-varying covariate effects. Data analysis is implemented via Markov chain Monte Carlo methods. The methodology is illustrated by a tentative re-analysis of Ohio lung cancer data 1968-1988. Two approaches that adjust for unmeasured spatial covariates, particularly tobacco consumption, are described. The first includes random effects in the model to account for unobserved heterogeneity; the second adds a simple urbanization measure as a surrogate for smoking behaviour. The Ohio data set has been of particular interest because of the suggestion that a nuclear facility in the southwest of the state may have caused increased levels of lung cancer there. However, we contend here that the data are inadequate for a proper investigation of this issue.
Gaussian conditional autoregressions have been widely used in spatial statistics and Bayesian image analysis, where they are intended to describe interactions between random variables at fixed sites in Euclidean space. The main appeal of these distributions is in the Markovian interpretation of their full conditionals. Intrinsic autoregressions are limiting forms that retain the Markov property. Despite being improper, they can have advantages over the standard autoregressions, both conceptually and in practice. For example, they often avoid difficulties in parameter estimation, without apparent loss, or exhibit appealing invariances, as in texture analysis. However, on small arrays and in nonlattice applications, both forms of autoregression can lead to undesirable second-order characteristics, either in the variables themselves or in contrasts among them. This paper discusses standard and intrinsic autoregressions and describes how the problems that arise can be alleviated using Dempster's (1972) algorithm or an appropriate modification. The approach represents a partial synthesis of standard geostatistical and Gaussian Markov random field formulations. Some nonspatial applications are also mentioned.
Markov chain Monte Carlo (MCMC) methods have been used extensively in statistical physics over the last 40 years, in spatial statistics for the past 20 and in Bayesian image analysis over the last decade. In the last five years, MCMC has been introduced into significance testing, general Bayesian inference and maximum likelihood estimation. This paper presents basic methodology of MCMC, emphasizing the Bayesian paradigm, conditional probability and the intimate relationship with Markov random fields in spatial statistics. Hastings algorithms are discussed, including Gibbs, Metropolis and some other variations. Pairwise difference priors are described and are used subsequently in three Bayesian applications, in each of which there is a pronounced spatial or temporal aspect to the modeling. The examples involve logistic regression in the presence of unobserved covariates and ordinal factors; the analysis of agricultural field experiments, with adjustment for fertility gradients; and processing of low-resolution medical images obtained by a gamma camera. Additional methodological issues arise in each of these applications and in the Appendices. The paper lays particular emphasis on the calculation of posterior probabilities and concurs with others in its view that MCMC facilitates a fundamental breakthrough in applied Bayesian modeling.
SUMMARY Markov chain Monte Carlo (MCMC) algorithms, such as the Gibbs sampler, have provided a Bayesian inference machine in image analysis and in other areas of spatial statistics for several years, founded on the pioneering ideas of Ulf Grenander. More recently, the observation that hyperparameters can be included as part of the updating schedule and the fact that almost any multivariate distribution is equivalently a Markov random field has opened the way to the use of MCMC in general Bayesian computation. In this paper, we trace the early development of MCMC in Bayesian inference, review some recent computational progress in statistical physics, based on the introduction of auxiliary variables, and discuss its current and future relevance in Bayesian applications. We briefly describe a simple MCMC implementation for the Bayesian analysis of agricultural field experiments, with which we have some practical experience.
Tests for clustering of rare diseases investigate whether an observed pattern of cases in one or more geographical regions could reasonably have arisen by chance alone, bearing in mind the variation in background population density. In contrast, tests for the detection of clusters are concerned with screening a large region for evidence of individual 'hot spots' of disease but without any preconception about their likely locations; the results of such tests may form the basis for subsequent small area investigation, statistical or non-statistical, but will rarely be an end in themselves. The main intention of the paper is to describe and illustrate a new technique for the identification of small clusters of disease. A secondary purpose is to discuss some common pitfalls in the application of tests of clustering to epidemiological data.
The assessment of statistical significance by Monte Carlo simulation may be costly in computer time. This paper looks at a number of ways of calculating exact Monte Carlo p-values by sequential sampling. Such p-values are shown to have properties similar to those obtained by sampling with a fixed sample size. Both standard and generalized Monte Carlo procedures are discussed and, in particular, sequential method is proposed for dealing with situations in which values can only be conveniently generated using a Markov chain, conditioned to pass through the observed data.
Journal Article A candidate's formula: A curious result in Bayesian prediction Get access JULIAN BESAG JULIAN BESAG Department of Mathematical Sciences, University of DurhamDurham DH1 3LE, U.K. Search for other works by this author on: Oxford Academic Google Scholar Biometrika, Volume 76, Issue 1, March 1989, Page 183, https://doi.org/10.1093/biomet/76.1.183 Published: 01 March 1989 Article history Received: 01 April 1988 Revision received: 01 October 1988 Published: 01 March 1989
SUMMARY Simple Monte Carlo significance testing has many applications, particularly in the preliminary analysis of spatial data. The method requires the value of the test statistic to be ranked among a random sample of values generated according to the null hypothesis. However, there are situations in which a sample of values can only be conveniently generated using a Markov chain, initiated by the observed data, so that independence is violated. This paper describes two methods that overcome the problem of dependence and allow exact tests to be carried out. The methods are applied to the Rasch model, to the finite lattice Ising model and to the testing of association between spatial processes. Power is discussed in a simple case.
The paper describes and provides examples of four different applications of the use of neighbouring plot values in the analysis of agricultural field experiments. The topics include the use of check plots to control environmental variation in large unreplicated variety trials; spatial models, based on first differences, to accommodate fertility effects in trials that have some degree of replication; adjustment for interplot competition using a simultaneous-equations formulation; and the analysis of trials in which plot values are affected by the particular treatments on neighbouring plots.
SUMMARY A continuous two-dimensional region is partitioned into a fine rectangular array of sites or “pixels”, each pixel having a particular “colour” belonging to a prescribed finite set. The true colouring of the region is unknown but, associated with each pixel, there is a possibly multivariate record which conveys imperfect information about its colour according to a known statistical model. The aim is to reconstruct the true scene, with the additional knowledge that pixels close together tend to have the same or similar colours. In this paper, it is assumed that the local characteristics of the true scene can be represented by a non-degenerate Markov random field. Such information can be combined with the records by Bayes' theorem and the true scene can be estimated according to standard criteria. However, the computational burden is enormous and the reconstruction may reflect undesirable large-scale properties of the random field. Thus, a simple, iterative method of reconstruction is proposed, which does not depend on these large-scale characteristics. The method is illustrated by computer simulations in which the original scene is not directly related to the assumed random field. Some complications, including parameter estimation, are discussed. Potential applications are mentioned briefly.
Starting from a suitable sequence of auto-Poisson lattice schemes, it is shown that (almost) any purely inhibitory pairwise-interaction point process can be obtained in the limit. Further pairwise-interaction processes are obtained as limits of sequences of auto-logistic lattice schemes.
D. M. Titterington合作论文数Department of Statistics2
Vladimir Batagelj合作论文数University of Ljubljana
FMF - Department of Mathematics and
IMFM - Institute of Mathematics, Physics and Mechanics
Department for theoretical computer science2