This chapter covers a mix of topics including splines, functional data, extreme values and density estimation. Being associated with continuous time monitoring processes, functional data are usually smooth curves or surfaces and are often treated as realizations of underlying random functions. The basic philosophy of functional data is to consider the observed curves as single entities, rather than only as a sequence of individual observations. Comprehensive surveys of statistical techniques for analyzing functional data can be found in Ramsay and Silverman and Ferraty and Vieu. Goldsmith study Diffusion Tensor Imaging (DTI) metrics of multiple sclerosis (MS) patients over multiple clinical visits. The data consist of 100 subjects, aged between 21 and 70 years at first visit. The number of visits per subject ranges from 2 to 8, and a total of 340 visits were recorded. Most statistical modeling is concerned with the mean …
This paper describes a method for creating a confidence interval for the ratio of rates using the score statistic. This non-iterative and easy to apply procedure produces confidence intervals that are suitable for use with Poisson data and simulation results indicate that it is close to the nominal level for a wide range of scenarios.
Meta-analysis is now a standard statistical tool for assessing the overall strength and interesting features of a relationship, on the basis of multiple independent studies. There is, however, recent acknowledgement of the fact that in many applications responses are rarely uniquely determined. Hence there has been some change of focus from a single response to the analysis of multiple outcomes. In this paper we propose and evaluate three Bayesian multivariate meta-analysis models: two multivariate analogues of the traditional univariate random effects models which make different assumptions about the relationships between studies and estimates, and a multivariate random effects model which is a Bayesian adaptation of the mixed model approach. Our preferred method is then illustrated through an analysis of a new data set on parental smoking and two health outcomes (asthma and lower respiratory disease) in children.
This paper proposes a template for modelling complex datasets that integrates traditional statistical modelling approaches with more recent advances in statistics and modelling through an exploratory framework. Our approach builds on the well-known and long standing traditional idea of 'good practice in statistics' by establishing a comprehensive framework for modelling that focuses on exploration, prediction, interpretation and reliability assessment, a relatively new idea that allows individual assessment of predictions.The integrated framework we present comprises two stages. The first involves the use of exploratory methods to help visually understand the data and identify a parsimonious set of explanatory variables. The second encompasses a two step modelling process, where the use of non-parametric methods such as decision trees and generalized additive models are promoted to identify important variables and their modelling relationship with the response before a final predictive model is considered. We focus on fitting the predictive model using parametric, non-parametric and Bayesian approaches.This paper is motivated by a medical problem where interest focuses on developing a risk stratification system for morbidity of 1,710 cardiac patients given a suite of demographic, clinical and preoperative variables. Although the methods we use are applied specifically to this case study, these methods can be applied across any field, irrespective of the type of response.
To increase the reliability of a particular diagnosis or similar type of binary evaluation, a common approach is to use multiple assessors. Through the use of Bayes's theorem, we show that one can compute the reliability of a "median agreement method" when there are three assessors, even if the information on individual assessors is limited. We show, perhaps counter-intuitively, that in some circumstances a seemingly greater reliability among assessors actually implies a greater rate of misclassification. The approach is exemplified through an investigation of the reliability of radiological diagnosis of opacities of the lung associated with exposure to asbestos. This provides a good example of the need for care in defining the conditioning events involved in discussing "reliability of diagnosis," and the differences between specificity and sensitivity on the one hand and predictive values or reliability on the other.
The maximum-likelihood (ML) approach is a powerful tool for reconstructing molecular phylogenies.In conjunction with the Kishino-Hasegawa test, it allows direct comparison of alternative evolutionary hypotheses.A commonly occurring outcome is that several trees are not significantly different from the ML tree, and thus there is residual uncertainty about the correct tree topology.We present a new method for producing a majority-rule consensus tree that is based on those trees that are not significantly less likely than the ML tree.Five types of consensus trees are considered.These differ in the weighting schemes that are employed.Apart from incorporating the topologies of alternative trees, some of the weighting schemes also make use of the differences between the log likelihood estimate of the ML tree and those of the other trees and the standard errors of those differences.The new approach is used to analyze the phylogenetic relationship of psbA proteins from four free-living photosynthetic prokaryotes and a chloroplast from green plants.We conclude that the most promising weighting scheme involves exponential weighting of differences between the log likelihood estimate of the ML tree and those of the other trees standardized by the standard errors of the differences.A consensus tree that is based on this weighting scheme is referred to as a standardized, exponentially weighted consensus tree.The new approach is a valuable alternative to existing treeevaluating methods, because it integrates phylogenetic information from the ML tree with that of trees that do not differ significantly from the ML tree.
Markov chain Monte Carlo (MCMC) methods have been used extensively in statistical physics over the last 40 years, in spatial statistics for the past 20 and in Bayesian image analysis over the last decade. In the last five years, MCMC has been introduced into significance testing, general Bayesian inference and maximum likelihood estimation. This paper presents basic methodology of MCMC, emphasizing the Bayesian paradigm, conditional probability and the intimate relationship with Markov random fields in spatial statistics. Hastings algorithms are discussed, including Gibbs, Metropolis and some other variations. Pairwise difference priors are described and are used subsequently in three Bayesian applications, in each of which there is a pronounced spatial or temporal aspect to the modeling. The examples involve logistic regression in the presence of unobserved covariates and ordinal factors; the analysis of agricultural field experiments, with adjustment for fertility gradients; and processing of low-resolution medical images obtained by a gamma camera. Additional methodological issues arise in each of these applications and in the Appendices. The paper lays particular emphasis on the calculation of posterior probabilities and concurs with others in its view that MCMC facilitates a fundamental breakthrough in applied Bayesian modeling.