
Summary. When evaluating the performance of discordancy tests both the type II error and the power function contain spurious and non-spurious elements as the test may not identify the contaminant observation. The probabilities that are associated with identifying as discordant a good observation and a contaminant observation are defined as spurious power and non - spurious power respectively. This nomenclature conveys the role of each as performance criteria for tests of discordancy. A general discussion of discordancy test performance criteria is presented, as is the case of tests for multiple outliers. Finally, the behaviours of the Grubbs single-outlier and multiple-outlier tests are studied using simulation.
Summary. Discrete data assumed to be generated by independent Poisson distributions are subject to censoring processes which determine the incomplete classification that is reported in several practical situations. The paper describes a probabilistic model that can explain incomplete data referring to partitions of the original set of categories. We overcome the model's lack of identifiability by assuming a missingness at random censoring mechanism. On the basis of the subsequent statistical model the results required to fit further non-informative censoring models by maximum likelihood methodology are obtained. This preliminary analysis opens the way to the analysis of structural models for Poisson expected rates. The paper describes how to test the fit of structural models for the Poisson rates, whether taken individually or simultaneously with special censoring models. Then, the maximum likelihood methodology is specialized to the analysis of strictly linear and log-linear models in such a way that its computational implementation opens up for any count table. The methods developed throughout the work are illustrated with a breast cancer data set.
Summary. We consider a class of ratio–product estimators for estimating a finite population mean. The asymptotically optimum estimator in the class is identified, along with its approximate mean-square error. This estimator requires prior knowledge of the parameter C = ρ C y / C x , where ρ is the correlation coefficient between the study variate y and the auxiliary variate x , and C y and C x are coefficients of variation of y and x respectively. If C is unknown in advance, then it can be replaced by its consistent estimate , with the resulting estimator known as an ‘estimator based on the estimating optimum’. It is shown that, to the first order of approximation, both estimators have the same mean-square error, and that they are generally more efficient than the usual ratio and product estimators.
Summary. Hoeffding's test of bivariate independence and its asymptotic equivalent due to Blum, Kiefer and Rosenblatt are well known to be consistent against all dependence alternatives. However, the two tests, which are often treated as interchangeable, are rarely used in data analysis mainly because their finite sample null distributions are unavailable, and little is known about their operating characteristics. In this paper the conventional wisdom regarding the equivalence of these tests and their distributions is examined by first tabulating their null distributions for sample sizes n =5,6,…,25,30,…,50,60,…,100, and then studying their power functions empirically. The power functions are compared with those of the commonly used methods based on the product moment correlation, the rank correlation and Kendall's τ , for bivariate normal and log-normal populations, as well as a variety of dependence models such as the well-known copulas due to Morgenstern, Gumbel, Plackett, Marshall and Olkin, Raftery, Clayton and Frank. It is seen that the Blum, Kiefer and Rosenblatt test is generally preferable in terms of power against positive dependence alternatives and that the conventional wisdom deserves a revision.
Summary. SAS software is often used for statistical simulations. The paper demonstrates a simple and effective approach for performing simulations in SAS. The method that is usually used generates data sets and performs the calculations sequentially in time, using the macro %DO loop to execute each cycle. The more efficient approach is to generate one large data set that includes all the individual data sets from each cycle as subsets and to perform the calculations on each subset in one pass, using the BY command. The paper presents an example of both methods of simulation, with their SAS codes. Both programs give the same numerical results, but the BY approach is 80 times faster than the macro %DO approach.
Summary. The paper focuses on the problem of data heaping that arises when measurements are recorded to varying degrees of precision. The work is motivated by an application in foetal medicine where measurements obtained from ultrasound images are rounded to varying numbers of decimal places causing heaping at integer values. We demonstrate the dangers of ignoring heaping before presenting a case-study of the ultrasound measurements. A mixture model, in which the different components represent different levels of rounding, is used for the heaping process. We illustrate a range of graphical posterior predictive checks to assess the fit of the model and we explore some extensions of the model. We adopt a Bayesian approach implemented by using the Gibbs sampler.
Summary. A normally distributed vector response variable is considered, related to a scalar explanatory variable through a linear regression model. The calibration data, i.e. data obtained on the response variable corresponding to known values of the explanatory variable, are to be used for making inferences concerning unknown values of the explanatory variable. The purpose of the paper is a detailed investigation of interval estimation and hypothesis testing problems that arise in this context. Such problems are addressed in the scenario of both single use and multiple use of the calibration data. In the single-use situation, the calibration data are used to make inferences concerning a single unknown value of the explanatory variable. In the multiple-use scenario, the calibration data are used to make inferences concerning a sequence of unknown values of the explanatory variable. A one-sided hypothesis testing problem is addressed and test procedures are developed in the context of both single use and multiple use of the calibration data. Since the test statistic has a distribution that depends on some unknown parameters, a parametric bootstrap procedure is advocated to carry out the test. The parametric bootstrap procedure can be adopted for interval estimation as well. The performance of the parametric bootstrap procedure is numerically investigated and is found to be quite satisfactory. Some examples are used to motivate the problems and to illustrate the applicability of our results. The overall conclusion is that the parametric bootstrap is a simple and satisfactory approach for making inferences in the calibration problem.