Lin (2014) developed a framework of the method of the sample-moment-based density approximant, for estimating the probability density function of microdata based on noise multiplied data. Theoretically, it provides a promising method for data users in generating the synthetic data of the original data without accessing the original data; however, technical issues can cause problems implementing the method. In this paper, we describe a software package called MaskDensity14, written in the R language, that uses a computational approach to solve the technical issues and makes the method of the sample-moment-based density approximant feasible. MaskDensity14 has applications in many areas, such as sharing clinical trial data and survey data without releasing the original data.
There are many areas of science and engineering where research and decision making are performed using computer models. These computer models are usually deterministic and may take minutes, hours or days to produce an output for a single value of the model inputs. Fitting mixtures of experts of computer models where the expert components use different values of the computer model parameters is considered. The efficient calibration of such models using emulators, which are fast statistical surrogates for the computer model, is discussed. It is argued that mixtures of experts are often insightful for describing model discrepancy and ways in which the computer model can be improved. This is not a strength of standard approaches to the statistical analysis of computer models where a certain ''best input'' assumption is usually made and model discrepancy is often described through a stationary Gaussian process prior on the discrepancy function. Application of the framework is presented for a dynamic hydrological rainfall-runoff model in which the mixture approach is helpful for highlighting model deficiencies.
A framework, the sample-moment-based density approximant, for estimating the probability density function based on noise multiplied data was proposed in Lin (2014). Based on the framework, an R package, MaskDensity10.R, is built in this paper. The package is available from http://www.uow.edu.au/ ∼yanxia/Confidential data analysis/. The framework is developed for continuous univeriates (see Lin, 2014). With the techniques of nonparametric smoothing and K-means clustering integrated, MaskDensity10.R can be used for estimating the mass functions of categorical variables. The same R package, MaskDensity10.R, can be used by the data agency to create the masked data set, as well as used by the end-user to obtain the approximation of the density function of the original data set based on masked data. Simulation studies and real life data applications of MaskDensity10.R are presented in this paper. The risk of disclosure in the application of the R package to microdata is discussed, particularly for category data.
Background: Emotional impairments are important determinants of functional outcome in psychosis, and current treatments are not particularly effective. Modafinil is a wake-promoting drug that has been shown to improve emotion discrimination in healthy individuals and attention and executive function in schizophrenia. We aimed to establish whether modafinil might have a role in the adjuvant treatment of emotional impairments in the first episode of psychosis, when therapeutic endeavor is arguably most vital.Methods: Forty patients with a first episode of psychosis participated in a randomized, double-blind, placebo-controlled crossover design study testing the effects of a single dose of 200 mg modafinil on neuropsychological performance. Emotional functions were evaluated with the emotional face recognition test, the affective go-no go task, and the reward and punishment learning test. Visual analogue scales were used throughout the study to assess subjective mood changes.Results: Modafinil significantly improved the recognition of sad facial expressions (z = 2.98, p = .003). In contrast, there was no effect of modafinil on subjective mood ratings, on tasks measuring emotional sensitivity to reward or punishment, or on interference of emotional valence on cognitive function, as measured by the affective go-no go task.Conclusions: Modafinil improves the analysis of emotional face expressions. This might enhance social function in people with a first episode of psychosis.
There is a well-recognized need to develop Bayesian computational methodologies that scale well to large data sets. Recent attempts to develop such methodology have often focused on two approaches—variational approximation and advanced importance sampling methods. This note shows how importance sampling can be viewed as a variational approximation, achieving a pleasing conceptual unification of the two points of view. We consider a particle representation of a distribution as defining a certain parametric model and show how the optimal approximation (in the sense of minimization of a Kullback–Leibler divergence) leads to importance sampling type rules. This new way of looking at importance sampling has the potential to generate new algorithms by the consideration of deterministic choices of particles in particle representations of distributions.
We consider Markov chain Monte Carlo (MCMC) computational schemes intended to minimize the number of evaluations of the posterior distribution in Bayesian inference when the posterior is computationally expensive to evaluate. Our motivation is Bayesian calibration of computationally expensive computer models. An algorithm suggested previously in the literature based on hybrid Monte Carlo and a Gaussian process approximation to the target distribution is extended in three ways. First, we consider combining the original method with tempering schemes in order to deal with multimodal posterior distributions. Second, we consider replacing the original target posterior distribution with the Gaussian process approximation, which requires less computation to evaluate. Third, we consider in the context of tempering schemes the replacement of the true target distribution with the approximation in the high temperature chains while retaining the true target in the lowest temperature chain. This retains the correct target distribution in the lowest temperature chain while avoiding the computational expense of running the computer model in moves involving the high temperatures. Application of our methodology is considered to calibration of a rainfall-runoff model where multimodality of the parameter posterior is observed.
Laplace approximation is one commonly used approach to the calculation of difficult integrals arising in Bayesian inference and the analysis of random effects models. Here we outline a procedure which is an extension of the Laplace approximation and which attempts to find changes of variable for which the integrand becomes approximately a product of one-dimensional functions. When the integrand is a product of one-dimensional functions, an approximation to the integral can be obtained using one-dimensional quadrature. The approximation is exact for a broader class of functions than the ordinary Laplace approximation and can be applied when the integrand is not smooth at the mode. As an illustration of this last point we consider calculation of marginal likelihoods for smoothing parameter selection in the lasso.
Model selection is an important activity in modern data analysis and the conventional Bayesian approach to this problem involves calculation of marginal likelihoods for different models, together with diagnostics which examine specific aspects of model fit. Calculating the marginal likelihood is a difficult computational problem. Our article proposes some extensions of the Laplace approximation for this task that are related to copula models and which are easy to apply. Variations which can be used both with and without simulation from the posterior distribution are considered, as well as use of the approximations with bridge sampling and in random effects models with a large number of latent variables. The use of a t-copula to obtain higher accuracy when multivariate dependence is not well captured by a Gaussian copula is also discussed.
A simple climate model is used to calculate the benefit, over time, of geosequestration Of C02 that would otherwise be released to the atmosphere. The analysis is performed relative to two reference cases. The first case is defined by a C02 concentration profile leading to stabilisation at 500 ppm. The second case is defined by 'business -as -usual' (IS92a) C02 emissions until 2100. The benefits are considered in terms of incremental change (per unit of displaced emission) in temperature and its rate of change, concentrating on the period to 2200. An automatic differentiation procedure has proved a convenient way of performing the calculations. The 'temperature benefit' of avoided carbon emission is found to be of order I mK/GtC on the time-scale of decades to centuries. This result is model-specific and would scale in proportion to the climate sensitivity of the model. Because of non-linearities in carbon-climate processes, the results have a small dependence (of order 10-20%) on the future emission scenario with a rather smaller contribution to uncertainty arising from model calibration uncertainties that reflect uncertainties in the 20th century carbon budget. Analysis over the longer term, to 2500, considers the effect of leakage of geologically stored C02 to the atmosphere, and shows that even at 0.1% per annum leakage, about half the climate benefit remains after 500 years. (c) 2008 Elsevier Ltd. All rights reserved.
Contrary to conventional belief, it turns out that in some problem instances of moderate size, fixed temperature simulated annealing algorithms based on a heuristic formula for determining the optimal temperature can be superior to algorithms based on cooling. Such a heuristic formula, however, often seems elusive. In practical cases considered we include instances of traveling salesman, quadratic assignment, and graph partitioning problems, where we obtain results that compare favorably to the ones known in the literature.
A sizable part of the theoretical literature on simulated annealing deals with a property called convergence, which asserts that the simulated annealing chain is in the set of global minimum states of the objective function with probability tending to 1. However, in practice, the convergent algorithms are considered too slow, whereas a number of nonconvergent ones are usually preferred. We attempt a detailed analysis of various temperature schedules. Examples will be given of when it is both practically and theoretically justified to use boiling, fixed temperature, or even fast cooling schedules which have a small probability of reaching global minima. Applications to traveling salesman problems of various sizes are also given.
In 2010 the Faculty of Science at the University of Wollongong (UOW) decided to inform its admissions policy by analysing the results in the introductory chemistry subject CHEM101. The data was analysed by members of the School of Mathematics and Applied Statistics using a generalised linear model. This paper discusses the model and its implications. One of the conclusions is that the level of mathematics studied for the Higher School Certificate (HSC) is a better predictor of performance in CHEM101 than either the HSC Chemistry mark or the student's Australian Tertiary Admission Rank. As a result of the analysis in this article, the Faculty of Science at UOW changed its early admission procedures. A pdf file of a presentation by the first author on this material and an interactive passing probability calculator based on the model are available at [11] and [10] respectively.
Yan-Xia Lin合作论文数School of Mathematics and Applied Statistics
University of Wollongong1