We propose a general robust prediction framework, termed conformity-based projective prediction (CPP), that integrates Bayesian predictive modeling with ideas from conformity-based conformal prediction. Rather than assessing conformity through residual-based scores, the CPP criterion defines conformity distributionally: a candidate value for a future response is considered conforming to the extent that its inclusion in the data leaves the leave-one-out predictive distributions of the observed responses undisturbed. The framework requires only that the leave-one-out and swapped predictive distributions are available in closed form and that the swapped predictive mean is differentiable in the candidate value. Under these conditions, we establish a general bounded-influence proposition and a general local convexity lemma, and prove that CPP dominates any plug-in predictor with unbounded influence in asymptotic variance under ε-contamination models. When the posterior mean is linear in the observations – as in Gaussian linear models, basis-expansion regression, and Gaussian process regression – the swapped predictive mean is affine in the candidate value, yielding closed-form or one-dimensional optimization solutions and an efficient rank-two computational update; all general theoretical results specialize to explicit corollaries in this setting. Simulation experiments and two data analyses under the Gaussian linear model illustrate the finite-sample advantages of the proposed method, confirming the theoretical predictions across contamination levels, sample sizes, and predictor dimensions.
The paper develops Bernstein von-Mises Theorem under hierarchical g -priors for linear regression models. The results are obtained both when the error variance is known, and also when it is unknown. An inverse gamma prior is attached to the error variance in the later case.
Small area estimation models are typically based on the normality assumption of response variables. More recently, attention has been drawn to the transformation of the original variables to justify the assumption of normality. Variance stabilizing transformation of observation serves the dual purpose of reaching closer to normality, as well as known variance of the transformed variables in contrast to the assumption of known variances of the original variables, the latter needed to avoid non-identifiability. However, the existing literature on the topic ignores a certain bias introduced in the seemingly correct back transformation. The present paper rectifies this deficiency by introducing asymptotically unbiased empirical Bayes (EB) estimators of small area means. Mean squared errors (MSEs) and estimated MSEs of such estimators are provided. The theoretical results were accompanied with simulations and data analysis. A somewhat surprising phenomenon is a finding which connects one of our results to the natural exponential family quadratic variance function (NEF-QVF) family of distributions introduced by Morris (1982,1983).
Global-local priors, often also referred to as shrinkage priors, have proved to be a very effective tool for the analysis of high dimensional data under sparsity. Asymptotic theoretical properties of such priors, studied under various scenarios, are now available in the literature. However, to our knowledge, theoretical guarantees of such priors provided so far, involve the assumption of known sample variance. The present paper relaxes this assumption, and carries out the analysis with a prior assigned to the error variance as well. In the process, some new tail bounds for shrinkage factors are developed, and these results are then utilized in providing asymptotic minimax rates for the posterior means of the parameters of interest.
We introduce a new small area predictor when the Fay-Herriot normal error model is fitted to a logarithmically transformed response variable, and the covariate is measured with error. This framework has been previously studied by Mosaferi et al. (2023). The empirical predictor given in their manuscript cannot perform uniformly better than the direct estimator. Our proposed predictor in this manuscript is unbiased and can perform uniformly better than the one proposed in Mosaferi et al. (2023). We derive an approximation of the mean squared error (MSE) for the predictor. The prediction intervals based on the MSE suffer from coverage problems. Thus, we propose a non-parametric bootstrap prediction interval which is more accurate. This problem is of great interest in small area applications since statistical agencies and agricultural surveys are often asked to produce estimates of right skewed variables with covariates measured with errors. With Monte Carlo simulation studies and two Census Bureau's data sets, we demonstrate the superiority of our proposed methodology.
There is a rich literature for modeling binary and polychotomous responses. However, existing methods are inadequate for handling combinatorial responses, where each response is an integer array under additional constraints. Such data are increasingly common in modern applications, such as surveys collected under skip logic, event propagation on a network, and observed matching in ecology. Ignoring the combinatorial structure leads to biased estimation and prediction. The fundamental challenge is the lack of a link function that connects a linear or functional predictor with a probability respecting the combinatorial constraints. In this article, we propose a novel augmented likelihood that views combinatorial response as a deterministic transform of a continuous latent variable. We specify the transform as the maximizer of integer linear program, and characterize useful properties such as dual thresholding representation. When taking a Bayesian approach and considering a multivariate normal distribution for the latent variable, our method becomes a direct generalization to the celebrated probit data augmentation, and enjoys straightforward computation via Markov chain Monte Carlo. We provide theoretical justification, including consistency and applicability, at an interesting intersection between duality and probability. We demonstrate the effectiveness of our method through simulations and a data application on the seasonal matching between waterfowl.
The Inverse-Wishart (IW) distribution is a standard and popular choice of priors for covariance matrices and has attractive properties such as conditional conjugacy. However, the IW family of priors has crucial drawbacks, including the lack of effective choices for non-informative priors. Several classes of priors for covariance matrices that alleviate these drawbacks, while preserving computational tractability, have been proposed in the literature. These priors can be obtained through appropriate scale mixtures of IW priors. However, in the era of increasing dimensionality, the posterior consistency of models that incorporate such priors has not been investigated. We address this issue for the multi-response regression setting (q responses, n samples) under a wide variety of IW scale mixture priors for the error covariance matrix. Posterior consistency and contraction rates for both the regression coefficient matrix and the error covariance matrix are established in the "large q, large n" setting under mild assumptions on the true data-generating covariance matrix and relevant hyperparameters. In particular, the number of responses qn is allowed to grow with n, but with qn = o(n). Also, some results related to the inconsistency of the posterior distribution and posterior mean for qn/n -> gamma, where gamma is an element of (0,infinity) are provided.
The paper considers a multiple testing problem of multivariate normal means under sparsity. First, the Bayes risk of the multivariate Bayes oracle is derived. Then, a hierarchical Bayesian approach is taken with global-local shrinkage priors, where the global parameter is either treated as a tuning parameter or is given a specific prior. The method is shown to attain an asymptotic Bayes optimal under sparsity (ABOS) property. Finally, an empirical Bayes procedure is proposed which involves estimation of the global shrinkage parameter. The approach is also shown to lead to the ABOS property.
We consider a small area estimation model under square-root transformation in the presence of functional measurement error. When measurement error is present, the Bayes predictor can no longer be used as it depends on the covariates even if parameters are known. Therefore suitable replacements are called for, and we propose a predictor that only depends on observed responses and data obtained from a large secondary survey. Moreover, some estimation methods of unknown parameters are considered. In the simulations section, we evaluate the performance using the mean squared prediction error (MSPE) and discuss several scenarios in terms of the number of areas and the sample sizes in a large secondary survey.
A random discrete probability distribution known as the GEM( α ) distribution, named after Griffiths, Engen and McCloskey, has found many applications for instance in population genetics, in mathematical ecology, in several probabilistic problems, in computer science etc. It reappeared under its current name “stick breaking distribution” in Sethuraman (Statist. Sinica, 4, 639–650, 1994) in the constructive definition of Dirichlet processes, which in turn has found numerous applications in Bayesian nonparametrics. The “invariant under size-biased property” (ISBP) of this distribution goes back to the Ph. D. dissertation of McCloskey (1965) and has been established by later authors (e.g. Donnelly J. Appl. Probab., 28, 321–335, 1991; Ongaro J. Statist. Plann. Inference, 128, 123–148, 2005; Pitman Adv. in Appl. Probab., 28, 525–539, 1996) by appealing to the Poisson Dirichlet distribution and random exchangeable permutations of the sets of integers. This paper gives a self contained proof of the ISBP property of stick breaking distribution using just its mixed moments. These mixed moments are also useful in nonparametric posterior estimation of the variance, skewness and kurtosis under Dirichlet process priors.
The paper develops Bernstein von Mises Theorem under hierarchical $g$ -priors for linear regression models. The results are obtained both when the error variance is known, and also when it is unknown. An inverse gamma prior is attached to the error variance in the later case. The paper also demonstrates some connection between the total variation and $\alpha$-divergence measures.
The paper obtains posterior Cramer-Rao bounds for arbitrary parameter vector in the spirit of van Trees. The relationship with the classical Cramer-Rao bound is also discussed.
Small area estimation is gaining increasing popularity among survey statisticians. Since the direct estimates of small areas usually have large standard errors, model-based approaches are often adopted to borrow strength across areas. The models often use covariates to link different areas and random effects to account for the additional variation. In the classic Fay-Herriot model, the random effects are assumed to have independent normal distributions with a shared variance. Recent studies showed that random effects are not necessary for all areas, so global-local priors have been introduced in Tang et al. [ 26 ] to effectively characterize the sparsity in random effects. This article introduces global-local priors in the context of small area estimation where the area level random effects exhibit a spatial structure. This generalizes the findings of Tang et al. [ 26 ] where independence of the area level effects is assumed. Our findings are illustrated via both simulation and real data examples. AMS Subject Classification: 62D05, 62M30
The paper addresses asymptotic estimation of normal means under sparsity. The primary focus is estimation of multivariate normal means where we obtain exact asymptotic minimax error under global-local shrinkage prior. This extends the corresponding univariate work of Ghosh and Chakrabarti (2017). In addition, we obtain similar results for the Dirichlet-Laplace prior as considered in Bhattacharya et al. (2015). Also, following van der Pas et al. (2017), we have been able to derive credible sets for multivariate normal means under global-local priors.
We study full Bayesian procedures for high-dimensional linear regression. We adopt data-dependent empirical priors introduced in Martin et al. (Bernoulli 23(3):1822–1847, 2017). In their paper, these priors have nice posterior contraction properties and are easy to compute. Our paper extend their theoretical results to the case of unknown error variance . Under proper sparsity assumption, we achieve model selection consistency, posterior contraction rates as well as Bernstein von-Mises theorem by analyzing multivariate t-distribution.
Statistical modeling and inference problems with sample sizes substantially smaller than the number of available covariates are challenging. Chakraborty et al. (2012) did a full hierarchical Bayesian analysis of nonlinear regression in such situations using relevance vector machines based on reproducing kernel Hilbert space (RKHS). But they did not provide any theoretical properties associated with their procedure. The present paper revisits their problem, introduces a new class of global-local priors different from theirs, and provides results on posterior consistency as well as posterior contraction rates
We consider a small area estimation model under square-root transformation in the presence of functional measurement error. When measurement error is present, the Bayes predictor can no longer be used as it depends on the covariates even if parameters are known. Therefore suitable replacements are called for, and we propose a predictor that only depends on observed responses and data obtained from a large secondary survey. Moreover, some estimating methods of unknown parameters are considered. In the simulations section, We evaluate the performance using the mean squared prediction error (MSPE) and discuss several scenarios in terms of the number of areas and the sample size in a large secondary survey.
Divergence between two distributions has been of statistical interest for more than a century, beginning with Karl Pearson with his famous chisquare test. The paper revisits some of the well-known density divergence measures, and studies their interrelationship. In addition, it is demonstrated how Scheffe’s pointwise density convergence implies convergence of distributions, based on different divergence measures. Journal of Statistical Research 2022, Vol. 56, No. 1, pp. 1-10
The arc-sin transformation has long been used as a variance stabilizer for the binomial sample proportion arising out of binary data. The natural back -transformed function is useful for returning an estimate to the original scale of the parameter of interest. However, it is known that such a transformation leads to bias when estimating the original parameter of interest. In this study, we find explicit asymptotic bias-adjusted empirical Bayes (EB) estimators for binomial sample pro-portions in the context of small area estimation. We obtain an explicit second-order correct approximation of the mean squared errors (MSEs) of such estimators, as well as second-order correct estimators of these MSEs. Moreover, the proposed EB esti-mators and corresponding MSE estimators outperform their competitors in terms of the bias and variance, as demonstrated in a simulation study. We apply our methodology to real data associated with Coronavirus Disease 2019 (COVID-19) for each prefecture in Japan.
This paper aims to examine the characteristics of the posterior distribution of covariance/precision matrices in a "large $p$, large $n$" scenario, where $p$ represents the number of variables and $n$ is the sample size. Our analysis focuses on establishing asymptotic normality of the posterior distribution of the entire covariance/precision matrices under specific growth restrictions on $p_n$ and other mild assumptions. In particular, the limiting distribution turns out to be a symmetric matrix variate normal distribution whose parameters depend on the maximum likelihood estimate. Our results hold for a wide class of prior distributions which includes standard choices used by practitioners. Next, we consider Gaussian graphical models which induce sparsity in the precision matrix. Asymptotic normality of the corresponding posterior distribution is established under mild assumptions on the prior and true data-generating mechanism.