In statistical inference, it is rarely realistic that the hypothesized statistical model is well-specified, and consequently it is important to understand the effects of misspecification on inferential procedures. When the hypothesized statistical model is misspecified, the natural target of inference is a projection of the data generating distribution onto the model. We present a general method for constructing valid confidence sets for such projections, under weak regularity conditions, despite possible model misspecification. Our method builds upon the universal inference method and is based on inverting a family of split-sample tests of relative fit. We study settings in which our methods yield either exact or approximate, finite-sample valid confidence sets for various projection distributions. We study rates at which the resulting confidence sets shrink around their target of inference and complement these results with a simulation study and a study of causal discovery using a linear causal model with the CausalEffectPairs dataset.
In statistical inference, it is rarely realistic that the hypothesized statistical model is well-specified, and consequently it is important to understand the effects of misspecification on inferential procedures. When the hypothesized statistical model is misspecified, the natural target of inference is a projection of the data generating distribution onto the model. We present a general method for constructing valid confidence sets for such projections, under weak regularity conditions, despite possible model misspecification. Our method builds upon the universal inference method of Wasserman et al. (2020) and is based on inverting a family of split-sample tests of relative fit. We study settings in which our methods yield either exact or approximate, finite-sample valid confidence sets for various projection distributions. We study rates at which the resulting confidence sets shrink around the target of inference and complement these results with a simulation study.
The world ocean plays a key role in redistributing heat in the climate system and hence in regulating Earth’s climate. Yet statistical analysis of ocean heat transport suffers from partially incomplete large-scale data intertwined with complex spatio-temporal dynamics, as well as from potential model misspecification. We present a comprehensive spatio-temporal statistical framework tailored to interpolating the global ocean heat transport using in-situ Argo profiling float measurements. We formalize the statistical challenges using latent local Gaussian process regression accompanied by a two-stage fitting procedure. We introduce an approximate Expectation-Maximization algorithm to jointly estimate both the mean field and the covariance parameters, and refine the potentially under-specified mean field model with a debiasing procedure. This approach provides data-driven global ocean heat transport fields that vary in both space and time and can provide insights into crucial dynamical phenomena, such as El Niño & La Niña, as well as the global climatological mean heat transport field, which by itself is of scientific interest. The proposed framework and the Argo-based estimates are thoroughly validated with state-of-the-art multimission satellite products and shown to yield realistic subsurface ocean heat transport estimates.
We propose Bayesian semiparametric mixed effects models with measurement error to analyze the literature data collected from multiple studies in a meta‐analytic framework. We explore this methodology for risk assessment in cadmium toxicity studies, where the primary objective is to investigate dose‐response relationships between urinary cadmium concentrations and ‐microglobulin. In the proposed model, a nonlinear association between exposure and response is described by a Gaussian process with shape restrictions, and study‐specific random effects are modeled to have either normal or unknown distributions with Dirichlet process mixture priors. In addition, nonparametric Bayesian measurement error models are incorporated to flexibly account for the uncertainty resulting from the usage of a surrogate measurement of a true exposure. We apply the proposed model to analyze cadmium toxicity data imposing shape constraints along with measurement errors and study‐specific random effects across varying characteristics, such as population gender, age, or ethnicity.
Model misspecification can compromise valid inference in conventional quantile regression models. To address this issue, we consider two flexible model extensions for high-dimensional data. The first is a Bayesian quantile regression approach with variable selection, which uses a sparse signal shrinkage prior on the high-dimensional regression coefficients. The second extension robustifies conventional parametric quantile regression methods by including observation specific mean shift terms. Since the number of outliers is assumed to be small, the vector of mean shifts is sparse, which again motivates the use of a sparse signal shrinkage prior. Specifically, we exploit the horseshoe+ prior distribution for variable selection and outlier detection in the high-dimensional quantile regression models. Computational complexity is alleviated using fast mean field variational Bayes methods, and we compare results obtained by variational methods with those obtained using Markov chain Monte Carlo (MCMC).
The Bayesian spectral analysis model (BSAM) is a powerful tool to deal with semiparametric methods in regression and density estimation based on the spectral representation of Gaussian process priors.The bsamGP package for R provides a comprehensive set of programs for the implementation of fully Bayesian semiparametric methods based on BSAM.Currently, bsamGP includes semiparametric additive models for regression, generalized models and density estimation.In particular, bsamGP deals with constrained regression models with monotone, convex/concave, S-shaped and U-shaped functions by modeling derivatives of regression functions as squared Gaussian processes.bsamGP also contains Bayesian model selection procedures for testing the adequacy of a parametric model relative to a non-specific semiparametric alternative and the existence of the shape restriction.To maximize computational efficiency, we carry out posterior sampling algorithms of all models using compiled Fortran code.The package is illustrated through Bayesian semiparametric analyses of synthetic data and benchmark data.
This paper presents a variational Bayes approach to a semiparametric regression model that consists of parametric and nonparametric components. The assumed univariate nonparametric component is represented with a cosine series based on a spectral analysis of Gaussian process priors. Here, we develop fast variational methods for fitting the semiparametric regression model that reduce the computation time by an order of magnitude over Markov chain Monte Carlo methods. Further, we explore the possible use of the variational lower bound and variational information criteria for model choice of a parametric regression model against a semiparametric alternative. In addition, variational methods are developed for estimating univariate shape-restricted regression functions that are monotonic, monotonic convex or monotonic concave. Since these variational methods are approximate, we explore some of the trade-offs involved in using them in terms of speed, accuracy and automation of the implementation in comparison with Markov chain Monte Carlo methods and discuss their potential and limitations.