In the field of approximate Bayesian inference, expectation propagation (EP) is an often overlooked counterpart to its older Laplace approximation and variational Bayes cousins, perhaps owing to a lack of theory (especially convergence guarantees) and a higher implementation overhead, where derivations need to be hand-crafted to the specific distributions being approximated. However, when EP is carefully implemented, it is often more accurate than these alternative approaches. With this in mind, the purpose of this review paper is to describe and to consolidate the current state of research on EP at a high level focusing on examples and applications. Our aim is for this broad-based, practical EP guide to encourage its novel application to both existing and new Bayesian problems, where it has the opportunity to dramatically extend the cutting edge in performance.
In this paper, we introduce a new probability distribution, the Lasso distribution. We derive several fundamental properties of the distribution, including closed-form expressions for its moments and moment-generating function. Additionally, we present an efficient and numerically stable algorithm for generating random samples from the distribution, facilitating its use in both theoretical and applied settings. We establish that the Lasso distribution belongs to the exponential family. A direct application of the Lasso distribution arises in the context of an existing Gibbs sampler, where the full conditional distribution of each regression coefficient follows this distribution. This leads to a more computationally efficient and theoretically grounded sampling scheme. To facilitate the adoption of our methodology, we provide an R package, BayesianLasso, available on CRAN, implementing the proposed methods. Our findings offer new insights into the probabilistic structure underlying the Lasso penalty and provide practical improvements in Bayesian inference for high-dimensional regression problems.
We consider large scale ordinal response models with crossed random effects. We use a Gaussian variational approximation simplifying the associated optimization problem using the delta method, and other approximations to calculate both point estimates and standard errors while maintaining computational scalability. We apply our methodology on large scale recommender systems and are able to fit the proposed model with millions of observations within minutes on a standard laptop computer. Experiments show that estimation using this methodology scales close to linearly, and a factor better than the Penalized Quasi-Likelihood (PQL), with the number of observations, while maintaining an accuracy close to the Laplace Approximation (LA) with only slightly more underestimated variance parameters compared to LA in these large-scale settings. Our method thereby enables estimation of accurate, large scale ordinal models with crossed random effects.
Mixed-effects regression models represent a useful subclass of regression models for grouped data; the introduction of random effects allows for the correlation between observations within each group to be conveniently captured when inferring the fixed effects. At a time where such regression models are being fit to increasingly large datasets with many groups, it is ideal if (a) the time it takes to make the inferences scales linearly with the number of groups and (b) the inference workload can be distributed across multiple computational nodes in a numerically stable way, if the dataset cannot be stored in one location. Current Bayesian inference approaches for mixed-effects regression models do not seem to account for both challenges simultaneously. To address this, we develop an expectation propagation (EP) framework in this setting that is both scalable and numerically stable when distributed for the case where there is only one grouping factor. The main technical innovations lie in the sparse reparameterisation of the EP algorithm, and a moment propagation (MP) based refinement for multivariate random effect factor approximations. Experiments are conducted to show that this EP framework achieves linear scaling, while having comparable accuracy to other scalable approximate Bayesian inference (ABI) approaches.
We introduce a new class of priors for Bayesian hypothesis testing, which we name "cake priors". These priors circumvent Bartlett's paradox (also called the Jeffreys-Lindley paradox); the problem associated with the use of diffuse priors leading to nonsensical statistical inferences. Cake priors allow the use of diffuse priors (having one's cake) while achieving theoretically justified inferences (eating it too). We demonstrate this methodology for Bayesian hypotheses tests for scenarios under which the one and two sample t-tests, and linear models are typically derived. The resulting Bayesian test statistic takes the form of a penalized likelihood ratio test statistic. By considering the sampling distribution under the null and alternative hypotheses we show for independent identically distributed regular parametric models that Bayesian hypothesis tests using cake priors are Chernoff-consistent, i.e., achieve zero type I and II errors asymptotically. Lindley's paradox is also discussed. We argue that a true Lindley's paradox will only occur with small probability for large sample sizes.
Many approximate Bayesian inference methods assume a particular parametric form for approximating the posterior distribution. A Gaussian distribution provides a convenient density for such approaches; examples include the Laplace, penalized quasi-likelihood, Gaussian variational, and expectation propagation methods. Unfortunately, these all ignore potential posterior skewness. The recent work of Durante et al. [Skewed Bernstein-von Mises theorem and skew-modal approximations; 2023. ArXiv preprint arXiv:2301.03038.] addresses this using skew-modal (SM) approximations, and is theoretically justified by a skewed Bernstein-von Mises theorem. However, the SM approximation can be impractical to work with in terms of tractability and storage costs, and uses only local posterior information. We introduce a variety of matching-based approximation schemes using the standard skew-normal distribution to resolve these issues. Experiments were conducted to compare the performance of this skew-normal matching method (both as a standalone approximation and as a post-hoc skewness adjustment) with the SM and existing Gaussian approximations. We show that for small and moderate dimensions, skew-normal matching can be much more accurate than these other approaches. For post-hoc skewness adjustments, this comes at very little cost in additional computational time.
In this work, we propose a novel approximated collapsed variational Bayes approach to model selection in linear regression. The approximated collapsed variational Bayes algorithm offers improvements over mean field variational Bayes by marginalizing over a subset of parameters and using mean field variational Bayes over the remaining parameters in an analogous fashion to collapsed Gibbs sampling. We have shown that the proposed algorithm, under typical regularity assumptions, (a) includes variables in the true underlying model at an exponential rate in the sample size, or (b) excludes the variables at least at the first order rate in the sample size if the variables are not in the true model. Simulation studies show that the performance of the proposed method is close to that of a particular Markov chain Monte Carlo sampler and a path search based variational Bayes algorithm, but requires an order of magnitude less time. The proposed method is also highly competitive with penalized methods, expectation propagation, stepwise AIC/BIC, BMS, and EMVS under various settings. Supplementary materials for the article are available online.
Many approximate Bayesian inference methods assume a particular parametric form for approximating the posterior distribution. A multivariate Gaussian distribution provides a convenient density for such approaches; examples include the Laplace, penalized quasi-likelihood, Gaussian variational, and expectation propagation methods. Unfortunately, these all ignore the potential skewness of the posterior distribution. We propose a modification that accounts for skewness, where key statistics of the posterior distribution are matched instead to a multivariate skew-normal distribution. A combination of simulation studies and benchmarking were conducted to compare the performance of this skew-normal matching method (both as a standalone approximation and as a post-hoc skewness adjustment) with existing Gaussian and skewed approximations. We show empirically that for small and moderate dimensional cases, skew-normal matching can be much more accurate than these other approaches. For post-hoc skewness adjustments, this comes at very little cost in additional computational time.
We introduce and develop moment propagation for approximate Bayesian inference. This method can be viewed as a variance correction for mean field variational Bayes which tends to underestimate posterior variances. Focusing on the case where the model is described by two sets of parameter vectors, we develop moment propagation algorithms for linear regression, multivariate normal, and probit regression models. We show for the probit regression model that moment propagation empirically performs reasonably well for several benchmark datasets. Finally, we discuss theoretical gaps and future extensions. In the supplementary material we show heuristically why moment propagation leads to appropriate posterior variance estimation, for the linear regression and multivariate normal models we show precisely why mean field variational Bayes underestimates certain moments, and prove that our moment propagation algorithm recovers the exact marginal posterior distributions for all parameters, and for probit regression we show that moment propagation provides asymptotically correct posterior means and covariance estimates.
The last decade has witnessed intense interest in how people perceive the minds of other entities (humans, non-human animals, and non-living objects and forces) and how this perception impacts behavior. Despite the attention paid to the topic, the psychological structure of mind perception-that is, the underlying properties that account for variance across judgements of entities-is not clear and extant reports conflict in terms of how to understand the structure. In the present research, we evaluated the psychological structure of mind perception by having participants evaluate a wide array of human, non-human animal, and non-animal entities. Using an entirely within-participants design, varied measurement approaches, and data-driven analyses, four studies demonstrated that mind perception is best conceptualized along a single dimension.
AIM:To analyse the effects of maternal diabetes mellitus (DM) and body mass Index (BMI) on central and peripheral fat accretion of large for gestational age (LGA) offspring.METHODS:This retrospective study included LGA fetuses (n = 595) with ultrasound scans at early (19.23 ± 0.68 weeks), mid (28.98 ± 1.62 weeks) and late (36.20 ± 1.59 weeks) stages of adipogenesis and measured abdominal (AFT) and mid-thigh (TFT) fat as surrogates for central and peripheral adiposity. Women were categorised according to BMI and DM status [pre-gestational (P-DM; n = 59), insulin managed (I-GDM; n = 132) and diet managed gestational diabetes (D-GDM; n = 29)]. Analysis of variance and linear regressions were applied.RESULTS:AFT and TFT did not differ significantly between BMI categories (normal, overweight and obese). In contrast, AFT was significantly higher in pregnancies affected by D-GDM compared to non-DM pregnancies from mid stage (0.44 mm difference, p = 0.002) and for all DM categories in late stage of adipogenesis (≥ 0.49 mm difference, p < 0.008). Late stage TFT accretion was higher than controls for P-DM and I-GDM but not for D-GDM (0.67 mm difference, p < 0.001; 0.49 mm difference, p = 0.001, 0.56 mm difference, p = 0.22 respectively). In comparison to the early non-DM group with an AFT to TFT ratio of 1.07, the I-GDM group ratio was 1.25 (p < 0.001), which normalised by 28 weeks becoming similar to control ratios.CONCLUSIONS:DM, independent of BMI, was associated with higher abdominal and mid-thigh fat accretion in fetuses. Use of insulin improved central to peripheral fat ratios in fetuses of GDM mothers.
Motivation High parameter histological techniques have allowed for the identification of a variety of distinct cell types within an image, providing a comprehensive overview of the tissue environment. This allows the complex cellular architecture and environment of diseased tissue to be explored. While spatial analysis techniques have revealed how cell-cell interactions are important within the disease pathology, there remains a gap in exploring changes in these interactions within the disease process. Specifically, there are currently no established methods for performing inference on cell localisation changes across images, hindering an understanding of how cellular environments change with a disease pathology. Results We have developed the spicyR R package to perform inference on changes in the spatial localisation of cell types across groups of images. Application to simulated data demonstrates a high sensitivity and specificity. We demonstrate the utility of spicyR by applying it to a type 1 diabetes imaging mass cytometry dataset, revealing changes in cellular associations that were relevant to the disease progression. Ultimately, spicyR allows changes in cellular environments to be explored under different pathologies or disease states. Availability and Implementation R package freely available at http://bioconductor.org/packages/release/bioc/html/spicyR.html and shiny app implementation at http://shiny.maths.usyd.edu.au/spicyR/ Contact ellis.patrick@sydney.edu.au Supplementary information Code for reproducing key figures available at https://github.com/nickcee/spicyRPaper .
In this chapter we consider fitting several different nonstandard flexible regression models using variational Bayes (VB). More specifically we focus on fitting generalised additive models with penalised splines for complicating factors including outliers, heteroscedastic noise, overdispersed count data and missing covariates. For these models the complications arising mean that standard software is not readily available to fit them. We start from the point of standard VB methodology and show some simple tricks can be brought to bear when such methods are not easily applied. Our VB implementation in R is often orders of magnitude faster than state-of-the-art Markov Chain Monte Carlo (MCMC) implementation via stan, often with minor loss of accuracy in the estimation of posterior quantities.
Concerted examination of multiple collections of single-cell RNA sequencing (RNA-seq) data promises further biological insights that cannot be uncovered with individual datasets. Here we present scMerge, an algorithm that integrates multiple single-cell RNA-seq datasets using factor analysis of stably expressed genes and pseudoreplicates across datasets. Using a large collection of public datasets, we benchmark scMerge against published methods and demonstrate that it consistently provides improved cell type separation by removing unwanted factors; scMerge can also enhance biological discovery through robust data integration, which we show through the inference of development trajectory in a liver dataset collection.
Background: Differences in cell-type composition across subjects and conditions often carry biological significance. Recent advancements in single cell sequencing technologies enable cell-types to be identified at the single cell level, and as a result, cell-type composition of tissues can now be studied in exquisite detail. However, a number of challenges remain with cell-type composition analysis - none of the existing methods can identify cell-type perfectly and variability related to cell sampling exists in any single cell experiment. This necessitates the development of method for estimating uncertainty in cell-type composition. Results: We developed a novel single cell differential composition (scDC) analysis method that performs differential cell-type composition analysis via bootstrap resampling. scDC captures the uncertainty associated with cell-type proportions of each subject via bias-corrected and accelerated bootstrap confidence intervals. We assessed the performance of our method using a number of simulated datasets and synthetic datasets curated from publicly available single cell datasets. In simulated datasets, scDC correctly recovered the true cell-type proportions. In synthetic datasets, the cell-type compositions returned by scDC were highly concordant with reference cell-type compositions from the original data. Since the majority of datasets tested in this study have only 2 to 5 subjects per condition, the addition of confidence intervals enabled better comparisons of compositional differences between subjects and across conditions. Conclusions: scDC is a novel statistical method for performing differential cell-type composition analysis for scRNA-seq data. It uses bootstrap resampling to estimate the standard errors associated with cell-type proportion estimates and performs significance testing through GLM and GLMM models. We have made this method available to the scientific community as part of the scdney package (Single Cell Data Integrative Analysis) R package, available from https://github.com/SydneyBioX/scdney.
Expectation propagation is a general prescription for approximation of integrals in statistical inference problems. Its literature is mainly concerned with Bayesian inference scenarios. However, expectation propagation can also be used to approximate integrals arising in frequentist statistical inference. We focus on likelihood-based inference for binary response mixed models and show that fast and accurate quadrature-free inference can be realized for the probit link case with multivariate random effects and higher levels of nesting. The approach is supported by asymptotic calculations in which expectation propagation is seen to provide consistent estimation of the exact likelihood surface. Numerical studies reveal the availability of fast, highly accurate and scalable methodology for binary mixed model analysis. Supplementary materials for this article are available online.
We provide several examples of Bayesian semiparametric regression analysis via the Infer.NET package for approximate deterministic inference in Bayesian models. The examples are chosen to encompass a wide range of semiparametric regression situations. Infer.NET is shown to produce accurate inference in comparison with Markov chain Monte Carlo via the BUGS package, but to be considerably faster. Potentially, this contribution represents the start of a new era for semiparametric regression, where large and complex analyses are performed via fast Bayesian inference methodology and software, mainly being developed within Machine Learning.
Class labels are required for supervised learning but may be corrupted or missing in various applications. In binary classification, for example, when only a subset of positive instances is labeled whereas the remaining are unlabeled, positive-unlabeled (PU) learning is required to model from both positive and unlabeled data. Similarly, when class labels are corrupted by mislabeled instances, methods are needed for learning in the presence of class label noise (LN). Here we propose adaptive sampling (AdaSampling), a framework for both PU learning and learning with class LN. By iteratively estimating the class mislabeling probability with an adaptive sampling procedure, the proposed method progressively reduces the risk of selecting mislabeled instances for model training and subsequently constructs highly generalizable models even when a large proportion of mislabeled instances is present in the data. We demonstrate the utilities of proposed methods using simulation and benchmark data, and compare them to alternative approaches that are commonly used for PU learning and/or learning with LN. We then introduce two novel bioinformatics applications where AdaSampling is used to: 1) identify kinase-substrates from mass spectrometry-based phosphoproteomics data and 2) predict transcription factor target genes by integrating various next-generation sequencing data.
MOTIVATION:Genes act as a system and not in isolation. Thus, it is important to consider coordinated changes of gene expression rather than single genes when investigating biological phenomena such as the aetiology of cancer. We have developed an approach for quantifying how changes in the association between pairs of genes may inform the outcome of interest called Differential Correlation across Ranked Samples (DCARS). Modelling gene correlation across a continuous sample ranking does not require the dichotomisation of samples into two distinct classes and can identify differences in gene correlation across early, mid or late stages of the outcome of interest.RESULTS:When we evaluated DCARS against the typical Fisher Z-transformation test for differential correlation, as well as a typical approach testing for interaction within a linear model, on real TCGA data, DCARS significantly ranked gene pairs containing known cancer genes more highly across several cancers. Similar results are found with our simulation study. DCARS was applied to 13 cancers datasets in TCGA, revealing several distinct relationships for which survival ranking was found to be associated with a change in correlation between genes. Furthermore, we demonstrated that DCARS can be used in conjunction with network analysis techniques to extract biological meaning from multi-layered and complex data.AVAILABILITY AND IMPLEMENTATION:DCARS R package and sample data are available at https://github.com/shazanfar/DCARS. Publicly available data from The Cancer Genome Atlas (TCGA) was used using the TCGABiolinks R package. Supplementary Files and DCARS R package is available at https://github.com/shazanfar/DCARS.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.