Knowledge concerning spatio-temporal distributions of populations is a prerequisite for successful conservation and management of migratory animals. Achieving cost-effective monitoring of large-scale movements is often difficult due to lack of effective and inexpensive methods. Taiga bean goose Anser fabalis fabalis and tundra bean goose A. f. rossicus offer an excellent example of a challenging management situation with harvested migratory populations. The subspecies have different conservation statuses and population trends. However, their distribution overlaps during migration to an unknown extent, which, together with their similar appearance, has created a conservation-management dilemma. Gaussian process (GP) models are widely adopted in the field of statistics and machine learning, but have seldom been applied in ecology so far. We introduce the R package gplite for GP modelling and use it in our case study together with birdwatcher observation data to study spatio-temporal differences between bean goose subspecies during migration in Finland in 2011-2019. We demonstrate that GP modelling offers a flexible and effective tool for analysing heterogeneous data collected by citizens. The analysis reveals spatial and temporal distribution differences between the two bean goose subspecies in Finland. Taiga bean goose migrates through the entire country, whereas tundra bean goose occurs only in a small area in south-eastern Finland and migrates later than taiga bean goose. Synthesis and applications. Within the studied bean goose populations, harvest can be targeted at abundant tundra bean goose by restricting hunting to south-eastern Finland and to the end of the migration period. In general, our approach combining citizen science data with GP modelling can be applied to study spatio-temporal distributions of various populations and thus help in solving challenging management situations. The introduced R package gplite can be applied not only to ecological modelling, but to a wide range of analyses in other fields of science.
Variable selection, or more generally, model reduction is an important aspect of the statistical workflow aiming to provide insights from data. In this paper, we discuss and demonstrate the benefits of using a reference model in variable selection. A reference model acts as a noise-filter on the target variable by modeling its data generating mechanism. As a result, using the reference model predictions in the model selection procedure reduces the variability and improves stability, leading to improved model selection performance. Assuming that a Bayesian reference model describes the true distribution of future data well, the theoretically preferred usage of the reference model is to project its predictive distribution to a reduced model, leading to projection predictive variable selection approach. We analyse how much the great performance of the projection predictive variable is due to the use of reference model and show that other variable selection methods can also be greatly improved by using the reference model as target instead of the original data. In several numerical experiments, we investigate the performance of the projective prediction approach as well as alternative variable selection methods with and without reference models. Our results indicate that the use of reference models generally translates into better and more stable variable selection.
Adaptive importance sampling is a class of techniques for finding good proposal distributions for importance sampling. Often the proposal distributions are standard probability distributions whose parameters are adapted based on the mismatch between the current proposal and a target distribution. In this work, we present an implicit adaptive importance sampling method that applies to complicated distributions which are not available in closed form. The method iteratively matches the moments of a set of Monte Carlo draws to weighted moments based on importance weights. We apply the method to Bayesian leave-one-out cross-validation and show that it performs better than many existing parametric adaptive importance sampling methods while being computationally inexpensive.
A salient approach to interpretable machine learning is to restrict modeling to simple models. In the Bayesian framework, this can be pursued by restricting the model structure and prior to favor interpretable models. Fundamentally, however, interpretability is about users' preferences, not the data generation mechanism; it is more natural to formulate interpretability as a utility function. In this work, we propose an interpretability utility, which explicates the trade-off between explanation fidelity and interpretability in the Bayesian framework. The method consists of two steps. First, a reference model, possibly a black-box Bayesian predictive model which does not compromise accuracy, is fitted to the training data. Second, a proxy model from an interpretable model family that best mimics the predictive behaviour of the reference model is found by optimizing the interpretability utility function. The approach is model agnostic-neither the interpretable model nor the reference model are restricted to a certain class of models-and the optimization problem can be solved using standard tools. Through experiments on real-word data sets, using decision trees as interpretable models and Bayesian additive regression models as reference models, we show that for the same level of interpretability, our approach generates more accurate models than the alternative of restricting the prior. We also propose a systematic way to measure stability of interpretabile models constructed by different interpretability approaches and show that our proposed approach generates more stable models.
This paper discusses predictive inference and feature selection for generalized linear models with scarce but high-dimensional data. We argue that in many cases one can benefit from a decision theoretically justified two-stage approach: first, construct a possibly non-sparse model that predicts well, and then find a minimal subset of features that characterize the predictions. The model built in the first step is referred to as the \emph{reference model} and the operation during the latter step as predictive \emph{projection}. The key characteristic of this approach is that it finds an excellent tradeoff between sparsity and predictive accuracy, and the gain comes from utilizing all available information including prior and that coming from the left out features. We review several methods that follow this principle and provide novel methodological contributions. We present a new projection technique that unifies two existing techniques and is both accurate and fast to compute. We also propose a way of evaluating the feature selection process using fast leave-one-out cross-validation that allows for easy and intuitive model size selection. Furthermore, we prove a theorem that helps to understand the conditions under which the projective approach could be beneficial. The benefits are illustrated via several simulated and real world examples.
Variable selection for Gaussian process models is often done using automatic relevance determination, which uses the inverse length-scale parameter of each input variable as a proxy for variable relevance. This implicitly determined relevance has several drawbacks that prevent the selection of optimal input variables in terms of predictive performance. To improve on this, we propose two novel variable selection methods for Gaussian process models that utilize the predictions of a full model in the vicinity of the training points and thereby rank the variables based on their predictive relevance. Our empirical results on synthetic and real world data sets demonstrate improved variable selection compared to automatic relevance determination in terms of variability and predictive performance.
The accuracy of an integral approximation via Monte Carlo sampling depends on the distribution of the integrand and the existence of its moments. In importance sampling, the choice of the proposal distribution markedly affects the existence of these moments and thus the accuracy of the obtained integral approximation. In this work, we present a method for improving the proposal distribution that applies to complicated distributions which are not available in closed form. The method iteratively matches the moments of a sample from the proposal distribution to their importance weighted moments, and is applicable to both standard importance sampling and self-normalized importance sampling. We apply the method to Bayesian leave-one-out cross-validation and show that it can significantly improve the accuracy of model assessment compared to regular Monte Carlo sampling or importance sampling when there are influential observations. We also propose a diagnostic method that can estimate the convergence rate of any Monte Carlo estimator from a finite random sample.
Gaussian graphical models are used for determining conditional relationships between variables. This is accomplished by identifying off-diagonal elements in the inverse-covariance matrix that are non-zero. When the ratio of variables (p) to observations (n) approaches one, the maximum likelihood estimator of the covariance matrix becomes unstable and requires shrinkage estimation. Whereas several classical (frequentist) methods have been introduced to address this issue, Bayesian methods remain relatively uncommon in practice and methodological literatures. Here we introduce a Bayesian method for estimating sparse matrices, in which conditional relationships are determined with projection predictive selection. Through simulation and an applied example, we demonstrate that the proposed method often outperforms both classical and alternative Bayesian estimators with respect to frequentist risk and consistently made the fewest false positives.We end by discussing limitations and future directions, as well as contributions to the Bayesian literature on the topic of sparsity.
Antenna line alarms are a significant indicator for failing components in LTE base stations. By predicting such alarms it is possible to prepare for system repairs or for system resets in advance, which reduces the maintenance downtime for LTE operators when a hardware failure actually occurs. In this paper we present results from an antenna line alarm prediction method in a real LTE network. The prediction model is based on the behavior of a large set of performance counters (PMs) collected in a timeseries in the LTE base stations. We create the alarm prediction model by using the random forest search method and we reach and accuracy of 67-72% From this, we pinpoint the most influential PMs for predicting different types of alarms and we discuss how to use these PMs in future research.
In high-dimensional prediction problems, where the number of features may greatly exceed the number of training instances, fully Bayesian approach with a sparsifying prior is known to produce good results but is computationally challenging. To alleviate this computational burden, we propose to use a preprocessing step where we first apply a dimension reduction to the original data to reduce the number of features to something that is computationally conveniently handled by Bayesian methods. To do this, we propose a new dimension reduction technique, called iterative supervised principal components (ISPC), which combines variable screening and dimension reduction and can be considered as an extension to the existing technique of supervised principal components (SPCs). Our empirical evaluations confirm that, although not foolproof, the proposed approach provides very good results on several microarray benchmark datasets with very affordable computation time, and can also be very useful for visualizing high-dimensional data.
Gaussian graphical models are used for determining conditional relationships between variables. This is accomplished by identifying off-diagonal elements in the inverse-covariance matrix that are non-zero. When the ratio of variables (p) to observations (n) approaches one, the maximum likelihood estimator of the covariance matrix becomes unstable and requires shrinkage estimation. Whereas several classical (frequentist) methods have been introduced to address this issue, fully Bayesian methods remain relatively uncommon in practice and methodological literatures. Here we introduce a Bayesian method for estimating sparse matrices, in which conditional relationships are determined with projection predictive selection. With this method, that uses Kullback-Leibler divergence and cross-validation for neighborhood selection, we reconstruct the inverse-covariance matrix in both low and high-dimensional settings. Through simulation and applied examples, we characterized performance compared to several Bayesian methods and the graphical lasso, in addition to TIGER that similarly estimates the inverse-covariance matrix with regression. Our results demonstrate that projection predictive selection not only has superior performance compared to selecting the most probable model and Bayesian model averaging, particularly for high-dimensional data, but also compared to the the Bayesian and classical glasso methods. Further, we show that estimating the inverse-covariance matrix with multiple regression is often more accurate, with respect to various loss functions, and efficient than direct estimation. In low-dimensional settings, we demonstrate that projection predictive selection also provides competitive performance. We have implemented the projection predictive method for covariance selection in the R package GGMprojpred
The horseshoe prior has proven to be a noteworthy alternative for sparse Bayesian estimation, but as shown in this paper, the results can be sensitive to the prior choice for the global shrinkage hyperparameter. We argue that the previous default choices are dubious due to their tendency to favor solutions with more unshrunk coefficients than we typically expect a priori. This can lead to bad results if this parameter is not strongly identified by data. We derive the relationship between the global parameter and the effective number of nonzeros in the coefficient vector, and show an easy and intuitive way of setting up the prior for the global parameter based on our prior beliefs about the number of nonzero coefficients in the model. The results on real world data show that one can benefit greatly -- in terms of improved parameter estimates, prediction accuracy, and reduced computation time -- from transforming even a crude guess for the number of nonzero coefficients into the prior for the global parameter using our framework.
The goal of this paper is to compare several widely used Bayesian model selection methods in practical model selection problems, highlight their differences and give recommendations about the preferred approaches. We focus on the variable subset selection for regression and classification and perform several numerical experiments using both simulated and real world data. The results show that the optimization of a utility estimate such as the cross-validation (CV) score is liable to finding overfitted models due to relatively high variance in the utility estimates when the data is scarce. This can also lead to substantial selection induced bias and optimism in the performance evaluation for the selected model. From a predictive viewpoint, best results are obtained by accounting for model uncertainty by forming the full encompassing model, such as the Bayesian model averaging solution over the candidate models. If the encompassing model is too complex, it can be robustly simplified by the projection method, in which the information of the full model is projected onto the submodels. This approach is substantially less prone to overfitting than selection based on CV-score. Overall, the projection method appears to outperform also the maximum a posteriori model and the selection of the most probable variables. The study also demonstrates that the model selection can greatly benefit from using cross-validation outside the searching process both for guiding the model size selection and assessing the predictive performance of the finally selected model.
The horseshoe prior has proven to be a noteworthy alternative for sparse Bayesian estimation, but has previously suffered from two problems. First, there has been no systematic way of specifying a prior for the global shrinkage hyperparameter based on the prior information about the degree of sparsity in the parameter vector. Second, the horseshoe prior has the undesired property that there is no possibility of specifying separately information about sparsity and the amount of regularization for the largest coefficients, which can be problematic with weakly identified parameters, such as the logistic regression coefficients in the case of data separation. This paper proposes solutions to both of these problems. We introduce a concept of effective number of nonzero parameters, show an intuitive way of formulating the prior for the global hyperparameter based on the sparsity assumptions, and argue that the previous default choices are dubious based on their tendency to favor solutions with more unshrunk parameters than we typically expect a priori. Moreover, we introduce a generalization to the horseshoe prior, called the regularized horseshoe, that allows us to specify a minimum level of regularization to the largest values. We show that the new prior can be considered as the continuous counterpart of the spike-and-slab prior with a finite slab width, whereas the original horseshoe resembles the spike-and-slab with an infinitely wide slab. Numerical experiments on synthetic and real world data illustrate the benefit of both of these theoretical advances.
The authors present a detailed analysis of the asymptotic frequentist properties of credible sets derived from posteriors with normal-linear measurement models and horseshoe priors. Although we disagree with the claim that “In Bayesian practice credible balls are nevertheless used as if they were confidence sets”, the results in the paper are important for identifying where the horseshoe priors are fragile asymptotically, and hence particularly dangerous in the non-asymptotic regimes more typical of the applied problems where sparse models are needed. One clarification we believe is warranted is that the horseshoe family of prior distributions does not encode sparsity as is typically interpreted. Instead of partitioning parameters into those that are zero and non-zero, the horseshoe priors actually separate parameters into those that are resolvable by measurements and those that are not. In particular, as with any model the horseshoe priors cannot be interpreted outside of the context of a particular likelihood (Gelman et al., 2017). Consequently the statement that “τ can be interpreted as the proportion of nonzero parameters, up to a logarithmic factor” is not quite true. Piironen and Vehtari (2017b; 2017c) demonstrate that the effects of τ in horseshoe priors are intimately related to the measurement variability σ, even for the simple normal-linear measurement model. Figure 6 of Piironen and Vehtari (2017c), for example, clearly illustrates that rescaling the data changes the impact of the horseshoe prior unless τ is scaled by σ, even with an oracle prior information about the true number of significant parameters, p0 = pn. In particular, the resolution threshold √ 2 log(n/pn) arising in the paper implicitly assumes that the measurement variability σ is equal to 1,