In this work, we introduce ShinyDataMatcher, a user-friendly R Shiny application designed to support the integration of survey data through statistical matching. The tool enables practitioners to import, explore and process survey data, harmonize variables, select appropriate matching variables, and apply a wide range of macro- and micro-level matching methods without writing any code. The application guides the user through the full workflow of a matching exercise, from data preparation to the creation of a synthetic matched dataset, and includes diagnostic tools for assessing matching quality. To illustrate its capabilities, we present an application based on the Italian Household Budget Survey (HBS) and the Survey on Household Income and Wealth (SHIW), where the goal is to fuse income and expenditure information and to construct a Social Accounting Matrix (SAM). The example also highlights how repeated random hot-deck imputations can be used to account for the additional uncertainty induced by statistical matching. Overall, ShinyDataMatcher provides a transparent and accessible environment for exploring, prototyping, and implementing statistical matching procedures, lowering the technical barriers that often limit their use in applied and official-statistics contexts.
Bayesian spatio-temporal models are particularly effective in disease mapping, enabling the smoothing of relative risks and the monitoring of their temporal and spatial variations. Motivated by the study of mortality risk in the elderly population across Italian provinces, we propose an extended model that also accounts for seasonal variations. These seasonal patterns prove highly informative for understanding the impact of major events, such as the COVID-19 pandemic and heatwaves. The inclusion of seasonality increases both the number of interaction effects and the overall complexity of the model. Additionally, specifying meaningful prior distributions for the variance parameters of random effects is challenging due to the influence of model design and correlation structures. To ensure balanced prior variability among model components, we adopt the Design and Structure Dependent (DSD) prior specification approach, which appears to be an intuitive prior elicitation strategy, even for practitioners. Analysis of the results from the proposed model provides valuable insights into the phenomenon under study, including a quantification of the impacts of the first and second waves of COVID-19 and the differing effects of unusually warm summer seasons across regions.
In this paper we revise a popular alternative for estimating Poisson regression models in a Bayesian framework and discuss possible pitfalls tied to data features. The MCMC algorithms are based on augmenting the model via the introduction of auxiliary variables. This leads to a model linear in the regression parameters with errors following a Gumbel or a log-Gamma distribution depending on the augmentation strategy. Such distributions are approximated by a Gaussian Mixture in order to favor standard MCMC sampling with Gibbs steps after augmentation. We show situations when such an approximation deteriorates and causes non-convergence of the algorithm, discussing how this can be detected while the algorithm is running.
Bayesian hierarchical Poisson models are an essential tool for analyzing count data. However, designing efficient algorithms to sample from the posterior distribution of the target parameters remains a challenging task. Auxiliary mixture sampling algorithms have been proposed to this aim. They involve two steps of data augmentation: the first leverages the theory of Poisson processes, and the second approximates the residual distribution of the resulting model through a mixture of Gaussian distributions. In this way, an approximate Gibbs sampler can be implemented. This strategy is particularly beneficial for latent Gaussian models, as it allows one to exploit the sparsity of the precision matrix associated with the random effects and to efficiently incorporate linear constraints. In this paper, we focus on the accuracy of the approximation step, highlighting scenarios where the mixture fails to represent accurately the true underlying distribution, leading to a lack of convergence in the algorithm. We outline key features to monitor, in order to assess if the approximation performs as intended. Building on this, we propose a robust version of the auxiliary mixture sampling algorithm. Our approach includes mechanisms for detecting approximation failures and introduces an enhanced approximation of the right tail of the auxiliary variable distribution, supplemented by a Metropolis-Hastings correction step when needed. Finally, we evaluate the proposed algorithm together with the original mixture sampling algorithms on both simulated and real datasets.
This paper examines the spatio-temporal evolution of individuals aged 80 and over across Italian provinces from 2010 to 2022. From a methodological point of view, we employ a well-established Bayesian model for spatio-temporal disease mapping, with particular attention to model parameterization, the application of appropriate constraints, and the specification of priors. For the latter, we use design- and structure-dependent priors that account for both the design and structure matrices involved in defining the prior distributions of the random effects.
Many common correlation structures assumed for data can be described through latent Gaussian models. When Bayesian inference is carried out, it is required to set the prior distribution for scale parameters that rules the model components, possibly allowing to incorporate prior information. This task is particularly delicate and many contributions in the literature are devoted to investigating such aspects. We focus on the fact that the scale parameter controls the prior variability of the model component in a complex way since its dispersion is also affected by the correlation structure and the design. To overcome this issue that might confound the prior elicitation step, we propose to let the user specify the marginal prior of a measure of dispersion of the model component, integrating out the scale parameter, the structure and the design. Then, we analytically derive the implied prior for the scale parameter. Results from a simulation study, aimed at showing the behavior of the estimators sampling properties under the proposed prior elicitation strategy, are discussed. Lastly, some real data applications are explored to investigate prior sensitivity and allocation of explained variance among model components.
Statistical matching involves integrating different data sources with different units and a set of common variables. Applying statistical matching methods presents challenges for practitioners, including data pre-processing and the application of different statistical matching methods. To address this, we develop an R-Shiny application to simplify the integration of different datasets for users and enhance efficiency in statistical matching tasks. In this paper, we delineate the various features of the application, encompassing tabs for importing new data or selecting from available surveys, choosing the matching variable, conducting data wrangling, employing diverse statistical matching methods, and using the matched dataset for regression analysis.
In the last two decades, significant research efforts have been dedicated to addressing the issue of spatial confounding in linear regression models. Confounding occurs when the relationship between the covariate and the response variable is influenced by an unmeasured confounder associated with both. This results in biased estimators for the regression coefficients reduced efficiency, and misleading interpretations. This article aims to understand how confounding relates to the parameters of the data generating process. The sampling properties of the regression coefficient estimator are derived as ratios of dependent quadratic forms in Gaussian random variables: this allows us to obtain exact expressions for the marginal bias and variance of the estimator, that were not obtained in previous studies. Moreover, we provide an approximate measure of the marginal bias that gives insights of the main determinants of bias. Applications in the framework of geostatistical and areal data modeling are presented. Particular attention is devoted to the difference between smoothness and variability of random vectors involved in the data generating process. Results indicate that marginal covariance between the covariate and the confounder, along with marginal variability of the covariate, play the most relevant role in determining the magnitude of confounding, as measured by the bias.
The problem of computing the distribution of quadratic forms in normal variables has a long tradition in the statistical literature. Well-established numerical algorithms that deal with this task rely on the inversion of Fourier transforms or series representations. In this article, the Mellin transform is proposed as a tool to compute both the density and the cumulative distribution functions of a positive definite quadratic form: an outline of the numerical algorithm is presented, providing details on the error analysis. The algorithm's characteristics allow us to propose an efficient way to compute the random variables' quantiles. From the theoretical point of view, the analytic properties of the Mellin transform are exploited to provide a novel representation of the distribution of the ratio of independent quadratic forms as a mixture of beta random variables of the second kind. Moreover, algorithms are proposed for computations related to ratios of both independent and dependent quadratic forms. The methods are tested and compared to popular numerical algorithms in terms of computational times and accuracy. The R package QF implementing all the proposed algorithms is also made available. Supplementary materials for this article are available online.
Estimating poverty and inequality parameters for small sub-populations with adequate precision is often beyond the reach of ordinary survey-weighted methods because of small sample sizes. In small area estimation, survey data and auxiliary information are combined, in most cases using a model. In this paper, motivated by the analysis of EU-SILC data for Italy, we target the estimation of a selection of poverty and inequality indicators, that is mean, headcount ratio and quintile share ratio, adopting a Bayesian approach. We consider unit-level models specified on the log transformation of a skewed variable (equivalized income). We show how a finite mixture of log-normals provides a substantial improvement in the quality of fit with respect to a single log-normal model. Unfortunately, working with these distributions leads, for some estimands, to the non-existence of posterior moments whenever priors for the variance components are not carefully chosen, as our theoretical results show. To allow the use of moments in posterior summaries, we recommend generalized inverse Gaussian distributions as priors for variance components, guiding the choice of hyperparameters.
The availability of computational tools that allow to easy implement MCMC algorithms, makes the Bayesian approach attractive to deal with such inferential issues. However, it is known that the posterior moments of relevant quantities may not exist under the log-normal model, and quantiles are among them. As known, quantiles are crucial, for example, in estimating extreme values, assessing the exposure to pollutants in occupational health, and environmental monitoring.
The analysis of variance, and mixed models in general, are popular tools for analyzing experimental data in psychology. Bayesian inference for these models is gaining popularity as it allows to easily handle complex experimental designs and data dependence structures. When working on the log of the response variable, the use of standard priors for the variance parameters can create inferential problems and namely the non-existence of posterior moments of parameters and predictive distributions in the original scale of the data. The use of the generalized inverse Gaussian distributions with a careful choice of the hyper-parameters is proposed as a general purpose option for priors on variance parameters. Theoretical and simulations results motivate the proposal. A software package that implements the analysis is also discussed. As the log-transformation of the response variable is often applied when modelling response times, an empirical data analysis in this field is reported.
The log-normal distribution is very popular for modeling positive right-skewed data and represents a common distributional assumption in many environmental applications. Here we consider the estimation of quantiles of this distribution from a Bayesian perspective. We show that the prior on the variance of the log of the variable is relevant for the properties of the posterior distribution of quantiles. Popular choices for this prior, such as the inverse gamma, lead to posteriors without finite moments. We propose the generalized inverse Gaussian and show that a restriction on the choice of one of its parameters guarantees the existence of posterior moments up to a prespecified order. In small samples, a careful choice of the prior parameters leads to point and interval estimators of the quantiles with good frequentist properties, outperforming those currently suggested by the frequentist literature. Finally, two real examples from environmental monitoring and occupational health frameworks highlight the improvements of our methodology, especially in a small sample situation.
We consider the estimation of the relative median poverty gap (RMPG) at the level of Italian provinces by using data from the European Union Survey on Income and Living Conditions. The overall sample size does not allow reliable estimation of income-distribution-related parameters at the provincial level; therefore, small area estimation techniques must be used. The specific challenge in estimating the RMPG is that, as it summarizes the income distribution of the poor, samples for estimating it for small subpopulations are even smaller than those available in other parameters. We propose a Bayesian strategy where various parameters summarizing the distribution of income at the provincial level are modelled by means of a multivariate small area model. To estimate the RMPG, we relate these parameters to a distribution describing income, namely the generalized beta distribution of the second kind. Posterior draws from the multivariate model are then used to generate draws for the distribution's area-specific parameters and then of the RMPG defined as their functional.
Quantile and M-quantile regression have been applied successfully to small area estimation within the frequentist approach. Quantile regression is applied in the same context but from a Bayesian perspective. Joint modelling of the quantile function is considered, adopting a non parametric assumption on the data generating process that nonetheless explicitly includes the normal distribution as a special case. A specification of the random part of the model that is simple and consistent with the predictive aim of small area estimation is proposed. Although the main output of the method is the estimation of the whole quantile function, estimators of the small area means based on the integration of the quantile function are proposed and discussed. A simulation exercise is used to assess the frequentist properties of these proposed predictors, that result at least as efficient as frequentist small area estimators based on quantile regression in scenarios characterized by the presence of outliers. The proposed method is illustrated using data from the European survey on Income and Living Conditions (EU-SILC).
SummaryWe consider the estimation of the relative median poverty gap (RMPG) at the level of Italian provinces by using data from the European Union Survey on Income and Living Conditions. The overall sample size does not allow reliable estimation of income-distribution-related parameters at the provincial level; therefore, small area estimation techniques must be used. The specific challenge in estimating the RMPG is that, as it summarizes the income distribution of the poor, samples for estimating it for small subpopulations are even smaller than those available in other parameters. We propose a Bayesian strategy where various parameters summarizing the distribution of income at the provincial level are modelled by means of a multivariate small area model. To estimate the RMPG, we relate these parameters to a distribution describing income, namely the generalized beta distribution of the second kind. Posterior draws from the multivariate model are then used to generate draws for the distribution's area-specific parameters and then of the RMPG defined as their functional.