Projective outlier-robust M-quantile-based small area estimators can be substantially biased when the sample data contain representative outliers. In this article we propose two new predictive type bias corrected versions of these estimators for continuous and discrete outcomes. Given both area level and individual level outliers in the population, these new estimators are more efficient than the robust-predictive and robust-projective estimators that have been proposed in the small area estimation literature. We also propose two estimators of the prediction mean-squared error of these estimators: one based on Taylor linearization and the other based on a new semi-parametric bootstrap method. We summarize the empirical evidence for these theoretical results in this article, while in the supplementary material we describe in more detail how the properties of these M-quantile-based small area estimators have been assessed in model-based and design-based simulations, as well as in a realistic application focusing on estimation of average income and unemployment rates for local labor market areas in Italy. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
The paper introduces a novel framework for small area estimation based on spatio-temporal M-quantile regression. The proposed approach extends the Geographically Weighted Regression by incorporating both spatial and temporal weighting schemes, and integrates them with the M-quantile modelling to effectively capture local distributional features across space and time. The resulting predictors are specifically designed for out-of-sample prediction in small domains and are accompanied by analytical estimators of their mean squared error. The methodology is evaluated through extensive simulation studies, demonstrating strong robustness to spatio-temporal dependence and the presence of outliers at both unit and area levels. An application to county-level air quality data in the United States (2016-2023) highlights the predictive performance and practical relevance of the proposed methods.
M-quantile (MQ) regression provides a robust and flexible alternative to mixed models for small area estimation. However, several theoretical aspects remain underexplored. In this paper, a parametric bootstrap method is proposed to approximate the distributions of area-specific MQ coefficients and applied to adjust the bias in the mean squared error (MSE) estimation of predictors for population means. The unified bias adjustment method, based on the laws of total expectation and variance, is general and can be applied to any MSE estimator that neglects the uncertainty in predicting MQ coefficients. Simulation experiments evaluate the performance of the adjusted MSE estimators under different scenarios, including those with atypical values. A real-world case study illustrates the practical relevance of the proposed methodology.
In recent years, addressing Educational Poverty (EP) has emerged as a pressing issue on the political agendas of several countries, recognized as a new social challenge requiring urgent attention. To assess EP in Italy, the Italian National Statistical Institute introduced a multidimensional composite index known as the EPI. Considering the same dimensions of this index and the relationships among them, we employ a quantile approach to have a comprehensive understanding of the relationships among the variables. Our results suggest that the EPI dimensions play a different role at different points of the index distribution.
In Industry 4.0 factories, innovative prediction tools are adopted so that data can be systematically processed into information that can explain uncertainties and support decisions. Predictive manufacturing systems begin with acquiring data from monitored assets using appropriate sensors to extract various signals. These signals can then be integrated with historical data into extensive datasets containing a multitude of variables. Consequently, addressing the challenge of reducing dimensionality becomes of paramount importance. Dimension reduction techniques such as partial least squares (PLS) have recently gained attention to deal with the problem of big datasets with a large number of correlated variables. Standard PLS approaches confine the estimation to examining only average effects, resulting in an insufficient portrayal. In this paper, we combine the standard PLS technique with M-quantile regression. The proposed approach aims at offering a more comprehensive view of the effect of various dimensions on the degradation of etching equipment in the microchip fabrication process.
Small Area Estimation (SAE) methods are used to obtain reliable estimates of finite population descriptive quantities of interest when domain sample sizes are too small to provide adequate precision for direct domain estimators. Standard SAE methods are based on linear mixed models (LMM) with area-specific random effects. However, due to recent advances in software and hardware capabilities, it is more and more common to have high-dimensional datasets with many potential predictors highly correlated. Therefore, dimension reduction techniques such as partial least squares (PLS) have recently gained attention to deal with these problems. Conventional PLS approaches do not allow to explicitly address the hierarchical dependence. For this reason, in this paper, we combine the standard Fay-Herriot model with a PLS technique for estimating averages at the small area level. The performance of the proposed predictors is empirically assessed in model-based simulations.
In small area estimation, it is a smart strategy to rely on data measured over time. However, linear mixed models struggle to properly capture time dependencies when the number of lags is large. Given the lack of published studies addressing robust prediction in small areas using time-dependent data, this research seeks to extend M-quantile models to this field. Indeed, our methodology successfully addresses this challenge and offers flexibility to the widely imposed assumption of unit-level independence. Under the new model, robust bias-corrected predictors for small area linear indicators are derived. Additionally, the optimal selection of the robustness parameter for bias correction is explored, contributing theoretically to the field and enhancing outlier detection. For the estimation of the mean squared error (MSE), a first-order approximation and analytical estimators are obtained under general conditions. Several simulation experiments are conducted to assess the performance of the new predictors and MSE estimators, as well as the optimal selection of the robustness parameter. Finally, an application to the Spanish Living Conditions Survey data illustrates the usefulness of the proposed predictors.
This study examines the relationship between gender and academic performance across different quantiles among students enrolled in a 3-year STEM (Science, Technology, Engineering, and Mathematics) degree program in Italy. We make use of a unique dataset of linked administrative records, provided through an agreement with the Italian Ministry of University and Research (MUR). The statistical modeling of earned credits presents challenges posed by the discrete and often irregular nature of the observed distribution and the hierarchical structure of our data, which demand an estimation strategy that extends beyond the simplicity of quantile regression. We implement a methodology based on the jittering approach for counts and penalized fixed effects in order to deal with these two distinct extensions over standard quantile regression.
In recent years, the need to address educational poverty (EP) has become a pressing concern on the political agenda of several countries, and this issue has been recognized as a novel social challenge that demands immediate attention. To measure the level of EP in Italy, based on an original proposal by the Italian National Statistical Institute, a composite index known as the Educational Poverty Index (EPI) has been defined and used by several authors. In this paper, we focus on the same set of dimensions included in the multidimensional EPI and the relationships among them; we also employ a hierarchical composite model to measure EP in Italian regions while simultaneously taking gender into account as an additional factor. However, employing this method limits the estimation solely to average effects, resulting in an insufficient portrayal. Therefore, we also use quantile composite-based path modelling, which offers a comprehensive view of the relationships among the variables. Our results suggest that the dimensions of the EPI play distinct roles at different points in the index distribution. Moreover, the results tend to exhibit different patterns at the lower and higher quantiles of the index distribution according to gender.
Zero-inflated data often suffer from non-observation or measurement inaccuracies, which complicate prediction. While mixtures of generalized linear mixed models have been widely used for these outcomes, their accuracy relies heavily on strong parametric assumptions, making them sensitive to outliers in small areas. To address this, a robust modeling framework is introduced based on an extension of the M-quantile approach. The proposed methodology accommodates both zero-inflated and standard M-quantile models, offering greater resistance to outliers and misspecification. Asymptotic properties are established, and robust predictors with bias correction are derived, together with analytical expressions for their mean squared error. Model-based simulations highlight improvements over existing methods in settings with atypical data. An application to the Spanish Living Conditions Survey illustrates the practical advantages of the proposed methodology.
Estimating economic poverty indicators at the local level is essential for well-targeted data-driven welfare policies. However, Italy is a country characterized by strong geographical heterogeneity represented by unequal price levels among different areas, and computing poverty indicators with a national monetary poverty threshold can be misleading. This work proposes a novel approach to estimate monetary poverty incidence at the provincial level in Italy considering the different price levels within national boundaries. To account for local price variation, Spatial Price Indices (SPIs) are computed using scanner data on retail prices. The SPIs are estimated in two ways, referring to the mean local prices and using the 20th percentile of the prices. These two kinds of SPIs are used to adjust the national poverty line when computing the poverty incidence at the provincial level using Small Area Estimation (SAE) models. Our findings suggest that adjusting the national poverty line using the SPIs to compute a monetary poverty index can modify the poverty mapping results from the map produced with the traditional national poverty line that ignores price differences.
Spatial data are becoming increasingly accessible to urban scientists, but these data are often prone to measurement error. Motivated by the analysis of the Milan (Italy) apartment market heterogeneity, we propose a semiparametric approach to adjust for the presence of measurement error in the covariates when estimating M-quantile regression. The M-quantile approach helps explain the heterogeneity across individual units, preserving robustness and efficiency in the estimates. The model’s parameters are estimated within a penalised likelihood framework and an analytical expression is proposed to estimate standard errors. Asymptotic properties of estimates are also provided.
In this study, we proposed a new method for estimating the sensitivity of enterprises in Italy to the United Nation's sustainable development goals at the provincial level using web-scraping data (a nonprobability sample) because this value is not surveyed by the Italian National Institute of Statistics. The proposed method used a probability sample to reduce the selection bias of estimates obtained from the nonprobability sample in the context of small area estimation and integrated nonprobability and probability samples using a double robust estimator that combined (i) propensity weighting to improve the representativeness of the nonprobability sample and (ii) a statistical model to predict the units that were not in the nonprobability sample. A bootstrap procedure for estimating variance was also proposed. To validate the proposed method, a Monte Carlo simulation was performed. Results showed that the proposed method allowed the correction of bias from the nonprobability sample while maintaining a good level of estimate reliability.
This paper proposes an M-quantile regression approach to address the heterogeneity of the housing market in a modern European city. We show how M-quantile modelling is a rich and flexible tool for empirical market price data analysis, allowing us to obtain a robust estimation of the hedonic price function whilst accounting for different sources of heterogeneity in market prices. The suggested methodology can generally be used to analyse nonlinear interactions between prices and predictors. In particular, we develop a spatial semiparametric M-quantile model to capture both the potential nonlinear effects of the cultural environment on pricing and spatial trends. In both cases, nonlinearity is introduced into the model using appropriate bases functions. We show how the implicit price associated with the variable that measures cultural amenities can be determined in this semiparametric framework. Our findings show that the effect of several housing attributes and urban amenities differs significantly across the response distribution, suggesting that buyers of lower-priced properties behave differently than buyers of higher-priced properties.
Sample surveys on income and living conditions rarely give credible estimates of poverty indicators at sub-regional and local level. This explains the importance of Small Area Estimation (SAE) methods for measuring poverty at the local level. In this chapter, the reader is introduced to SAE for obtaining reliable estimates of poverty indicators at the local level when survey data are not sufficient (e.g., due to a lack of precision or a complete lack of data). Standard SAE methods and some new developments that allow the use of big data for estimating poverty at the local level are presented.
In this paper, we extend the linear M-quantile random intercept model (MQRE) to discrete data and use the proposed model to evaluate the effect of selected covariates on two count responses: the number of generic medical examinations and the number of specialised examinations for health districts in three regions of central Italy. The new approach represents an outlier-robust alternative to the generalised linear mixed model with Gaussian random effects and it allows estimating the effect of the covariates at various quantiles of the conditional distribution of the target variable. Results from a simulation experiment, as well as from real data, confirm that the method proposed here presents good robustness properties and can be in certain cases more efficient than other approaches.
The estimation of population agricultural indicators at sub-national local level (such as municipalities) is important to provide useful insights for suitable policy interventions. In many cases, information collected from national surveys allows estimation only for larger regions and the direct estimates of agricultural statistics cannot be produce with an adequate level of precision at the sub-national local level. Consequently, small area estimation (SAE) techniques could be implemented to obtain more precise estimator at local level, which will be used by policy makers. This chapter presents a review of the most important area-level models and shows the benefit of considering the spatial information to provide estimates of agricultural and rural statistics at a local level. The presented models are fitted with R to estimate the mean agrarian surface area used for the production of grape at the municipality level in Tuscany using data set based on the Italian Agricultural Census of the year 2000 for the Italian region of Tuscany.
The COVID-19 pandemic has heavily hit international economy giving a particular setback to the tourism sector. Between March and May 2020, during the first lockdown, and between October and December of the same year, during the second lockdown, a questionnaire was administrated in Italy, Greece and Great Britain. Through the questionnaire, people’s feelings and expectations of their desire to take a vacation were collected regarding the period of constraint due to the new coronavirus and the possible end of the pandemic, or the first government approved travel openings. In particular, the question of how long it would take to decide on a holiday, the type and duration, after the period of constriction due to the coronavirus was over, was asked. Both surveys, in the two different lockdown periods, showed the potential desire of tourists to leave relatively quickly, and to take forms of domestic tourism, characterized by small and short-lived trips. The favorite destination being the seaside.
Using the Programme for International Student Assessment (PISA) 2015 data for Italy, this paper offers a complete overview of the relationship between test anxiety and school performance by studying how anxiety affects the performance of students along the overall conditional distribution of mathematics, literature and science scores. We aim to indirectly measure whether higher goals increase test anxiety, starting from the hypothesis that high-skilled students generally set themselves high goals. We use an M-quantile regression approach that allows us to take into account the hierarchical structure and sampling weights of the PISA data. There is evidence of a negative and statistically significant relationship between test anxiety and school performance. The size of the estimated association is greater at the upper tail of the distribution of each score than at the lower tail. Therefore, our results suggest that high-performing students are more affected than low-performing students by emotional reactions to tests and school-work anxiety.
M-quantile random-effects regression represents an interesting approach for modelling multilevel data when the interest of researchers is focused on the conditional quantiles. When data are based on complex survey designs, sampling weights have to be incorporate in the analysis. A pseudo-likelihood approach for accommodating sampling weights in the M-quantile random-effects regression is presented. The proposed methodology is applied to the Italian sample of the "Program for International Student Assessment 2015" survey in order to study the gender gap in mathematics at various quantiles of the conditional distribution. Findings offer a possible explanation of the low share of females in "Science, Technology, Engineering and Mathematics" sectors.