In this study, we proposed a new method for estimating the sensitivity of enterprises in Italy to the United Nation's sustainable development goals at the provincial level using web-scraping data (a nonprobability sample) because this value is not surveyed by the Italian National Institute of Statistics. The proposed method used a probability sample to reduce the selection bias of estimates obtained from the nonprobability sample in the context of small area estimation and integrated nonprobability and probability samples using a double robust estimator that combined (i) propensity weighting to improve the representativeness of the nonprobability sample and (ii) a statistical model to predict the units that were not in the nonprobability sample. A bootstrap procedure for estimating variance was also proposed. To validate the proposed method, a Monte Carlo simulation was performed. Results showed that the proposed method allowed the correction of bias from the nonprobability sample while maintaining a good level of estimate reliability.
SummaryUnemployment rate estimates for small areas are used to efficiently support the distribution of services and the allocation of resources, grants and funding. A Fay–Herriot type model is the most used tool to obtain these estimates. Under this approach out‐of‐sample areas require some synthetic estimates. As the geographical context is extremely important for analysing local economies, in this paper, we allow for area random effects to be spatially correlated. The spatial model parameters are estimated by a marginal likelihood method and are used to predict in‐sample as well as out‐of‐sample areas. Extensive simulation experiments are used to assess the impact of the auto‐regression parameter and of the rate of out‐of‐sample areas on the performance of this approach. The paper concludes with an illustrative application on real data from the Italian Labour Force Survey in which the estimation of the unemployment rate in each Local Labour Market Area is addressed.
In-work poverty has risen to become a key feature of European societies. In 2017, the percentage of workers at risk of low pay in Italy reached an estimated 25
Sample surveys on income and living conditions rarely give credible estimates of poverty indicators at sub-regional and local level. This explains the importance of Small Area Estimation (SAE) methods for measuring poverty at the local level. In this chapter, the reader is introduced to SAE for obtaining reliable estimates of poverty indicators at the local level when survey data are not sufficient (e.g., due to a lack of precision or a complete lack of data). Standard SAE methods and some new developments that allow the use of big data for estimating poverty at the local level are presented.
Citizen Data and Citizen Science are undoubtedly a challenge and an opportunity for Official Statistics. The paper follows the evolution in the production of statistics and indicators and gives some indications for the use of Citizen data in the production of indicators for monitoring of SDGs achievements.
The small area estimation (SAE) theory is widely used when local or domain-specific reliable estimates based on survey data are needed. Small area model-based estimates use a model that links the response variable to some auxiliary information borrowing strength from the related areas. When geographical information on the areas of interest is available, the specification of a spatial area level model can increase the estimates’ efficiency, depending on available auxiliary data. In this article, we first review the most popular area level spatial models, and we then compare their performance under two alternative scenarios of auxiliary information availability to estimate the average equivalized household income in Italian Local Labour Market Areas (LLMAs) using the EU-SILC (European Union Statistics on Income and Living Conditions) survey data. Our findings suggest that the spatial information can “fill the gap” when the covariates do not have a high predictive power, a crucial result when there is lack of auxiliary data. AMS Subject Classification: 62D05, 62G05, 62H11
In Italy, a crucial anti-poverty policy “Reddito di Cittadinanza” (RdC), a measure of guaranteed minimum income, was introduced in April 2019. We aim to evaluate the targeting of the RdC policy at the local level, as aggregated analyses could mask important misalignments between the share of beneficiaries of the RdC and the share of poor households. To measure the poverty share in the local areas of interest, two main indicators to capture and monitor poverty are used in Europe: the At-Risk-of-Poverty Rate based on the EU Statistics on Income and Living Conditions survey and the Absolute Poverty Index based on consumption data collected through the Household and Budget Survey. To obtain reliable estimates of these indicators at the local level, it is necessary to introduce small area estimation models that allow the use of data from different sources. We apply a bivariate Fay and Herriot model to provide reliable estimates of absolute and relative poverty for the assessment of RdC policy targeting in the 59 areas represented by the region by degree of urbanisation level in Italy. The degree of urbanisation is indeed a key geographical variable in the study of the poverty phenomenon. Our results suggest that the RdC policy implemented at the national level shows heterogeneous targeting performance at the local level, excluding large shares of poor households from the program. These findings yield a set of policy implications for improving the targeting of the measure.
In recent decades, the measurement and evaluation of important social and natural phenomena has significantly evolved, with many traditional measurements based on single variables increasingly being replaced by multidimensional approaches. One key aspect of these approaches is the development of composite indexes, usually real-value functions of multiple achievements of a group of units. The achievements in each of the selected dimensions are generally synthesised through one or more variables, often referred to as indicators. When indicators are obtained through an estimation process, it is crucial to understand if and how their estimation error – for example, sampling error – affects the resulting composite index. This paper presents a methodology based on a parametric bootstrap technique that evaluates to what extent uncertainty in indicators affects the reliability of the aggregate composite index. The method is applied to four composite indexes measuring the environmental performances of Italian regions based on real population and survey data. To our knowledge, this is the first attempt to measure the impact of indicators’ sampling error on composite indexes. If adequately generalised, our methodology could be used in the presence of measurement errors, non-response issues, or other kinds of non-sampling errors.
This chapter presents the data and the indicators used to define the Educational Poverty (EP) multi-dimensional measure. From the policymaking perspective, the results of our analysis demonstrate the actual consistency and dimensions of EP at a useful local level, which in turn, suggest, for example, a revision of local governments' normal policies employed to reverse EP. From a methodological point of view, taking a fuzzy perspective on EP attains particular significance in our multi-dimensional approach. Indeed, it overcomes the issue of transforming each single indicator used in the analysis into a simple dichotomous variable to indicate the percentage of people who are deprived with respect to each specific aspect, as implied by the approach taken by Mazziotta-Pareto. Although it is good practice to do so, this procedure requires information on the primary sampling units, rotational groups and strata, which are not available in the user database of the aspects of everyday life (AVQ) survey.
After 12 years of EMOS experience it is time to open the discussion on the future of EMOS. This papers briefly describes the experience from the perspective of the Universities, trying also to describe the needs and role of the NSIs, Banks and other possible actors to join the network, and unlock the future. EMOS should reload (or evolute) to stay current and attractive. Statistical ’thinking’ evolved and a major change and challenge for EMOS is to pick up this trend in its cooperation with the universities.
The estimation of population agricultural indicators at sub-national local level (such as municipalities) is important to provide useful insights for suitable policy interventions. In many cases, information collected from national surveys allows estimation only for larger regions and the direct estimates of agricultural statistics cannot be produce with an adequate level of precision at the sub-national local level. Consequently, small area estimation (SAE) techniques could be implemented to obtain more precise estimator at local level, which will be used by policy makers. This chapter presents a review of the most important area-level models and shows the benefit of considering the spatial information to provide estimates of agricultural and rural statistics at a local level. The presented models are fitted with R to estimate the mean agrarian surface area used for the production of grape at the municipality level in Tuscany using data set based on the Italian Agricultural Census of the year 2000 for the Italian region of Tuscany.
Composite indicators (CIs) are frequently employed to measure complex, multidimensional phenomena such as well-being, sustainability, or work quality. It is nowadays well studied, that the multiple decisions made in the construction process of building the CI strongly affect the result. These construction decisions range from the selection of a set of sub-indicators, to the standardization method, the weighting and the choice of the aggregation function. Sensitivity analysis is an established tool to analyze and quantify the impact of the choices made in this regard. Less attention has been paid to the impact of data quality on CIs: CIs are usally based on sub-indicators that are estimated from sample surveys and the resulting aggregated measure can only be as good as the underlying data. The uncertainty in the CI due to sampling and possible non-sampling errors is, however, frequently neglected. With this report we aim to fill this gap. To do so, we perform a sensitivity analysis for an example CI for working quality, that includes the selection of a sampling design and the sampling itself as possible sources of variability. Further, we propose a parametric bootstrap-based apporach to estimate the standard error of a CI and apply it to the illustrating example of a CI on environmental perfomance.
Official statistics are collected and produced by national statistical institutions (NSIs) based upon standardized questionnaire forms and a priori designed survey frame. Although the response to NSIs' surveys is mandatory for respondent units, increasing disaffection in replying to official surveys is a common trend across many advanced countries. This work explores the possibility to use Citizen-Generated Data (CGD) as a new information source for the compilation of official statistics. CGD represent a unique and still unexploited data source that share some key characteristics with Big Data, while they present some specific features in terms of information relevance and data generating process. Given the relevance of CGD to reduce the information gap between the demand and supply of new or more robust Sustainable Development Goals (SDG) indicators, the experimental setting to assess the data quality of CGD refers to different ways to integrate official statistics and CGD. Istat collects CGD within the framework of a pilot survey focused on key SDG indicators, and the appropriate methodological approach to assess data quality for official statistics is defined according to different data integration modalities.
Reducing the inequality between member states is a target the European Union (EU) has set itself in its treaties and monitors through its cohesion reports. This has become a much debated and researched issue over the last decade. One result emerges clearly in the debate: this goal is far to be reached without a deeper studying of inequalities within each Member States. This paper describes a general approach to the previous issue proposing a set of statistical methods that can be applied in the almost totality of the EU countries. It is based on data from European current sample surveys on consumption expenditure (Household Budget Survey) and on income and living conditions (EUSILC - European Survey on Income and Living Conditions). It uses the most popular poverty indicator from the Laeken set, the At Risk of Poverty Rate (ARPR) or Head Count Ratio (HCR). The examples are built on Italian data. The sub national level used is defined on the basis of the NUTS classification used by Eurostat.
The objective of this study is to investigate whether the quality of educational services and the university’s institutional image influence students’ overall satisfaction with their university experience as well as the possible consequences of these relationships on students’ loyalty. In particular, in today’s increasingly competitive higher education environment, such concepts have become of strategic concern in both public and private universities. To explain the complex system of relationships among these constructs, several hypotheses were formulated and tested through a structural equation model. Data were collected through a web questionnaire handed out to 14,870 students enrolled at the University of Pisa. The results provide valuable insight and show that teaching and lectures and teaching and course organization are the main determinants of students’ satisfaction and students’ loyalty among the more academic components of the educational service. Furthermore, the crucial role played by university image is worth noting, both for its direct and indirect effects on students’ satisfaction as well as on students’ loyalty and on teaching and lectures.
The importance of computing poverty measures at sub-national level is nowadays widely attested. Local poverty indicators are relevant both for a detailed planning of the policy actions against poverty and social exclusion, and for the citizens to evaluate their effects. However, there are still open problems to compute adequate sub-national poverty indicators. They refer to: 1) the definition of poverty lines; 2) the methods for accounting the spatial variation of the cost of living to make comparisons in ‘real terms’ between different areas; 3) the use of Small Area Estimation methods when the sample size is not enough to obtain accurate estimates of the indicators at local level. In this paper, we discuss the issues above by presenting some analyses on the impact of using different poverty lines on the value of the poverty rate for the 20 Italian Regions, which represent a planned domain of study in Italy. Then, we estimate the poverty rate for the 110 Italian Provinces, unplanned domains in Italy, by using specific parametric models and SAE methods. The key results highlight strong differences in the territorial distribution of the poverty rate by using national versus subnational specific poverty lines. The effect of the heterogeneity of the general spatial price indexes on the poverty rates seems instead less important in comparison with the relevant territorial differences in the cost of housing. Moreover, the different methods of estimation of poverty rates at local level provides interesting first results and indicates the route for further research to improve the methods of estimation of poverty at the sub-regional level.
While the contributions on the organized crime and Mafia environments are many, there is a lack of empirical evidence on the firm’s decision to resist to extortion. Our case study is based on Addiopizzo , an NGO that, from 2004, invites firms to refuse requests from the local Mafia and to join a public list of “non-payers”. The research is based on a dataset obtained linking the current administrative archives maintained by the chambers of commerce and the list updated by the NGO. The objective of this paper is twofold: first, to gather sound data on the characteristics of the Addiopizzo joiners; second to model the probability to join Addiopizzo by a two-level logistic regression model. We find that the resilience behavior is likely to be the result of both individual (firm) and environmental factors. In particular, we find that firm’s total assets, firm’s age and being in the construction sector are negatively correlated with the probability of joining AP, while a higher level of human capital embodied in the firm and a higher number of employees are positively correlated. Among the district-level variables, we find that the share of district’s population is negatively correlated with the probability to join, while a higher level of socio-economic development, including education levels, are positively correlated.