Abstract Background As populations are aging globally and healthcare systems are transitioning from being centred around single disease to patient-centred approach, precise and flexible identification of the multimorbidity patterns allows sound measures to prevent adverse health outcomes and better allocate healthcare resources. Methods Combining probabilistic approach of graphical model and intuitive visibility of network analysis, we use administrative health data of individuals aged 50 and above residing in Emilia-Romagna region (northern Italy) in 2011 and followed up to 2019 (N = 1,010,571) to investigate multimorbidity patterns and their impact across time. Results Four consistent multimorbidity patterns are identified across sex, across age groups (50–59, 60–69, 70–79, 80 +), and across time-points (2011, 2016 and 2019), which consists of cardiovascular, neuropsychiatric, respiratory-digestive and metabolic-pain pattern. This finding suggests plausible existence of stable population multimorbidity structure for public health monitoring. We also bring evidence of the better performance of multimorbidity patterns in mortality modelling compared to traditional methods. While neuropsychiatric pattern is the leading group causing mortality at older ages, higher education and living in suburban areas provide consistent protective effect against mortality, unplanned hospitalization and length of stay for all sexes and age groups. We focus on the gatekeeper diseases as potential targets for intervention. They are not necessarily the diseases with highest prevalence, but we demonstrate that an early detection of these diseases can contribute to improving three different health outcomes at older ages. Conclusion For each sex and age group above 50, we provide systematically the network of chronic diseases, the corresponding multimorbidity patterns, and the estimated impacts of multimorbidity, and thus a comprehensive understanding of multimorbidity at older ages. The observed consistent multimorbidity structure at population level and the identification of target diseases for intervention that goes beyond the prevalence-based approach can open new doors for more directions shaping public health policy.
This article illustrates how to measure the heterogeneity of spatial data presenting a finite number of categories via computation of spatial entropy. The R package SpatEntropy contains functions for the computation of entropy and spatial entropy measures. The extension to spatial entropy measures is a unique feature of SpatEntropy. In addition to the traditional version of Shannon's entropy, the package includes Batty's spatial entropy, O'Neill's entropy, Li and Reynolds' contagion index, Karlstrom and Ceccato's entropy, Leibovici's entropy, Parresol and Edwards' entropy and Altieri's entropy. The package is able to work with both areal and point data. This paper is a general description of SpatEntropy, as well as its necessary theoretical background, and an introduction for new users.
Entropy indices are commonly used to evaluate the heterogeneity of spatially arranged data by exploiting various approaches capable of including spatial information. Unfortunately, in practical studies, difficulties can arise regarding both the availability of computational tools for fast and easy implementation of these indices and guidelines supporting the correct interpretation of the results. The present work addresses such issues for the most known spatial entropy measures: the approach based on area partitions, the one based on distances between observations, and the decomposable spatial entropy. The newly released version of the R package SpatEntropy is introduced here and we show how it properly supports researchers in real case studies. This work also answers practical questions about the spatial distribution of nesting sites of an endangered species of gorillas in Cameroon. Such data present computational challenges, as they are marked points in continuous space over an irregularly shaped region, and covariates are available. Several aspects of the spatial heterogeneity of the nesting sites are addressed, using both the original point data and a discretised pixel dataset. We show how the diversity of the nesting habits is related to the environmental covariates, while seemingly not affected by the interpoint distances. The issue of scale dependence of the spatial measures is also discussed over these data. A motivating example shows the power of the SpatEntropy package, which allows for the derivation of results in seconds or minutes with minimum effort by users with basic programming abilities, confirming that spatial entropy indices are proper measures of diversity.
Entropy measures are standard tools in environmental and ecological sciences to describe the heterogeneity of data. This paper reviews a selection of spatial entropy indices, some of which are very recent, suitable to deal with spatial data on variables presenting a finite number of categories. A special focus is given on biodiversity data, but methods can be applied to any other environmental phenomena. The new R package SpatEntropy is here introduced to compute spatial entropy measures in practice. The extension from traditional entropy measures to their spatial version is a unique feature of the package, which is able to work with both areal and point data. A practical part is also presented, where two types of environmental data are considered, regarding trees biodiversity and urban expansion, respectively. The package SpatEntropy is run over these dataset and shown to represent a user-friendly and helpful tool for new package users.
Entropy is a measure of heterogeneity widely used in applied sciences, often when data are collected over space. Recently, a number of approaches has been proposed to include spatial information in entropy. The aim of entropy is to synthesize the observed data in a single, interpretable number. In other studies the objective is, instead, to use data for entropy estimation; several proposals can be found in the literature, which basically are corrections of the estimator based on substituting the involved probabilities with proportions. In this case, independence is assumed and spatial correlation is not considered. We propose a path for spatial entropy estimation: instead of correcting the global entropy estimator, we focus on improving the estimation of its components, i.e. the probabilities, in order to account for spatial effects. Once probabilities are suitably evaluated, estimating entropy is straightforward since it is a deterministic function of the distribution. Following a Bayesian approach, we derive the posterior probabilities of a multinomial distribution for categorical variables, accounting for spatial correlation. A posterior distribution for entropy can be obtained, which may be synthesized as wished and displayed as an entropy surface for the area under study.
A very recent proposal of a set of entropy measures for spatial data, based on building pairs of realizations, allows to split the data heterogeneity that is usually assessed via Shannon's entropy into two components: spatial mutual information, identifying the role of space, and spatial residual entropy, measuring heterogeneity due to other sources. A further decomposition into partial terms deeply investigates the role of space at specific distance ranges. The present work proposes improvements to the method and adds relevant results proving that the new set of spatial entropies satisfies a list of desirable properties. We extend the methodology to sets of realizations greater than pairs. We also show that the approach is more general, better performing and more interpretable than the most popular proposals in the literature, thanks to the property of additivity and a new way of computing entropy that explicitly discards the order within sets. A novel procedure for building the necessary quantities for computations is also provided. A comparative study illustrates the superior performance of the new set of measures over representative spatial configurations. Practical questions are answered by means of a case study on land use data.
The work presents some preliminary results of a research project on variations and determinants of secondary sex ratios in Italy. Variations in sex ratios at birth is still an active research field and several studies focus on this topic. Early works stated that the human sex ratio at birth was universally stable, without significant fluctuations across time and space. However, in the last decades various authors directly challenged these conclusions. After reviewing these studies on the historical trends of the sex ratio at birth, long term tendencies in Italy are analyzed. We collect historical and contemporary series of sex ratios at birth, covering more the 160 years, and compare them with those of different European countries. In this work, studies on the main determinant of long- and short-term trends are briefly reviewed, taking into account findings and results from different kinds of disciplines. The variations in the sex ratio at birth mentioned in the literature are assessed by using birth series at the national and regional level. Compared with other industrialized countries, sex ratio trends in Italy occurred with some delays. Intrauterine and fetal mortality appeared to be a key factor in sex ratio variations.
The concept of entropy, firstly introduced in information theory, rapidly became popular in many applied sciences via Shannon's formula to measure the degree of heterogeneity among observations. A rather recent research field aims at accounting for space in entropy measures, as a generalization when the spatial location of occurrences ought to be accounted for. The main limit of these developments is that all indices are computed conditional on a chosen distance. This work follows and extends the route for including spatial components in entropy measures. Starting from the probabilistic properties of Shannon's entropy for categorical variables, it investigates the characteristics of the quantities known as residual entropy and mutual information, when space is included as a second dimension. This way, the proposal of entropy measures based on univariate distributions is extended to the consideration of bivariate distributions, in a setting where the probabilistic meaning of all components is well defined. As a direct consequence, a spatial entropy measure satisfying the additivity property is obtained, as global residual entropy is a sum of partial entropies based on different distance classes. Moreover, the quantity known as mutual information measures the information brought by the inclusion of space, and also has the property of additivity. A thorough comparative study illustrates the superiority of the proposed indices.
In this article, we present a model-based framework to estimate the educational attainments of students in latent groups defined by unobservable or only partially observed features that are likely to affect the outcome distribution, as well as being interesting to be investigated. We focus our attention on the case of students in the first year of the upper secondary schools, for which the teachers’ suggestion at the end of their lower educational level toward the subsequent type of school is available. We use this information to develop latent strata according to the compliance behavior of students simplifying to the case of binary data for both counseled and attended school (i.e., academic or technical institute). We consider a likelihood-based approach to estimate outcome distributions in the latent groups and propose a set of plausible assumptions with respect to the problem at hand. In order to assess our method and its robustness, we simulate data resembling a real study conducted on pupils of the province of Bologna in year 2007/2008 to investigate their success or failure at the end of the first school year.
In this paper, the problem of combining information from different data sources is considered. We focus our attention on spatially misaligned data, where available information (typically counts or rates from administrative sources) refers to spatial units that are different from the ones of interest. A hierarchical Bayesian perspective is considered, as proposed by Mugglin et al. in 2000, to provide a fully model-based approach in an inferential, and not only descriptive, sense. In particular, explanatory covariates are arranged to be modeled according to spatial correlations through a conditionally autoregressive prior structure. In order to assess model performance and its robustness we generate artificial data inspired by a real study and a simulation exercise is then carried out.
In this article, we aim at assessing hierarchical Bayesian modeling for the analysis of multiple exposures and highly correlated effects in a multilevel setting. We exploit an artificial data set to apply our method and show the gains in the final estimates of the crucial parameters. As a motivating example to simulate data, we consider a real prospective cohort study designed to investigate the association of dietary exposures with the occurrence of colon-rectum cancer in a multilevel framework, where, e.g., individuals have been enrolled from different countries or cities. We rely on the presence of some additional information suitable to mediate the final effects of the exposures and to be arranged in a level-2 regression to model similarities among the parameters of interest (e.g., data on the nutrient compositions for each dietary item).
Multilevel modeling is a recently new class of statistical methods to handle nested data. Mainly thanks to the wide range of applicability and the great increase of statistical softwares, in the last decades multilevel modeling has enjoyed an explosion of published papers and books in both methodological and application field. Currently, there is a need to not only develop the research on multilevel approach for the analysis of complex data, but also to have instructions to properly address the usage. This work aims at summarizing methodological aspects related to multilevel models, illustrating good-practices, advantages, and limits by reviewing applications in various fields, such as socio-economic, educational, health, and medical sciences. We further focus our attention on the latest advances of multilevel modeling towards, e.g., the inclusion of latent variables and the Bayesian approach.
In the paper, we investigate the effects of family characteristics on the achievement of students in the first year of the upper secondary schools of the province of Bologna. In particular, we focus our attention on the number of siblings as potential causal factor influencing the outcome. We employ a matching strategy based on propensity score to create treatment groups, corresponding to the values of the factor under study, with the same distribution of observed covariates. As a result, students are stratified in blocks according to the propensity score to obtain estimates of the average treatment effect using nearest neighbour matching. In order to further compare the achievements of students of upper secondary schools in the city of Bologna with those in the other towns of the province, we show that valid inference is assured by controlling for family characteristics whose influence on the outcome has been previously assessed.
In this paper, we present a model-based framework to estimate the educational attainments of students in latent groups defined by unobservable features which are likely to affect the outcome. We focus our attention on the case of students in the first year of the upper secondary schools, for which the teachers’ suggestion at the end of the lower educational level toward the subsequent type of school is available. We exploit this information to develop latent strata according to the compliance behavior of students simplifying, as a first attempt, to binary data for both counselled and attended type of school (e.g., academic or technical institute). We consider a likelihood-based approach to estimate outcome distributions in the latent groups and propose a set of plausible assumptions with respect to the problem at hand. In order to assess our method and its robustness, we simulate data to resemble a real study conducted on pupils of the province of Bologna in year 2007/2008 to investigate their success or failure at the end of the first school year.
The paper deals with the analysis of the effects of multiple exposures on the occurrenceof a disease in observational case-control studies. We consider the case of multilevel data, with subjects nested in spatial clusters. As a result, we often face problems of small and sparse data, along with correlations among the exposures and the observations, which both invalidate the results from the ordinary analyses. A hierarchical Bayesian model is here proposed to manage the within-cluster dependence and the correlation among the exposures. We assign prior distributions on the crucial parameters by exploiting additional information at different levels and by making suitable assumptions according to the problem at hand. The model is conceived to be applied to a real multi-centric study aiming at investigating the association of dietary exposures with colon-rectum cancer occurrence. Compared with results obtained with conventional regressions, the hierarchical Bayesian model is shown to yield great gains in terms of more consistent and less biased estimates. Thanks to its flexibility, this approach represents a powerful statistical tool to be adopted in a wide range of applications. Moreover, the specification of more realistic priors may facilitate and extend the use of Bayesian solutions in the epidemiological field.
We consider a new approach to deal with non ignorable non response on an outcome variable, in a causal inference framework. Assuming that a binary instrumental variable for non response is available, we provide a likelihood-based approach to identify and estimate heterogeneous causal effects of a binary treatment on specific latent subgroups of units, named principal strata, defined by the non response behavior under each level of the treatment and of the instrument. We show that, within each stratum, non response is ignorable and respondents can be properly compared by treatment status. In order to assess our method and its robustness when the usually invoked assumptions are relaxed or misspecified, we simulate data to resemble a real experiment conducted on a panel survey which compares different methods of reducing panel attrition.
The pattern of longevity in the Italian north-eastern region of Emilia Romagna was investigated at the municipality level, considering a modified version of the centenarian rate (CR) in two different periods (1995-1999 and 2005-2009). Due to the rareness of such events in small areas, spatio-temporal modelling was used to tackle the random variations in the occurrence of long-lived individuals. This approach allowed us to exploit the spatial proximity to smooth the observed data, as well as controlling for the effects of a set of covariates. As a result, clusters of areas characterised by extreme indexes of longevity could be identified and the temporal evolution of the phenomenon depicted. A persistence of areas of lower and higher occurrences of long-lived subjects was observed across time. In particular, mean and median values higher than the regional ones, showed up in areas belonging to the provinces of Ravenna and Forli-Cesena, on one side spreading out along the Adriatic coast and, on the other stretching into the Apennine municipalities of Bologna and Modena. Further, a longitudinal perspective was added by carrying out a spatial analysis including the territorial patterns of past mortality. We evaluated the effects of the structure of mortality on the cohort of long-lived subjects in the second period. The major causes of death were considered in order to deepen the analysis of the observed geographical differences. The circulatory diseases seem to mostly affect the presence of long-lived individuals and a prominent effect of altitude and population density also emerges.
The paper deals with the analysis of the effects of multiple exposures on the occurrence of a disease in observational case-control studies. We consider the case of multilevel data, with subjects nested in spatial clusters. As a result, we often face problems of small and sparse data, along with correlations among the exposures and the observations, which both invalidate the results from the ordinary analyses. A hierarchical Bayesian model is here proposed to manage the within-cluster dependence and the correlation among the exposures. We assign prior distributions on the crucial parameters by exploiting additional information at different levels and by making suitable assumptions according to the problem at hand. The model is conceived to be applied to a real multi-centric study aiming at investigating the association of dietary exposures with colon-rectum cancer occurrence. Compared with results obtained with conventional regressions, the hierarchical Bayesian model is shown to yield great gains in terms of more consistent and less biased estimates. Thanks to its flexibility, this approach represents a powerful statistical tool to be adopted in a wide range of applications. Moreover, the specification of more realistic priors may facilitate and extend the use of Bayesian solutions in the epidemiological field.
In the paper, the effects of subsidies to Tuscan handicraft firms are evaluated; the study is affected by missing outcome values, which cannot be assumed missing at random. We tackle this problem within a causal inference framework. By exploiting Principal Stratification and the availability of an instrument for the missing mechanism, we conduct a likelihood-based analysis, proposing a set of plausible identification assumptions. Causal effects are estimated on (latent) subgroups of firms, characterized by their response behavior.
In this paper, we investigate the pattern of longevity during the last 15 years in Emilia Romagna, a North-Eastern region of Italy, at a municipality level. We consider a specific index of extreme longevity based on people aged 95 and over in two different periods (1995-1999 and 2005-2009). Spatio-temporal modeling is used to tackle at both periods the random variations in the occurrence of people 95+, due to the increasingly rareness of such events, especially in small areas. This method exploits the spatial proximity and the consequent interaction of the geographical areas to smooth the observations, as well as to control for the effects of a set of regressors. As a result, clusters of areas characterized by high and low indexes of longevity are well identified and the temporal evolution of the phenomenon can be depicted. In a parallel analysis, we consider the past levels of mortality on the same cohort of individuals reaching 95 years and over in the second period and when they were aged 80-89 and 90-99. Within this longitudinal framework, the longevity outcome is modeled by a spatial regression. The area-specific structures of mortality are included as regressors, whose effects represent the causal link between the occurrence of people 95+ and the causes of death in the same cohort.