Defining a 'hat-matrix' for a model is essential in many model diagnostic procedures, as it acts as an orthogonal projector from the observation space to the model space. In this paper, we introduce a unique Hat-matrix for the class of hierarchical generalised linear models (HGLMs), which includes, as a special case, the subclass of generalised linear mixed models (GLMMs). We provide a practical discussion on interpreting the hat matrix values in HGLMs across various settings, aimed at assisting practitioners in model diagnostics. Additionally, we propose two new empirical thresholds to identify high-leverage observations and clusters. We demonstrate the advantages of using these empirical thresholds over the traditional approach with a simulation study. Lastly, we present an application to real data to illustrate the effectiveness of our proposed methodology in real-world scenarios.
In this paper, we propose a Bayesian approach for spatial causal inference based on combining spatial propensity scoring with Integrated Nested Laplace Approximation. The method models both local and spillover exposure effects via multiple likelihoods and treats counterfactuals as missing data, allowing inference also for non-Gaussian outcomes. We validated the proposed method through simulations and an application to U.S. county-level cancer data, demonstrating the critical importance of properly accounting for spatial dependence when drawing causal conclusions from geostatistical data. Our results show that the proposed method achieves MCMC-comparable accuracy with substantially reduced computational time.
This paper explores the concept of consensus in the context of ternary preferences, an extension of dichotomous preference approvals, where alternatives are classified into three categories: acceptable, neutral, and unacceptable.We propose a novel distance-based measure to quantify consensus among voters and introduce a method for calculating the marginal contribution of each voter to the overall consensus, drawing parallels to the Banzhaf value in cooperative game theory. To handle large voter groups, we also present an estimation procedure based on sampling techniques to derive the marginal contributions. We performed comprehensive simulation studies to validate the statistical properties and computational efficiency of the proposed approach. Finally, empirical analyses using data from the Italian National Institute of Statistics (ISTAT) and the Balkan Barometer highlight its practical applicability.
The effective integration of real and synthetic clinical data in multiple languages is essential to advance healthcare research. In this study, we propose a statistical framework that leverages cross-lingual embeddings to validate semantic alignment between authentic Italian EHRs and synthetic English clinical notes. Using two state-of-the-art models, E5 and BGE, we encode the texts and employ Fuzzy C-Means clustering along with multidimensional scaling to assess their semantic coherence. Our analysis reveals distinct language-specific patterns alongside robust cross-lingual alignment, highlighting the promise of synthetic data augmentation in mitigating resource scarcity.
Textual analysis has gained significant interest in medical research, particularly for automated patient diagnosis based on clinical narratives. While traditional approaches often focus on associational methods, this paper explores the application of causal forests to analyze textual data from electronic health records (EHRs), aiming to identify causal relationships between specific words and the likelihood of receiving certain medical diagnoses. Utilizing the MIMIC-III dataset, we assess how linguistic factors influence diagnosis probabilities for three conditions: diabetes, hypothyroidism, and adrenal gland disorders. Our findings reveal significant causal links between certain clinical terms and diagnosis probabilities, emphasizing the potential of causal inference techniques to improve the analysis of language in clinical narratives. Additionally, we uncover heterogeneity in treatment effects, demonstrating that specific words can identify high-risk patient subgroups. This study highlights the importance of integrating causal inference in natural language processing within healthcare settings.
Text analysis has become increasingly common in medical research, especially for tasks like patient diagnosis based on medical notes. However, most existing approaches do not account for causal relationships between words and diagnoses. This paper proposes a causal approach using the MIMIC-III dataset to identify words or word pairs that causally affect the probability of receiving a specific diagnosis. We employ causal forests to assess the impact of individual linguistic factors on patient outcomes while adjusting for potential confounders. Our analysis reveals significant causal relationships between specific terms in clinical notes and the presence of hypothyroidism diagnosis.
Understanding comorbidity patterns is crucial for improving patient outcomes and optimising healthcare strategies. In this study, we propose an approach to detect comorbidities of two diseases from clinical discharge notes. To account for the complexities of textual data, we summarise the information through propensity scores, which represent the probability of receiving a certain diagnosis conditional on the extracted text. These scores are then used as covariates in a logistic regression model to explore the association between diseases. Specifically, we compare models trained on TF-IDF weighted document-term matrices and text embeddings, employing LASSO regression, XGBoost, and multilayer perceptrons (MLP). Our results, obtained by applying this method to study the association between diabetes and Chronic Kidney Disease, demonstrate the potential of Natural Language Processing (NLP) and machine learning techniques in advancing observational healthcare research.
Data from music streaming has gained increasing attention since it allowsstudying music preferences across diverse cultures and different periods of time.Indeed, the study of “music and emotion” is crucial for understanding the psychologicalrelationship between human sentiments and music. The temporal studyof musical emotions provides beneficial insights into the analysis of the mood oflisteners during periods of particular relevance and stress (e.g., the COVID-19pandemic). This study performs music streaming data analysis to retrieve themusical emotions of the top Italian streamed songs during the pandemic. To thisend, we propose two new indices for measuring anger and joy in songs. We suggesta procedure for clustering music streaming data: the DISTATIS procedureand Partitioning Around Medoids (PAM) clustering algorithm are sequentiallyapplied to identify intervals of time sharing similar sentiments. Finally, we employthe proposed procedure to investigate the relationship between the evolutionof the pandemic spread and sentiments extracted from songs. The results show that music streaming data analysis allow identifying fiveclusters of time intervals sharing similar sentiments,
Ranking and rating methods for preference data result in a different underlying organization of data that can lead to manifold probabilistic approaches to data modelling. As an alternative to existing approaches, two new flexible probability distributions are discussed as a modelling framework: the Discrete Beta and the Shifted Beta-Binomial . Through the presentation of three real-world examples, we demonstrate the practical utility of these distributions. These illustrative cases show how these novel distributions can effectively address real-world challenges, with a particular focus on data derived from surveys concerning environmental issues. Our analysis highlights the new distributions’ capability to capture the inherent structures within preference data, offering valuable insights into the field.
Plastic pollution has been extensively documented in the marine food web, but targeted studies focusing on the relationship between microplastic ingestion and fish trophic niches are still limited. In this study we investigated the frequency of occurrence and the abundance of micro- and mesoplastics (MMPs) in eight fish species with different feeding habits from the western Mediterranean Sea. Stable isotope analysis (δ13C and δ15N) was used to describe the trophic niche and its metrics for each species. A total of 139 plastic items were found in 98 out of the 396 fish analysed (25%). The bogue revealed the highest occurrence with 37% of individuals with MMPs in their gastrointestinal tract, followed by the European sardine (35%). We highlighted how some of the assessed trophic niche metrics seem to influence MMPs occurrence. Fish species with a wider isotopic niche and higher trophic diversity were more probable to ingest plastic particles in pelagic, benthopelagic and demersal habitats. Additionally, fish trophic habits, habitat and body condition influenced the abundance of ingested MMPs. A higher number of MMPs per individual was found in zooplanktivorous than in benthivore and piscivorous species. Similarly, our results show a higher plastic particles ingestion per individual in benthopelagic and pelagic species than in demersal species, which also resulted in lower body condition. Altogether, these results suggest that feeding habits and trophic niche descriptors can play a significant role in the ingestion of plastic particles in fish species.
Preference data represent a particular type of ranking data where a group of people gives their preferences over a set of alternatives. Within this framework, distance-based decision trees represent a non-parametric tool for identifying the profiles of subjects giving a similar ranking. This paper aims at detecting, in the framework of (complete and incomplete) ranking data, the impact of the differently structured weighted distances for building decision trees. By means of simulations, we will compute the impact of higher/lower homogeneity in groups and different weighting structures both on splitting and on consensus ranking. The distances that will be used satisfy Kemeny’s axioms and, accordingly, a modified version of the rank correlation coefficient τ _x , proposed by Emond and Mason, will be proposed and used for rank aggregation and class label in the tree leaves.
Preference-approval structures combine preference rankings and approval voting for declaring opinions over a set of alternatives. In this paper, we propose a new procedure for clustering alternatives in order to reduce the complexity of the preference-approval space and provide a more accessible interpretation of data. To that end, we present a new family of pseudometrics on the set of alternatives that take into account voters’ preferences via preference-approvals. To obtain clusters, we use the Ranked k -medoids (RKM) partitioning algorithm, which takes as input the similarities between pairs of alternatives based on the proposed pseudometrics. Finally, using non-metric multidimensional scaling, clusters are represented in 2-dimensional space.
Digital music distribution is increasingly powered by automated mechanisms that continuously capture, sort and analyze large amounts of Web-based data. This paper deals with the management of songs audio features from a statistical point of view. In particular, it explores the data catching mechanisms enabled by Spotify Web API and suggests statistical tools for the analysis of these data. Special attention is devoted to songs popularity and a Beta model including random effects is proposed in order to give the first answer to questions like which are the determinants of popularity? The identification of a model able to describe this relationship, the determination within the set of characteristics of those considered most important in making a song popular is a very interesting topic for those who aim to predict the success of new products.
Label Ranking (LR) is an emerging non-standard supervised classification problem with practical applications in different research fields. The Label Ranking task aims at building preference models that learn to order a finite set of labels based on a set of predictor features. One of the most successful approaches to tackling the LR problem consists of using decision tree ensemble models, such as bagging, random forest, and boosting. However, these approaches, coming from the classical unweighted rank correlation measures, are not sensitive to label importance. Nevertheless, in many settings, failing to predict the ranking position of a highly relevant label should be considered more serious than failing to predict a negligible one. Moreover, an efficient classifier should be able to take into account the similarity between the elements to be ranked. The main contribution of this paper is to formulate, for the first time, a more flexible label ranking ensemble model which encodes the similarity structure and a measure of the individual label importance. Precisely, the proposed method consists of three item-weighted versions of the AdaBoost boosting algorithm for label ranking. The predictive performance of our proposal is investigated both through simulations and applications to three real datasets.
A preference–approval on a set of alternatives consists of a weak order on that set and, additionally, a cut-off line that separates acceptable and unacceptable alternatives. In this paper, we propose a new method for defining the distance between preference–approvals taking into account jointly the disagreements in preferences and approvals for each pair of alternatives. The proposed distance is compared to the existing distance functions to deal with clustering problems. Specifically, we prove that our metric improves the estimated clusters in terms of both stability and accuracy.
We propose an iterative algorithm to select the smoothing parameters in additive quantile regression, wherein the functional forms of the covariate effects are unspecified and expressed via B-spline bases with difference penalties on the spline coefficients. The proposed algorithm relies on viewing the penalized coefficients as random effects from the symmetric Laplace distribution, and it turns out to be very efficient and particularly attractive with multiple smooth terms. Through simulations we compare our proposal with some alternative approaches, including the traditional ones based on minimization of the Schwarz Information Criterion. A real-data analysis is presented to illustrate the method in practice.
Preference data are a particular type of ranking data where some subjects (voters, judges,...) express their preferences over a set of alternatives (items). In most real life cases, some items receive the same preference by a judge, thus giving rise to a ranking with ties. An important issue involving rankings concerns the aggregation of the preferences into a “consensus”. The purpose of this paper is to investigate the consensus between rankings with ties, taking into account the importance of swapping elements belonging to the top (or to the bottom) of the ordering (position weights). By combining the structure of τ x proposed by Emond and Mason (J Multi-Criteria Decis Anal 11(1):17–28, 2002) with the class of weighted Kemeny-Snell distances, a position weighted rank correlation coefficient is proposed for comparing rankings with ties. The one-to-one correspondence between the weighted distance and the rank correlation coefficient is proved, analytically speaking, using both equal and decreasing weights.
The introduction of oral disease-modifying therapies (DMTs) for relapsing–remitting multiple sclerosis (RRMS) changed algorithms of RRMS treatment. To compare the effectiveness of treatment with dimethyl fumarate (DMF) and teriflunomide (TRF) in a large multicentre Italian cohort of RRMS patients. Patients with RRMS who received treatment with DMF and TRF between January 1st, 2012 and December 31st, 2018 from twelve MS centers were identified. The events investigated were “time-to-first-relapse”, “time-to-Magnetic-Resonance-Imaging (MRI)-activity” and “time-to-disability-progression”. 1445 patients were enrolled (1039 on DMF, 406 on TRF) and followed for a median of 34 months. Patients on TRF were older (43.5 ± 8.6 vs 38.8 ± 9.2 years), with a predominance of men and higher level of disability (p < 0.001 for all). Patients on DMF had a higher number of relapses and radiological activity (p < .05) at baseline. Time-varying Cox-model for the event “time-to-first relapse” revealed that no differences were found between the two groups in the first 38 months of treatment (HRt < 38DMF = 0.73, CI = 0.52 to 1.03, p = 0.079). When the time-on-therapy exceeds 38 months patients on DMF had an approximately 0.3 times lower relapse hazard risk than those who took TRF (HRt>38DMF = 3.83, CI = 1.11 to 13.23, p = 0.033). Both DMTs controlled similarly MRI activity and disability progression. Patients on DMF had higher relapse-free survival time than TRF group after the first 38 months on therapy.
Understanding the response of species to disturbance and the ability to recover is crucial for preventing their potential collapse and ecosystem phase shifts. Explosive submarine activity, occurring in shallow volcanic vents, can be considered as a natural pulse disturbance, due to its suddenness and high intensity, potentially affecting nearby species and ecosystems. Here, we present the response of Posidonia oceanica, a long-lived seagrass, to an exceptional submarine volcanic explosion, which occurred in the Aeolian Archipelago (Italy, Mediterranean Sea) in 2002, and evaluate its resilience in terms of time required to recover after such a pulse event. The study was carried out in 2011 in the sea area off Panarea Island, in the vicinity of Bottaro Island by adopting a back-dating methodological approach, which allowed a retrospective analysis of the growth performance and stable carbon isotopes (δ13C) in sheaths and rhizomes of P. oceanica, during a 10-year period (2001–2010). After the 2002 explosion, a trajectory shift towards decreasing values for both growth performance and δ13C in sheaths and rhizomes was observed. The decreasing trend reversed in 2004 when recovery took place progressively for all the analysed variables. Full recovery of P. oceanica occurred 8 years after the explosive event with complete restoration of all the variables (rhizome growth performance and δ13C) by 2010. Given the ecological importance of this seagrass in marine coastal ecosystems and its documented large-scale decline, the understanding of its potential recovery in response to environmental changes is imperative.