Cluster analysis is widely used in educational settings to gain insights into student learning. To justify their choice of clustering approach, authors often draw on methods used in earlier studies that they perceive to be similar. This can sometimes come at the expense of selecting a clustering method better suited to addressing their own study's goals. We argue that the selection of an appropriate clustering method should be closely connected to the context of the study. We demonstrate our argument through the use of the Open University Learning Analytics Dataset (OULAD), a well-established open access educational data set. Through a review of studies citing the OULAD, we identify five possible motivations for clustering educational data and then focus on two of these: early identification of at-risk students and identification of similar groups of learners. We discuss the educational context behind the two motivations selected, describe which variable subsets from the OULAD might best align with the specific research questions that they motivate, and illustrate how a preferred clustering method may be highly influenced and driven by specifics of the research question. For example, the desire for an expression of uncertainty in cluster membership allocation. Fully reproducible R code is provided.
INTRODUCTION:Cost-effectiveness modelling often requires the extrapolation of survival data from clinical trials. The choice of extrapolation method is often uncertain and can have a profound effect on the results. We propose an approach that incorporates external evidence (eg, from published studies, registries, or elicited clinical opinion) into the extrapolation process to reduce uncertainty. This method can be applied to adjust previously fitted parametric curves as a secondary analysis, without access to the underlying patient-level data. Adjusted curves can be easily exported and used in spreadsheet-based models. METHODS:Standard parametric survival curves are fitted to time-to-event data using maximum likelihood estimation (MLE). These are combined with external evidence on expected cohort-level survival at a future time point(s), formulated as a probability distribution, to generate adjusted survival curves that simultaneously incorporate both observed data and external evidence. Parameter estimation uses importance sampling and multivariate normal approximations of the likelihood. We apply our method to a case study of survival extrapolation from immuno-oncology. RESULTS:Our method resulted in extrapolated survival predictions that were more closely aligned with the external evidence compared with the standard (MLE-based) approach. The incorporation of external evidence decreased the between-distribution variance (reduced structural uncertainty) and for most distributions also decreased within-distribution variance (reduced parametric uncertainty). CONCLUSION:Our extrapolation method can reduce uncertainty when valid external evidence is available. Only MLE-based parameter estimates are required to implement our method; thus, secondary model users such as health technology assessment bodies can adjust survival extrapolations from existing cost-effectiveness models without access to patient-level data. Implementation is straightforward and computationally efficient, and outputs are easily incorporated into existing cost-effectiveness models.
Generative approaches to clustering provide information on geometric properties of clusters, whereas discriminative approaches provide boundaries between clusters. Ideas from both approaches are incorporated to present a fully unsupervised, probabilistic, and discriminative clustering method via a regularized mutual information objective function, wherein a mixture of mixtures of Gaussian and uniform distributions is used for formulation of the conditional model. Automatic selection of the number of components is established with the introduction of the regularizing term and a merge step, similar to those applied in reversible jump Markov chain Monte Carlo methods used in Bayesian clustering. Consequently, the turtle shell method – a fully unsupervised clustering method capable of estimating non-linear boundary lines, automatically selecting the number of components, and capturing intuitive clusters in the presence of data abnormalities such as noise and/or irregular cluster shapes – is introduced. We test this method on various simulated and real datasets commonly explored in clustering research, and extend the analysis to datasets arising from flow cytometry experiments.
Abstract Organ transplant recipients are at increased risk of skin cancer, particularly squamous cell carcinoma (SCC), due to long-term immunosuppression. The Skin and UV Neoplasia Transplant Risk Assessment Calculator (SUNTRAC) stratifies organ transplant recipients by skin cancer risk to guide initial surveillance, while recent British Society for Skin Care in Immunosuppressed Individuals (BSSCII) guidance provides recommendations on surveillance frequency. We aimed to evaluate skin cancer incidence, assess alignment of dermatological surveillance with SUNTRAC risk stratification, and compare observed follow-up practices with BSSCII recommendations in a cohort of kidney transplant recipients (KTR). A retrospective review was conducted of 281 adult KTRs under active immunosuppression at a single tertiary centre. Clinicopathological data were collected between June and August 2025. Skin cancers were confirmed histologically. Cumulative incidence, time to first and subsequent skin cancers, and surveillance intervals were analysed and compared with SUNTRAC and BSSCII recommendations. Overall, 107 of 281 KTRs (38.1%) developed skin cancer, with cumulative incidences of 13.4%, 24.8%, 44.7% and 64.2% at 5, 10, 20 and 30 years post-transplant, respectively. Ten-year cumulative incidence differed significantly by SUNTRAC risk category (19.2%, 45.0% and 85.7% in medium-, high- and very high-risk groups; P < 0.001). Median time to first skin cancer decreased with increasing risk (11.1, 4.9 and 0.8 years, respectively). Timing of initial dermatological surveillance did not consistently align with SUNTRAC recommendations, particularly in lower-risk patients (median 3.3 vs. 10 years). Restricting analysis to transplants from 2020 demonstrated earlier initial surveillance across all risk groups, although misalignment persisted. Following a first skin cancer, 57.8% developed subsequent lesions within 5 years. In patients who developed SCC, follow-up occurred significantly earlier than recommended (median 4 vs. 12 months; P < 0.001). SUNTRAC effectively stratified skin cancer risk in this cohort; however, surveillance practices did not consistently align with guideline recommendations. Formal integration of risk-based surveillance may optimize resource utilization while maintaining patient safety.
Staphylococcus aureus bloodstream infections are a significant risk for hemodialysis patients, who would significantly benefit from a preventative vaccine. To-date S. aureus vaccine trials have failed, in part due to lack of consideration of pre-existing immune imprints in relevant patient cohorts. Using a machine learning algorithm this study revealed, prior S. aureus exposure in a cohort of 180 hemodialysis patients, was associated with an increase in circulating S. aureus antigen-specific exTh17 cells and reduced antigen-specific IL-10 producing CD4 cells and circulating NK cells and Vδ2 cells. It provides novel insights into pre-existing S. aureus immunity in hemodialysis patients that could inform next generation vaccine development.
Sleep difficulties in children are heterogeneous in presentation, yet conventional assessment tools like the Children's Sleep Habits Questionnaire (CSHQ) reduce this complexity to a single cumulative score, obscuring distinct patterns of sleep disturbance that require different interventions. Latent Class Regression (LCR) models offer a principled approach to identify subgroups with shared sleep behaviour profiles whilst incorporating predictors of group membership, but Bayesian inference for these models has been hindered by computational challenges and the absence of variable selection methods. We propose a fully Bayesian framework for LCR that uses Pólya-Gamma data augmentation, enabling efficient sampling of regression coefficients. We extend this framework to include variable selection for both predictors and item responses: predictor variable selection via latent inclusion indicators and item selection through a partially collapsed approach. Through simulation studies, we show that the proposed methods yield accurate parameter estimates, resolve identifiability issues arising in full models and successfully identify informative predictors and items while excluding noise variables. Applying this methodology to CSHQ data from 148 children reveals distinct latent subgroups with different sleep behaviour profiles, anxious nighttime sleepers, short/light sleepers and those with more pervasive sleep problems, with each carrying distinct implications for intervention. Results also highlight the predictive role of Autism Spectrum Disorder diagnosis in subgroup membership. These findings demonstrate the limitations of conventional CSHQ scoring and illustrate the benefits of a probabilistic subgroup-based approach as an alternative for understanding paediatric sleep difficulties.
ABSTRACTModel‐based clustering is a statistical approach to cluster analysis, which has been successfully deployed in a number of domains due to its principled framework, clear assumptions, and adaptability. For these reasons, there has been substantial interest in applying model‐based clustering methods to flow cytometry and mass cytometry data. The identification of relevant cell populations is a crucial step in the analysis of cytometry data for immunological research. Technological advances have led to a rapid increase in the dimensionality and complexity of cytometry data, prompting significant interest in the use of clustering algorithms in place of traditional manual data analysis techniques for cell population identification. This article highlights how model‐based clustering methods, such as mixture models, have been adapted to meet the many interesting and unusual challenges that present themselves to the researcher when analyzing flow and mass cytometry data. These innovations demonstrate that there is considerable potential for further methodological development and collaboration between the cytometry and model‐based clustering research communities.
Attribution modelling lies at the heart of marketing effectiveness, yet most existing approaches depend on user-level path data, which are increasingly inaccessible due to privacy regulations and platform restrictions. This paper introduces a Causal-Driven Attribution (CDA) framework that infers channel influence using only aggregated impression-level data, avoiding any reliance on user identifiers or click-path tracking. CDA integrates temporal causal discovery (using PCMCI) with causal effect estimation via a Structural Causal Model to recover directional channel relationships and quantify their contributions to conversions. Using large-scale synthetic data designed to replicate real marketing dynamics, we show that CDA achieves an average relative RMSE of 9.50
The presence of outliers can prevent clustering algorithms from accurately determining an appropriate group structure within a data set. We present outlierMBC, a model-based approach for sequentially removing outliers and clustering the remaining observations. Our method identifies outliers one at a time while fitting a multivariate Gaussian mixture model to data. Since it can be difficult to classify observations as outliers without knowing what the correct cluster structure is a priori, and the presence of outliers interferes with the process of modelling clusters correctly, we use an iterative method to identify outliers one by one. At each iteration, outlierMBC removes the observation with the lowest density and fits a Gaussian mixture model to the remaining data. The method continues to remove potential outliers until a pre-set maximum number of outliers is reached, then retrospectively identifies the optimal number of outliers. To decide how many outliers to remove, it uses the fact that the squared sample Mahalanobis distances of Gaussian distributed observations are Beta distributed when scaled appropriately. outlierMBC chooses the number of outliers which minimises a dissimilarity between this theoretical Beta distribution and the observed distribution of the scaled squared sample Mahalanobis distances. This means that our method both clusters the data using a Gaussian mixture model and implements a model-based procedure to identify the optimal outliers to remove without requiring the number of outliers to be pre-specified. Unlike leading methods in the literature, outlierMBC does not assume that the outliers follow a known distribution or that the number of outliers can be pre-specified. Moreover, outlierMBC performs strongly compared to these algorithms when applied to a range of simulated and real data sets.
Objective ANCA-associated vasculitis (AAV) is a relapsing-remitting disease, resulting in incremental tissue injury. The gold-standard relapse definition (Birmingham Vasculitis Activity Score, BVAS>0) is often missing or inaccurate in registry settings, leading to errors in ascertainment of this key outcome. We sought to create a computable phenotype (CP) to automate retrospective identification of relapse using real-world data in the research setting.Methods We studied 536 patients with AAV and >6 months follow-up recruited to the Rare Kidney Disease registry (a national longitudinal, multicentre cohort study). We followed five steps: (1) independent encounter adjudication using primary medical records to assign the ground truth, (2) selection of data elements (DEs), (3) CP development using multilevel regression modelling, (4) internal validation and (5) development of additional models to handle missingness. Cut-points were determined by maximising the F1-score. We developed a web application for CP implementation, which outputs an individualised probability of relapse.Results Development and validation datasets comprised 1209 and 377 encounters, respectively. After classifying encounters with diagnostic histopathology as relapse, we identified five key DEs; DE1: change in ANCA level, DE2: suggestive blood/urine tests, DE3: suggestive imaging, DE4: immunosuppression status, DE5: immunosuppression change. F1-score, sensitivity and specificity were 0.85 (95% CI 0.77 to 0.92), 0.89 (95% CI 0.80 to 0.99) and 0.96 (95% CI 0.93 to 0.99), respectively. Where DE5 was missing, DE2 plus either DE1/DE3 were required to match the accuracy of BVAS.Conclusions This CP accurately quantifies the individualised probability of relapse in AAV retrospectively, using objective, readily accessible registry data. This framework could be leveraged for other outcomes and relapsing diseases.
Objectives This study aims to describe the data structure and harmonisation process, explore data quality and define characteristics, treatment, and outcomes of patients across six federated antineutrophil cytoplasmic antibody-associated vasculitis (AAV) registries. Methods Through creation of the vasculitis-specific Findable, Accessible, Interoperable, Reusable, VASCulitis ontology, we harmonised the registries and enabled semantic interoperability. We assessed data quality across the domains of uniqueness, consistency, completeness and correctness. Aggregated data were retrieved using the semantic query language SPARQL Protocol and Resource Description Framework Query Language (SPARQL) and outcome rates were assessed through random effects meta-analysis. Results A total of 5282 cases of AAV were identified. Uniqueness and data-type consistency were 100% across all assessed variables. Completeness and correctness varied from 49%–100% to 60%–100%, respectively. There were 2754 (52.1%) cases classified as granulomatosis with polyangiitis (GPA), 1580 (29.9%) as microscopic polyangiitis and 937 (17.7%) as eosinophilic GPA. The pattern of organ involvement included: lung in 3281 (65.1%), ear-nose-throat in 2860 (56.7%) and kidney in 2534 (50.2%). Intravenous cyclophosphamide was used as remission induction therapy in 982 (50.7%), rituximab in 505 (17.7%) and pulsed intravenous glucocorticoid use was highly variable (11%–91%). Overall mortality and incidence rates of end-stage kidney disease were 28.8 (95% CI 19.7 to 42.2) and 24.8 (95% CI 19.7 to 31.1) per 1000 patient-years, respectively. Conclusions In the largest reported AAV cohort-study, we federated patient registries using semantic web technologies and highlighted concerns about data quality. The comparison of patient characteristics, treatment and outcomes was hampered by heterogeneous recruitment settings.
BACKGROUND:Antineutrophil cytoplasmic antibody (ANCA)-associated vasculitis is a heterogenous autoimmune disease. While traditionally stratified into two conditions, granulomatosis with polyangiitis (GPA) and microscopic polyangiitis (MPA), the subclassification of ANCA-associated vasculitis is subject to continued debate. Here we aim to identify phenotypically distinct subgroups and develop a data-driven subclassification of ANCA-associated vasculitis, using a large real-world dataset. METHODS:In the collaborative data reuse project FAIRVASC (Findable, Accessible, Interoperable, Reusable, Vasculitis), registry records of patients with ANCA-associated vasculitis were retrieved from six European vasculitis registries: the Czech Registry of ANCA-associated vasculitis (Czech Republic), the French Vasculitis Study Group Registry (FVSG; France), the Joint Vasculitis Registry in German-speaking Countries (GeVas; Germany), the Polish Vasculitis Registry (POLVAS; Poland), the Irish Rare Kidney Disease Registry (RKD; Ireland), and the Skåne Vasculitis Cohort (Sweden). We performed model-based clustering of 17 mixed-type clinical variables using a parsimonious mixture of two latent Gaussian variable models. Clinical validation of the optimal cluster solution was made through summary statistics of the clusters' demography, phenotypic and serological characteristics, and outcome. The predictive value of models featuring the cluster affiliations were compared with classifications based on clinical diagnosis and ANCA specificity. People with lived experience were involved throughout the FAIRVASVC project. FINDINGS:A total of 3868 patients diagnosed with ANCA-associated vasculitis between Nov 1, 1966, and March 1, 2023, were included in the study across the six registries (Czech Registry n=371, FVSG n=1780, GeVas n=135, POLVAS n=792, RKD n=439, and Skåne Vasculitis Cohort n=351). There were 2434 (62·9%) patients with GPA and 1434 (37·1%) with MPA. Mean age at diagnosis was 57·2 years (SD 16·4); 2006 (51·9%) of 3867 patients were men and 1861 (48·1%) were women. We identified five clusters, with distinct phenotype, biochemical presentation, and disease outcome. Three clusters were characterised by kidney involvement: one severe kidney cluster (555 [14·3%] of 3868 patients) with high C-reactive protein (CRP) and serum creatinine concentrations, and variable ANCA specificity (SK cluster); one myeloperoxidase (MPO)-ANCA-positive kidney involvement cluster (782 [20·2%]) with limited extrarenal disease (MPO-K cluster); and one proteinase 3 (PR3)-ANCA-positive kidney involvement cluster (683 [17·7%]) with widespread extrarenal disease (PR3-K cluster). Two clusters were characterised by relative absence of kidney involvement: one was a predominantly PR3-ANCA-positive cluster (1202 [31·1%]) with inflammatory multisystem disease (IMS cluster), and one was a cluster (646 [16·7%]) with predominantly ear-nose-throat involvement and low CRP, with mainly younger patients (YR cluster). Compared with models fitted with clinical diagnosis or ANCA status, cluster-assigned models demonstrated improved predictive power with respect to both patient and kidney survival. INTERPRETATION:Our study reinforces the view that ANCA-associated vasculitis is not merely a binary construct. Data-driven subclassification of ANCA-associated vasculitis exhibits higher predictive value than current approaches for key outcomes. FUNDING:European Union's Horizon 2020 research and innovation programme under the European Joint Programme on Rare Diseases.
IntroductionCost-effectiveness modeling often requires extrapolation of survival data from clinical trials over a long-term horizon. The choice of extrapolation method is often uncertain and can have a profound impact on the results. We propose a novel Bayesian approach towards incorporating external information (e.g., registry data or clinical opinion) into the extrapolation process as a means of reducing this uncertainty.Methods Standard parametric survival curves are fitted to immature time-to-event data using maximum likelihood estimation (MLE). Separately, external information on expected cohort-level survival at a future time point is used to specify a prior probability distribution. These are combined to generate posterior distributions of survival curve extrapolations that simultaneously incorporate both observed data and external information. This is done using importance sampling and multivariate normal approximations of the likelihood and posterior distributions; it requires only summary model parameter estimates (and not patient-level data). We apply our method to analyze survival data from the KEYNOTE-426 trial of pembrolizumab+axitinib in advanced renal cell carcinoma.ResultsThe method was implemented in R, and outputs survival curve parameters that are compatible with cost-effectiveness models developed in other software (e.g., Microsoft Excel). In all examples considered, our method resulted in extrapolated survival predictions that were more closely aligned with the external information compared with the standard (MLE-based) approach. Incorporation of external information decreased between-distribution variance (reduced structural uncertainty), and generally also decreased within-distribution variance as well (reduced parameter uncertainty). Results were comparable with those obtained from the method of Cooney and White, which uses similar ideas but requires full patient-level data and is more computationally complex.ConclusionsThe extrapolation method we describe can reduce uncertainty when valid external information is available. Only MLE-based parameter estimates are required to implement our method, thus secondary model users such as HTA agencies can adjust survival extrapolations from existing cost-effectiveness models without access to patient-level data. Implementation is straightforward and computationally efficient, and outputs are easily incorporated into existing cost-effectiveness models.
Introduction: Modelling of relative treatment effects is an important aspect to consider when extrapolating the long-term survival outcomes of treatments. Flexible parametric models offer the ability to accurately model the observed data, however, the extrapolated relative treatment effects and subsequent survival function may lack face validity. Methods: We investigate the ability of change-point survival models to estimate changes in the relative treatment effects, specifically treatment delay, loss of treatment effects and converging hazards. These models are implemented using standard Bayesian statistical software and propagate the uncertainty associate with all model parameters including the change-point location. A simulation study was conducted to assess the predictive performance of these models compared with other parametric survival models. Change-point survival models were applied to three datasets, two of which were used in previous health technology assessments. Results: Change-point survival models typically provided improved extrapolated survival predictions, particularly when the changes in relative treatment effects are large. When applied to the real world examples they provided good fit to the observed data while and in some situations produced more clinically plausible extrapolations than those generated by flexible spline models. Change-point models also provided support to a previously implemented modelling approach which was justified by visual inspection only and not goodness of fit to the observed data. Conclusions: We believe change-point survival models offer the ability to flexibly model observed data while also modelling and investigating clinically plausible scenarios with respect to the relative treatment effects.
Atopic dermatitis (AD) is an inflammatory skin condition with a childhood prevalence of up to 25%. Microbial dysbiosis is characteristic of AD, with Staphylococcus aureus the most frequent pathogen associated with disease flares and increasingly implicated in disease pathogenesis. Therapeutics to mitigate the effects of S. aureus have had limited efficacy and S. aureus-associated temporal disease flares are synonymous with AD. An alternative approach is an anti-S. aureus vaccine, tailored to AD. Experimental vaccines have highlighted the importance of T cells in conferring protective anti-S. aureus responses; however, correlates of T cell immunity against S. aureus in AD have not been identified. We identify a systemic and cutaneous immunological signature associated with S. aureus skin infection (ADS.aureus) in a pediatric AD cohort, using a combined Bayesian multinomial analysis. ADS.aureus was most highly associated with elevated cutaneous chemokines IP10 and TARC, which preferentially direct Th1 and Th2 cells to skin. Systemic CD4+ and CD8+ T cells, except for Th2 cells, were suppressed in ADS.aureus, particularly circulating Th1, memory IL-10+ T cells, and skin-homing memory Th17 cells. Systemic γδ T cell expansion in ADS.aureus was also observed. This study suggests that augmentation of protective T cell subsets is a potential therapeutic strategy in the management of S. aureus in AD.
The 41st volume of the Journal of Classification begins with an article, by Karlsson and Hössjer, that marks an important contribution to Bayesian set-valued classification, and the second contribution is an erratum thereto.In the third paper, Hou-Liu and Browne contribute to the rapidly growing literature on model-based clustering by considering nested Gaussian clusters via orthogonal intrinsic variable subspaces.The fourth paper is an interesting contribution by McLaughlin, Franczak, and Kashlak, who developed a parsimonious family of contaminated shifted asymmetric Laplace distributions for clustering.The fifth paper, by Lim, introduces a new non-parametric polytomous attributes diagnostic classification method that can be used with a variety of cognitive diagnosis models.In the sixth paper of this issue, Chen, Peng, Nie, and Kong contribute to the literature on the search for low-dimensional clustering subspaces.The seventh paper, by Gaye, Ka Diongue, Sylla, Diarra, Diallo, Talla, and Loucoubar, proposes an approach for supervised classification of high-dimensional, correlated, data.This approach is then illustrated via a genomic dataset comprising over 700,000 single nucleotide polymorphisms for 445 individuals.In the penultimate paper, Nam, Mun, Jo, and Kim combine a weighted support vector machine and Gaussian mixture-based oversampling to develop an approach for the supervised classification of infrequent events.In the final article of this issue, de Rooij considers restricted multidimensional unfolding within a multinomial logistic framework for supervised classification.
John Kelleher合作论文数School of Computing,
Dublin Institute of Technology,3