Haemorrhage is the leading preventable cause of trauma death, primarily through ischaemic consequences that current treatments cannot adequately address. We combined human transcriptomic data (n=458) with a controlled porcine model of haemorrhagic shock to identify treatment-responsive molecular mechanisms. Using latent factorisation, we prioritised distinct molecular signatures of the human shock response, including stress signalling, neutrophil activation, and cytotoxic lymphocyte programmes. We assessed the behaviour of these pathways in the experimental porcine system, revealing that shock-initiated immune trajectories are not immutable: blood resuscitation normalised maladaptive transcriptomic changes whilst noradrenaline exacerbated them. While resuscitation modulated neutrophil, heat-shock, and barrier defence programmes, interferon and coagulation pathways were neither mortality-predictive nor treatment-responsive. A factor representing p38-MAPK/AP-1 stress signalling emerged as the dominant mortality-predictive pathway. An in silico small molecule screen identified p38 inhibitors as leading candidates for reversing shock-induced transcriptomic signatures. Our framework identifies modifiable pathways in trauma shock, prioritising p38-MAPK inhibition for therapeutic development and providing a systematic approach for trauma drug repurposing. One sentence summary : Human-porcine cross-analysis reveals treatment-modifiable molecular signatures of trauma shock, identifying potential therapeutic strategies.
Sjögren's disease (SjD) remains a major unmet medical challenge, characterised by biological complexity, patient heterogeneity, and the absence of curative treatments. To advance mechanistic understanding and support therapeutic discovery, we developed a comprehensive Molecular Interaction Map (MIM). Differential expression analyses were conducted on peripheral blood samples from SjD patients and healthy controls across three datasets (GSE51092, UKPSSR, PRECISESADS), identifying 1,625 differentially expressed genes (DEGs), of which 25 were shared across all datasets. Nine common DEGs were linked to interferon signalling, reinforcing its pivotal role in SjD pathogenesis. Pathway enrichment analysis revealed 137 pathways, 43 of which were integrated into the MIM alongside literature-derived knowledge. The resulting SjD Map, freely available at https://sjdmap.elixir-luxembourg.org/ , encompasses 829 molecular entities connected by 598 interactions by transcriptomic data and by curated evidence. This first comprehensive SjD Map provides an integrative framework for visualising pathways, overlaying omics data, and exploring therapeutic opportunities.
BACKGROUND:Type 2 diabetes (T2D) presents a significant global health challenge, and its prevalence is expected to rise. T2D is associated with diverse complications, including cardiovascular diseases (CVD), end-stage renal disease (ESRD), and mental health disorders like depression and anxiety. Previous research has demonstrated sex differences in the metabolic risk profile, clinical characteristics, and care related to T2D and its complications, including mental health conditions. In this study, we aimed to describe the trajectories of hypertension, cardiovascular diseases, ESRD and mental health conditions in women and men from different age groups with T2D. METHODS AND FINDINGS:Using individual-level anonymised linked primary and secondary electronic health record (EHR) data from the United Kingdom (UK) Clinical Practice Research Datalink (CPRD) GOLD, we carried out an observational cohort study of 28,720 individuals with incident T2D between 2010 and 2020. We applied competing risk semi-Markov multi-state analysis to investigate and understand the different transitions to cardiovascular disease, ESRD, and mental health conditions following the onset of T2D in women and men. During a median follow-up of 6.8 years, 39% of the cohort moved to another health state. Men constituted 62% of the cohort, with a median age of 54 (IQR 46, 64) at T2D onset, while women were typically older at diagnosis (mean = 56, IQR = 45, 67). Disease trajectories differed significantly by sex: women were more likely than men to be diagnosed with mental health conditions post-T2D, (8.1% versus 5.3%), while men were more likely to be diagnosed with CVD/ESRD (8.4% versus 5.3%) and hypertension (21.1% versus 20.2%) trajectories than women. Women showed higher average of health service utilisation and long-term prescribing than men. Women experienced a later onset of events compared to men. Both women and men who had a combination of T2D and a mental health condition experienced premature mortality compared to other trajectories. Hypertension was the most frequent event following T2D onset across all age groups and sexes. As a study limitation, some of the observed differences in mental health diagnoses between women and men may reflect differences in healthcare-seeking behaviour rather than true differences in the underlying proportions of these conditions. CONCLUSIONS:Our study highlights sex-specific differences in disease trajectories following T2D onset, emphasising the need for stratified healthcare interventions. Women exhibited greater vulnerability to mental health conditions post-T2D, while men showed higher proportion of cardiovascular and renal complications, and had earlier onset of most events compared to women. Both sexes experienced earlier mortality when T2D was combined with a mental health condition. Current UK medical guidelines lack sex-specific prevention strategies and management for T2D and comorbidities, especially mental health conditions. Our findings draw attention to the importance of tailored management strategies that consider sex, age, and the complex interplay of physical and mental health in T2D care.
ABSTRACT Background Including public contributors in the development of artificial intelligence (AI) systems in healthcare research is growing, however, traditional methods of participation fail to engage people from minoritised groups. This work explores how we can utilise art‐based methods to involve the perspectives of those not previously included in AI development. Methods We collaborated with a East London‐based organisation to involve people not previously included in research to contribute to a study on multiple long‐term conditions (MLTCs) and polypharmacy. Patient and public involvement and engagement (PPIE) contributors all had lived experience of MLTCs and represented a range of different ages, genders, socio‐demographic backgrounds and multilingual abilities. We ran a series of six workshops that used different visual arts methods; ceramics, collage, body mapping and AI‐generated images, to create research priorities and to inform AI development. Findings The arts‐based methods served as a platform for communication which supported PPIE contributors to develop multiple research priorities, for example the impact of the lack of routine appointments on MLTCs. Through these workshops PPIE contributors also highlighted concepts that are important to consider during AI model development, such as utilising local housing data and considering bias. Visual images and art helped to facilitate different forms of communication, whilst being fun and engaging and provided a way to make abstract AI concepts more tangible whilst building AI literacy. Conclusions Arts‐based methods were a useful tool to make involvement in research more accessible for under‐represented communities in the development of AI tools in healthcare research. There is a need for more inclusive participatory approaches as the use of AI in healthcare and research increases. Patient or Public Contribution Working with staff and interpreters from a local community‐based charity, Social Action for Health, we invited 22 PPIE contributors from under‐represented communities in Tower Hamlets who had no previous experience of PPIE research. PPIE contributors developed the research priorities for a large academic consortia and helped create a community art exhibition to highlight their artwork. Additionally, two experienced PPIE contributors from the wider AI‐Multiply study assisted with the preparation of this manuscript.
Type 2 diabetes (T2D) precipitates diabetic cardiomyopathy (dbCM), a condition characterized by chronic inflammation, metabolic dysregulation and impaired cardiac performance. Here we show that the glucokinase activator AZD1656, originally developed for glycemic control but later identified to have immunomodulatory effects, reverses cardiac dysfunction and metabolic remodeling in dbCM. In obese, hyperglycemic db/db mice with diastolic dysfunction, 6 weeks of AZD1656 treatment improved myocardial performance, reduced infarct size and enhanced post-ischaemic recovery. Integrated metabolic, functional and histological analyses revealed restoration of mitochondrial metabolism and attenuation of fibrosis. Mechanistically, AZD1656 remodeled the cardiac immune landscape by promoting infiltration of regulatory T cells. These findings demonstrate a link between cardiac inflammation and metabolic remodeling in dbCM and highlight that modulation of immune cells and metabolism can protect the diabetic heart. Targeting immunometabolic pathways may therefore offer a therapeutic strategy to alleviate cardiac dysfunction and reduce infarct vulnerability in T2D.
Evidence has shown that lipoprotein(a) (Lp[a]) is an independent, causal, genetic risk factor for cardiovascular disease (CVD) that promotes the progression of high-risk, vulnerable atherosclerotic plaque phenotypes. Systems biology integrates multiomics datasets to study linear and nonlinear relationships to enhance understanding of the molecular patterns of disease. One such example is the Genetic Loci and the Burden of Atherosclerotic Lesions (GLOBAL) study, which utilizes multiomics profiling to unravel the molecular signatures of Lp(a)-driven CVD. Using deep phenotyping of coronary atherosclerosis by coronary computed tomography angiography, whole-genome sequencing for genetic analysis, and evaluation of thousands of omics measurements and circulating biomarkers, it is possible to describe the atherogenic milieu associated with Lp(a)-driven CVD. By leveraging the multiomic evaluation of Lp(a)-driven coronary phenotypes, we can begin to translate these findings into real-world strategies for earlier recognition of distinct Lp(a)-driven CVD, which may contribute to improved risk mitigation strategies in clinical practice.
Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across disciplines. We conducted an inductive thematic analysis of 13 semi-structured interviews with participants in early stages of AI-in-healthcare research consortia in the UK. Our findings identify that participants needed to adapt both the tools used for sharing and the information communicated according to their audience, particularly when working with those with a clinical or patient perspective. We identify the novelty of participating in AI research, how AI knowledge is shared, and the inclusion of clinician and patient stakeholder perspectives as key areas within collaborative AI practices in healthcare. These findings highlight that bringing AI into the mix can introduce new obstacles to interdisciplinary work.
Chronic Kidney Disease (CKD) is a global health challenge, affecting 5-10% of the population, with a significant burden on healthcare systems. Early prediction of CKD progression from stage III to stage V is crucial to enable timely interventions. Traditional predictive methods rely on biochemical markers and demographic factors, but are often limited by issues such as missing data and reliance on structured inputs. This study explores the potential of several encoder-based language models, to predict CKD progression using a cohort from the Clinical Practice Research Datalink (CPRD) GOLD database. We applied both Full Fine-Tuning (FFT) and Parameter-Efficient Fine-Tuning (PEFT) with LoRA to pre-trained models, comparing them against traditional machine learning algorithms such as Random Forest and XGBoost. Our results show that fine-tuned models, particularly dmis-lab/biobert-v1.1-FFT, outperform traditional models in predicting CKD progression, with an AUC of 0.7787, precision of 0.7261, and accuracy of 0.7045. Although LoRA-based models are more computationally efficient, they consistenly exhibit lower performance. These findings suggest that fine-tuned encoder models hold significant potential for improving CKD progression prediction. However, there is still room for further enhancement in their accuracy and applicability in clinical settings.
Type 2 Diabetes (T2D) can lead to diabetic cardiomyopathy (dbCM), which is characterised by chronic, systemic inflammation, disrupted metabolism and impaired cardiac function. However, whether cardiac inflammation is present in dbCM and causally linked to metabolic remodelling remains unknown. AZD1656 (AZD), an activator of glucokinase, was postulated to provide glycaemic control in T2D by acting on in the pancreas and liver. However, AZD failed to control hyperglycaemia in clinical trials. Nevertheless, testing of the drug in COVID-19 T2D patients as part of the ARCADIA trial indicated an immunomodulatory effect. Therefore, we used the db/db mouse model of dbCM and an integrated in vivo and ex vivo experimental approach to examine the effects of AZD on cardiac functional and metabolic disturbances and inflammation. 20-week db/db mice displaying the features of human dbCM (obesity, hyperglycaemia and diastolic dysfunction) treated for six weeks with AZD showed improved metabolic remodelling, attenuated diastolic dysfunction, reduced infarct size and improved functional post-ischemic recovery compared to untreated dbCM, alongside an improved cardiac immunophenotype, including reduced T cell-mediated fibrosis and B cell infiltration. Therefore, targeting of immunometabolism may offer a new therapeutic approach to treat cardiac dysfunction and metabolic dysregulation and reduce infarct size in dbCM.
BACKGROUND:The classification of Sjögren's disease partly relies on focus score grading from a minor salivary gland biopsy. Expert regrading of the focus score leads to disease reclassification in half of cases. This study aimed to leverage machine learning to automatically classify the focus score and Sjögren's disease to identify new histological disease subtypes based on minor salivary gland biopsy. METHODS:This retrospective cohort study included minor salivary gland biopsy scanned haematoxylin and eosin slides from six expert centres (three centres in the UK and one each in Greece, Portugal, and France) of the European H2020 NECESSITY consortium. Participants with sicca but without Sjögren's disease and patients with Sjögren's disease and a focus score of either at least 1 or less than 1 where included. All patients with Sjögren's disease fulfilled the American College of Rheumatology-European League Against Rheumatism 2016 criteria. A deep learning model was trained on slides from five centres and validated on slides from the sixth centre. The primary outcome was the area under the receiver operator curve (AUROC) to classify the focus score and Sjögren's disease. Shapley values, an explainable machine learning technology, were computed to identify histological patterns driving the model's classification. People with lived experience of Sjögren's disease were involved in the decision to fund this research and in the dissemination of the findings. FINDINGS:The study was conducted between Oct 13, 2021, and Sept 5, 2024, and included 545 participants with a mean age of 54·2 (SD 13·5); 490 (90%) were female and 55 (10%) were male. After external validation, the model had an AUROC of 0·88 (95% CI 0·82-0·94) for the focus score classification task and an AUROC of 0·89 (0·82-0·94) for Sjögren's disease classification. The performance of Sjögren's disease classification for patients who were negative for anti-Sjögren's syndrome-related antigen A was 0·92 (0·87-1·00). Of histological patterns identified by the model, a new pattern of CD8+ T cells around acinar epithelial cells was associated with Sjögren's disease diagnosis. INTERPRETATION:This study showed that deep learning can reliably classify the focus score and Sjögren's disease using minor salivary gland biopsy exclusively. The study identified that CD8+ T-cell infiltration in acini was associated with Sjögren's disease. Further studies are needed to validate the models. FUNDING:Société Française de Rhumatologie, European Alliance of Associations for Rheumatology.
Patient and Public Involvement and Engagement (PPIE) is critical in the development and application of Artificial Intelligence (AI) in healthcare research to ensure that outcomes align with patients' and the public's needs. However, current PPIE practices often limit involvement to reactive tasks such as reviewing documents and providing plain English summaries. Whilst important, this approach can sideline PPIE from influencing key research decisions. Consequently, PPIE interactions often fail to adequately reach and influence everyday decision makers. On AI and big data research projects, these decisions are often made by Early Career Researchers (ECRs) who play a vital role in the day-to-day research process. After realising these limitations, and to address them, the NIHR-funded AI MULTIPLY consortium introduced twice-monthly "ECRs meet PPIE" sessions. These sessions began in May 2024 and enabled ECRs to present and discuss work in progress and gain targeted input from PPIE members during early phases of research, such as research direction, data and variable selection. By integrating PPIE at this stage, the project aimed to improve the relevance and impact of the healthcare research but also provide ECRs with essential skills in public engagement. At time of writing, 12 sessions have been conducted. Through ethnographic observations integrated with internal surveys, the findings show how the sessions were developed, overcame challenges, and helped to embed PPIE contributors' voices into an AI-in-healthcare project. Based on our findings we have identified 5 recommendations for other large interdisciplinary consortia to strengthen the contribution of PPIE to everyday decision-making in research.
Evidence-based precision medicine strategies do not currently exist to guide the choice of biologics in the treatment of psoriasis. As a result, a costly and arduous trial-and-error approach is often adopted. Artificial intelligence has the potential to improve personalization through the prediction of treatment outcomes using real-world data, such as that within the British Association of Dermatologists Biologics and Immunomodulators Register (BADBIR). We aimed to develop an explainable machine learning (ML) model to predict biologic drug discontinuation in a biologic-naive psoriasis cohort using BADBIR data. BADBIR data (2007–2024) were engineered to enable readability. Adult biologic-naive patients across all biologic cohorts with > 6 months of follow-up data were included. Recruitment centres representing 10% of the overall cohort were randomly separated for external validation (model testing). The residual cohort was then randomly split for model training (80%) and internal validation (20%, for hyperparameter tuning). Random forest modelling was applied for imputation of missing data. Only clinical data at baseline prior to biologic initiation were used for model training to enhance future clinical utilization. The performance of several ML (XG-Boost, AdaBoost, random forest) and deep learning (simple and recurrent neural networks) algorithms was evaluated. External validation was performed with a cross-validation leave-group-out approach of individual recruitment centres. SHAP (SHapley Additive exPlanations) and permutation feature importance values were generated to understand model predictions. In total, 10 806 patients were included, in the cohorts for training (n = 7722), internal validation (n = 1930) and external validation (for final model testing: nine centres, n = 1154). Most patients (n = 7290, 67%) discontinued initial biologic therapy within their follow-up duration (median 6.6 years). Within the discontinuation cohort, adalimumab (originator and biosimilars, 57%) was most prescribed. Higher proportions of female patients (43% vs. 37%) and patients with psoriatic arthritis (21% vs. 17%) and scalp psoriasis (59% vs. 51%) were noted in the discontinuation vs. the continuation cohort, respectively. AdaBoost, an ensemble ML model, outperformed other evaluated models with regards to area under the receiver operating characteristic curve (AUROC). Model testing predicted discontinuation of biologic therapy with (mean, 95% confidence interval) precision 0.85 (0.83–0.88), recall 0.80 (0.78–0.83), F1 score 0.82, AUROC 0.76 (0.71–0.78) and area under the precision recall curve (AUPRC) 0.83 (0.81–0.86). Performance metrics following testing with cross-validation [mean (SD)] were precision 0.79 (0.09), recall 0.69 (0.2), F1 score 0.74 (0.16), AUROC 0.71 (0.06) and AUPRC 0.75 (0.11). The features contributing most significantly to model performance were initial biologic drug, baseline Psoriasis Area and Severity Index, patient age, recruitment centre and baseline white cell count. In conclusion, AdaBoost represents an explainable, ML model with potential clinical utility to predict treatment outcomes of patients with psoriasis using real-world registry data. Future work will investigate discontinuation risk across a range of individual biologic therapies.
The development of biologically interpretable and explainable models remains a key challenge in computational pathology, particularly for multistain immunohistochemistry (IHC) analysis. We present BioX-CPath, an explainable graph neural network architecture for whole slide image (WSI) classification that leverages both spatial and semantic features across multiple stains. At its core, BioX-CPath introduces a novel Stain-Aware Attention Pooling (SAAP) module that generates biologically meaningful, stain-aware patient embeddings. Our approach achieves state-of-the-art performance on both Rheumatoid Arthritis and Sjogren's Disease multistain datasets. Beyond performance metrics, BioX-CPath provides interpretable insights through stain attention scores, entropy measures, and stain interaction scores, that permit measuring model alignment with known pathological mechanisms. This biological grounding, combined with strong classification performance, makes BioX-CPath particularly suitable for clinical applications where interpretability is key. Source code and documentation can be found at: https://github.com/AmayaGS/BioX-CPath.
This study evaluates the generalisation capabilities of state-of-the-art histopathology foundation models on out-of-distribution multi-stain autoimmune Immunohistochemistry datasets. We compare 13 feature extractor models, including ImageNet-pretrained networks, and histopathology foundation models trained on both public and proprietary data, on Rheumatoid Arthritis subtyping and Sjogren's Disease detection tasks. Using a simple Attention-Based Multiple Instance Learning classifier, we assess the transferability of learned representations from cancer H E images to autoimmune IHC images. Contrary to expectations, histopathology-pretrained models did not significantly outperform ImageNet-pretrained models. Furthermore, there was evidence of both autoimmune feature misinterpretation and biased feature importance. Our findings highlight the challenges in transferring knowledge from cancer to autoimmune histopathology and emphasise the need for careful evaluation of AI models across diverse histopathological tasks. The code to run this benchmark is available at https://github.com/AmayaGS/ImmunoHistoBench.
The co-occurrence of multiple long-term conditions (MLTC), or multimorbidity, in an individual can reduce their lifespan and severely impact their quality of life. Exploring the longitudinal patterns, e.g. clusters, of disease accrual can help better understand the genetic and environmental drivers of multimorbidity, and potentially identify individuals who may benefit from early targeted intervention. We introduce probabilistic modelling of onset times, or , for clustering and forecasting MLTC trajectories. seamlessly learns from incomplete and unreliable disease trajectories that is commonplace in Electronic Health Records but often ignored in existing longitudinal clustering methods. We analyse data from 150,000 individuals in the UK Biobank and identify 50 clusters showing patterns of disease accrual that have also been reported by some recent studies. We further discuss the forecasting capabilities of the model given the history of disease accrual.
Understanding the temporal properties of longitudinal data is critical for identifying trends, predicting future events, and making informed decisions in any field where temporal data is analysed, including health and epidemiology, finance, geosciences, and social sciences. Traditional time-series analysis techniques often fail to capture the complexity of irregular temporal patterns present in such data. To address this gap, we introduce bursty_dynamics, a Python package that enables the quantification of bursty dynamics through the calculation of the Burstiness Parameter (BP) and Memory Coefficient (MC). In temporal data, BP and MC provide insights into the irregularity and temporal dependencies within event sequences, shedding light on complex patterns of disease aetiology, human behaviour, or other information diffusion over time. An event train detection method is also implemented to identify clustered events occurring within a specified time interval, allowing for more focused analysis with reduced noise. With built-in visualisation tools, bursty_dynamics provides an accessible yet powerful platform for researchers to explore and interpret the temporal dynamics of longitudinal data. This paper outlines the core functionalities of the package, demonstrates its applications in diverse research domains, and discusses the advantages of using BP, MC, and event train detection for enhanced temporal data analysis.
Study Funding Global Genomics Group, LLC is the sponsor of the GLOBAL clinical study. Editing support was provided by Susannah Thornhill, Ph.D. (BOLDSCIENCE Ltd.), and was funded by Novartis Pharmaceuticals Corporation, East Hanover, NJ, USA, in accordance with GPP 20. Background/Synopsis Lipoprotein(a) [Lp(a)] is the most prevalent genetic cause of atherosclerotic coronary artery disease, and proprotein convertase subtilisin/kexin type 9 (PCSK9) is a genetically and clinically validated target. Little is known about the relationship between Lp(a) isoform size, serum levels of PCSK9, and coronary atherosclerosis. Objective/Purpose We hypothesized that circulating PCSK9 levels are associated with circulating levels of small Lp(a) (≤24 Kringle IV type 2 repeats) and coronary atherosclerosis. Methods The GLOBAL study (NCT01738828) enrolled patients referred for coronary computed tomography angiography (CCTA). Circulating PCSK9 was measured using ELISA (R&D Systems, MN). Circulating Lp(a) (nmol/L) was measured using isoform-independent ELISA. Kringle IV type 2 repeats and the percentage of each of the two Lp(a) isoforms were measured using a validated western blot technique (Northwest Lipid Metabolism and Diabetes Research Laboratories, WA). Circulating levels of small and large Lp(a) particles were calculated based on total Lp(a) and the percent contribution of each isoform. Coronary atherosclerosis was identified and quantified using CCTA in a core laboratory. First, we used linear regression to assess associations between PCSK9 and Lp(a) as continuous variables. Second, we compared serum levels of Lp(a) across PSCK9 quartiles using ANOVA. Third, the prevalence of coronary atherosclerosis was assessed based on serum PCSK9 and Lp(a) quartiles. Results We enrolled 340 patients: 53% female, mean age 55.6±9.8 years. Increasing levels of circulating PCSK9 were associated with increasing levels of Lp(a) when PCSK9 was assessed as a continuous variable (rho=0.11; p=0.038; Figure, Panel A) and when PCSK9 was assessed by quartiles (16±20.61, 24.05±29.5, 23.75±24.76, and 45±59.3 nmol/L, respectively; ANOVA p=0.043; Figure, Panel D). This association was seen for small Lp(a) particles (rho=0.122; p=0.026; Figure, Panels B and E) but was not seen for large Lp(a) particles (rho=0.055; p=0.32; Figure, Panels C and F). In patients in the highest Lp(a) quartile, the prevalence of coronary atherosclerosis across PCSK9 quartiles was 38.9%, 50.0%, 78.9%, and 77.8% (Figure, Panel G). Conclusions Increasing levels of circulating PCSK9 are associated with higher levels of circulating Lp(a), driven by the small Lp(a) particles; this is associated with the increasing prevalence of coronary atherosclerosis. These results are consistent with a mechanistic hypothesis that the PCSK9/low-density lipoprotein receptor axis may be involved in the metabolism of small but not large Lp(a) particles, and small Lp(a) particles are associated with coronary atherosclerosis. In patients with elevated Lp(a), high circulating PCSK9 levels are associated with a very high prevalence of atherosclerosis, representing a high-risk group.
Artificial Intelligence for Multiple Long-term conditions (AIM): A consensus statement from the NIHR AIM consortia Hajira Dambha-Miller1, Andrew Farmer2, Krishnarajah Nirantharakumar3, Thomas Jackson4 Christopher Yau5,6,7, Lauren Walker8 Iain Buchan9, Sarah Finer10, Michael R Barnes10, Nick J Reynolds11, Gyuchan Thomas Jun12, Satheesh Gangadharan13, Simon Fraser 14 and Bruce Guthrie8 1. Primary Care Research Centre, University of Southampton, Southampton, UK. 2. Nuffield Department of Primary Care Health Sciences, University of Oxford, Oxford, UK. 3. Institute of Applied Health Research, University of Birmingham 4. Institute of Inflammation and Ageing, University of Birmingham 5. Nuffield Department of Women's & Reproductive Health, University of Oxford, Oxford, UK 6. Nuffield Department of Population Health, University of Oxford, Oxford, UK 7. Health Data Research, London, UK 8. Institute of Systems, Molecular and Integrative Biology, University of Liverpool, Liverpool, UK 9. Institute of Population Health, University of Liverpool, Liverpool, UK 10. Faculty of Medicine and Dentistry, Queen Mary University of London, London, UK 11. Institute of Translational and Clinical Medicine, Newcastle University Medical School, Newcastle upon Tyne, UK 12. School of Design and Creative Arts, Loughborough University, UK 13. Leicestershire Partnership NHS Trust, UK 14. Department of Public Health, University of Southampton 15. Usher Institute, University of Edinburgh Correspondence: Dr Hajira Dambha-Miller, Primary Care Research Centre, University of Southampton, Southampton, SO16 5ST, Email: H.Dambha-Miller@soton.ac.uk Abstract Recent advances in causal machine learning and wider artificial intelligence (AI) methods could provide new insights into the natural histories and potential prevention of clusters of multiple long-term conditions or multimorbidity (MLTC-M). When combined with expertise in clinical practice, applied health research and social science, there is potential to systematically identify and map new clusters of disease, understand the trajectories of patients with these conditions throughout their life course, predict serious adverse outcomes, optimise therapies and consider the influence of wider determinants such as environmental, behavioural and psychosocial factors. The National Institute of Health Research (NIHR) recently funded multidisciplinary consortia to bring together AI specialists, experts in big data and MLTC-M in the first and second waves of this new programme. The so-called AIM consortia of researchers will spearhead the use of artificial intelligence methods and develop insights for the identification and subsequent prevention of MLTC-M. This consensus agreement is aimed at facilitating a community of learning within the AIM consortia, promoting cooperation, transparency and rigour in our approaches while maintaining high methodological standards and consistency in defining and reporting within our research. In bringing together these research collaborations, there is also an opportunity to foster shared learning, synergies and rapidly compare and validate new AI approaches across our respective studies. This step is critical to implementation on the pathway to patient and public benefit. Scope and aim This statement was developed by the first and second wave of the NIHR AIM consortia and includes representatives across thirteen universities from Edinburgh, Birmingham, Oxford, Southampton, Nottingham, Kent, Manchester, St. Andrews, Liverpool, Newcastle, QMUL, Loughborough and UCL. The multidisciplinary collaborations include front-line primary, secondary and social care staff; researchers in primary, secondary, and social care; health informatics and data science (including AI) experts; epidemiologists; qualitative researchers; statisticians; clinical and health services researchers, geographers; health economists; sociologists, human factors design researchers and public contributors. A summary of each study represented within the AIM consortia and their respective aims is included below. The agreements reached are entirely those of the consortia; sponsors and funders have had no role in the development or reporting of this statement. After initial discussions on the need for such a statement in our individual projects, we met to refine ideas and reach an agreement on item inclusion. We acknowledged variations in aims and purpose of research but found shared interest and overlapping aims with regards to the development of MLTC-M clusters that could then be externally validated across our respective studies. We collectively agreed on the need for a priori consensus on definitions of MLTC-M, clustering variables, outcomes, managing data requests, as well transparency in reporting and methodological approaches to permit meaningful comparison between findings and validation that in turn, could contribute toward more rapid translation to patient benefit. Summary of studies included within wave 1 and 2 of the AIM consortia * The development and validation of population clusters for integrating health and social care: A mixed-methods study on Multiple Long-Term Conditions (AIM-Cluster) + Principal Investigators: Hajira Dambha-Miller and Andrew Farmer + Aim: To develop and validate population clusters that consider health and social care determinants and subsequent need for people with MLTC-M using data-driven AI methods compared to expert-driven approaches, followed by evaluation of cluster trajectories and their association with health outcomes and costs. * + Study Design: AI methods, Delph and Qualitative: Semi-automated based on shallow machine learning clustering using expert crafted features calculated from raw data; Fully-automated based on deep artificial neural networks and explainable AI techniques + Databases: CRPD Aurum and Gold, SAIL, ELSA + Follow up: 10 years + Outcomes: For each cluster; Development of additional MLTC-M, All-cause mortality, Cause-mortality, Worsening frailty, Health and social care costs * OPTIMising therapies, disease trajectories and AI-assisted clinical management for patients Living with complex multimorbidity (OPTIMAL study) + Principal Investigators; Thomas Jackson and Krishnarajah Nirantharakumar + Aim: To integrate multimodal primary care, secondary care, qualitative, and prescribing data to develop an AI tool with the aim of optimising clinical decision making in patients with cMM. Study Design: AI methods and Mixed method + Databases: CPRD Aurum/GOLD, PIONEER, INSIGHT, Scottish data (Linked Tayside and Fife Database and Greater Glasgow and Clyde) + Follow-up: Up to 10 years + Outcomes: Trajectories of distinct phenotypes in clusters and time-points for prevention, Predictive algorithms for future morbidities, Predictive algorithm for choice of medication in the context of multimorbidity, Discovery of medications for repurposing * Artificial Intelligence and Multimorbidity: Clustering in Individuals, Space and Clinical Context (AIM-CISC) + Principal Investigators; Bruce Guthrie + Aim: To use artificial intelligence and state-of-the-art data science, social science and health service research methods to understand clustering of morbidities within individuals, within communities, and in key clinical contexts. + Study Design: AI methods and Mixed methods + Databases: : CPRD Aurum, Lothian DataLoch, UK Biobank, Generation Scotland, Scottish Longitudinal Study, Lothian Data + Follow-up: Up to 10 years + Outcomes: Objective 1: morbidity clusters, Objective 2: morbidity clusters from objective 1, Objective 3: morbidity clusters from objective 1 and a priori multimorbidty groups, Objective 4: serious adverse events of various kinds (eg delirium, falls, increasing dependency, hospital admission, care home admission, death). * DynAIRx: AIs for dynamic prescribing optimisation and care integration in multimorbidity (DynAIRx) + Principal Investigators; Lauren Walker and Iain Buchan + Aim: To extract information from scattered clinical records on how health and medication change over time, apply AI and statistical approaches to predict risk of poor outcomes and develop visualisations that overlay these combined data to enable easier review of medications. We shall pilot this approach in GP prescribing audit and feedback systems, and co-design future medicines optimisation decision support systems with patients and prescribers. + Study Design : AI methods and qualitative engagement + Databases: Structured data from care records across Cheshire & Merseyside, Greater Manchester, and Yorkshire & Humber, covering a population of ~11m using the www.cipha.nhs.uk blueprint that links population health management with care workflow, so can provide data for AI test/train and feedback actionable information to prescribers. + Follow-up: Up to 12 years + Outcomes: Patterns and clusters of conditions, medications, tests and clinical contacts preceding adverse events identified by AI in key high-risk groups (older people with frailty; coexisting physical and mental health problems; complex multimorbidity and problematic polypharmacy). Build the patterns into biostatistical causal inference and prediction of (clustered) clinical outcomes, Advance visualisation of longitudinal summaries of multi-provider care records overlain combined with key features from AI-learned patterns/structures, AI-augmented multimorbidity information built into an existing prescribing audit and feedback system – creating a learning system for medicines optimisation * AI-Multiply - Using artificial intelligence (AI) to characterize the dynamic inter-relationships between MUltiple Long-term condiTIons and PoLYpharmacy and across diverse UK populations and inform health care pathways (AI-Multiply) * + Principal Investigators; Nick Reynolds and Michael Barnes + Aim: To characterise MLTC-M and polypharmacy (MLTC-M-PP) trajectories and define the interrelationships between MLTC-M clusters, polypharmacy, inequalities and healthcare outcomes over the life course + Study Design: AI methods and mixed methods + Databases: + Follow-up: Up to 10 years + Outcomes: Polypharmacy, All-cause mortality, Premature mortality,Healthcare utilisation, MLTC-M inflection point trial emulation Interdisciplinary, theory-informed guidance, A defined and operational PPI strategy influencing study outcomes.Refined pipelines of data engineering, AI and ML methodology. Defined MLTC-M-PP clusters, replicated cross-methods and cross-data sets. External validation of MLTC-M-PP clusters.Defined drivers of MLTC-M-PP clusters including influence of intersectional factors.Theoretically informed guidance for future AI-in-health collaborations, based on interdisciplinary understanding of ideas, assumptions and knowledge creation. New interventions to improve long-term treatment of patients, for testing in pragmatic trials.Identification of a partner to develop MLTC-M-PP clinical decision support tools and identification of potential 'tipping’ points. * Enhanced training of early career researchers and increased critical mass in interdisciplinary MLTC research. Data-driven machine-learning aided stratification and management of multiple long-term conditions in adults with intellectual disabilities (DECODE) * + Principal Investigators; Gyuchan Thomas Jun and Satheesh Gangadharan + Aim: To apply machine learning approaches to identify clusters and trajectories of MLTCs in people with intellectual disabilities and to utilise them to develop actionable insights for effective care coordination to improve the health and wellbeing of people with intellectual disabilities + Study Design: AI methods and mixed methods + Databases: + Follow-up: Up to 20 years + Outcomes: New knowledge on clusters and trajectories of MLTCs and a model of effective care coordination for people with intellectual disabilities * Multidisciplinary Ecosystem to study Lifecourse Determinants and Prevention of Early-onset Burdensome Multimorbidity (MELD-B) * + Principal Investigators; Simon Fraser and Nisreen Alwan + Aim: To safely deliver an Artificial Intelligence (AI)-enhanced epidemiological analytic system in which optimal lifecourse time points and targets for prevention of early-onset, burdensome multimorbidity are identified through multidisciplinary synthesis and analysis of birth cohorts and electronic health records, and disseminated to key stakeholders + Study Design: AI methods and mixed methods + Databases: + Follow-up: Birth to age 65 for the primary analyses + Outcomes: A suite of burdensomeness/complexity indicators. Novel, burdensome MLTC-M clusters, Summary of early-life and wider determinants of burdensome cluster, Summary of critical time points and targets for intervention, Models of potential ‘preventable moments’ of early, burdensome MLTC-M.Co-produced public health implementation recommendations. Defining MLTC-M Researchers acknowledged the substantial variations within the literature in defining MLTC-M and in the conditions that might be included under this terminology.[1] [2]Moreover, each consortium may have a slightly different clinical or public health focus meriting inclusion of a diverse range of conditions [ref]. Acknowledging these challenges, the consortia agreed to conform to the definition of MLTC-M set out by Guthrie et al., (paper in submission) and to include a minimum set of conditions within our data requests based on the 59 core conditions summarised in figure 1 below. Not all these conditions will be relevant to every study and indeed researchers will include additional conditions. However, agreement on the availability of a minimum set of conditions by every consortium will facilitate future external validation across studies. Figure 1: Core conditions to be included within MLTC-M (Guthrie et al., in submission) Body System (based on ICD-10 chapters) Cardiovascular: Always include: Stroke, Coronary Artery Disease, Heart Failure, Peripheral Artery Disease Usually include (unless good reason not to in your context): Heart Valve Disorder, Arrythmia, Venous Thromboembolic Disease,Aneurysm, Hypertension (treated and untreated) Body System (based on ICD-10 chapters) : Metabolic and endocrine Always include: Diabetes,Addison’s Disease,Cystic Fibrosis Usually include (unless good reason not to in your context):) Thyroid Disorders Body System (based on ICD-10 chapters) : Respiratory Always include: Chronic Obstructive Pulmonary Disease,Asthma Usually include (unless good reason not to in your context):) Bronchiectasis Body System (based on ICD-10 chapters) : Neurological Always include: Parkinson’s, Epilepsy,Multiple Sclerosis,Paralysis Usually include (unless good reason not to in your context): Transient Ischaemic Attack,Peripheral Neuropathy,Chronic Primary Pain Body System (based on ICD-10 chapters) : Mental and behavioural Always include: Dementia, Schizophrenia Usually include (unless good reason not to in your context):) Depression, Anxiety, Bipolar Disorder,Drug/Alcohol misuse, Eating disorder, Autism, Post-traumatic stress disorder Body System (based on ICD-10 chapters) : Cancers Always include: Solid Organ Cancer, Haematological Cancers, Metastatic Cancers Usually include (unless good reason not to in your context):) Melanoma, Benign Cerebral Tumours that can cause disability Body System (based on ICD-10 chapters) : Musculoskeletal Always include: Connective Tissue Disease Usually include (unless good reason not to in your context):) Osteoarthritis, Long-term musculoskeltal problems due to injury, Osteoporosis, Gout Body System (based on ICD-10 chapters) : Digestive Always include: Chronic Liver Disease Inflammatory Bowel Disease Usually include (unless good reason not to in your context):) Chronic Pancreatitis, Peptic Ulcers Body System (based on ICD-10 chapters) : Urogenital Always include: Chronic Kidney Disease, End-stage Kidney Disease Usually include (unless good reason not to in your context):) Endometriosis, Chronic urinary tract infection Body System (based on ICD-10 chapters) : Haematological Always include: none Usually include (unless good reason not to in your context):) Anaemia (including pernicious anaemia and sickle cell disease) Body System (based on ICD-10 chapters) : Eye Always include: None Usually include (unless good reason not to in your context):) Vision Impairment that cannot be corrected Body System (based on ICD-10 chapters) : Ear Always include: none Usually include (unless good reason not to in your context):) Hearing Impairment that cannot be corrected, Meniere’s Disease Body System (based on ICD-10 chapters) : Infections Always include: HIV/AIDS Usually include (unless good reason not to in your context):) Chronic Lyme Disease,Tuberculosis,Post-acute COVID-19 Body System (based on ICD-10 chapters) : Congenital Always include: None Usually include (unless good reason not to in your context):) Congenital Disease and Chromosomal abnormalities Protocols, data requests and coding Study protocols and analysis plans will be made widely available, and wherever possible the consortia will include within data requests and approvals, opportunities for unspecified replication for other studies within the NIHR AIM programme of research[3]. This is an essential step to achieving the overall objective of cross-collaboration replication and validation. Code-sets and analytical code should also be made available. We constructed clinical code lists using a rigorous process, which involved reviewing existing code lists (e.g., CALIBER and Cambridge CPRD codes), applying comprehensive search terms to identify codes and finally review by clinicians and where necessary a consensus process to produce the final list of codes. The process is documented for transparency and was developed using the DExtER code builder. Initial code list was generated by the University of Birmingham for the 59 agreed conditions and shared freely across the consortia. Individual groups will include additional codes where appropriate according to their projects. Analytical code will be [4,5]available through the HDR-UK phenotype library. Researchers will also consider The FAIR Guiding Principles for scientific data management and stewardship, [6] and utilise appropriate reporting standards such as STROBE, RECORD or TRIPOD. [4,7]This will facilitate study consistency and rigour, improve transparency and reduce ambiguity in both methods and reporting. Data variables The consortia agreed to include a core minimum list of variables in their data requests. Individual study aims, objectives and methodologies will vary and thus not all variables will be included in all analytical models. Additional variables could be required accordingly. Depending on study aims, variables might be used as exposures, covariates or outcomes and thus we have purposely not specified these here. The inclusion of a minimum set of clustering variables means that different research teams will be able to test their algorithms and validate research findings across datasets. We have agreed on the following core groups of variables but acknowledge that they may be measured, reported and defined differently across datasets and by each consortium. Clear reporting and explanations of variables will be included to permit meaningful interpretation and comparison between measures across consortia: 1. Sociodemographic variables 1. + Age + Sex / Gender + Ethnicity + Deprivation 2. Clinical variables (a clear justification will be provided for their inclusion) 3. Health and social care service utilisation 4. Drugs (specify generic name, class of drugs and purpose) 5. Mortality (specify if cause-specific or all-cause) 6. Frailty (specify measure/s used) AI fairness The consortia acknowledge ethical concerns around bias and ‘unfairness’ that occur in AI systems and the impact this can have on the misuse of predictive models in decision making.[6] We acknowledge ethical frameworks that could be applied in the production and deployment of AI methods. Share learning and conclusions We agree to have regular engagement as a community of practice aiming to share learning, discuss problems and find collective solutions to the wider challenges of research that uses AI and big data such as those around data governance limitations, data sharing for the purposes of external validation or methodological limitations around incomplete records. We recognise that further work in this new field is necessary, specifically around how best to effectively foster research collaboration and sharing across groups and individual databases, that encourage external validation of novel findings, maintain transparency and promote cooperation with the purpose of patient benefit. Competing Interests: CY declares the receipt of remuneration from the Medicines and Healthcare products Regulatory Agency for consultancy work in artificial intelligence. IB is Chief Data Scientist advisor for AstraZeneca. Funding: All the authors of this report have received funding from the National Institute for Health Research (Artificial Intelligence for Multiple Long-Term Conditions (AIM). The views expressed in this publication are those of the author(s) and not necessarily those of the NHS, the National Institute for Health Research or the Department of Health and Social Care. Contributors: All authors contributed equally Transparency declaration: This manuscript is an honest, accurate, and transparent account; no important aspects have been omitted. Ethical approval: Not applicable Data sharing: No other data is available for this consensus statement References: 1. Johnston MC, Crilly M, Black C, et al. Defining and measuring multimorbidity: a systematic review of systematic reviews. European Journal of Public Health 2019;29:182–9. doi:10.1093/EURPUB/CKY098 2. Payne RA, Mendonca SC, Elliott MN, et al. Development and validation of the Cambridge Multimorbidity Score. CMAJ 2020;192:E107–14. doi:10.1503/CMAJ.190757 3. Artificial Intelligence for Multiple Long-Term Conditions (AIM) - Research Specification | NIHR. https://www.nihr.ac.uk/documents/nihr-artificial-intelligence-for-multiple-long-term-conditions-aim-clusters-call-research-specification-finalised/24646 (accessed 24 May 2022). 4. Benchimol EI, Smeeth L, Guttmann A, et al. The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) Statement. PLOS Medicine 2015;12:e1001885. doi:10.1371/JOURNAL.PMED.1001885 5. Collins GS, Reitsma JB, Altman DG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD Statement. BMC Medicine 2015;13:1–10. doi:10.1186/S12916-014-0241-Z/TABLES/1 6. Wilkinson MD, Dumontier M, Aalbersberg IjJ, et al. Comment: The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 2016;3:1–9. doi:10.1038/sdata.2016.18 7. Elm E von, Altman DG, Egger M, et al. Strengthening the reporting of observational studies in epidemiology (STROBE) statement: guidelines for reporting observational studies. BMJ 2007;335:806–8. doi:10.1136/BMJ.39335.541782.AD
Rates of Hospital Readmission (HR), defined as unplanned readmission within 30 days of discharge, have been increasing over the years, and impose an economic burden on healthcare services worldwide. Despite recent research into predicting HR, few models provide sufficient discriminative ability. Three main drawbacks can be identified in the published literature: (i) imbalance in the target classes (readmitted or not), (ii) not including demographic and lifestyle predictors, and (iii) lack of interpretability of the models. In this work, we address these three points by evaluating class balancing techniques, performing a feature selection process including demographic and lifestyle features, and adding interpretability through a combination of SHapley Additive exPlanations (SHAP) and Accumulated Local Effects (ALE) post hoc methods. Our best classifier for this binary outcome achieves a UAC of 0.849 using a selection of 1296 features, extracted from patients’ Electronic Health Records (EHRs) and from their sociodemographics profiles. Using SHAP and ALE, we have established the importance of age, the number of long-term conditions, and the duration of the first admission as top predictors. In addition, we show through an ablation study that demographic and lifestyle features provide even better predictive capabilities than other features, suggesting their relevance toward HR.
Gap junctional communication is required for cellular coordination in health and disease. Gap junction channels are composed of connexins and allow direct intercellular communication. Connexins also play an important role in processes such as wound healing, with connexin 43 (Cx43) being downregulated at the plasma membrane of cells at the wound edge during the initial stages of the normal wound healing process. Connexins must be located on the plasma membrane in order to form gap junctions; therefore, it is important to identify proteins involved in the regulation of this. Previously, in a study of human papillomavirus (HPV)-positive tumour cells, we reported that the gap junction protein Cx43 is a physiologically relevant binding partner of the human homologue of Drosophila Discs large (Dlg1). Dlg1 is a member of the membrane associated-guanylate kinase (MAGUK) scaffolding protein family that is known to control cell shape and polarity, and intracellular trafficking. Here we show, for the first time, that Cx43 interacts with Dlg1 in uninfected keratinocytes in vitro and in normal human epithelial tissues. Depletion of Dlg1 in keratinocytes did not alter Cx43 transcription but was associated with an overall reduction in Cx43 protein levels. The loss of Dlg1 in keratinocytes was further associated with a loss of Cx43 at the plasma membrane and a reduction in intercellular communication. This decrease in Cx43 levels was partially reversed through use of a lysosomal inhibitor; however, a proteasomal inhibitor resulted in no difference to Cx43 levels, suggesting that, in the absence of Dlg1, Cx43 is degraded through the lysosomal but not the proteasomal pathway. This correlated with an increased association between Cx43 and LAMP2, a marker of lysosomes in Dlg1 knockdown cells. In the absence of Dlg1, Cx43 localized to the Golgi but not the endoplasmic reticulum. Although this result could indicate that Dlg1 is required for stable Cx43 membrane location, a study of Cx43 and Dlg1 during the wound healing process in HaCaT cells (a spontaneously immortalized HPV-negative epithelial cell line), showed the proteins co-localizing as they relocated from the plasma membrane to the cytoplasm between 4 and 16 h postwounding before returning to the plasma membrane at 24 h postwounding. This may suggest that Dlg1 is involved in Cx43 trafficking. The levels of both proteins were reduced between 0 and 4 h postwounding and then increased between 4 and 16 h, before returning to normal levels at 24 h postwounding. Taken together, our data suggest that Dlg1 is crucial for promoting plasma membrane localization of Cx43 in keratinocytes.