The integration of artificial intelligence (AI) and machine learning (ML) in healthcare has significantly advanced predictive analytics. Acute Kidney Injury (AKI), a condition marked by sudden kidney function loss, necessitates early intervention to improve patient outcomes. This study introduces a novel approach that leverages knowledge graph embeddings (KGE) to enhance the predictive accuracy of ML models for AKI detection. Knowledge graphs (KGs) model entities and their interrelations in a graph format, integrating heterogeneous data sources to provide a comprehensive view of complex biological systems. The AI contribution lies in the application of embedding techniques that transform these graphs into continuous vector spaces, improving the ability to capture semantic similarities and infer new relationships within the data.On the engineering side, we applied this AI-driven approach to healthcare by leveraging the Medical Information Mart for Intensive Care (MIMIC) III dataset to construct a KG, generate embeddings, and incorporate them into ML models for AKI prediction. This engineering application aims to demonstrate the utility of KGEs in clinical settings. Our approach involved extracting patient features, generating KGE, and training various models, such as those utilizing complex, translational, holographic, and multiplicative embeddings. The models were evaluated through ranking-based and binary classification protocols. The transparency of the AI models enhances their trustworthiness in clinical practice, and the findings underscore the need for continued collaboration with clinicians to refine these techniques and ensure successful deployment in healthcare.
The emergence of the Metaverse as a new platform for human interaction could revolutionize several sectors. This paper provides a new “Metaverse Knowledge Graph” (MateKG) framework to enhance user interaction, data management, and knowledge representation in the virtual environment. The building and interacting MetaKGs is proposed to allow for efficient and successful data management, organization, and exploitation. The novel framework comprises several layers: infrastructure, data gathering, Knowledge Graph (KG), user interaction, evaluation and enhancement, and applications. We discuss several approaches, storage options, ontology designs for KG creation, and data-collecting sources. We look at customization strategies, accessibility concerns for the user interaction layer, and user interface design. Furthermore, we showcase case studies and MetaKG applications in e-commerce, smart cities, healthcare, and entertainment. We provide a starting point for creating reliable and efficient MetaKGs that serve a range of applications and domains.
Knowledge graph embedding (KGE) models are designed for the task of link prediction, which aims to infer missing triples by learning representations for entities and relations. While KGE models excel at ranking-based link prediction, the critical issue of probability calibration has been largely overlooked, resulting in uncalibrated estimates that limit their adoption in high-stakes domains where trustworthy predictions are essential. Addressing this is challenging, as we demonstrate that existing calibration methods are ill-suited to KGEs, often significantly degrading the essential ranking performance they are meant to support. To overcome this, we introduce the KGE Calibrator (KGEC), the first probability calibration method tailored for KGE models to enhance the trustworthiness of their predictions. KGEC integrates three key techniques: a Jump Selection Strategy that improves efficiency by selecting the most informative instances while filtering out less significant ones; Multi-Binning Scaling, which models different confidence levels separately to increase capacity and flexibility; and a Wasserstein distance-based calibration loss that further boosts calibration performance. Extensive experiments across multiple datasets demonstrate that KGEC consistently outperforms existing calibration methods in terms of both effectiveness and efficiency, making it a promising solution for calibration in KGE models.
In this study, we evaluate the effectiveness of foundational artificial intelligence (AI) models, particularly large language models (LLMs), in comparison to traditional machine learning methods for predicting tumor relapse in patients with non-small-cell lung cancer (NSCLC). With a high recurrence risk in NSCLC, early and accurate prediction is essential for improving patient outcomes and guiding treatment decisions. Our analysis utilizes a dataset of 1,348 patients, examining the performance of traditional machine learning models such as Random Forest, alongside cutting-edge LLMs like Mistral-7B, LLaMA-7B, Falcon-7B, and GPT-based models. While the Random Forest model slightly outperforms Mistral-7B in precision-recall for relapse prediction, the comparable results suggest that both approaches offer valuable insights for early relapse detection. This study underscores the potential of integrating classical machine learning with foundational AI models to enhance predictive accuracy in cancer prognosis, providing pathways for more personalized medical interventions.
Retrieval-Augmented Generation (RAG) has gained significant attention from many researchers as an effective solution to address the hallucination issue of Foundational Models (FMs), particularly Large Language Models (LLMs). Although the RAG framework is considered a successful approach for enhancing LLMs by providing a suitable retrieval mechanism to obtain appropriate external knowledge, it still has limitations in acquiring high-quality knowledge from diverse data sources. The complementary integration of RAG and data spaces is proposed to exploit RAG’s capabilities within data spaces. Data spaces provide RAG with the ability to obtain diverse and high quality data sources from several data providers under secure data-sharing mechanisms and direct data exchange negotiations. At the same time, RAG enhances the support services of data spaces. In this paper, we present a high-level architecture for RAG data space models (RAG-DSMs) with a unified lifecycle for RAG and data spaces, highlight the possible challenges of the proposed integration while presenting potential opportunities. Moreover, we present two use cases for leveraging RAG-DSMs in the mobility and health domains.
The lightning development of artificial intelligence (AI) has revolutionized healthcare, helping significant improvements in various applications. This paper provides a comprehensive review of foundation models in healthcare, highlighting their transformative potential in areas such as diagnostics, personalized treatment, and operational efficiency. We argue the key capabilities of these models, including their ability to process diverse data types such as medical images, clinical notes, and structured health records. Regardless their assurance, difficulties remain, including data privacy concerns, bias in AI algorithms, and the need for extensive computational resources. Our analysis identifies emerging trends and future directions, emphasizing the importance of ethical AI deployment, improved interoperability over healthcare systems, and the development of more robust, domain-specific models. Future research should focus on enhancing model interpretability, ensuring equitable access, and fostering collaboration between AI developers and healthcare professionals to maximize the advantages of these technologies.
Retrieval-Augmented Generation (RAG) has gained significant attention from many researchers as an effective solution to address the hallucination issue of Foundational Models (FMs), particularly Large Language Models (LLMs). Although the RAG framework is considered a successful approach for enhancing LLMs by providing a suitable retrieval mechanism to obtain appropriate external knowledge, it still has limitations in acquiring high-quality knowledge from diverse data sources. The complementary integration of RAG and data spaces is proposed to exploit RAG's capabilities within data spaces. Data spaces provide RAG with the ability to obtain diverse and high quality data sources from several data providers under secure data-sharing mechanisms and direct data exchange negotiations. At the same time, RAG enhances the support services of data spaces. In this paper, we present a high-level architecture for RAG data space models (RAG-DSMs) with a unified lifecycle for RAG and data spaces, highlight the possible challenges of the proposed integration while presenting potential opportunities. Moreover, we present two use cases for leveraging RAG-DSMs in the mobility and health domains.
The recent developments in Neurso-Symbolic AI (NESyAI) for healthcare predictive models and explainable techniques have demonstrated immense potential in transforming medical diagnostics and treatment. This chapter dives into the integration of neural networks with symbolic reasoning to enhance both the precision and interpretability of AI systems in healthcare. By harmonising the strength of these approaches, NeSY AI can analyse complex medical data, make accurate diagnoses, and develop effective treatment plans personalised to patients. Applications such as drug repurposing, lung cancer tumour recurrence prediction, and the classification of chronic kidney disease caused from clinical notes exemplify the practical benefits of this hybrid AI approach. Detailed case studies and the methodologies illustrate how NeSY can create more accurate, interpretable and trustworthy examination systems. This dual potential not only boosts patient outcomes but also encourages clinician acceptance, accelerating the integration of AI into mainstream medical fields. Because of these advancements, NeSY holds the promise of significantly enhancing healthcare delivery and patient care.
Recently, there has been more interest in Decentralized Data-Sharing (DDS) because of the introduction of Dataspace 4.0. DDS is becoming increasingly popular as a safe, open, and effective way for many parties to data-sharing. Unlike conventional, centralized methods, DDS has several benefits, such as better knowledge exchange, higher accessibility and interoperability, and data privacy and security. The paper covers DDS’s advantages, including improved resilience, higher security, increased privacy, and improved interoperability. DDS gives people and organizations more control and ownership over their data while reducing the dangers of centralized data management. In this survey, we highlight promising technologies for DDS in Dataspace 4.0, including Federated Learning (FL), blockchain, decentralized file systems, semantic web and knowledge representation, and Peer-to-Peer (P2P) networks. We highlight the challenges, opportunities, and future directions of technology enabling further DDS advancement in Industry 4.0.
Motivation: Low-stage lung cancer is known to recur unpredictably, and patients receiving various treatment methods like radiation, chemotherapy, and immunotherapies have been seen to respond very differently. Identifying a priori if a patient is going to relapse or not could make a difference in terms of saving lives and personalized care offered. In this work, we provide an answer to the following research question: Is it possible to enhance the machine learning (ML) of the estimated probability of relapse in early-stage non-small-cell lung cancer (NSCLC) patients with aneuploidy imputation scores?Results: To predict recurrence in 1,348 early-stage (I-II) NSCLC patients, we train graph ML models utilizing the Spanish pulmonary cancer group knowledge graph enriched with triples from pathway imputation. ML models trained on Knowledge graph data enriched with triples from pathway score imputation present an 82% Precision and 91% Specificity in predicting relapse over 200 patients from a held-out test set. ML models trained using graphs data could prove useful supplemental tool in the TNM classification systems and improve a lung cancer patient's prognosis.
Chronic kidney disease affects over 800 million people globally and has many etiologies, including hypertension, diabetes, and IgA nephropathy. Effective management relies on precise medical documentation to monitor disease progression and treatment efficacy. Artificial intelligence, specifically BERT-based models, has shown promise in enhancing disease classification from clinical notes. In Ireland, clinical notes from hemodialysis and nephrology outpatients are stored in the Kidney Disease Clinical Patient Management System (KDCPMS). This study uses DistilBERT with 2,947 unstructured clinical notes from the KDCPMS to classify individual patient notes as stating that the patient has or does not have IgA Nephropathy, Diabetes, and Hypertension. The training pipeline involved data extraction, anonymization, pre-processing, annotation, and model training. On the test set the IgA Nephropathy classifier achieved high precision (99%) and recall (100%), effectively identifying positive cases with minimal errors. The Diabetes classifier had precision of 76% and recall of 90%. The Hypertension classifier had high precision (91%) and medium recall (72%). The performance of the classifiers was further evaluated against a clinical validation dataset of 112 patients created through a full chart review by clinicians. The IgA, Diabetes and Hypertension classifiers had precision of 83%, 61% and 73% respectively and recall of 83%, 73% and 83% respectively. The results highlight the models' performance on the different classification problems in both the test set and a more clinically focused dataset. The results emphasize the importance of ongoing refinement and expert collaboration to enhance clinical relevance and applicability.
The recurrence of low-stage lung cancer poses a challenge due to its unpredictable nature and diverse patient responses to treatments. Personalized care and patient outcomes heavily rely on early relapse identification, yet current predictive models, despite their potential, lack comprehensive genetic data. This inadequacy fuels our research focus—integrating specific genetic information, such as pathway scores, into clinical data. Our aim is to refine machine learning models for more precise relapse prediction in early-stage non-small cell lung cancer. To address the scarcity of genetic data, we employ imputation techniques, leveraging publicly available datasets such as The Cancer Genome Atlas (TCGA), integrating pathway scores into our patient cohort from the Cancer Long Survivor Artificial Intelligence Follow-up (CLARIFY) project. Through the integration of imputed pathway scores from the TCGA dataset with clinical data, our approach achieves notable strides in predicting relapse among a held-out test set of 200 patients. By training machine learning models on enriched knowledge graph data, inclusive of triples derived from pathway score imputation, we achieve a promising precision of 82% and specificity of 91%. These outcomes highlight the potential of our models as supplementary tools within tumour, node, and metastasis (TNM) classification systems, offering improved prognostic capabilities for lung cancer patients. In summary, our research underscores the significance of refining machine learning models for relapse prediction in early-stage non-small cell lung cancer. Our approach, centered on imputing pathway scores and integrating them with clinical data, not only enhances predictive performance but also demonstrates the promising role of machine learning in anticipating relapse and ultimately elevating patient outcomes.
Almost every community Question-Answering (cQA) platform has the pressing need of enhancing user experience by presenting dedicated displays, connecting potential answerers with open questions and revitalizing the material in their archives. In doing so, it is crucial to understand the profile of their community members, especially as it relates to their demographics. In this realm, variables such as age and gender have shown to be particularly promising for managing content. For instance, they make it easier to connect questions posted by one generation that are more likely to be answered by individuals from the previous generation. This paper advances the current body of knowledge in this area by exploring the performance of nineteen frontier transformer-based models (e.g., BERT and ELECTRA) on age recognition across a large-scale collection of cQA members. In effect, the best encoder (LongFormer) finished with an accuracy of 78.61% (F1-Score of 0.7424) by taking full-questions and answers into account. Unlike gender recognition, our outcomes do not show a noticeable difference between cased and uncased models. But on the other hand, they confirm that the transition from one age group to the other is smooth, and thus boundary individuals pose a tough challenge to discriminant models built on top of frontier machine learning approaches.
The study addresses customer churn, a major issue in service-oriented sectors like telecommunications, where it refers to the discontinuation of subscriptions. The research emphasizes the importance of recognizing customer satisfaction for retaining clients, focusing specifically on early churn prediction as a key strategy. Previous approaches mainly used generalized classification techniques for churn prediction but often neglected the aspect of interpretability, vital for decision-making. This study introduces explainer models to address this gap, providing both local and global explanations of churn predictions. Various classification models, including the standout Gradient Boosting Machine (GBM), were used alongside visualization techniques like Shapley Additive Explanations plots and scatter plots for enhanced interpretability. The GBM model demonstrated superior performance with an 81% accuracy rate. A Wilcoxon signed rank test confirmed GBM’s effectiveness over other models, with the p-value indicating significant performance differences. The study concludes that GBM is notably better for churn prediction, and the employed visualization techniques effectively elucidate key churn factors in the telecommunications sector.
When the global pandemic struck in 2020, most countries established task forces to meet a challenge that impacted governmental resources. It became apparent that data, intelligence gathering, and both modelling and predictive capabilities were required. While artificial intelligence (AI) based solutions had already begun to emerge within the public sector, the Covid-19 pandemic accelerated this process. In particular, modelling of case numbers with the development of predictive algorithms. The development of AI solutions for public sector organizations is inherently multidisciplinary. This is crucial to understanding how solutions can be developed, outputs understood, and the benefits and risks measured. Furthermore, the development of AI solutions often requires data which may not be accessible from a single location. In the case of Covid-19 modelling, data must be extracted from multiple locations to construct data assets. In this research, a collaborative approach to developing machine learning expertise for the public sector is presented. Using Covid-19 as a case study, the role of different government sectors when building data assets is examined along with the use of standard data models, and how this type of cooperation led to the development of a pipeline for data assets to underpin AI solutions for the public sector.
Lung cancer is one of the leading health complications causing high mortality worldwide. The relapsing behavior of medically treated early-stage lung cancer makes this disease even more complicated. Thus predicting such relapse using a data-centric approach provides a complementary perspective for clinicians to understand the disease. In this preliminary work, we explored off-the-shelf survival models to predict the relapse of early-stage lung cancer patients. We analyzed the survival models on a cohort of 1348 early-stage non-small cell lung cancer (NSCLC) patients in different timestamps. Using the prediction explanation model SHAP (SHapley Additive exPlanations), we further explained the best-performing survival model's predictions. Our explainable predictive model is a potential tool for oncologists that address an unmet clinical need for post-treatment patient stratification based on the relapse hazard.
For online social networks, demographic analysis is absolutely essential for improving their services in many ways. It is instrumental in understanding their different audiences, members and competitors. As well as that, it is pivotal in designing effective personalization and contextualization strategies, especially for displaying and creating better content. There is, for this reason, a great bulk of research into how demographic variables are characterized and how they impact online platforms such as Facebook and Twitter. But surprisingly, only a handful of works delve into their characterization and effect on community Question-Answering (cQA) websites. In this particular context, the subject of age demographics remains largely unexplored. This paper takes the lead on interpreting automatic age recognition on CQAs (a.k.a. age screening) as a regression task. To this effect, it compares state-of-the-art graph-based neural network regression and embedding models on a massive activity-graph encompassing ca. 16 and 837 million nodes (members) and edges, respectively. For this study, a large-scale subset of ca. 657,000 community fellows was automatically associated with their age via aligning their profile texts with a limited number of linguistic patterns. In short, our results show that Node2vec significantly outperforms other embeddings regardless of the regression model used for casting predictions. When this embedding is combined with Artificial Neural Network Regressions, we obtained our best configuration scoring a Root Mean Square Error (RMSE) of 8.39. An interesting qualitative feature of this embedding space is that age-based centroid vectors tend to form a trail ordered by age. Lastly, our outcomes also signal that activity graph based models can rival its counterparts based on image and textual inputs, paving the way for constructing effective multi-modal approaches.
PURPOSE:Stratifying patients with cancer according to risk of relapse can personalize their care. In this work, we provide an answer to the following research question: How to use machine learning to estimate probability of relapse in patients with early-stage non-small-cell lung cancer (NSCLC)? MATERIALS AND METHODS:For predicting relapse in 1,387 patients with early-stage (I-II) NSCLC from the Spanish Lung Cancer Group data (average age 65.7 years, female 24.8%, male 75.2%), we train tabular and graph machine learning models. We generate automatic explanations for the predictions of such models. For models trained on tabular data, we adopt SHapley Additive exPlanations local explanations to gauge how each patient feature contributes to the predicted outcome. We explain graph machine learning predictions with an example-based method that highlights influential past patients. RESULTS:Machine learning models trained on tabular data exhibit a 76% accuracy for the random forest model at predicting relapse evaluated with a 10-fold cross-validation (the model was trained 10 times with different independent sets of patients in test, train, and validation sets, and the reported metrics are averaged over these 10 test sets). Graph machine learning reaches 68% accuracy over a held-out test set of 200 patients, calibrated on a held-out set of 100 patients. CONCLUSION:Our results show that machine learning models trained on tabular and graph data can enable objective, personalized, and reproducible prediction of relapse and, therefore, disease outcome in patients with early-stage NSCLC. With further prospective and multisite validation, and additional radiological and molecular data, this prognostic model could potentially serve as a predictive decision support tool for deciding the use of adjuvant treatments in early-stage lung cancer.
Amidst prevailing healthcare challenges, a dynamic solution emerges, fusing knowledge graph technology, clinical trials optimization, dataspace integration, and AI innovation.This unified approach tackles issues like limited patient insights, suboptimal trial designs, and imprecise treatments.By interlinking diverse data through knowledge graphs, this method illuminates disease trends, therapeutic efficacies, and patient prognoses.AI techniques, especially machine learning, contribute predictive power by unveiling hidden patterns for accurate diagnostics, prognostics, and personalized treatments.This multidisciplinary fusion transforms clinical trials, enhancing comprehensiveness and precision through real-world data analysis and subgroup identification.In reshaping healthcare, this proposition aims to accelerate treatment personalization, elevate therapeutic efficacy, and empower informed medical decisions, encompassing the essence of 'Advancing Healthcare through Innovation: Knowledge Graphs, Clinical Trials, Dataspace, and AI'.
Vít Novácek合作论文数National University of Ireland, Galway9
Mathieu D’Aquin合作论文数Knowledge Media institute (KMi) of the Open University in Milton Keynes, UK3