Fast Healthcare Interoperability Resources (FHIR), developed by Health Level Seven International (HL7), has emerged as the leading healthcare data standard to address persistent barriers in interoperability, fragmented exchange, and inconsistent data harmonization. As health systems worldwide undergo digital transformation, FHIR offers a flexible framework for integrating electronic health records, analytics platforms, and decision-support tools. Its growth has been accelerated by policy mandates such as the 21st Century Cures Act, as well as the availability of application programming interfaces (APIs), software development kits (SDKs), and web standards. Globally, FHIR has been adopted or piloted by national health systems in the United States, United Kingdom, Canada, and Australia, and incorporated into World Health Organization data initiatives, underscoring its role in global digital health strategy. Documented outcomes of this review include comprehensive mapping of FHIR applications across clinical, research, and public health domains; identification of adoption barriers and enablers; insights into integration with generative AI and large language models for predictive modeling, automated documentation, and decision support; and guidance for future innovations such as blockchain-enabled infrastructure and cloud-native scalability. Nonetheless, challenges remain, including uneven implementation, workforce training gaps, scalability limitations, and unresolved concerns around privacy, security, and regulatory compliance. This synthesis provides actionable insights for providers, researchers, policymakers, and developers to advance global health interoperability.
Haematoxylin and Eosin (H&E) staining and ImmunoFluo-rescence (IF) imaging are usually performed for tissue analysis. H&E staining is preferred in some downstream tasks but are time-consuming and can degrade tissue during washing. We propose a deep learning approach using a U-Net to generate synthetic H&E images from their IF markers counterpart. Since we consider this a deterministic mapping problem, we compare against the traditional GAN approach in literature, finding comparable quality yet reduced training costs. Using a colorectal cancer tissue dataset with 50+ markers from CODEX technology, we conducted an ablation study showing that DRAQ5 and Hoechst strongly influence synthetic H&E quality. Comparing subsets of markers, we assessed synthetic H&E images in downstream tasks like tissue segmentation, demonstrating the model's applicability and enabling explain-ability by quantifying marker relationships to H&E staining.
BackgroundAdolescent idiopathic scoliosis (AIS) is the most common type of scoliosis, affecting 1-4% of adolescents. The Scoliosis Research Society-22R (SRS-22R), a health-related quality-of-life instrument for AIS, has allowed orthopedists to measure subjective patient outcomes before and after corrective surgery beyond objective radiographic measurements. However, research has revealed that there is no significant correlation between the correction rate in major radiographic parameters and improvements in patient-reported outcomes (PROs), making it difficult to incorporate PROs into personalized surgical planning.MethodsThe objective of this study is to develop an artificial intelligence (AI)-enabled surgical planning and counseling support system for post-operative patient rehabilitation outcomes prediction in order to facilitate personalized AIS patient care. A unique multi-site cohort of 455 pediatric patients undergoing spinal fusion surgery at two Shriners Children's hospitals from 2010 is investigated in our analysis. In total, 171 pre-operative clinical features are used to train six machine-learning models for post-operative outcomes prediction. We further employ explainability analysis to quantify the contribution of pre-operative radiographic and questionnaire parameters in predicting patient surgical outcomes. Moreover, we enable responsible AI by calibrating model confidence for human intervention and mitigating health disparities for algorithm fairness.ResultsThe best prediction model achieves an area under receiver operating curve (AUROC) performance of 0.86, 0.85, and 0.83 for individual SRS-22R question response prediction over three-time horizons from pre-operation to 6-month, 1-year, and 2-year post-operation, respectively. Additionally, we demonstrate the efficacy of our proposed prediction method to predict other patient rehabilitation outcomes based on minimal clinically important differences (MCID) and correction rates across all three-time horizons.ConclusionsBased on the relationship analysis, we suggest additional attention to sagittal parameters (e.g., lordosis, sagittal vertical axis) and patient self-image beyond major Cobb angles to improve surgical decision-making for AIS patients. In the age of personalized medicine, the proposed responsible AI-enabled clinical decision-support system may facilitate pre-operative counseling and shared decision-making within real-world clinical settings.
Polygenic risk scores (PRSs) hold promise in their potential translation into clinical settings to improve disease risk prediction. An important consideration in integrating PRSs into clinical settings is to gain an understanding of how to identify which subpopulations of individuals most benefit from PRSs for risk prediction. In this study, using the UK Biobank dataset, we trained logistic regression models to predict the 10 year incident risk of myocardial infarction, breast cancer, and schizophrenia using either just clinical features or clinical features combined with PRSs. For each disease, we identified the top 10% subgroup with the greatest magnitude of improvement in risk prediction accuracy attributed to PRSs in the multi-modal model. Using up to similar to 3.6 k demographic, lifestyle, diagnostic, lab, and physical measurement features from the UK Biobank dataset of similar to 500 k individuals, we characterized these subgroups based on various clinical, lifestyle, and demographic characteristics. The incident cases in the top 10% subgroup for each disease represent distinct phenotypes that differ from other cases and that are strongly correlated with genetic predisposition. Our findings provide insights into disease subtypes and can encourage future studies aimed at classifying these individuals to enhance the targeting of polygenic risk scoring in practice.
Objective: To develop a clinical decision support tool that can predict cardiovascular disease (CVD) risk with high accuracy while requiring minimal clinical feature input, thus reducing the time and effort required by clinicians to manually enter data prior to obtaining patient risk assessment. Results: In this study, we propose a robust feature selection approach that identifies five key features strongly associated with CVD risk, which have been found to be consistent across various models. The machine learning model developed using this optimized feature set achieved state-of-the-art results, with an AUROC of 91.30%, sensitivity of 89.01%, and specificity of 85.39%. Furthermore, the insights obtained from explainable artificial intelligence techniques enable medical practitioners to offer personalized interventions by prioritizing patient-specific high-risk factors. Conclusion: Our work illustrates a robust approach to patient risk prediction which minimizes clinical feature requirements while also generating patient-specific insights to facilitate shared decision-making between clinicians and patients.
With the recent advancement of novel biomedical technologies such as high-throughput sequencing and wearable devices, multi-modal biomedical data ranging from multi-omics molecular data to real-time continuous bio-signals are generated at an unprecedented speed and scale every day. For the first time, these multi-modal biomedical data are able to make precision medicine close to a reality. However, due to data volume and the complexity, making good use of these multi-modal biomedical data requires major effort. Researchers and clinicians are actively developing artificial intelligence (AI) approaches for data-driven knowledge discovery and causal inference using a variety of biomedical data modalities. These AI-based approaches have demonstrated promising results in various biomedical and healthcare applications. In this review paper, we summarize the state-of-the-art AI models for integrating multi-omics data and electronic health records (EHRs) for precision medicine. We discuss the challenges and opportunities in integrating multi-omics data with EHRs and future directions. We hope this review can inspire future research and developing in integrating multi-omics data with EHRs for precision medicine.
Despite the myriad peer-reviewed papers demonstrating novel Artificial Intelligence (AI)-based solutions to COVID-19 challenges during the pandemic, few have made a significant clinical impact, especially in diagnosis and disease precision staging. One major cause for such low impact is the lack of model transparency, significantly limiting the AI adoption in real clinical practice. To solve this problem, AI models need to be explained to users. Thus, we have conducted a comprehensive study of Explainable Artificial Intelligence (XAI) using PRISMA technology. Our findings suggest that XAI can improve model performance, instill trust in the users, and assist users in decision-making. In this systematic review, we introduce common XAI techniques and their utility with specific examples of their application. We discuss the evaluation of XAI results because it is an important step for maximizing the value of AI-based clinical decision support systems. Additionally, we present the traditional, modern, and advanced XAI models to demonstrate the evolution of novel techniques. Finally, we provide a best practice guideline that developers can refer to during the model experimentation. We also offer potential solutions with specific examples for common challenges in AI model experimentation. This comprehensive review, hopefully, can promote AI adoption in biomedicine and healthcare.
Expert microscopic analysis of cells obtained from frequent heart biopsies is vital for early detection of pediatric heart transplant rejection to prevent heart failure. Detection of this rare condition is prone to low levels of expert agreement due to the difficulty of identifying subtle rejection signs within biopsy samples. The rarity of pediatric heart transplant rejection also means that very few gold-standard images are available for developing machine learning models. To solve this urgent clinical challenge, we developed a deep learning model to automatically quantify rejection risk within digital images of biopsied tissue using an explainable synthetic data augmentation approach. We developed this explainable AI framework to illustrate how our progressive and inspirational generative adversarial network models distinguish between normal tissue images and those containing cellular rejection signs. To quantify biopsy-level rejection risk, we first detect local rejection features using a binary image classifier trained with expert-annotated and synthetic examples. We converted these local predictions into a biopsy-wide rejection score via an interpretable histogram-based approach. Our model significantly improves upon prior works with the same dataset with an area under the receiver operating curve (AUROC) of 98.84% for the local rejection detection task and 95.56% for the biopsy-rejection prediction task. A biopsy-level sensitivity of 83.33% makes our approach suitable for early screening of biopsies to prioritize expert analysis. Our framework provides a solution to rare medical imaging challenges currently limited by small datasets.
Personalized medicine plays an important role in treatment optimization for COVID-19 patient management. Early treatment in patients at high risk of severe complications is vital to prevent death and ventilator use. Predicting COVID-19 clinical outcomes using machine learning may provide a fast and data-driven solution for optimizing patient care by estimating the need for early treatment. In addition, it is essential to accurately predict risk across demographic groups, particularly those underrepresented in existing models. Unfortunately, there is a lack of studies demonstrating the equitable performance of machine learning models across patient demographics. To overcome this existing limitation, we generate a robust machine learning model to predict patient-specific risk of death or ventilator use in COVID-19 positive patients using features available at the time of diagnosis. We establish the value of our solution across patient demographics, including gender and race. In addition, we improve clinical trust in our automated predictions by generating interpretable patient clustering, patient-level clinical feature importance, and global clinical feature importance within our large real-world COVID-19 positive patient dataset. We achieved 89.38% area under receiver operating curve (AUROC) performance for severe outcomes prediction and our robust feature ranking approach identified the presence of dementia as a key indicator for worse patient outcomes. We also demonstrated that our deep-learning clustering approach outperforms traditional clustering in separating patients by severity of outcome based on mutual information performance. Finally, we developed an application for automated and fair patient risk assessment with minimal manual data entry using existing data exchange standards.
Recent advances in artificial intelligence (AI) have sparked interest in developing explainable AI (XAI) methods for clinical decision support systems, especially in translational research. Although using XAI methods may enhance trust in black-box models, evaluating their effectiveness has been challenging, primarily due to the absence of human (expert) intervention, additional annotations, and automated strategies. In order to conduct a thorough assessment, we propose a patch perturbation-based approach to automatically evaluate the quality of explanations in medical imaging analysis. To eliminate the need for human efforts in conventional evaluation methods, our approach executes poisoning attacks during model retraining by generating both static and dynamic triggers. We then propose a comprehensive set of evaluation metrics during the model inference stage to facilitate the evaluation from multiple perspectives, covering a wide range of correctness, completeness, consistency, and complexity. In addition, we include an extensive case study to showcase the proposed evaluation strategy by applying widely-used XAI methods on COVID-19 X-ray imaging classification tasks, as well as a thorough review of existing XAI methods in medical imaging analysis with evaluation availability. The proposed patch perturbation-based workflow offers model developers an automated and generalizable evaluation strategy to identify potential pitfalls and optimize their proposed explainable solutions, while also aiding end-users in comparing and selecting appropriate XAI methods that meet specific clinical needs in real-world clinical research and practice.
In pediatric heart transplantation, manual annotations with interob-server and intraobserver variability among cardiovascular pathology experts lead to significant disagreements about the severity of rejection. Artificial intelligence (AI)-enabled computational pathology usually requires large-scale manual annotations of gigapixel whole-slide images (WSIs) for effective model training. To address these challenges, we develop and validate an AI-enabled rare disease detection framework for automating heart transplant rejection detection from whole-slide images of pediatric patients. Specifically, we conduct a novel dataset cartography with data maps and training dynamics to map and diagnose the augmented samples, exploring the model behavior on individual instances during model training. Extensive experiments on internal and external patient cohorts have demonstrated the feasibility of both tile-level and biopsy-level detection. The proposed data-efficient learning framework may support seamless scalability to real-world rare disease detection without the burden of iterative expert annotations.
World-renowned pediatric patient care in scoliosis, craniofacial, orthopedic, and other life-altering conditions is provided at the international Shriners Children's hospital system. The impact of scoliosis can be extreme with significant curvature of the spine that often progresses during childhood periods of growth and development. Gauging the impact of treatment is vital throughout the diagnostic and treatment process and is achieved using radiographic imaging and patient reported feedback surveys. Surgeons from multiple clinical centers have amassed a wealth of patient data from more than 1,000 scoliosis patients. However, these data are difficult to access due to data heterogeneity and poor interoperability between complex hospital systems. These barriers significantly decrease the value of these data to improve patient care. To solve these challenges, we create a generalizable multi-site and multi-modality cloud infrastructure for managing the clinical data of multiple diseases. First, we establish a standardized and secure research data repository using the Fast Health Interoperability Resources (FHIR) standard to harmonize multi-modal clinical data from different hospital sites. Additionally, we develop a SMART-on-FHIR application with a user-friendly graphical user interface (GUI) to enable non-technical users to access the harmonized clinical data. We demonstrate the generalizability of our solution by expanding it to also facilitate craniofacial microsomia and pediatric bone disease imaging research. Ultimately, we present a generalized framework for multi-site, multimodal data harmonization, which can efficiently organize and store data for clinical research to improve pediatric patient care.
According to the World Health Organization, Artificial Intelligence (AI) technology may assist in COVID-19 management. However, existing image segmentation using AI suffers from a lack of accuracy and explainability, which prevents its adoption in actual clinical practice. In this paper, we investigated an attention-based image segmentation method for COVID-19 CT imaging with enhanced interpretation capabilities. Specifically, we developed U-Net architecture-based for segmentation with attention coefficients to produce a salient feature map. We use the DICE score and accuracy to perform a comprehensive model evaluation. We compared to other well-known methods such as Light U-Net, COPLE-Net, and Res U-Net and demonstrated that attention U-Net is superior for COVID-19 segmentation tasks in terms of performance and explainability. We also developed the tool as a web-application with a graphic user interface with the goal to translate this AI-driven clinical decision-support system for real-world clinical use.
As of May 15th, 2022, the novel coronavirus SARS-COV-2 has infected 517 million people and resulted in more than 6.2 million deaths around the world. About 40% to 87% of patients suffer from persistent symptoms weeks or months after their original infection. Despite remarkable progress in preventing and treating acute COVID-19 conditions, the clinical diagnosis of long-term COVID remains difficult. In this work, we use free-text clinical notes and natural language processing (NLP) techniques to explore long-term COVID effects. We first obtain free-text clinical notes from 719 outpatient encounters representing patients treated by physicians at Emory Clinic to detect patterns in patients with long-term COVID symptoms. We apply state-of-the-art NLP frameworks to automatically identify patients with long-term COVID effects, achieving 0.881 recall (sensitivity) score for note-level prediction. We further interpret the prediction outcomes and discuss potential phenotypes. Our work aims to provide a data-driven solution to identify patients who have developed persistent symptoms after acute COVID infection. With this work, clinicians may be able to identify patients who have long-term COVID symptoms to optimize treatment.
Digital Twins (DTs) and Virtual Reality (VR) have recently gained significant popularity. The term digital twin refers to a virtual representation of a real-world environment. DTs are known for providing intuitive, accurate, and real-time visualizations of complex physical systems. In this paper, we present a novel, cloud-based framework for creating digital twin environments using VR within a cloud environment. We demonstrate our framework with two real-world case studies inspired by clinical insights obtained from experts in post-stroke rehabilitation. In the case of rehabilitation, sessions within virtual scenarios mimicking patients’ everyday life may significantly improve their rehabilitation quality. Inspired by biofeedback, we used a digital twin in VR to allow patients to enhance hand stability using real-time data streaming. Our second case study demonstrated the utility of our approach for allowing patients to more comfortably navigate within their homes by generating a virtual twin of an apartment. Our work is designed to transcend the boundary between reality and virtuality by letting users interact directly with the real and virtual worlds simultaneously. To achieve real-time synchronization between the real and the virtual environments, we established an infrastructure blueprint that features efficient sensor data collection, streaming, processing, and storage in a secure cloud environment. Data may then be accessed from any site, directly from the cloud, to update the virtual environment. In the end, we show that our infrastructure is generalizable across clinically-relevant tasks. We hope to extend our work to other areas that would benefit from a digital twin-based VR space.
Shriners Children's (SHC) is a hospital system whose mission is to advance the treatment and research of pediatric diseases. SHC success has generated a wealth of clinical data. Unfortunately, barriers to healthcare data access often limit data-driven clinical research. We decreased this burden by allowing access to clinical data via the standardized data access standard called FHIR (Fast Healthcare Interoperability Resources). Specifically, we converted existing data in the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) standard into FHIR data elements using a technology called OMOP-on-FHIR. In addition, we developed two applications leveraging the FHIR data elements to facilitate patient cohort curation to advance research into pediatric musculoskeletal diseases. Our work enables clinicians and clinical researchers to use hundreds of currently available open-sourced FHIR applications. Our successful implementation of OMOP-on-FHIR within a large hospital system will accelerate advancements in pediatric disease treatment and research.
This paper reports an interpretable automated grading system for diabetic retinopathy using color fundus images. First, we develop shallow learners as baselines. Second, we pre-train deep neural networks to extract high-dimensional features and complex patterns from fundus images and utilize ensemble models to do automatic grading. Then we develop several explainable artificial intelligence models to visualize the extracted deep features and to interpret the predicted outcomes. We investigate the robustness of our system over two publicly available diabetic retinopathy fundus imaging datasets. In addition, we displayed both local and global explainable results to further illustrate the clinical decision-making process with deep models. The innovations of our work include (1) using ensemble models to boost the performance of diabetic retinopathy grading system, and (2) providing transparency of ensemble models using explainable artificial intelligence. The result has shown the potential to improve the effectiveness and accessibility of diabetic retinopathy screening in clinical practice and research settings.
There is a perennial need to identify novel, effective therapeutic agents to combat rising infections. Recently, prediction of therapeutic targets to decrease the impact of COVID-19 has posed an urgent challenge requiring innovative solutions. Successful identification of novel drug-target combinations may greatly facilitate drug development. To meet this need, we developed a COVID-19 drug target prediction model using machine learning approaches to quickly identify drug candidates for 18 COVID-19 protein targets. Specifically, we analyzed the performance of three prediction models to predict drug-target docking scores, which represents the strength of interactions between ligands and proteins. Docking scores were predicted for 300,457 molecules on 18 different COVID-19 related protein docking targets. Our proposed approach achieved a competitive performance with $\mathrm{R}^{2}$=0.69,MAE=0.285, MSE=0.627. In addition, we identify chemical structures associated with stronger binding affinities across target binding sites. We believe our work could potentially save pharmaceutical companies significant resources, especially during the early stages of drug development.
Bio-marker identification for COVID-19 remains a vital research area to improve current and future pandemic responses. Innovative artificial intelligence and machine learning-based systems may leverage the large quantity and complexity of single cell sequencing data to quickly identify disease with high sensitivity. In this study, we developed a novel approach to classify patient COVID-19 infection severity using single-cell sequencing data derived from patient BronchoAlveolar Lavage Fluid (BALF) samples. We also identified key genetic biomarkers associated with COVID-19 infection severity. Feature importance scores from high performing COVID-19 classifiers were used to identify a set of novel genetic biomarkers that are predictive of COVID-19 infection severity. Treatment development and pandemic reaction may be greatly improved using our novel big-data approach. Our implementation is available on https://github.com/aekanshgoel/COVID-19_scRNAseq.
Alzheimer's Disease (AD) is an irreversible and progressive neurodegenerative disorder with three stages: cognitively normal (CN), mild cognitive impairment (MCI), and clinical dementia. Progression and stage prediction of dementia plays an important role in prognosis and treatment. In this work, we developed a multi-modal AD progress prediction model that integrates magnetic resonance imaging (MRI) and electronic health record (EHR) to classify patients into three stages: CN, MCI, and AD. We trained deep auto-encoder to extract features from EHR data, and ResNet and 3D U-Net for MRI imaging data. We developed an entropy-based weighted sum classification method to integrate the classification results from each individual modality to generate final prediction. We experimented on Alzheimer's Disease Neuroimaging Initiative (ADNI) data to demonstrate that the multi-modality integration model outperforms single modality models in accuracy, precision, recall, and F1 scores. In addition, our model achieves competitive performance in comparison with other state-of-the-art multimodality integration methods on AD progression prediction.