
While variations in electronic health record (EHR) use among healthcare professionals and across sociodemographic patient groups have been studied, the temporal EHR use by healthcare teams remains underexplored. This study aims to develop metrics to quantify temporal EHR use by teams, measure racial variations in these patterns, and assess their associations with patient outcomes using large-scale EHR audit logs. For each patient's care episode, we used audit logs to construct a sequence of healthcare professionals who were involved in the patient's care and performed actions in the patient's EHR. Two professionals are considered neighbors in the sequence if they performed consecutive EHR actions or if their actions overlapped in time. We then identified temporal patterns as k-tuples-subsequences of healthcare professionals of length k extracted from the patient's sequence. Term Frequency-Inverse Document Frequency (TF-IDF) was used to quantify the importance of each k-tuple within a patient's sequence, treating each patient sequence as a document and k-tuples as words. Proportional-odds (PO) logistic regression models evaluated racial variations in k-tuples, adjusting for baseline factors including body mass index, type 2 diabetes mellitus, surgery type, age, sex, Charlson Comorbidity Index, hypertension, ischemic heart disease, depression, surgery duration, and surgery time. PO models also assessed associations between k-tuples and postoperative length of stay (PLOS), controlling for these baseline variables. A total of 1,691 patient sequences were analyzed, generating 487, 1,870, and 4,408 unique 3-, 4-, and 5-tuples, respectively. The odds of a longer PLOS increased by 0.7% to 5.5% for each unit increase in the TF-IDF of 12 k-tuples (3 <= k <= 5), particularly those involving consecutive EHR actions performed by multiple anesthesiologists (false discovery rate [FDR] p < 0.05). Conversely, nine k-tuples (3 <= k <= 5), particularly those involving consecutive EHR actions performed by resident physicians, anesthesiologists, and physicians, were associated with shorter PLOS, with odds decreasing by 1.3% to 6.1% per unit increase in their TF-IDF (FDR p < 0.05). Variations in the TF-IDF of four k-tuples (3 <= k <= 5) involving consecutive EHR actions performed by physicians and anesthesiologists, or by resident physicians and nurse anesthetists, were observed between White and Black patients (odds ratios ranging from 1.009 to 1.023, FDR p < 0.05); importantly, these variations were not associated with PLOS. Our findings suggest that specific temporal patterns of EHR use by healthcare teams are associated with patient outcomes. Further research is needed to investigate the reasons behind k-tuples that are positively or negatively associated with PLOS to inform guidelines aimed at improving patient outcomes.
Occupational health is an aspect often neglected in our society. Illness, injuries, and fatalities in the workplace have remarkable consequences from both societal as well as economical point of view. Fitness-for-work judgment is critical to deem whether workers are fit, unfit, or fit with restrictions for a specific job. Identifying the key factors and how these impact the workers’ life quality is paramount.In this study, we harness machine learning, pattern mining and causal learning to first identify key impacting factors and patterns associated with fitness-for-work decisions in the maritime domain, and then to extract and use a causal model to enable a what-if analysis on the identified factors aimed at predicting the outcome of potential improvement actions.The current results highlight the prominent role of factors like age, smoking, and audiometry values in the judgment in relation to the job type. We inferred the 10 most frequent patterns, validated against decision tree rules and whose significance is confirmed by the Relative Risk and Odds Ratio. These are then used to query an automatically extracted causal model to simulate potential improvement actions as well as to support root cause analysis. We used real-world data from two maritime companies according to their health protocols. The dataset is made publicly available, contributing to the relatively scarce empirical body of knowledge on occupational health.
Accurate predictive reasoning is a cornerstone of biomedical decision-making, particularly in precision oncology, where elucidating the intricate relationships between disease-risk genes and biological processes is critical. This study presents a novel "Thought Graph" methodology, an advancement of the Tree of Thoughts framework, to systematically generate and refine biological process representations derived from gene sets while addressing the trade-off between specificity and uncertainty. Balancing these factors is essential for robust and interpretable gene set analyses, as it accounts for the complexity, variability, and overlapping functions of biological pathways. Furthermore, we introduce a quantitative metric that integrates specificity and uncertainty, thereby enhancing the rigor and transparency of the inference process. Using a subset of the Gene Ontology database, we evaluate the effectiveness of our system in generating biologically meaningful terms that accurately describe the underlying biological processes of gene sets. We compare its performance against a domain-specific tool (GSEA) and five LLM baselines across multiple metrics. Our system achieves the highest cosine similarity (64.00%) and specificity percentile (96.40%), highlighting its capacity to generate terms closely aligned with human annotations while maintaining a balance between specificity and accuracy. By advancing the artificial intelligence driven analyses, this work facilitates more informed decision-making in biomedical research, precision oncology, and related fields.
Video-based patient behavior analysis offers deep insight into patient activities, adherence, risk assessment, and clinical decision-making. By integrating computer vision and machine learning, healthcare providers can monitor and analyze patient behaviors, improving quality of care, recovery progress, and treatment. In this paper, we propose a method to systemati- cally monitor patient behavior by analyzing interactions between a patient (in various postures) and clinically significant objects. Our methodology includes a CNN-based person/object detection & tracking, posture classification, ambulation assessment, depth estimation, and interaction analysis. A customized clinically relevant image dataset is constructed for training and validation. Our platform has achieved promising performance, i.e. an average F1 score of 98% for posture classification and an average accuracy of 90% for interaction detection.
Advancements in treating infectious diseases have improved in the past two decades, but the COVID-19 pandemic showed how quickly diseases can spread in today’s interconnected world. Computational epidemiology, using artificial intelligence and model simulations, helps experts analyse and control the spread of illnesses. This objective can be achieved using two different modelling paradigms: (i) macro-simulations , and (ii) micro-simulations. The former aims to characterize a system’s behaviour from a macroscopic perspective, utilizing mathematical methods like deterministic or stochastic processes to depict disease evolution and population dynamics. By concentrating on average or aggregated metrics, macro-simulations are especially useful for analysing broad trends and long-term implications. Differently, the latter aim at describing a system’s behaviour from a microscopic level in terms of its components and their interactions. Agent-Based Models are a common approach in this paradigm, representing individual agents with specific characteristics and decision-making rules. These models allow researchers to simulate complex, heterogeneous populations and capture localized or emergent phenomena, such as super-spreading events or the effects of targeted interventions [1] .
The enrichment of healthcare data coupled with advancements in computational capabilities has driven the adoption of artificial intelligence in healthcare. However, these approaches, when implemented without fairness considerations, may exacerbate existing disparities, leading to inequitable resource allocation and diagnostic inaccuracies across different demographic groups. This study introduces a novel approach via gradient projection to multi-attribute fairness optimization in healthcare AI, optimizing fairness across multiple demographic attributes and predictive performance concurrently. The approach aligns conflicting optimization objectives by projecting each gradient onto the normal plane of the other. Our method also enhances interpretability by elucidating the adjustments made during optimization, providing insights into trade-offs between fairness and accuracy.
Autism is the fastest-growing developmental condition, and as a spectrum disorder, individuals vary in severity in communication and behavior. The rapid development of artificial intelligence (AI) and machine learning (ML) presents opportunities to leverage such technology to support autistic individuals in multiple contexts, including education, healthcare, and disability research. However, very limited research has focused on developing AI or ML-based solutions for autistic individuals largely due to the scarcity of autism-related datasets. In order for autistic people to equally benefit from the breakthroughs in AI, the data representation gap in AI and ML must be addressed. The project proposes an innovative human expert and AI combined approach to generate synthetic data about autistic individuals. Through this approach, we are developing the first dataset of case studies of autistic individuals to effectively identify evidence-based practices (EBPs) to address the unique needs of autistic individuals. The project improves the representation and inclusion of autistic individuals in the rapidly evolving field of ML and AI through both methodological innovations (i.e., the expert/AI combined data generation approach) and practical contributions (i.e., the EBP dataset that will be made publicly accessible). The innovative data generation method and supporting materials developed in this project can be generalized to individuals with other disabilities.
Language models have significant potential to improve clinical workflows, accelerate research, and enhance patient care in hospitals. However, privacy constraints, limited compatibility with diverse IT systems, and the absence of a holistic approach for managing language models within internal networks hinder broader adoption. JAVIS tackles these challenges by offering a secure, scalable, and modular framework for Large Language Models (LLMs) and Vision-Language Models (VLMs), fully operating on private hospital networks. Its features include high-throughput data labeling (text, images, and audio), a no-code LLM training interface, an auto-labeling module for named entity recognition (NER) tasks, and distributed deployment of LLMs and VLMs on multi-GPU infrastructures that provide an internal network chat service. This poster outlines JAVIS’s architecture, key features, and results demonstrating robust performance in data labeling, training, and large-scale inference.
Automatic sleep stage classification is an important task to assist experts to perform diagnosis of sleep-related disorders. In supervised learning setting, the availability of a large amount of labeled training data is often the bottleneck for training the sleep stage classification model. Furthermore, the advent of multi-channel EEG signal from in-home EEG devices for sleep monitoring also pose a challenge of designing a label-efficient model suitable for such data. This work proposes a label-efficient approach leveraging self-supervised learning for performing 5 sleep stage classification suitable for multichannel in-home EEG device signal. Our work demonstrates the effectiveness of employing contrastive learning technique on unlabeled EEG data to learn the prominent features. We fine-tuned the features learned by contrastive learning to perform sleep stage classification using a limited amount of labeled data. The experimental results suggest that the proposed approach is more effective than training a traditional end-to-end sleep stage classification model without contrastively learned features. Specially, our approach is suitable in the situation where a large amount of unlabeled data is available and a small amount of labeled data is provided. Furthermore, our results suggest that the proposed method predicts with higher confidence than the end-to-end model.
In this paper, we systematically investigate the possibility, quality, and potential usage of generating synthetic free-text medical records, such as discharge summaries, admission notes, and doctor correspondences, using the Masked Language Modelling (MLM) strategy. Our system is designed to preserve the critical information of the records while introducing significant diversity and minimising re-identification risk. The system incorporates a de-identification component that uses Philter to mask Protected Health Information (PHI), followed by a Medical Entity Recognition (NER) model to retain key medical information. We explore various masking ratios and mask-filling techniques to balance the trade-off between diversity and fidelity in the synthetic outputs without affecting overall readability. Our results demonstrate that the system can produce high-quality synthetic data with significant diversity while achieving a HIPAA-compliant PHI recall rate of 0.96 and a low re-identification risk of 0.035. Furthermore, downstream evaluations using an NER task reveal that the synthetic data can be effectively used to train models with performance comparable to those trained on real data. The flexibility of the system allows it to be adapted for specific use cases, making it a valuable tool for privacy-preserving data generation in medical research and healthcare applications. We host our models and data publicly at https://github.com/HECTA-UoM/MLM4SynMed
The Frailty Index (FI) is a well-established clinical tool for assessing elderly health. Recent efforts have aimed to refine FI for diverse settings or simplify its collection, enabling more efficient population screening. Motivated by this clinical challenge, this is the first study that aims to enhance both index efficiency and practicality by proposing a data pipeline composed of: (i) a nonlinear feature selection method to identify the most relevant variables for index prediction; (ii) Multi-Objective Symbolic Regression, a symbolic machine learning technique to generate simplified index formulas that replicate the original index’s values, distribution, and risk stratification ability using fewer variables and (iii) rigorous model evaluation through calibration, temporal correlation and associations with related clinical outcomes. We tailored our approach to streamlining the 53-item FI for people living with HIV (PLWH), in use at the Modena HIV Metabolic Clinic, utilizing electronic health records from a public hospital clinic providing data on about 4,800 patients. Several reduced FI (rFI) formulas were derived, with the simplest model, rFI(16), relying on just 16 indicators. Validation confirmed that the rFIs successfully replicate FI statistical properties and screening power also maintaining consistent relations with other geriatric outcomes. By facilitating index distillation from retrospective data, our method offers broad adaptability to various clinical case studies.
Efficient and safe handover of surgical instruments is crucial for optimizing operating room performance and ensuring patient safety. Traditional observational methods, such as manual annotation and direct supervision, often lack the ability to quantitatively analyze these critical interactions, leading to gaps in understanding and workflow inefficiencies. This paper introduces SurgiGard, a novel approach that leverages multimodal machine learning with the CLIP model and Neo4j graph database to precisely detect and analyze instrument handovers during surgical procedures. We developed a comprehensive dataset by collecting and annotating video footage from real surgeries, using the CLIP model to identify key events of instrument possession and transfer. By integrating these detections into a graph-based framework, SurgiGard quantifies handover frequencies, directionality, and patterns, revealing hidden bottlenecks and opportunities for procedural improvements. Our results demonstrate that this method not only accurately captures surgical handovers but also uncovers new insights into team dynamics, paving the way for targeted enhancements in surgical training and protocols. The SurgiGard methodology presents a promising step toward enhancing safety, communication, and efficiency in the operating room.
Traditional drug discovery is a long and costly process with a low success rate. Drug repurposing (DR) offers a promising alternative by leveraging existing drugs with known safety profiles. Network-based DR models have gained attention for their ability to integrate diverse types of biological and clinical data. However, effectively combining such heterogeneous data remains a significant challenge. We propose Gene-Drug-Disease-Bidirectional Encoder Representations from Transformers (GDD-BERT), an AI model designed to identify drug-gene-disease interactions, aiding drug repositioning. It utilizes BERTwalk, a novel embedding technique based on BERT, to generate vector representations of nodes in biological networks by analyzing complex relationships through random walks. These embeddings are combined into drug-gene-disease triplets, which are used in a classifier to predict potential interactions, such as a drug targeting a protein or a protein’s link to a disease. GDD-BERT efficiently reconstructs missing interactions and excels in sparse networks, outperforming traditional methods in discovering new therapeutic applications for drugs.
With recent advances in Deep Learning (DL) models, the healthcare domain has seen an increased adoption of neural networks for clinical diagnosis, monitoring, and prediction. Deep Learning models have been developed for various tasks using 1D (one-dimensional) time-series signals. Time-series healthcare data, typically collected through sensors, have specific structures and characteristics such as frequency and amplitude. The nature of these features, including varying sampling rates that depend on the instruments used for sensing, poses challenges in handling them. Electrocardiograms (ECG), a class of 1D time-series signals representing the electrical activity of the heart, have been used to develop heart condition classification decision support systems. The sampling rate of these signals, influenced by different ECG instruments as well as their calibrations, can greatly impact the learning functions of deep learning models and subsequently, their decision outcomes. This hinders the development and deployment of generalized, DL-based ECG classifiers that can work with data from a variety of ECG instruments, particularly when the sampling rate of the training data remains unknown to users. Moreover, DL models are not designed to recognize the sampling rate of the testing data on which they are being deployed, further complicating their effective application across diverse clinical settings. In this study, we investigated the effect of different sampling rates of time-series ECG signals on DL-based ECG classifiers. To the best of our knowledge, this is the first work to understand how varying sampling rates affect the performance of DL-based models for classifying 1D time-series ECG signals. Through our comprehensive experiments, we showed that accuracy can drop by as much as 20% when the training and testing sampling rates are different. We provide visual explanations to understand the differences in learned model features through activation maps when the sampling rates for training and testing data are different. We also investigated potential strategies to address the challenges posed by different sampling rates: (i) transfer learning, (ii) resampling, and (iii) training a DL model using ECG data at different sampling rates.
Healthcare staff in a hospital ward setting typically monitor patients by taking vital sign observations at regular intervals, usually every 6 - 12 hours for routine observations, but more frequently for critical patients. Patients on similar schedules have been shown to be regularly batched in a practice called ‘ward rounds’, but to what extent healthcare staff manage observations on ‘non-routine’ intervals independently to those scheduled for the subsequent ward round has yet to be established. This study examines the Time-To-Next-Observation (TTNO) for vital sign observations with planned schedules (e.g., 1 hour) as random variables defined by their hazard distribution functions. Joint distribution functions sampled from any pair of TTNO distributions could be used to calculate the probability that any two observations on set schedules will happen collectively. However, it is clear that this model fails to capture underlying dependency structures seen in empirical results. We propose a copula approach to extract and quantify pairwise nonlinear relationships for all standard observation intervals across 20 study wards. This study showed that most wards operate with significant levels of dependency between observation schedules, largely reflecting broad ward characteristics referenced in other works, yet present a deeper level of insight into individual ward operations. Understanding current levels of dependency between regular observation scheduling and routine operations has the potential to become essential knowledge for ward stakeholders when designing staff resource strategies.
This study introduces a co-learning style semi-supervised regression (SSR) to predict future health checkup results. Co-learning is a significant approach within SSR, which builds a model based on labeled samples and estimates target response assigned to unlabeled samples as pseudo-labels, which are then used as labeled samples during training. Recently, this approach has emphasized the importance of enhancing pseudo-labels’ quality. However, conventional SSR generates pseudo-labels using only given labeled and unlabeled samples, limiting their quality. Therefore, incorporating auxiliary data and diverse perspectives is crucial for generating reliable pseudo-labels.We present a new co-learning style SSR using auxiliary data from health checkup results accumulated across different age groups. Using optimal transport (OT) theory, we learn a map of the unpaired distributions of health checkup results from different age groups. This OT map assigns target responses to unlabeled samples as pseudo-labels. Our approach assumes that change in health checkup results over time can be inferred by comparing them collected in different age groups despite being different individuals. By integrating pseudo-labels assigned from labeled and unlabeled samples and ones assigned using our OT map from an auxiliary view in an optimally weighted manner, we improve the quality of pseudo-labels to enhance prediction accuracy. Experimental results using health checkup results from a large cohort study over fifteen years show that our method using auxiliary data with an OT map is effective in achieving higher accuracy than the existing co-learning approach on SSR.
Object detection is a crucial task in computer vision with applications in a plethora of fields such as autonomous driving, medical imaging, surveillance, and robotics. The effectiveness of object detection models heavily depends on the availability of large-scale annotated datasets, where objects are labeled with precise bounding boxes. However, the manual annotation of images to generate bounding boxes is a time-consuming and labor-intensive process, often requiring domain-specific expertise. In this paper, we propose a method to automatically generate bounding boxes from Class Activation Maps images, with the aim of generate object detection datasets annotated with bounding boxes derived from these heatmaps. The proposed approach leverages the localized information provided by Class Activation Maps to detect and isolate regions of interest, encapsulating them within bounding boxes. The proposed method has broad applicability across various domains, including medical imaging, cybersecurity, and industrial inspection, where annotated datasets are often scarce.
This study investigates the potential of large language models (LLMs) to leverage their extensive pre-trained knowledge for conducting affect classification using ambulatory data collected from smartphone sensors in support of precise mental health (MH) monitoring. We propose several prompting strategies to conduct a three-way classification of happiness and sadness levels: baseline, temporal method, personalization method, and two variations of a combination method to guide the learning process. These designs incorporate in-context learning, prompt engineering, and additional MH-related context to enhance model performance. Results indicate that the combination strategy outperforms the others in terms of balanced accuracy and macro F1 score for both the sadness and happiness tasks using longitudinal GPS, phone log, and accelerometer data. Additionally, we conduct a linguistic analysis of the LLM’s reasoning that led to its decision across these strategies. The findings suggest that prompting methods where the LLM reasoning incorporates language related to affect, analytical thinking and authenticity yield improved affect classification performance.
Forecasting migraine attacks before they occur can significantly enhance the quality of life for patients by allowing them to take preemptive measures to mitigate or prevent the onset of symptoms. In this study, we aim to predict migraine attacks by leveraging the largest real-world migraine dataset available to date, collected from a mobile app between 2016 and 2022, it encompasses approximately 43,000 users and 7 million daily records gathered throughout the entire 6-year period. We introduce TRACE (Temporal Recurrent Autoencoder for Concept Embedding), a novel approach developed to anticipate migraine attacks. On two held-out sets, using 2-day and 3-day lookback data, TRACE successfully predicts approximately 70% of migraine episodes before their onset, outperforming competing methods in terms of sensitivity, i.e., accurately predicting migraine days. TRACE maintains a comparable specificity of around 60%, i.e., accurately predicting no-headache days. To understand the predictive factors driving models performance, we conduct both global and local feature importance analyses. Global importance is assessed using ablation analysis and Random Forest, while local importance is evaluated on a per-patient basis using SHapley Additive exPlanations (SHAP). Our SHAP analysis show that while common migraine triggers exist, individual triggers can vary significantly, highlighting the need for personalized predictive tools.
The human gut microbiome plays a pivotal role in health and disease, influencing conditions ranging from gastrointestinal disorders to systemic diseases. The emergence of digital twin technology provides a transformative approach to understanding and predicting complex interactions between the gut microbiota and medications. This paper presents a comprehensive framework for creating a digital twin of the human gut microbiome, using data-driven methodologies to enable personalized medicine. The proposed system models the dynamics of the gut microbiome, predicts drug-microbe interactions using a Random Forest machine learning approach, and estimates the changes in the microbial population through mathematical modeling. By integrating genomic data, drug-specific chemical properties, and dose-response relationships, it simulates the microbiome responses to drugs, presenting results through interactive 3D visualizations and graphical interfaces. Performance evaluations highlight the model’s robustness, achieving a ROC-AUC of 0.97 and PR-AUC of 0.91, outperforming other machine learning approaches. This research underscores the potential of digital twins in advancing personalized gut health treatments, paving the way for precision medicine.