
Retinopathy is characterized by pathological alterations in the retina that can lead to partial or complete vision loss. It can result from a variety of causes, including diabetes, hypertension, and autoimmune disorders, such as Autoimmune Retinopathy, which is particularly challenging to detect and diagnose. Due to its visually similar features, Autoimmune Retinopathy is often mistaken for other retinal diseases that require distinct treatments. The objective of this work is to develop a system for detecting clinically relevant anomalies in retinal images using a Transformer-based Deep Learning model. By leveraging Transformers, the model captures subtle variations in retinal images, enhancing diagnostic accuracy. This approach aims to assist ophthalmologists in precise disease detection, improving early diagnosis and enabling personalized treatment strategies.
Breast cancer remains one of the most prevalent and life-threatening diseases worldwide, being the most frequently diagnosed cancer among women and the second leading cause of cancer-related mortality. Precise tumor segmentation is essential for breast cancer assessment, as it enables accurate estimation of tumor size, monitoring of disease progression, and evaluation of treatment effectiveness. Despite the importance of this task, the development of reliable automatic methods is hindered by the scarcity of fully annotated datasets, which makes manual labeling both time-consuming and subject to inter-observer variability. In this study, we propose an unsupervised 3D tumor segmentation method based on Fuzzy C-Means (FCM) clustering, specifically designed for volumetric Dynamic Contrast-Enhanced MRI (DCE-MRI) of the breast. Unlike supervised deep learning approaches, our method does not require manual annotations for training, making it especially valuable in scenarios with limited labeled data. The proposed pipeline combines preprocessing, region-of-interest extraction, and FCM-based clustering to generate accurate segmentation masks with minimal human intervention. We evaluated our approach using clinical data from the ACRIN-6698 dataset, comparing the automatic segmentations against expert manual annotations. The method achieved high performance across multiple metrics, including accuracy, precision, recall, specificity, Dice-Sorensen coefficient (DSC), and Jaccard index (IoU). These results demonstrate the feasibility of unsupervised clustering techniques for volumetric breast tumor segmentation, offering a promising alternative to supervised methods in clinical contexts where annotated data is limited.
Obstructive Sleep Apnea (OSA) management struggles from static retrospective adherence models which build their assessment on rigid compliance benchmarks. This paper describes a modular AI-based framework for CPAP (Continuous Positive Airway Pressure) personalization, which combines real-time digital biomarkers and dynamic therapystage clustering. Patients are stratified into four progressive stages, new, adherent, attempting and non-adherent, using unsupervised learning on the CPAP usage data, clinical history, and behavioral surveys. The system dynamically adjusts biomarker tracking such as HRV, ODI and RVO, produces weekly risk scores and triggers phase specific interventions such as alerts, education modules and physician reviewed actions. Unlike previous models, the framework continually improves clusters and provides interventions depending on physiological feedback and engagement history. Preliminary timescale review shows its ability of identifying dropout threat and adaptation of interventions in response to real time surveillance. The proposed system enables scalable, explainable, and medically integrated for biomarker-driven real-time CPAP personalization methods throughout the OSA care pathway.
Lung cancer remains one of the deadliest cancers and a major public health concern. Although numerous studies have identified various risk factors, further research is essential, particularly in the biological domain. Existing data sources compile biological information on lung cancer and its subtypes but differ in structure and format, complicating data extraction and integration for artificial intelligence (AI) models. Ontologies and semantic technologies address this challenge by enabling the construction of unified knowledge graphs that promote interoperability. Lung-CABO is an ontology specifically designed for lung cancer, supporting the creation of a knowledge graph for risk factor identification and AI applications. Its modular design allows expansion to integrate additional data, such as environmental factors, further enhancing its utility and reusability.
The circadian rhythm is essential for regulating physiological functions and is compromised in people with dementia. Objective measurements are needed to track circadian trend and stability, usually from actigraphy data. This paper explores the feasibility of applying the stability model on activity counts obtained from raw accelerometer, and discusses preliminary results limited to two participants included in a randomized controlled trial that investigates the effect of virtual darkness as a treatment.
Antimicrobial resistance (AMR) is a growing global health challenge that necessitates accurate computational methods for predicting resistant phenotypes from genomic data. In this study, we evaluate the performance of traditional machine learning (ML) models-Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and kNearest Neighbors (k-NN)—alongside deep learning (DL) approaches, including Multilayer Perceptron (MLP), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN). We also introduce TabTransformer as an additional model for AMR classification. Our models leverage k-mer-based features, resistance gene presence/absence, and SNP data across three feature sets. Our results show that LR, RF, and SVM consistently achieved an over 90% accuracy, while k -NN underperformed (77.78 %) due to sensitivity to high-dimensional feature spaces. Among deep learning models, RNN demonstrated superior performance in high-dimensional settings (97.22 % accuracy on 1024 features), while CNN and MLP experienced performance declines. TabTransformer also exhibited robust accuracy across different feature sets, outperforming CNN and MLP. Increasing the number of selected features beyond 256 did not significantly improve performance, emphasizing the importance of feature selection. Future work will explore nucleotide models and further integration of biological metadata to enhance model interpretability and prediction accuracy.
Breast cancer is one of the most common and lethal types of cancer worldwide, with millions of new cases diagnosed each year. Early detection is pivotal in improving patient outcomes and significantly increases the chances of successful treatment. While traditional detection methods such as mammography are effective, they can be invasive, costly, painful, and less applicable for younger women with denser breast tissue. In this context, infrared thermography emerges as a promising, non-invasive technique for breast cancer detection. However, analyzing these images presents challenges due to noise and irrelevant information that can interfere with accurate diagnosis. In this work, we propose a method for segmenting infrared breast images using the DeepLabV3+ Convolutional Neural Network (CNN). Our approach leverages the power of deep learning to precisely delineate breast regions, enabling more accurate feature extraction for subsequent classification tasks. Results achieved an average accuracy of 98.69%, an Intersection over Union (IoU) of 97.18%, and a precision of 98.48%, demonstrating a clear improvement over previous approaches, particularly in terms of segmentation quality, making our method a robust tool for enhancing automated breast cancer detection.
Using simulation for medical education improves clinical practice and creates a safe environment for students to develop skills. High-fidelity patient simulators, as one kind of simulation tool, use advanced life-like manikins. A high-fidelity simulation session includes pre-briefing, simulation scenarios, and debriefing. While pre-briefing and debriefing are widely recognized, the reflection-in-action during simulation is often overlooked despite its importance. Reflection-in-action is the ability to reflect on the student and adapt during the simulation. One major issue regarding reflection-in-action is the limited number of instructors at the university compared to the large number of students. In addition, the intensive time required for the simulations further constrains the capacity of the instructors. To address this critical gap, we propose using a proactive computing approach to provide real-time support for decisions in high-fidelity medical simulations. By implementing the proactive system and providing immediate adaptive feedback to students based on sensor data, we hope to assist teachers and improve learning outcomes.
Effective management of pediatric diabetes remains a clinical challenge, particularly due to the onset of lipodystrophy resulting from repeated insulin injections and inadequate rotation of injection sites. These complications adversely affect subcutaneous tissue integrity and insulin pharmacokinetics, ultimately compromising glycemic control. In this study, we introduce diPen, a novel smart case designed to integrate with commercial insulin pens and equipped with a dual-sensor system: an optical module for non-invasive lipodystrophy detection and an inertial measurement unit for monitoring injection-site rotation. To evaluate the feasibility of the proposed approach, a clinical study was conducted involving a pediatric cohort. The system employs a personalized machine learning pipeline, leveraging a leave-one-acquisition-out validation strategy to replicate realworld deployment scenarios, wherein newly acquired data from the same subject are evaluated without prior exposure during training. The system demonstrated promising performance in both detection and classification tasks, suggesting that di-Pen may represent a viable tool for enhancing insulin therapy through personalized, data-driven injection guidance and tissue health monitoring.
The identification of fraud, waste, and abuse in healthcare systems is critical to minimise financial losses incurred by medical aid companies. Traditional fraud detection methods are based on manual processes and often fail to identify hidden patterns within medical claims data. The complexity and scale of fraudulent activities require more advanced, data-driven approaches to detect. This paper develops a framework that employs unsupervised machine learning models to improve fraud detection from medical claims data for a South African medical scheme. The framework is a two-layer approach that employs six unsupervised anomaly detection algorithms, including isolation forest, one-class support vector machine, autoencoder, selforganising map, local outlier factor, and hierarchical densitybased spatial clustering of applications with noise. The first layer focuses on detecting global outliers, while the second layer refines these classifications using local outlier factors. The results show that the framework flags potential fraud, waste, and abuse and provides valuable insights into which medical practices should be investigated further. The paper highlights the potential of unsupervised learning for fraud detection in medical claims, reducing reliance on manual audits and improving detection accuracy.
Deep Learning approaches show promise for improving tomosynthesis reconstruction but require paired CT and tomosynthesis datasets that are difficult to obtain. This work presents ChestXsim, an open-source Python framework that enable the simulation of digital chest tomosynthesis (DCT) from chest CT data. Its modular design includes the preprocessing pipeline to adapt helical chest CT volumes to a standard tomosynthesis positioning and the simulation of polychromatic projections with noise modelling. Additionally, ChestXsim provides reconstruction techniques (FDK/SART) and leverages GPU acceleration through open-source ASTRA kernels. ChestXsim offers a fast and efficient method for simulating DCT from standard chest CT data, making it a valuable tool for generating paired datasets.
Breast cancer diagnosis using histopathological images is a challenging task due to the scarcity of annotated medical data, particularly for rare cancer stages. Traditional deep learning models struggle to generalize effectively in such low-data scenarios. To address this problem, we propose a few-shot classification framework for breast histopathological images based on metric-based learning. Our approach leverages Vision Transformers (ViTs) for feature extraction, capturing global contextual information better than conventional Convolutional Neural Networks (CNNs). Additionally, we integrate BLIP-2, a Vision Language Model (VLM), to incorporate manual text prompts and contextual textual descriptions, enhancing the model's interpretability and adaptability. The extracted visual and textual features are fused using a novel feature fusion module, and classified the samples based on cosine distance. We evaluated our approach on BreakHis and BACH datasets, showing its effectiveness in few-shot learning (FSL). Our model achieves 57.12% and around 89% in 5-shot setting, respectively, on the BACH and BreakHis datasets. As the number of support samples increases, the performance of the model improves. Our findings suggest that combining transformer-based architectures with VLMs enhances the performance of FSL based medical image classification systems. The code implementation of the methodology is available at MultiModal-FewShot
Accurately detecting pain in infants remains a complex challenge. Conventional deep neural networks used for analyzing infant cry sounds typically demand large labeled datasets, substantial computational power, and often lack interpretability. In this work, we introduce a novel approach that leverages OpenAI's vision-language model, GPT-4(V), combined with mel spectrogram-based representations of infant cries through prompting. This prompting strategy significantly reduces the dependence on large training datasets while enhancing transparency and interpretability. Using the USF-MNPAD-II dataset, our method achieves an accuracy of 83.33% with only 16 training samples, in contrast to the 4,914 samples required in the baseline model. To our knowledge, this represents the first application of few-shot prompting with vision-language models such as GPT-4o for infant pain classification.
This study presents the first phase of a transdisciplinary research project aimed at improving the design of visual feedback stimuli in neurofeedback (NFB) applications. While current NFB research has focused extensively on signal processing and feature extraction, limited attention has been given to the design and user experience of feedback stimuli. To address this gap, the research team conducted generative user research including site visits, expert consultations, and semistructured interviews with domain experts and previous NFB participants. Analysis of the collected data yielded a preliminary set of design requirements. User-centered requirements include minimizing cognitive load, enhancing attention and engagement, incorporating positive reinforcement, supporting a sense of agency, and providing clear instructions. Technical requirements include reducing artifacts, ensuring low-latency feedback, and promoting participant relaxation. These findings lay the groundwork for iterative design and evaluation phases, with the ultimate goal of delivering validated stimuli and design guidelines to the NFB research and clinical communities.
Type 1 diabetes is a chronic disease that results from insufficient insulin production by the pancreas. While artificial pancreas systems have emerged as an alternative therapy, current commercial devices do not account for physical activity in their control algorithms. A major issue is the scarcity of high-quality and unbiased datasets comprising physical activity records. This study investigates the application of transfer learning techniques and physical activity data to enhance glucose prediction. To this end, we propose three distinct methods: a graph topology model, a substitution model and a retraining method. A comprehensive analysis was conducted to assess accuracy, delay and clinical utility. The study revealed that the substitution model exhibited superior accuracy and reduced delay compared to the base model. The study compares the graph topology model and the retrained model, with the former proving to be the best for a PH of 30 minutes, obtaining an RMSE of 21.63 mg/dL compared to the RMSE of 21.85 mg/dL obtained with the retrained model. For a PH of 60 minutes, the graph topology model achieved an RMSE of 33.87 mg/dL in comparison to the 34.14 mg/dL of the retrained model. In the Error Grid analysis for a PH of 30 minutes, both the graph (98.56%) and the retrained model (98.41%) achieve results close to the acceptable range. However, with a PH of 60 minutes, the clinical utility drops significantly (96.4% vs. 96.21%). These findings underscore the importance of incorporating physical activity data and the need for further exploration of approaches that account for their impact.
This paper presents a novel approach to noninvasively assess different levels of perfusion using remote photoplethysmography (rPPG) from an RGB camera. In a proof-of-concept study involving 20 healthy participants, three degrees of stenosis were induced in the upper extremity by applying external pressure. The rPPG signal from the affected arm was correlated with a reference signal from the well-perfused control arm. The results show that correlation analysis of rPPG signals can distinguish between mild, moderate, and full stenosis. This supports the potential of rPPG as a cost-effective tool for detecting perfusion disorders such as peripheral artery disease (PAD).
Cognitive impairment is a growing public health concern, with early detection playing a crucial role in improving patient outcomes. The Montreal Cognitive Assessment (MoCA) is widely used for screening mild cognitive impairment (MCI) and early-stage dementia. However, traditional MoCA assessments require manual scoring by trained professionals, making the process labor-intensive, time-consuming, and susceptible to human error. To overcome these limitations, we propose an automated pipeline for MoCA score estimation using eye-gaze data and Vision Transformers (ViTs). Our approach leverages gaze-tracking technology to capture spatial and temporal eyemovement patterns during structured cognitive tasks, identifying subtle cognitive impairments that may otherwise go unnoticed. The raw gaze data is preprocessed and mapped onto taskrelevant image regions, where a pretrained ViT extracts highdimensional feature representations. To address inconsistencies in gaze sampling and improve temporal modeling, we introduce a time-aware positional embedding mechanism that enhances the model's ability to infer cognitive performance. These extracted features are then processed by a transformer-based classification model to predict MoCA scores with high accuracy. We validate our approach using a dataset collected from seven cognitive gaming sessions, demonstrating its effectiveness in automated cognitive assessment. The experimental results indicate that our method provides a reliable and efficient alternative to traditional MoCA evaluations, reducing dependency on human intervention while maintaining diagnostic accuracy.
This study proposes a classification framework for Alzheimer's Disease (AD) using statistical features from displacement vector fields and Jacobian determinants computed from structural magnetic resonance imaging (MRI). Images from 542 cognitively normal (CN), 341 mild cognitive impairment (MCI), and 245 AD individuals were analyzed. Groupwise registration and deformable coregistration quantified spatial deformations and local volumetric changes. Statistical moments from displacement vector fields and Jacobian determinants enabled regionspecific analysis. Stratified by sex, CN vs. AD classification achieved an AUC of 0.93 and 88.68 % accuracy for males, and an AUC of 0.94 with 87.89 % accuracy for females, demonstrating the efficacy of deformation-based biomarkers for AD diagnosis.
Event-Based Surveillance Systems (EBS) are crucial for detecting emerging public health threats. However, these systems face significant challenges, including overreliance on manual expert intervention, limited handling of heterogeneous textual data, etc. The Description-Detection Framework (DDF) addresses some of these limitations by leveraging PropaPhen (Core Propagation Phenomenon Ontology), UMLS, and OpenStreetMaps to detect suspicious health-related cases using spatiotemporal and textual data. However, DDF is restricted to detection and lacks the ability to classify the detected observations into meaningful categories. To adress this limitation, we propose to enhance DDF by incorporating a clustering-based classification process. This enhancement employs BioSTransformers, a pretrained biomedical language model built on Sentence Transformers trained on PubMed data, to compute semantic similarity between observations. By capturing domain-specific semantic relationships, BioSTransformers enables clustering that integrates biological semantics with spatiotemporal context, outperforming traditional methods from the literature in observation classification. Our proposed approach reduces the dependency on manual expert effort, improves the system's ability to process heterogeneous data, and enhances the accuracy and contextual relevance of case classification. The results demonstrate the potential of this method to advance EBS systems, providing a scalable and automated solution to public health surveillance challenges.
Deep learning models have achieved great success in medical imaging tasks. However, recent work on fairness in healthcare has shown these models can be biased, potentially leading to discriminatory treatment of patients based on demographic attributes such as race, gender, and age. Data bias, often resulting from imbalanced and non-representative datasets, can negatively impact model fairness. While aggregating data from multiple sources can help mitigate data bias, privacy concerns make the data-sharing process challenging. In this scenario, Federated Learning (FL) has emerged as a solution for the collaborative training of models without data sharing. This paper presents FairMed-FL, a methodology to assess fairness in FL for medical imaging tasks. By utilizing two public chest X-ray datasets partitioned by sex and age, we compared federated models trained with clients from single and multiple datasets against centralized models trained on each client's data. The results indicate that FL reduces performance discrepancies between demographic groups, enhances the performance of the worstperforming groups, and improves overall metrics compared to centralized approaches. These findings highlight its potential for promoting fairness in medical imaging.