
Understanding the neural mechanisms underlying emotional processing is critical for advancements in emotional neuroscience. This study explores the relationship between emotion and motion perception using Event-Related Potentials (ERPs) in a structured experimental setup. We incorporate subjective intensity ratings to enrich the data by capturing the subjective experiences of participants in response to emotional stimuli with implied motion and no motion. Thirty university students participated in the study, where EEG data was collected and analyzed using threshold-free cluster enhancement (TFCE) for time-domain analysis and Repeated Measures ANOVA for frequency-domain analysis. Furthermore, we developed a multimodal deep learning model to predict subjective intensity levels from EEG-derived features. This model leverages statistical, spectral, and auto covariance features, integrated through a transformer encoder layer, to enhance predictive capability. Our findings contribute to a deeper understanding of emotional processing in the brain and highlight the importance of incorporating subjective measures in neuroscience research.
Despite the renal biopsy being the gold standard for diagnosing glomerulonephritis, this practice remains inaccessible for many patients worldwide. Nephropathologists typically combine microscopy, immunohistology, transmission electron microscopy, clinical information, and genetic studies for diagnosis. However, variability in nephropathology evaluation has hindered its integration with emerging technologies and personalized medicine. This study proposes the use of deep learning to extract significant features to distinguish glomerulonephritis from PAS sections without other modalities. To test this hypothesis, various AI methods were tested for classifying 12 common glomerulonephritis diagnoses. Finally, a sequential classification was implemented, initially characterizing sclerosed and non-sclerosed glomeruli using Swin-Transformers, followed by classifying the non-sclerosed glomeruli into 12 types of glomerulonephritis using ConvNeXt. The first step achieved an average Balanced Accuracy of 97% and an AUC of 0.96. In the second step, a Balanced Accuracy considering up to the top3 of 79.5% and an avarage AUCs of 0.76 were achieved. This study establishes a baseline for this challenging classification task, demonstrating promising results even on single PAS glomerular crops.
The Receiver Operating Characteristic (ROC) curve is a critical tool for binary classification analysis in medicine, with the Area Under the ROC Curve (AUROC) serving as a widely accepted metric to evaluate the performance of binary classifiers. This study conducts a comprehensive review of the ROC curve with a focus on its utility in outlier identification. We introduce a novel scoring method to rank actual positives and actual negatives within a test set, according to their impact on AUROC degradation. We bridge the scoring system with the ROC curve analysis to quantify each data point's contribution to AUROC loss. Furthermore, we introduce the IMICS ROC Analyzer, a graphical user interface-based software, embedded with our innovative algorithms. Through the use of an open-source prostate cancer dataset, we illustrate the application of our algorithms for practical outlier detection in binary classification tasks. The IMICS ROC Analyzer enhances the field of precision medicine by allowing for measuring an individual's contributions (be it patients, lesions, or samples) to the overall AUROC, thus facilitating confidence measurement of Machine Learning (ML) classifiers for individual cases of interest in a cohort.
The critical shortage of medical professionals in low-resource countries, notably in Africa, hinders adequate health-care delivery. AI, particularly Multimodal Large Language models (MLLMs), can enhance the efficiency of healthcare systems by assisting in medical image analysis and diagnosis. However, the deployment of state-of-the-art MLLMs is limited in these regions due to the high computational demands that exceed the capabilities of consumer-grade GPUs. This paper presents a framework for optimizing MLLMs for resource-constrained environments. We introduce optimized medical MLLMs including TinyLLaVA-Med-F, a medical fine-tuned MLLM, and quantized variants (TinyLLa VA - Med- FQ4, Tiny LLa VA - Med- FQ8, LLa VA - Med-Q4, and LLaVA-Med-Q8) that demonstrate substantial reductions in memory usage without significant loss in accuracy. Specifically, TinyLLaVA-Med-FQ4 achieves the greatest reductions, lowering dynamic memory by approximately 89% and static memory by 90% compared to LLaVA-Med. Similarly, LLaVA-Med-Q4 reduces dynamic memory by 65% and static memory by 67% compared to state-of-the-art LLaVA-Med. These memory reductions make these models feasible for deployment on consumer-grade GPUs such as RTX 3050. This research underscores the potential for deploying optimized MLLMs in low-resource settings, providing a foundation for future developments in accessible AI-driven healthcare solutions.
Early identification of individuals at risk for Alzheimer's disease is essential to improve treatment effectiveness. Cerebrospinal fluid analyses and positron emission tomography (PET) scans are commonly used to detect the presence of beta-amyloid and tau, which are associated with an increased risk of conversion to Alzheimer's disease. However, these biomarker tests are expensive and involve invasive procedures. Researchers are working towards discovering easily measurable biomarkers to detect individuals at risk, but only a few have been identified thus far. There is a need to discover biomarkers that are cost-efficient and non-invasive to test. We propose a machine learning approach for discovering potential risk biomarkers of Alzheimer's disease through the analysis of physiological responses to the cognitively complex task of driving by using decision tree ensemble techniques. Though driving patterns in early Alzheimer's have been previously studied, physiological responses of cognitively normal seniors during driving remain unexplored. As a first step, we measure heart rate, electrodermal activity, and temperature responses to several driving events, such as right turns and roundabouts, of seniors with and without elevated PET beta-amyloid levels to explore the relationship between these physiological responses and amyloid level. Data were collected from 26 participants with elevated beta-amyloid and 28 without. We used four machine learning algorithms for classification: Random Forest, Extra Trees, AdaBoost, and XGBoost, and developed a novel methodology to extract significant features from these models. By doing so, we successfully identified five risk biomarkers most influential in differentiating the two groups with and without elevated beta-amyloid.
In the healthcare sector, the scarcity of data and privacy concerns present formidable challenges to the widespread adoption of machine learning. In the present-day scenario, Federated Learning (FL) emerges as a pivotal solution, fostering the rapid evolution of distributed machine learning paradigms while adeptly addressing the problem of data governance and privacy. It allows distributed clients to collaboratively train a global model by synchronizing their local updates without sharing private data. In recent years, federated learning and quantum computing have individually shown great promise to revolutionize various sectors, including healthcare, finance, and manufacturing, where privacy protection is paramount. In this article, we propose a communication-efficient Quantum Federated Learning (QFL) framework based on a variational circuit that enables clients to efficiently train and transmit quantum model parameters, thereby reducing communication rounds significantly and enhancing QFL performance using quantum natural gradient descent (QNGD) optimization. This paper demonstrates the feasibility of a QFL framework for predicting the presence of coronary heart disease, diagnosing whether a patient is suffering from diabetes or not, and differentiating malignant and benign cancer by distributing the UCI datasets unbalanced among healthcare institutions. The proposed framework has the potential to incorporate privacy, security, and the expedited processing of distributed data. QNGD outperformed classical GD by reducing communication rounds by a range of 5% to 60%. In addition to reducing the communication rounds by optimizing the QFL training algorithm and achieving quicker convergence, it also determines the important features regardless of the data imbalance among the clients.
With the growing adoption of computer-aided diagnostic and treatment recommendation systems in healthcare, it is essential to ensure both the accuracy and reliability of AI-enabled clinical decision support systems. In this study, we comprehensively examine existing model confidence calibration methods and propose an ensemble-based calibration approach for reliable predictions in clinical decision support systems (CDSSs). Specifically, we introduce an ENsemble-based Confidence-caLibrated deep neural network, ENCL-DNN, to improve respiratory disease screening using cough sounds. We also leverage local interpretable model-agnostic explanations to monitor the behavior of the CDSS, identifying the key features that contribute to its predictions and ensuring transparency in the diagnosis. By employing the ensemble-based calibration method, ENCL-DNN demonstrates superior performance on two publicly available respiratory audio datasets, Coswara and Cambridge, as evidenced by a 50% and a 28.74% reduction in Expected Calibration Error (ECE), respectively, compared to the uncalibrated baselines. Our experiments highlight the significance of well-calibrated deep neural networks in respiratory disease screening and the enhancement of reliability in mobile healthcare systems. By providing reliable and transparent predictions, ENCL-DNN has the potential to promote the wide adoption of AI-driven CDSSs and thereby improve patient outcomes through early diagnosis and intervention.
Oral cancer is the 13th most common cancer, affecting 380,000 people globally. The biggest challenge is that in its initial stage, cancer can go unnoticed until it reaches the most advanced, difficult-to-treat stages. Although a 90% survival rate is assured when diagnosed earlier, early-stage detection requires expensive periodic dental check-ups. Under these circumstances, converting a smartphone into a cancer screening tool has the potential to reduce cancer mortality. Mobile vision technology is a promising platform for early diagnosis of oral cancer. The aim of this research is to develop a telemedical mobile application for Oral Cancer Detection (OCD) using deep learning as a backend. This paper details experiments with various lightweight machine learning architectures, including MobileNetV3Large, which achieved an 84% accuracy, 86% sensitivity and 80% specificity on the test data. By incorporating the machine learning model into the smartphone app, users can capture or upload images for instant or offline screening, ensuring on-device processing that maintains privacy. This application promises to revolutionize oral healthcare accessibility and delivery.
Understanding working memory's genetic and neural bases is crucial for advancing cognitive neuroscience and identifying biomarkers for cognitive impairments, particularly in the older population. This study integrates SNP and neuroimaging data from the UK biobank to improve the classification of high vs. low working memory capacity and reveal genetic factors associated with brain structure. 1060 SNPs belonging to Protein-Protein Interaction networks of amyloid precursor protein and Aβ of Alzheimer's disease were integrated with latent features of whole brain gray matter density, extracted by a pre-trained CNN, via supervised contrastive learning. Our model effectively extracts latent representations of both modalities through enhancing genetic-imaging relation within individuals and within working memory groups, in contrast to across individuals and groups. Features derived from contrastive learning outperformed other baseline models in terms of classification. Sparse canonical correlation analysis was applied to the latent representations and uncovered significantly related genetic variants and brain regions. Genetic components highlight SNPs in genes FYN, RPL28, MAPT, enriched in the pathways of dendrite and synapse, among others. The linked brain regions support the cerebellum and striatum's role in cognitive functions. These findings provide new insights into the genetic and neural mechanisms underlying working memory, potentially guiding future research and therapeutic strategies for cognitive impairment.
The integration of model uncertainty quantification in clinical decision support systems, incorporating machine learning models, can augment the models’ reliability and robustness against domain shifts while also promoting user confidence and trust. In the present study, an uncertainty-informed active learning approach, leveraging Monte Carlo dropout for uncertainty estimation, is proposed towards the development of a deep learning model able to classify carotid ultrasound images as high-risk and low-risk for cardiovascular disease. An auxiliary dataset (CUBS) is employed for the initial model's development and fine-tuning as well as the optimization of the Monte Carlo dropout's hyperparameters. A dataset (87 B-mode ultrasound sequences) from ATTIKON hospital is subsequently utilized within the framework of active learning for model retraining based on the selection of the most informative samples according to the Monte Carlo dropout uncertainty estimation. In this context, the use of three active learning strategies is investigated, including uncertainty rank selection, pseudo-labeling for certain samples, and pseudo-labeling with variable sample weighting. The obtained results indicate that pseudo-labeling with variable sample weighting yields the best performance, achieving an AUC of 87.28% with only 21 annotated samples, which account for 30% of the total training data. Thus, this work provides evidence regarding the ability of uncertainty quantification and active learning to reduce labeling costs while maintaining model performance and enhancing the robustness and reliability of cardiovascular risk prediction models.
In this study, we focus on developing efficient calibration methods via Bayesian decision-making for the family of compartmental epidemiological models. The existing calibration methods usually assume the compartmental model is cheap in terms of its output and gradient evaluation, which may not hold in practice when extending them to more general settings. Therefore, we introduce model calibration methods based on a “graybox” Bayesian optimization (BO) scheme, more efficient calibration for general epidemiological models. This approach uses Gaussian processes as a surrogate to the expensive model, and leverages the functional structure of the compartmental model to enhance calibration performance. Additionally, we develop model calibration methods via a decoupled decision-making strategy for BO, which further exploits the decomposable nature of the functional structure. The calibration efficiencies of the multiple proposed schemes are evaluated based on various data generated by a compartmental model mimicking real-world epidemic processes. Experimental results demonstrate that our proposed graybox variants of BO schemes can further improve the calibration performance measured by the logarithm of mean square errors and achieve faster performance convergence in terms of BO iterations. We anticipate that the proposed calibration methods can be extended to enable fast calibration of more complex epidemiological models, such as the agent-based models.
Developing frameworks using high-dimensional magnetic resonance imaging (MRI) data to characterize underlying brain changes in neurological disorders is crucial and challenging. While deep learning models offer a better prediction, tracking automated higher-order explanations at the level of brain networks is harder in learned models. We introduce a novel constrained source-based salience (cSBS) framework to automatically learn and visualize multiple independently salient brain networks associated with clinical diagnostic assessments. This is achieved by performing active subspace learning (ASL) and spatially constrained independent component analysis (scICA) in the saliency space of trained convolutional neural networks (CNNs), such that the resultant components are interpretable in terms of brain network components from existing templates. By employing a robust analysis across repeated training scenarios for an Alzheimer's disease (AD) classification task, we visualize cSBS components via full-brain back-reconstruction. We show that the cSBS components and their corresponding loadings are consistent and relevant in terms of AD-related brain areas. Our approach is able to synthesize multiple objectives of utilization of high-dimensional MRI data for deep learning along with automated detection of low-dimensional representations of the consistently involved features in terms of intrinsically salient brain networks. Our framework of automated identification of consistent underlying brain subsystems associated with clinically observed assessments is an important step toward biomarker development for various clinically observed characteristics and disorders.
Scanpath prediction is crucial in the medical domain as it captures the visual attention patterns of experienced clini-cians, offering insights into diagnostic processes and enhancing training programs. Understanding where experts focus can lead to improved medical imaging interpretation and decision-making. However, scanpath prediction is extremely challenging due to the inherent noise in eye-tracking data, individual variability among clinicians, and the complexity of medical images. This work introduces a pioneering adaptation of the “Show, Attend and Tell” (SAT) [1] framework to analyze the gaze patterns of ophthalmologists on Optical Coherence Tomography (OCT) reports. Instead of using Convolutional Neural Networks (CNNs) for visual feature extraction, we integrate self-supervised learning through a Masked Autoencoder (MAE) [2]. The MAE re-constructs masked regions of OCT images, enabling the encoder to generate robust image representations despite limited labeling in medical imaging datasets. We trained separate LSTM models for each clinician to account for individual inspection patterns. The model demonstrated strong evaluation results, with the best-performing model achieving a ScanMatch score up to 0.5595 and Pearson correlation of up to 0.866 in predicting expert gaze on OCT reports. We showcase a downstream use-case of predicting the sequence of expert-fixated regions on an OCT report and visualizing these for ophthalmic resident education. Our findings highlight the framework's potential to enhance the understanding and emulation of expert-level diagnostic mechanisms' aiding in the explanation of AI-based predictions in the clinic and guiding novice residents in ophthalmic education, especially in resource-diverse environments with limited access to expert ophthalmologists or labeled datasets.
Parkinson's Disease (PD) is a neurological disease that progresses over time and causes severe motor symptoms. Therefore, treating PD requires constant patient monitoring, which may turn clinical practice overwhelming, preventing its practical implementation, and raising the need for patient monitoring outside the clinical setting. The iHandU system described in this paper fulfils this need by providing an objective way to quantify motor symptoms of PD in non-clinical settings. It integrates an innovative real-time assessment of the severity of motor symptoms based on signal processing and Machine Learning models that mimic the clinical severity classification scales used in practice and allows for a more continuous and personalized therapy planning and management by doctors, through the use of a web dashboard user-friendly interface. This system, recently tested at 5 patients' homes, has shown promising results as a PD patient management digital platform, reaching a usability score of 83.9% (A grade) based on the System Usability Scale (SUS). Such a level shows a strong alignment between user needs, expectations and functionalities. This study highlights the potential of the used system as a Patient Management Tool showing a case study from an ongoing clinical study. By giving additional information to the doctors with features beyond the semi-quantitative rating scales currently used, allowing a more optimized and continuous PD symptom management, it will be possible to advance PD management further.
This study presents a groundbreaking strategy for the homecare management of Obstructive Sleep Apnea (OSA). With an emphasis on patient empowerment through the integration of feedback management systems into sleep treatment, this study offers an innovative approach to the homecare management of obstructive sleep apnea (OSA). Fundamentally the research attempts to significantly increase, via tailored interventions, adherence to Continuous Positive Airway Pressure (CPAP) therapy. It accomplishes this by combining qualitative patient feedback with quantitative CPAP machine monitoring data to improve patient clustering and in turn treatment outcomes. The research methodology is comprehensive, encompassing various stages that include advanced patient grouping, continuous incorporation of new patient data, integration of feedback from surveys on both intervention and medical sleep, and a thorough cycle of interventions and evaluations. This iterative refinement process is essential as it allows for the dynamic updating of patient profiles and clustering based on evolving data and treatment responses. All these efforts are focused on fostering a more tailored approach to patient care. The creation of patient-centered treatment plans that maximize treatment efficacy by utilizing intervention repositories and data analytics is at the heart of this research. The study also explores how personalized care can improve CPAP adherence and underscores the need to tailor interventions and content based on patient feedback. This multi-layered approach aims at improving patient treatment adherence. It creates a more efficient patient-centered care model for individuals with OSA by continuously adapting and personalizing the interventions provided to each patient, thereby fostering a stronger relationship between patients and their treatments.
The gold standard for diagnosing dysphagia is the Videofluoroscopic Swallowing Study (VFSS). In patients with dysphagia, the invasion of food material into the airway is known as penetration-aspiration. Assessing this risk using VFSS is inherently subjective, with significant inter-patient and inter-rater variability. This article proposes an AI pipeline that introduces a novel approach in which bolus segmentation and airway detection are combinationally assessed to interpret frame-wise penetration-aspiration risk. The existing AI approaches rely on manual frame selection and overlook the clinical significance of bolus and airway. Additionally, addressing challenges posed by varying airway orientations, we develop an automated AI pipeline that tracks bolus and airway throughout VFSS videos. We curated a VFSS dataset and annotated one-third of the frames from 82 VFSS clips obtained from 40 patients due to a lack of benchmarks. Our approach involved comparing various segmentation models for bolus segmentation and fine-tuning object detection model for airway detection. The segmented bolus area and airway information are then processed to identify penetration-aspiration events. Our pipeline achieved a dice score of 0.80, a mean average precision of 0.93, and an accuracy of 89% in bolus segmentation, airway detection, and penetration-aspiration detection. Our pipeline could be effectively trained even with limited annotated frames. This saved clinicians time and also reduced the burden of manual annotation. These promising results have significant potential for assisting clinicians in assessing penetration-aspiration risk.
Cone beam computed tomography (CBCT) plays a vital role in the jaw lesions clinical diagnosis. However, different types of jaw lesions exhibit similar appearances in CBCT slices, while existing computer-aided diagnostic models, neither 2D slices with lack of distinctive features or 3D volume with highly redundant information, resulting in limited performance. For better detection of jaw lesions, we proposed a novel cross-view feature mining detection network based on reinforcement learning to adaptively extract the most characteristic slices from multi-views. Specifically, for every transverse plane slice in the CBCT image, policy network is designed to extract these corresponding sagittal and coronal slices with the most critical features for lesion detection. And then these slices are encoded and fused into the recognition branch which enhanced the overall performance. In our experiments, the proposed network reached detection recall of 79.7%, precision of 89.2%, and high average precision (AP) of 0.84 with an intersection-over-union (IoU) of 0.5. Quantitative results show that the proposed network is more effective than existing advanced approaches in the clinical detection and recognition of jaw lesions.
In recent years, the number of patients using continuous glucose monitoring (CGM) has increased. In addition to helping patients manage their disease, CGM produces time series data that can be used for integration in control algorithms, predictive models, and for retrospective analyses. Through feature extraction, many digital biomarkers can be derived from CGM. In this work, we provide a tool to extract features derived from the frequency domain. We first introduce a novel open-source Python library, CGM-Freq, for the analysis of CGM data in the frequency domain. We then test the library on real data. This work provides an open-source tool to further investigate the frequency domain of CGM signals.
EEG signal analysis and audio processing, though distinct in application, share inherent structural similarities in their data patterns. Recognizing this parallel, our study pioneers the application of two renowned audio processing models, PaSST and LEAF, to the realm of EEG signal classification. In our experiments, the adapted PaSST and LEAF models delivered exceptional performance on the Temple University Hospital Abnormal EEG Corpus (TUAB). Specifically, PaSST achieved an impressive accuracy of 95.7%, while LEAF registered 94.0%, both substantially outstripping previously established benchmarks. Such achievements underscore the potential of tapping into cross-domain models, particularly from the audio sector, for advancing EEG research. Notably, while these larger audio models brought about unparalleled results, maximizing their capabilities required addressing the limitations of available EEG data volume. Thus, we introduced innovative pre-training strategies derived from diverse datasets, further enhancing the performance efficacy. With these refinements, PaSST reached a landmark accuracy of 96.1% on the TUAB dataset, marking a significant stride forward in EEG signal processing. By leveraging the intrinsic resemblances between EEG and audio signals, we have successfully repurposed these audio models. We recommend further work devoted to the exploration of the transferability of machine learning audio techniques to healthcare time series tasks.
Radar has garnered great interest for remote health monitoring due to its ambient operation, effectiveness in the dark, and inability to make visual recordings of private scenes/faces. However, the current state-of-the-art in human activity recognition (HAR) focuses on the classification of persistent gaits, such as walking, and ignores the transitions between activities. The characterization of a person's ability to transition between postural states is highly individual and influenced by the person's physical and mental health. This paper presents a personalized, ethogram-based approach to HAR, which jointly characterizes the agility of transitions in addition to activity classification. We develop a multi-input multi-task learning (MIMTL) approach to simultaneously classify both human activity and agility. Our proposed approach yields accuracies of over 98% and 90% for the joint characterization tasks. Various interventions affecting gait are applied to show how the proposed approach can lead to agility-based detection of changes in gait.