
Sleep is essential for human survival, and accurate sleep staging has significant potential for real-world applications. While deep learning methods for automatic sleep staging have shown promise, several challenges remain: 1) how to efficiently extract multimodal representations with strong generalization; 2) how to leverage the similarities and differences in multimodal signals for accurate sleep staging; and 3) how to utilize abundant unlabeled data to enhance model practicality. To address these challenges, we propose MultiConsSleepNet, a multimodal consistency-based sleep staging network. It integrates a unimodal feature extractor and a multimodal consistency extractor to capture local-domain representations of electroencephalogram (EEG) and electrooculogram (EOG) signals, enhancing feature consistency both within and across modalities. We also incorporate self-supervised contrastive learning strategies for both unimodal and multimodal consistency learning, thereby improving model generalization and transferability. Experiments on three datasets demonstrate that MultiConsSleepNet achieves state-of-the-art performance with limited labeled data, effectively leveraging unlabeled data to enhance application value.Clinical relevance— This study highlights the potential of a multimodal deep learning approach for automated sleep staging, achieving high performance with only 10% labeled data. Specifically, it attains an average F1 score of 74.2% and accuracy of 79.6% in cross-dataset scenarios. This method provides a valuable tool for clinicians, enabling efficient and reliable sleep staging with minimal labeled data. It can improve the speed and consistency of sleep disorder diagnoses while reducing reliance on manual analysis, thereby enhancing clinical workflow efficiency.
Sleep staging is an essential step in sleep quality assessment and sleep disorder diagnosis. However, most neural network approaches to automatic sleep staging have focused on improving network architectures and data augmentation techniques, the latent potential of sleep staging labels has not been fully explored. Existing approaches often use the one-hot encoding method, which encodes each sleep stage as an independent and parallel label, fail to capture the prior information and relationships among sleep stages, thereby limiting their potential. In this study, we propose a novel label encoding method called Binary Carry-Forward Encoding (BCFE) for sleep staging classification. Unlike traditional approaches with one-hot encoding, we first decompose the 5-class sleep staging task into four binary classification tasks aligned with the logical progression of sleep stages. By introducing 4-dimensional binary vectors, BCFE embeds potentially prior information into the label structure. Predicted labels are generated through Human-Like Progression (HLP) inference, and training is optimized with a Masked Binary CrossEntropy (MBCE) loss. These methods enables the network to better utilize nuanced label characteristics, enhancing classification performance. The evaluation was conducted across combinations of three models (DeepSleepNet, AttnSleep and SleePyCo) and two datasets (ISRUC-S3 and Sleep-EDF). The results showed that our proposed method demonstrated significant improvements compared to the traditional One-Hot + Cross-Entropy method, particularly for the DeepSleepNet architecture on the ISRUC-S3 dataset. Notably, it achieved nearly 7% increase in performance (Accuracy: 73.7% → 80.3%; Macro-F1: 71.3% → 78.8%; kappa: 66.7% → 74.6%).
Shape registration between patient-specific organ geometries and endoscopic camera images is crucial for image-guided surgery. Although machine learning methods have been investigated for this task, collecting training datasets remains challenging because of the limited use of three-dimensional (3D) imaging in surgical settings. Offline learning with synthetic images generated from preoperative 3D-CT data has been proposed; however, ensuring robust learning under domain discrepancies between synthetic and real images remains a major challenge. In this study, we propose a diffusion-based offline learning strategy to register the shape a liver mesh in laparoscopic camera images. Within this framework, semantic organ labels serve as an image feature shared between synthetic and real-world images, thereby mitigating the domain gap and facilitating accurate registration. During model training, Gaussian noise is introduced into the registration parameters, and the visual changes in the 2D organ labels guide the training to predict the noise. Experimental results demonstrate that the prediction accuracy surpasses that of conventional approaches, producing image overlays that visualize tumors and vascular structures for intraoperative guidance.Clinical Relevance: The proposed 2D/3D registration model generates image overlays that facilitate the visualization of tumors and vascular structures in laparoscopic surgery, thereby enabling highly precise surgical guidance.
This study presents an automatic segmentation algorithm for accurately delineating lungs with pathological attenuations in 3D computed tomography (CT) images. Classical image processing methods are applied, including intrinsic image decomposition (IID) filtering and wavelet transform. Two contour refinement strategies—convex hull and corner detection—are evaluated, with the convex hull approach demonstrating superior performance, achieving a Dice similarity coefficient (DSC) of 98% against expert manual delineations. When compared to a state-of-the-art deep learning (DL) model (TotalSegmentatorV2), the classical approach remains competitive, achieving lower Hausdorff distance (HD), indicating fewer extreme segmentation errors. These results suggest that while DL methods provide high segmentation accuracy, classical approaches that combine 2D and 3D processing still offer advantages in mitigating outlier errors and ensuring interpretability. Additionally, the robustness and interpretability of classical methods make them ideal for generating accurate, well-annotated training datasets for DL models, enhancing performance, especially in lung disease contexts. Future work could explore hybrid models that leverage the strengths of both classical and DL-based techniques for robust lung segmentation.
Medical image classification plays an increasingly vital role in identifying various diseases by classifying medical images, such as X-rays, MRIs and CT scans, into different categories based on their features. In recent years, deep learning techniques have attracted significant attention in medical image classification. However, it is usually infeasible to train an entire large deep learning model from scratch. To address this issue, one of the solutions is the transfer learning (TL) technique, where a pre-trained model is reused for a new task. In this paper, we present a comprehensive analysis of TL techniques for medical image classification using deep convolutional neural networks. We evaluate six pre-trained models (AlexNet, VGG16, ResNet18, ResNet34, ResNet50, and InceptionV3) on a custom chest X-ray dataset for disease detection. The experimental results demonstrate that InceptionV3 consistently outperforms other models across all the standard metrics. The ResNet family shows progressively better performance with increasing depth, whereas VGG16 and AlexNet perform reasonably well but with lower accuracy. In addition, we also conduct uncertainty analysis and runtime comparison to assess the robustness and computational efficiency of these models. Our findings reveal that TL is beneficial in most cases, especially with limited data, but the extent of improvement depends on several factors such as model architecture, dataset size, and domain similarity between source and target tasks. Moreover, we demonstrate that with a well-trained feature extractor, only a lightweight feedforward model is enough to provide efficient prediction. As such, this study contributes to the understanding of TL in medical image classification, and provides insights for selecting appropriate models based on specific requirements.
The diagnosis of liver diseases such as fibrosis and steatosis has traditionally relied on invasive techniques like liver biopsy. This study investigates hepatic biomarkers derived from magnetic resonance imaging (MRI) to non-invasively detect and stratify liver fibrosis and steatosis. Parameters such as liver stiffness, fat fraction, T2 relaxation time, and diffusion coefficients (apparent diffusion coefficient -ADC- and intravoxel incoherent motion -IVIM-) were analyzed in a cohort of 27 patients suspected of metabolic dysfunction-associated steatotic liver disease (MASLD), using a 3T MR scanner. Results show ADC as a promising fibrosis biomarker with a significant negative correlation to liver stiffness (r = -0.6801, p-value = 0.0000953), while T2 relaxation time correlated positively with fat fraction (r = 0.6087, p-value = 0.0007536). No significant correlation was found between PDFF and fibrosis, and IVIM parameters showed limited predictive value. These findings highlight the potential of MRI biomarkers for non-invasive liver disease assessment.Clinical Relevance— The findings of this study highlight the potential of non-invasive MRI-derived biomarkers for the diagnosis and stratification of liver fibrosis and steatosis. By reducing reliance on invasive procedures such as liver biopsy, these biomarkers could improve clinical outcomes, enable earlier detection, and minimize the risks and economical costs associated with traditional diagnostic methods
Implantable neurostimulators are key devices for treating neurological disorders such as Parkinson’s disease and epilepsy by delivering electrical pulses to specific neural tissues. The electrochemical impedance at the electrode-tissue interface significantly influences therapeutic effects and impacts the quality of closed-loop brain signal acquisition. While the real-time and accurate impedance is difficult to measure in vivo, since reference electrodes cannot be implanted for a long time. This study aims to quantify the differences in electrochemical impedance spectroscopy of various electrode configurations by utilizing equivalent circuit models. In this work, 2-electrode, 3-electrode and bipolar systems of clinical deep brain stimulation electrodes were tested in vitro, and the in-body practicality were discussed. 2-electrode and 3-electrode model showed similar impedance spectroscopy results, supporting the possibility of using simpler configuration in vivo. However, the study reveals obvious impedance differences in bipolar mode, almost doubling those in monopolar stimulation. These findings provide essential methodological support for modeling tissue interfaces in implantable neural stimulators, ensuring safety and therapeutic effectiveness. It also has key clinical implications for the development of microelectrodes and the advancement of closed-loop therapies.Clinical Relevance— This provides a theoretical reference for the accurate evaluation of the long-term in vivo impedance of implantable electrodes.
This study evaluates the performance and usability of the EXOTIC upper limb exoskeleton, designed to assist individuals with severe disabilities in performing activities of daily living. The EXOTIC system, featuring a tongue-based control interface and intelligent computer vision assistance, was tested by six individuals with tetraplegia. They needed a minute to drink from a bottle or toggle a light switch and about 1.5 minutes to eat a strawberry or scratch their heads on average. The study highlights the potential of the system to improve independence for individuals with complete functional tetraplegia while revealing important design considerations such as adaptability to different lean angles in the wheelchair and increased stiffness of the fingers of potential users.Clinical relevance—This study presents a hybrid deep learning model for automated wrist fracture detection, which could assist clinicians by enhancing diagnostic accuracy, efficiency, and timeliness.
Wearables continue to expand their ability to sample more signals from the brain and body while reducing form factors and limiting on-device computing to meet real-world comfort and power constraints. Where evolving multimodal sensing platforms can combine electrophysiological (ExG), optical, chemical, and mechanical sensors for holistic brain-body state estimation, understanding of cross-modal coupling mechanisms remains limited. Here, we collect and analyze bio-signals in a brain-body coupling experiment using simultaneous electroencephalogram (EEG), electrocardiogram (ECG), and respiratory measurements. Our experimental paradigm contrasted listening to self-selected music, preceded by tempo-matched isochronous cues. The results show a clear entrainment of cardiac rhythms with the underlying beat of auditory stimuli for both metronome cues and subject-selected music. Magnitude-Squared coherence analysis showed frequency coupling across recorded modalities, with ECG, EEG, and breath, exhibiting peak coherence near heard or underlying musical beat or its harmonics. We present minimally pre-processed data to motivate future methodological explorations of multiscale brain-body entrainment to rhythmic stimuli.Clinical relevance—These findings have implications for designing music therapy and biofeedback interventions that could adapt to ongoing brain-body rhythms. Understanding brain-body entrainment mechanisms and validating simplified measurement approaches could enable more accessible and effective rhythmic interventions in clinical settings, particularly for conditions benefiting from audio-based therapies or cardiac rehabilitation.
Magnetic particle imaging (MPI) as a new imaging method which senses magnetic nanoparticles (MNPs) concentration, and the simultaneous temperature mapping is a promising direction. MPI temperature imaging is still in the phantom stage, and a key challenge in achieving in vivo temperature imaging through MPI is how to obtain in vivo calibration parameters for temperature measurement. We used MPS/MPI dual-mode system to obtain the in vitro and in vivo phase difference of MNPs, indirectly obtained the in vivo phase lag parameters of MPI, and calibration of temperature and phase is achieved without heating the animal body. Finally, we achieved simultaneous imaging of in vivo temperature and MNP concentration. For MPI temperature reconstruction, the mean temperature deviation can reach 1.3 ° C in vitro and 2.23 ° C in vivo.Clinical relevance— The accuracy of our results is comparable to the level of temperature measurement with MRI and ultrasound. This demonstrates the potential of MPI multiparameter imaging and opens up new directions for MPI applications, particularly in MPI-guided precision magnetothermal therapy.
Photoplethysmogram (PPG) reflects the pulsation of the heart and is closely related to blood pressure (BP), but its generation mechanism is significantly different from that of BP. While prior research has emphasized their correlation, particularly for BP estimation from PPG, this correlation is often disrupted by individual differences, environmental factors, and signal noise, complicating the understanding of PPG-based BP estimation mechanisms. Causal inference methods have gradually gained attention in the field of physiological signal analysis and can reveal the causal relationship between variables. In this study, we aim to investigate the causal relationship between PPG and BP. We first using ensemble empirical mode decomposition to decomposed the PPG signal into frequency components and constructed counterfactual sequences via causal decomposition. We then analyze the causal relationship between these components and BP with a causal discovery algorithm DirectLiNGAM. It revealed that IMF4 and IMF5, among 10 intrinsic mode functions (IMFs), exhibited a strong causal link with BP. Validated on 164 subjects from VitalDB, BP estimation using these causally related components achieved an estimation error of 10.65 mmHg and 5.75 mmHg for SBP and DBP, respectively, outperforming estimates from the original PPG signal. This indicates that specific components of PPG rather than the original PPG signal can be better casually related with BP.
Brain-Computer interface (BCI), which translates neural activities into commands for external devices, holds significant promise for clinical rehabilitation and assisted movement for individuals with motor disabilities. Among various BCI paradigms, the steady-state visual evoked potential (SSVEP) based BCI garnered considerable attention due to its relatively stable and high-speed communication capabilities. However, a notable portion of the population, referred to as BCI illiteracy, struggles to effectively control BCI systems due to their inability to generate or modulate the neural patterns required for interaction. To address this issue, we proposed a user-centered approach using neurofeedback training (NFT) to improve individual’s performance on SSVEP-BCI. As a result, after a five-day training period, significant improvements in SSVEP-BCI performance were only observed in the training group rather than the control group without training. Notably, some subjects initially determined as BCI-illiterate also gained effective control of the BCI system after training. Further analysis revealed that the improvement of SSVEP-BCI performance had a close link with increased power and inter-trial phase coherence of the SSVEP response, indicating that NFT successfully strengthened the user’s task-related neural responses. These findings highlight the potential of NFT as a user-centered intervention to improve BCI control performance, offering a promising pathway to address BCI illiteracy and promote the broader application of BCI systems.Clinical Relevance— This study proposes an effective approach to enhancing the controllability of SSVEP-BCI systems, addressing the critical issue of individual control limitations. The developed method demonstrates significant clinical potential for promoting SSVEP-BCI applications, particularly in facilitating communication and device control for patients with severe motor impairments, such as amyotrophic lateral sclerosis (ALS) and locked-in syndrome (LIS).
Ultrasound elastography can serve as an advanced point of care tool in the field of sports medicine, specifically, to assess the player's injury risk and status at the playing or training arena. In this work, we present a simulation study of our proposed external shaker-based ultrasound elastography for musculoskeletal application. The algorithm estimates are validated with ground truth finite element model (FEM) results. Preliminary experimental proof of concept study is also presented on a regular tissue-mimicking phantom. The simulation results suggest that it is feasible to estimate the shear modulus using the proposed external-shaker based shear wave elastography technique. Further, the results depict elastography estimates to be affected by the imaging probe alignment with respect to the muscle fiber orientation.Clinical Relevance— The results from present work are encouraging towards developing portable point of care ultrasound tool with shear wave elastography for sports medicine application.
Diabetes has become an increasingly severe problem in China in recent years. Although some methods have been conducted to promote diabetes prevention before, they have certain limitations. To better activate diabetes management in China, this research proposes an AI-driven web application that offers personalized suggestions for diabetes prevention in the Chinese population. Its core is fine-tuning a large language model based on a tailored training dataset containing conversation prompts about diabetes lifestyle prevention suggestions. Its novelty includes the training dataset building on a self-defined diabetes prevention guideline, focusing on the Chinese lifestyle and, particularly, dietary habits. Therefore, it helps bridge gaps in existing prevention tools, such as the lack of personalization and cultural relevance. The system enables users to interact with AI in real time to receive advice and download chatting histories as well. This study demonstrates the feasibility of using AI to enhance early prevention strategies, contributing to advancing AI applications in healthcare. Its future implications may include expanding the app’s features for broader lifestyle management, as well as integrating with a community feedback mechanism and relevant healthcare systems to improve Chinese diabetes prevention.Clinical relevance—This project might be of interest to practicing clinicians, since it had the potential to improve the quality and accessibility of current Chinese diabetes prevention. In addition, if people’s diabetes prevention self-management is improved, the burden of clinicians may be reduced to some extent.
Neuromuscular conditions arising from accidents, injuries, or genetic neurodegenerative conditions can significantly affect the patient’s physical and mental well-being, and tracking their progression and treatment often requires continuous monitoring. Herein, we propose a self-powered IoT-enabled wearable electromyography (EMG) solution that integrates flexible solar panels embedded in everyday clothes. This study presents a self-powered wearable EMG system that integrates flexible solar cells, hydrogel-based electrodes, and a microcontroller-driven circuit. The solar cells deliver 1.763 W, 1.7 W, and 0.95 W in cloudy, sunny, and indoor conditions, respectively, ensuring sufficient power for charging a rechargeable battery and supporting extended operation. The innovative hydrogel-based electrodes outperform traditional silver-silver chloride (Ag/AgCl) electrodes, offering enhanced signal quality, skin-friendliness, sustainability, and conformability for comfortable long-term use. The system employs a multitasking algorithm to process muscle activity in real-time and uses WiFi to transmit the data to the Arduino cloud for remote monitoring. This user-friendly, sustainable solution overcomes the limitations of conventional EMG systems by enabling uninterrupted monitoring in remote areas and diverse environments. Its versatile design benefits healthcare, rehabilitation, and sports applications.Clinical relevance-The extracted EMG feature profile and pattern enable a deeper quantitative assessment of the overall muscle state. For instance, each parameter could be used to track an aspect of the muscle activity. Healthcare providers are empowered to diagnose diseases more accurately by matching each disorder with a certain profile, monitor rehabilitation progress by observing changing trends in a certain parameter, or, similarly, assess sports performance.
Retinopathy of prematurity (ROP) manifests clinically through abnormal development of the retinal capillary, ischemia, proliferative retinopathy, and retinal detachment, making it one of the leading causes of childhood blindness. Diagnosis of ROP from images of the full retinal fundus is challenging due to the presence of small peripheral lesions and underdeveloped ocular structures in premature infants. Current methods for fundus image analysis, based on Convolutional neural networks (CNNs) and transformers, often struggle to capture long-range dependencies or are limited by the complex computational demands. To address these challenges, we propose a CNN-Mamba interaction fusion network, CMIFNet, for diagnosing ROP. This network consists of a CNN-based branch for local feature extraction and a bidirectional state space model (BSSM) based branch for global information integration. We developed an interactive attention fusion module (IAFM) to enhance feature interaction and facilitate attention-based integration between the two branches. Our approach demonstrates promising recognition performance on a clinical ROP dataset.
There is strong clinical evidence that patients with depression have a high probability to exhibit cardiovascular disease (CVD) and vice versa. Thus, it is important to accurately identify these patients to provide optimal management of the comorbid conditions. Although the existing literature focuses on the development of artificial intelligence (AI) models for the diagnosis of CVD and/or depression, there is not currently any reported tool or system which integrates such models for clinical practice. In this work, we present a cloud-based platform to enable the easier, accurate, and cost-effective diagnosis of CVD and depression. The cloud-based platform is an integrated cloud-enabled computing unit that provides the execution of artificial intelligence computing algorithms along with data exchange services by utilizing the REST (Representational State Transfer) architecture. The platform enables the seamless and transparent interfacing of AI models and applications for the end-users. During the development a variety of state-of-the-art technologies and architectural models were integrated through a Payara Application Server, the Python programming environment (version 3) and a MySQL database server. Java SDK 11 was used for developing the full-stack API of the user interfaces and the back-end logic including the REST interfaces. The platform is hosted on a Linux Virtual Machine (VM). The development resulted in a cost-effective, accurate and efficient tool for the risk stratification of depression and CVD.Clinical Relevance— This is a state-of-the-art cloud-based platform for the risk stratification of CVD and depression. Example: Cardiologists and psychiatrists can use this platform to identify patients with CVD and depression and then prescribe more detailed examinations.
Deep learning has shown strong potential for detecting depression from speech. However, speaker-specific traits and language differences can cause feature shifts, limiting the model’s ability to generalize across individuals and languages. To overcome this challenge, we propose a Multi-Domain Feature Alignment (MDFA) model that minimizes the impact of individual and language differences on depression-related features. Using gradient reversal techniques, our approach reduces the feature encoder’s sensitivity to speaker and language-specific traits, allowing the model to focus on depression-related patterns for more effective classification. Experimental results on datasets from different languages, including DAIC-WoZ and Androids, demonstrate that the MDFA model significantly improves generalization across individuals and enhances performance in cross-language tasks.
Aboriginal pregnant women and new mothers face an increased risk of mental health issues, often stemming from historical trauma, including violence and discrimination. These challenges could contribute to complex trauma and adverse perinatal outcomes, highlighting the need for culturally sensitive care. However, non-Aboriginal clinicians often face barriers due to limited cultural knowledge, exacerbated by other factors such as time constraints for training and reliance on one-time training. Large Language Models (LLMs)-based chatbots offer the potential to support self-directed learning and enhance clinicians’ self-efficacy through interactive question and answer. However, LLMs also pose challenges, including hallucinatory responses, outdated knowledge, fictitious information, unverifiable references, and difficulty handling domain-specific queries. In this study, we aim to mitigate these challenges by developing a specialized chatbot for improving Aboriginal perinatal mental health question-answering. The chatbot integrates Retrieval-Augmented Generation (RAG) with a semantic search engine, enabling it to retrieve verified external knowledge and provide more accurate, contextually relevant responses without frequent retraining. We evaluate its performance against a baseline GPT-3.5-turbo model and compare LLMs integrated with different RAG techniques to assess improvements in accuracy and reliability.Clinical Relevance— This study shows the potential of the specialized RAG LLM-based chatbot to improve domain-specific, clinically relevant, and on-demand question-answering support for clinicians. By providing accurate, verified information through interactive responses, it may help bridge knowledge gaps, support self-directed learning, and complement existing training.