The diagnosis and screening of colon polyps are essential for the early detection of colorectal cancer. Polyps can be identified through colonoscopies before becoming cancerous, making accurate detection and prompt intervention critical for colorectal health. A comprehensive evaluation of deep learning models using colonoscopy images and comparisons with state-of-the-art models is presented in this study. A total of 7900 still and video sequence images from the PolypGen multicenter data set were used to train cutting-edge object detection models, including YOLOv5, YOLOv7, YOLOv8, and F-RCNN + ResNet101. In terms of accuracy, precision, recall, and mAP, the YOLOv8x model achieved the best performance with an F1 score of 0.9058, accuracy of 0.949, precision of 0.863, and mAP@0.5. The robustness of the model was further confirmed across varying patient demographics and conditions using the external Kvasir data set. To enhance interpretability, the EigenCam explainable AI (XAI) technique was used, offering visual insights into the model's decision-making process by highlighting the most influential regions in the input images.
A method for accurately estimating physiological signals from video streams at a minimal cost holds immense value, particularly in pre-clinical health monitoring applications. This technique is particularly indispensable in scenarios where traditional sensors, such as finger photoplethysmography (PPG), are not viable, such as in cases involving burn victims, premature infants, or individuals with sensitive skin. Remote photoplethysmography (rPPG) is a process of estimating PPG signals using video streams instead of traditional sensors. rPPG has thus been seen as a promising alternative to traditional PPG. As an alternative to using PPG for estimating oxygen saturation (SpO2), we propose ROSE-Net. ROSE-Net, trained on clinical PPG, was tested on an external rPPG dataset, PURE. The model achieved a mean absolute error (MAE) of 1.20 and a root mean square error (RMSE) of 1.86 on clinical PPG. When tested on rPPG, it exhibited an MAE of 1.95 and an RMSE of 2.46 in PURE, an MAE of 0.77, and an RMSE of 0.96 in ARPOS. These results demonstrate the model’s ability to estimate SpO2 levels within acceptable margins when applied to rPPG data. Consequently, rPPG presents a viable approach for estimating SpO2 levels, paving the way for non-contact health tracking applications.
Hypertension influences cardiovascular diseases, such as heart attacks and strokes. Blood pressure (BP) monitoring is essential for detecting hypertension and assessing its consequences. BP was traditionally measured using stethoscopes and pressure cuffs, which had several limitations. Additionally, automated blood pressure machines are not always accurate. Blood pressure measurement can be conducted more accurately and sensitively through a novel, non-invasive, and automated method. In this paper, a hybrid classification-mapping model is proposed to estimate Systolic (SBP) and Diastolic (DBP) blood pressure using 155 subjects from the University of New South Wales non-invasive BP (NIBP) dataset. In addition to exploring new beat-related features derived from oscillometric waveforms (OW), our study employs eight distinct feature ranking techniques to optimize the performance of different machine learning classifiers (K Nearest Neighbor (KNN), Ensemble KNN, Ensemble Bagged Tree, and Support Vector Machine (SVM)). As a comparison to existing methods for estimating DBP, which report a Mean Absolute Error (MAE) of 3.42 ± 5.38 mmHg, our approach achieves remarkably comparable results for estimating SBP, with an MAE of 1.28 ± 2.27 mmHg. Considering our promising results, implementing our methodology could provide a more reliable and convenient way to monitor blood pressure via remote healthcare.
An injury, chronic illness, obesity, infection, and more can negatively affect the hip joint. Surgery and implant placement are the standard treatments for moderate to severe hip issues. These treatments, however, can alter a patient's gait patterns. Gait patterns must be assessed clinically by a qualified physician and a specialized examination is required to detect and monitor these changes. By contrast, Machine Learning (ML) techniques assist in diagnosing a wide variety of anomalies and illnesses. In addition to being extremely accurate, it reduces subjectivity in clinical expert evaluations. Gait anomalies can also be quickly identified and monitored inexpensively and quickly using ML. Three open-source datasets (GaitRec, Gutenberg, and Orthoload) were utilized in this study for the gait cycle conditions for healthy control, hip surgery, and hip implant patients. This study classifies individuals into two classes: Healthy control/Gait Abnormality and three classes: Healthy control/ Hip surgery/ Hip implant by gait cycle conditions using only vertical ground reaction forces (vGRFs) from these datasets, which consist of 3D GRFs. The essential steps in data preparation include filtering, denoising, normalizing, resampling, and augmenting. The purpose of these efforts was to improve the model performance in classification and reduce biases. We used several feature extraction techniques, focusing on excluding highly correlated features. The final analysis utilized five widely recognized feature selection algorithms (Minimum Redundancy Maximum Relevance (mRMR), Neighborhood Component Analysis (NCA), Multi-Cluster Feature Selection (MCFS), Chi-square, and Relief) to arrange the features systematically. Based on a comprehensive examination of five machine learning classifiers (k-nearest neighbor (KNN), artificial neural network (ANN), decision tree (DT), support vector machine (SVM), and Na & iuml;ve Bayes (NB)), The KNN classifier exhibited the highest level of accuracy. The two and three-class classification's overall accuracy, precision, sensitivity, and F1 score are 95.48%, 96.13%, 95.48%, 95.63% and 89.18%, 89.30%, 89.18, and 87.15%, respectively. With the proposed solution, clinicians can more easily identify gait abnormalities based on vertical ground reaction forces.
The Magnetohydrodynamic (MHD) effect on the bloodstream, induced by the static magnetic field of Magnetic Resonance Imaging (MRI) devices, distorts Electrocardiogram (ECG) components, and poses challenges to cardiac gating and monitoring during MRI examinations. Restoring ECG components from MHD-induced artifacts has always been a challenge, primarily due to the influence of numerous device-induced and physiological variables and the absence of paired ground truth clean ECG for validation. In this research, at first, we employ unpaired restoration of MHD-corrupted 12-lead ECG signals using one dimensional (1D) Cycle Generative Adversarial Networks (CycleGANs) for channel and patient orientation-specific restoration. Subsequently, a 1D-segmentation-based comprehensive restoration scheme is proposed based on the phase 1 outputs to simultaneously restore 12-channel ECG signals, irrespective of internal and external factors affecting ECG. By employing the proposed Magnetohydrodynamic Generative Adversarial Network (MHDGAN) and MHD-to-ECG Network (MHD2ECGNet) for unpaired and comprehensive restoration, we achieved an overall spectral correlation improvement of 92.56% and 77.64%, respectively. In R-peak detection, MHDGAN demonstrated an overall precision, recall, and F1-score of 98.14%, 94.12%, and 96.01%, respectively. The extracted heart rate (HR) and heart rate variability (HRV) from the restored waveforms closely match the ground truth, as confirmed through quantitative and qualitative assessments. The proposed framework has proven effective in restoring MHD-corrupted ECG waveforms, even in the absence of paired ground truth labels and amidst compounding distortions from multiple sources.
Breathing conditions affect a wide range of people, including those with respiratory issues like asthma and sleep apnea. Smartwatches with photoplethysmogram (PPG) sensors can monitor breathing. However, current methods have limitations due to manual parameter tuning and pre-defined features. To address this challenge, we propose the PPG2RespNet deep-learning framework. It draws inspiration from the UNet and UNet + + models. It uses three publicly available PPG datasets (VORTAL, BIDMC, Capnobase) to autonomously and efficiently extract respiratory signals. The datasets contain PPG data from different groups, such as intensive care unit patients, pediatric patients, and healthy subjects. Unlike conventional U-Net architectures, PPG2RespNet introduces layered skip connections, establishing hierarchical and dense connections for robust signal extraction. The bottleneck layer of the model is also modified to enhance the extraction of latent features. To evaluate PPG2RespNet’s performance, we assessed its ability to reconstruct respiratory signals and estimate respiration rates. The model outperformed other models in signal-to-signal synthesis, achieving exceptional Pearson correlation coefficients (PCCs) with ground truth respiratory signals: 0.94 for BIDMC, 0.95 for VORTAL, and 0.96 for Capnobase. With mean absolute errors (MAE) of 0.69, 0.58, and 0.11 for the respective datasets, the model exhibited remarkable precision in estimating respiration rates. We used regression and Bland-Altman plots to analyze the predictions of the model in comparison to the ground truth. PPG2RespNet can thus obtain high-quality respiratory signals non-invasively, making it a valuable tool for calculating respiration rates.
Brain magnetic resonance imaging (MRI) offers intricate soft tissue contrasts that are essential for diagnosing diseases and conducting neuroscience research. At 7 Tesla (7T) magnetic field intensity, MRI enables increased resolution, enhanced tissue contrast, and improved SNR, compared to MRI collected from the commonly employed 3 Tesla (3T) MRI scanners. However, the exorbitant expenses associated with 7T MRI scanners hinder their broad use in research and clinical facilities. Efforts are underway to develop algorithms that can generate 7T MRI from 3T MRI to achieve better image quality without the need for 7T MRI machines. In this study, we have adopted a cycle consistent generative adversarial network (CycleGAN)-based approach for 3T MRI to 7T MRI translation, and vice versa, using a recently published dataset of paired T1-weighted MR images collected at 3T and 7T from a total of ten subjects. Various CycleGAN architectures were experimented with and compared on this dataset. The best performing CycleGAN architecture successfully produced the reconstructed images with a high level of accuracy based on different quantitative and qualitative evaluation criteria. Utilizing a post-processing technique, the best performing model generated 7T MRI from 3T MRI with a structural similarity index measure (SSIM) of 83.80%, peak SNR (PSNR) of 26.25, normalized mean squared error (NMSE) of 0.0088 and normalized mean absolute error (NMAE) of 0.0630. Utilizing CycleGAN to convert images from 3T to 7T MRI has shown a substantial improvement in MRI resolution, setting the stage for advancements in more informative and precise diagnostic imaging.
Patients with hyperglycemia require routine glucose monitoring to effectively treat their condition. We have developed a lightweight wristband device to capture Photoplethysmography (PPG) signals. We collected PPG signals, demographic information, and blood pressure data from 139 diabetic (49.65%) and non-diabetic (50.35%) subjects. Blood glucose was estimated, and diabetic severity (normal, warning, and dangerous) was stratified using Mel frequency cepstral coefficients, time, frequency, and statistical features from PPG and their derivative signals along with physiological parameters. Bagged Ensemble Trees outperform other algorithms in estimating blood glucose level with a correlation coefficient of 0.90. The proposed model’s prediction was all in Zone A and B in the Clarke Error Grid analysis. The predictions are thus clinically acceptable. Furthermore, K-nearest neighbor model classified the severity levels with an accuracy of 98.12%. Furthermore, the proposed models were deployed in Amazon Web Server. The wristband is connected to an Android mobile application to collect real-time data and update the estimated glucose and diabetic severity every 10-seconds, which will allow the users to gain better control of their diabetic health.
A method to accurately estimate physiological signals from video streams at a minimal cost is invaluable. The importance of such a technique in pre-clinical health monitoring cannot be understated. Remote photoplethysmography (rPPG) can be used as a substitute for finger photoplethysmography (PPG) when such sensors are not recommended, such as for burn victims, premature babies, and patients with sensitive skin. Good quality rPPG signal that is highly correlated to finger PPG can be used to estimate many vital health signs. In this work, a shallow encoder-decoder architecture, LGI-rPPG-Net is proposed. The proposed model aims to produce highly correlated rPPG signals which can be substituted for finger PPG. In the reconstruction of rPPG, the model achieved a very good Pearson's Correlation Coefficient (PCC), Root Mean Squared Error (RMSE), and dynamic time warping distance of 0.862, 0.148, and 0.699, respectively. This highly correlated rPPG was compared to finger PPG by calculating heart rate from rPPG and finger PPG. The model achieved a PCC of 0.984 and RMSE, and MAE of 2.91, 1.51 beats per minute (BPM), respectively. LGI-rPPG-Net model with video streaming to predict rPPG can thus be used as a replacement for finger PPG where in-contact collection is not feasible.
In the context of effective disease management for hyperglycemia patients, regular monitoring of blood glucose levels is imperative. However, traditional glucose monitoring methods suffer from invasiveness, discomfort, and potential infection. This paper introduces an innovative approach that utilizes non-invasive data sources derived from wearable devices, namely Photoplethysmography (PPG), Electrodermal Activity (EDA), and skin temperature (ST), in combination with user-provided food logs. The proposed model, MMG-Net, uses the three waveforms along with food features extracted from food logs to estimate blood glucose levels. MMG-Net delivers exceptional performance metrics, achieving a Mean Absolute Error of 13.51 mg/dL, a Mean Absolute Percentage Error of 12.57 %, and a Root Mean Square Error of 17.26 mg/dL. Notably, MMG-Net outperforms existing solutions in the estimation of blood glucose levels, solidifying its status as an innovative approach. The model's clinical precision is substantiated through Clarke Error Grid analysis, with a remarkable 99.43 % of predictions falling within clinically acceptable ranges. This paper presents a substantial advancement in non-invasive blood glucose monitoring, offering a promising avenue for enhanced disease management among hyperglycemic patients with only wearable devices.
Electrical and mechanical equipment with rotating parts often face the challenge of early breakdown due to defects in the gears or rolling bearings. Automated industrial systems can be significantly impeded by this type of fault in revolving components because of manual fault detection and the additional time required for repairing and replacing them. This research presents GearFaultNet, a novel, lightweight 1D Convolutional Neural Network (CNN)-based network, designed to detect gearbox faults. GearFaultNet can be an effective measure for real-time detection of sudden shutdowns and can alleviate downtime and system losses in the industrial aspect. The proposed framework involves the integration of four-channel vibration data from different loading conditions, which are preprocessed in the temporal domain and fed to GearFaultNet to classify the gearbox’s condition as either Healthy or Broken. The developed lightweight deep learning network has achieved higher accuracy than those proposed in existing literature. The overall accuracy achieved by this framework is 94.04%. This shallow network can also be applied to estimate other mechanical faults in different machinery.
Tuberculosis (TB) is a chronic infectious lung disease, which caused the death of about 1.5 million people in 2020 alone. Therefore, it is important to detect TB accurately at an early stage to prevent the infection and associated deaths. Chest X-ray (CXR) is the most popularly used method for TB diagnosis. However, it is difficult to identify TB from CXR images in the early stage, which leads to time-consuming and expensive treatments. Moreover, due to the increase of drug-resistant tuberculosis, the disease becomes more challenging in recent years. In this work, a novel deep learning-based framework is proposed to reliably and automatically distinguish TB, non-TB (other lung infections), and healthy patients using a dataset of 40,000 CXR images. Moreover, a stacking machine learning-based diagnosis of drug-resistant TB using 3037 CXR images of TB patients is implemented. The largest drug-resistant TB dataset will be released to develop a machine learning model for drug-resistant TB detection and stratification. Besides, Score-CAM-based visualization technique was used to make the model interpretable to see where the best performing model learns from in classifying the image. The proposed approach shows an accuracy of 93.32% for the classification of TB, non-TB, and healthy patients on the largest dataset while around 87.48% and 79.59% accuracy for binary classification (drug-resistant vs drug-sensitive TB), and three-class classification (multi-drug resistant (MDR), extreme drug-resistant (XDR), and sensitive TB), respectively, which is the best reported result compared to the literature. The proposed solution can make fast and reliable detection of TB and drug-resistant TB from chest X-rays, which can help in reducing disease complications and spread.
The continuous monitoring of respiratory rate (RR) and oxygen saturation (SpO2) is crucial for patients with cardiac, pulmonary, and surgical conditions. RR and SpO2 are used to assess the effectiveness of lung medications and ventilator support. In recent studies, the use of a photoplethysmogram (PPG) has been recommended for evaluating RR and SpO2. This research presents a novel method of estimating RR and SpO2 using machine learning models that incorporate PPG signal features. A number of established methods are used to extract meaningful features from PPG. A feature selection approach was used to reduce the computational complexity and the possibility of overfitting. There were 19 models trained for both RR and SpO2 separately, from which the most appropriate regression model was selected. The Gaussian process regression model outperformed all the other models for both RR and SpO2 estimation. The mean absolute error (MAE) for RR was 0.89, while the root-mean-squared error (RMSE) was 1.41. For SpO2, the model had an RMSE of 0.98 and an MAE of 0.57. The proposed system is a state-of-the-art approach for estimating RR and SpO2 reliably from PPG. If RR and SpO2 can be consistently and effectively derived from the PPG signal, patients can monitor their RR and SpO2 at a cheaper cost and with less hassle.
Diabetic sensorimotor polyneuropathy (DSPN) leads to pain, diabetic foot ulceration (DFU), amputation, and death. The diagnosis of advanced DSPN to identify those at risk is key to preventing DFU and amputation. Alterations in foot pressure and temperature may help to detect DSPN and the risk of DFU. We have applied a robust machine-learning approach to identify patients with severe DSPN using standing foot temperature maps generated using temperature sensor data. A robust shallow operational neural network model DSPNet is proposed. The study utilized a labeled dataset from the University Hospital Magdeburg, Magdeburg, Germany, consisting of temperature sensor data from eight different points on the foot in seating and standing positions in patients with severe DSPN (n =25) and healthy controls (n =18). The proposed network achieved an F1 score of 90.3% for identifying patients with DSPN and outperformed current state-of-the-art deep-learning network methods. This is the first of its kind of research where the results confirm that temperature maps are not only effective in the detection of those at high risk of DFU but also in identifying patients with severe DSPN. Such sensors could easily be incorporated into smart insoles.
Hypospadias is a complex multifactorial disorder that is impacted by both environmental and genetic variables, with varying degrees of severity and significant long-term functional effects if not treated properly. Typically, surgical guidelines founded on essential anatomical features aid in the decision-making process for urethroplasty. In the years thereafter, these methods have been tuned to the point that they may be utilized in clinical practice via risk assessment models, increasing diagnostic accuracy and streamlining workflow using artificial intelligence. It is also true that artificial intelligence (AI) has changed many facets of contemporary life, but its full potential has not yet been realized in pediatric urology and hypospadiology. Insufficient vast, high-quality, and publicly available datasets from a variety of institutes and locales may help explain the paucity of studies attempting to apply ML to hypospadiology. In this chapter, we will define key AI concepts, present a brief overview of AI’s use in medicine, explore exciting new AI-powered medical tools, and describe how to put AI to work solving specific medical issues. (See Video 10.1).
Gait analysis is helpful for rehabilitation, clinical diagnoses, and sporting activities. Among the gathered signals, ground reaction forces (GRF) may be used for assisting doctors in recognizing and categorizing gait patterns using Machine-Learning methods. In this study, GaitRec and Gutenberg databases were used, where GaitRec contains 2645 gait disorder (GD) patients and 211 Healthy Controls (HCs), and the Gutenberg database has 350 HCs. The combined database has HCs and four GD classes: hip, knee, ankle, and calcaneus. GD is an abnormality in the hip, knee, or ankle joints, whereas Calcaneus gait is calcaneus fractures or ankle fusions. We pre-processed the GRF signals, applied different feature extraction techniques, removed the highly correlated features, and ranked the features using three feature selection algorithms. K-nearest neighbour model (KNN) showed the top performance in terms of accuracy in all experiments. Four different experimental schemes were pursued: (i) 6 binary classifications; (ii) 1 three-class classification; (iii) 2 four-class classifications; (iv) one five-class classification. We also compared the performance of vertical GRF with three-dimensional GRF. We found that using three-dimensional GRF increased the overall performance. Furthermore, it is found that time-domain and Wavelet features are among the most useful in identifying gait patterns. The findings show promising performance in automated gait disorder classification.
Magnetic resonance imaging (MRI) is commonly used in medical diagnosis and minimally invasive image-guided operations. During an MRI scan, the patient's electrocardiogram (ECG) may be required for either gating or patient monitoring. However, the challenging environment of an MRI scanner, with its several types of magnetic fields, creates significant distortions of the collected ECG data due to the Magnetohydrodynamic (MHD) effect. These changes can be seen as irregular heartbeats. These distortions and abnormalities hamper the detection of QRS complexes, and a more in-depth diagnosis based on the ECG. This study aims to reliably detect R-peaks in the ECG waveforms in 3 Tesla (T) and 7T magnetic fields. A novel model, Self-Attention MHDNet, is proposed to detect R peaks from the MHD corrupted ECG signal through 1D-segmentation. The proposed model achieves a recall and precision of 99.83% and 99.68%, respectively, for the ECG data acquired in a 3T setting, while 99.87% and 99.78%, respectively, in a 7T setting. This model can thus be used in accurately gating the trigger pulse for the cardiovascular functional MRI.
Diabetes mellitus (DM) is one of the most prevalent diseases in the world, and is correlated to a high index of mortality. One of its major complications is diabetic foot, leading to plantar ulcers, amputation, and death. Several studies report that a thermogram helps to detect changes in the plantar temperature of the foot, which may lead to a higher risk of ulceration. However, in diabetic patients, the distribution of plantar temperature does not follow a standard pattern, thereby making it difficult to quantify the changes. The abnormal temperature distribution in infrared (IR) foot thermogram images can be used for the early detection of diabetic foot before ulceration to avoid complications. There is no machine learning-based technique reported in the literature to classify these thermograms based on the severity of diabetic foot complications. This paper uses an available labeled diabetic thermogram dataset and uses the k-mean clustering technique to cluster the severity risk of diabetic foot ulcers using an unsupervised approach. Using the plantar foot temperature, the new clustered dataset is verified by expert medical doctors in terms of risk for the development of foot ulcers. The newly labeled dataset is then investigated in terms of robustness to be classified by any machine learning network. Classical machine learning algorithms with feature engineering and a convolutional neural network (CNN) with image-enhancement techniques are investigated to provide the best-performing network in classifying thermograms based on severity. It is found that the popular VGG 19 CNN model shows an accuracy, precision, sensitivity, F1-score, and specificity of 95.08%, 95.08%, 95.09%, 95.08%, and 97.2%, respectively, in the stratification of severity. A stacking classifier is proposed using extracted features of the thermogram, which is created using the trained gradient boost classifier, XGBoost classifier, and random forest classifier. This provides a comparable performance of 94.47%, 94.45%, 94.47%, 94.43%, and 93.25% for accuracy, precision, sensitivity, F1-score, and specificity, respectively.
Respiratory ailments are a very serious health issue and can be life-threatening, especially for patients with COVID. Respiration rate (RR) is a very important vital health indicator for patients. Any abnormality in this metric indicates a deterioration in health. Hence, continuous monitoring of RR can act as an early indicator. Despite that, RR monitoring equipment is generally provided only to intensive care unit (ICU) patients. Recent studies have established the feasibility of using photoplethysmogram (PPG) signals to estimate RR. This paper proposes a deep-learning-based end-to-end solution for estimating RR directly from the PPG signal. The system was evaluated on two popular public datasets: VORTAL and BIDMC. A lightweight model, ConvMixer, outperformed all of the other deep neural networks. The model provided a root mean squared error (RMSE), mean absolute error (MAE), and correlation coefficient (R) of 1.75 breaths per minute (bpm), 1.27 bpm, and 0.92, respectively, for VORTAL, while these metrics were 1.20 bpm, 0.77 bpm, and 0.92, respectively, for BIDMC. The authors also showed how fine-tuning a small subset could increase the performance of the model in the case of an out-of-distribution dataset. In the fine-tuning experiments, the models produced an average R of 0.81. Hence, this lightweight model can be deployed to mobile devices for real-time monitoring of patients.