Pulmonary diseases (PDs) are one of the third-largest causes of mortality worldwide. Recent developments on digital stethoscopes facilitate doctors in recording respiratory sounds (RSs) and monitoring various adventitious RSs, such as crackle, wheeze, or rhonchi, etc., in different breathing phases that are closely linked to disease-specific anatomical flaws. In contrast to the conventional approach, the present research first explores the prospect of utilizing breathing phase contextualized features from RSs for classifying the pulmonary diseases into obstructive (OPD) and restrictive (RPD) in a hierarchical fashion. Our proposed framework consists of four major stages: (a) pre-processing, (b) mel spectrogram-based time frequency response extraction, (c) categorizing the mel spectrograms into healthy or pathological classes by using our proposed self-organized operational neural network (SONN) architecture, thereafter, (d) these pathological RSs are transformed into equivalent breathing pattern signals, indicating volumetric flow of air during the inhalation and exhalation phase, which further used to find the onset and offset points of each breathing phases from the RSs and finally we classify these pathological cases into RPD or OPD, by employing the phase-specific features from the RSs through our developed dual-scale Kolmogorov-Arnold network (KAN) architecture. Experimental results illustrate that an overall accuracy (average) of 97.79% and 70.07% are achieved for stage-3, stage-4 classification using our proposed SONN and DS-KAN architecture. Furthermore, the utilization of the phase-specific or phase-contextualized features at stage-4 improves the class-wise accuracy of RPD and OPD significantly from 43.31%, 52.50% to 67.11%, 72.87%, respectively, which also suggests the efficacy of extracting such phase-contextualized features in classifying PDs.
The presence of motion artifact in the photoplethysmogram (PPG) signal makes it challenging for PPG-derived respiration (PDR) measurements in continuous health monitoring. In this work, a deep-learning framework for quality-aware PDR signal extraction, named QaPRExt, is described. This includes three key stages: pre-processing of the raw PPG signal, PPG snippet quality evaluation using a reservoir computing model, namely RCSQNet, and finally, PDR extraction using the proposed QaRExt module. QaPRExt was evaluated with a synthetic noisy signal using two noise insertion strategies, viz., additive and convolution, and the real-time noisy PPG signals with their corresponding reference respiration signals collected in the laboratory. The proposed framework achieves an average correlation of 0.91, 0.88 for PDR extraction and an average mean absolute error of 0.70, 0.10 breaths per minute, respectively, in estimating respiration rate from the noisy PPG segments for the 53 subjects of BIDMC and 30 subjects from the volunteers' database, respectively.
Heart murmurs are a prevalent indication of cardiovascular diseases and can offer valuable insights into early cardiac abnormalities. For the prompt diagnosis and treatment of any cardiac problems, their early identification and examination are essential for reducing the death risks. In this paper, we introduce the joint contrastive triplet loss for fusing the cross-domain and intra-domain embeddings obtained from the time and frequency signals to diagnosis of heart murmurs. For this, we examine the phonocardiogram (PCG) signals given in the circor digiscope heart sound database for three-class classification i.e., murmur present (HMP), murmur absent (HMA), and murmur unknown (HMU). The outcomes of the experiment demonstrate that the proposed work outperforms the existing works and attains an accuracy of 89.87% for three classes, along with additional performance measures such as 90.08% precision, 89.87% recall, and 89.89% F1-score.
Interstitial lung disease (ILD) represents a group of restrictive chronic pulmonary diseases that impair oxygen acquisition by causing irreversible changes in the lungs such as fibrosis, scarring of parenchyma, etc. ILD conditions are often diagnosed by various clinical modalities such as spirometry, high-resolution lung imaging techniques, crackling respiratory sounds (RSs), etc. In this letter, we develop a novel vision transformer (VIT)-based deep learning framework namely, ILD-VIT, to detect the ILD condition using the RS recordings. The proposed framework comprises three major stages: pre-processing, mel spectrogram extraction, and classification using the proposed VIT architecture using the mel spectrogram image patches. Experimental results using the publicly available BRACETS and KAUH databases show that our proposed ILD-VIT achieves an accuracy, sensitivity, and specificity of 84.86%, 82.67%, and 86.91%, respectively, for subject-independent blind testing. The successful onboard implantation of the proposed framework on a Raspberry-pi-4 microcontroller indicates its potential as a standalone clinical system for ILD screening in a real clinical scenario.
Non-invasive respiratory activity assessment, including airflow signal (AF)-derived vital extraction such as respiration rate (RR), tidal volume, expiratory flow rate, etc., and adventitious breathing event detection, are emerging research areas in continuous health monitoring. Recent studies have demonstrated a strong pathological correlation between AFs and respiratory sounds (RSs). In this work, for the first time, we present a unified deep learning framework, namely R2REst, for RR estimation by synthesizing equivalent electrical impedance tomography (EIT)-based AFs from RSs. The proposed framework comprises four major stages: pre-processing, mel spectrogram generation, mel spectrogram-vision transformer-based AF prediction, and lastly, RR estimation by analyzing the frequency spectrum of the predicted AF signal. Unlike prior works that utilizes bio-acoustic modalities other than RSs such as vocal, tracheal sounds in conjunction with pneumotachometry, or capnogram measurements, which are typically cumbersome to collect, our proposed R2REst uses publicly available RSs and AFs from the BRACETS dataset and achieves superior performance, with a mean square error and mean absolute error of 0.001, 0.003, and 0.010, 0.016 (in breaths per minute (BPM)) for tidal breathing followed by deep breathing (TBDB) and cough-speech (TBCS) induced cases, respectively.
Biometric authentication based on different physiological signals has attracted significant attention in the last decade due to advancements in wearable sensors and communication technologies apart from the traditional ways of recognition based on fingerprint, face. Recently, researchers have been allured by photoplethysmograph (PPG)-based biometric authentication owing to its non-invasiveness, low cost, and no use of adhesive, unlike widely used electrocardiogram (ECG)-based authentication. However, the identification accuracy (IA) severely deteriorates due to frequent motion artifacts. Further, it poses security issues due to few fiducial points and the compromise of live video of the subject in video-based PPG. Recently, few researchers have explored the use of seismocardiogram (SCG), another mechanical cardiac signal modality, for biometric authentication. However, these methods are unable to extract state-stable embeddings which impact the IA. To overcome, these issues, we propose a deep metric learning-based biometric authentication framework using SCGs. The proposed framework consists of the following stages: pre-processing, mel-spectrogram extraction, subject-specific-state-stable feature extraction using parameter-shared triplet neural network, embedding dictionary construction, and authentication using an intelligent cosine similarity-based authentication module. The proposed framework is evaluated using the only publicly available CEBS dataset under basal, music, and post-music states, and outperforms the existing works by achieving an IA and equal error rate (EER) of 99.79%, and 0.42%.
Respiratory disorders cause the death of around 8 million individuals globally. These disorders can be diagnosed by utilizing different clinical modalities, such as spirometry, FeNO test, and peak-flow measurements. However, chest auscultation-based respiratory sound (RS) examination with stethoscopes remains one of the most important clinical tools for identifying such disorders, since different adventitious RS cycles (ARSCs) correlate with various structural flaws of the lung. Therefore, an efficient identification of the ARSs using deep learning (DL)-based computerized algorithms, will help in the early detection of respiratory disorders. However, the RSs are highly susceptible to different types of interferences when collected in a real-world clinical setting, which might impair the precision of the ARSC classification algorithms and increase the chances of misdiagnosis. This work aims to investigate the effect of various auscultation-hindering noises, such as heart sounds (HSs), and background ambient noises on the automated classification of ARSCs using the DL algorithms. In this article, we have performed the ARSC classification experiments under various noisy circumstances by feeding the mel spectrogram representations into different pretrained audio neural networks (PANNs). Experimental results show that by using the proposed approach, we have achieved an ICBHI score (Mean [95% confidence interval]) of 81.14 ([80.18%-82.10%]), 77.08 ([74.36%-79.81%]), and 70.69 ([69.10%-72.28%]), while classifying the clean or reference RSs, RSs corrupted with HS, and hospital ambient noises, respectively.
Chronic obstructive pulmonary disease (COPD) is one of the most severe respiratory diseases, which can be diagnosed by several clinical modalities such as spirometric measures, lung function tests, parametric response mapping, wheezing events of lung sounds, etc. Since lung sounds are related to the respiratory irregularities caused by pulmonary illnesses, examining these sounds is more effective for identifying respiratory issues. In this paper, we propose a triplet time-frequency representation (TFR) driven multi-head self-organized operational neural network (MHSONN) for efficient detection of COPD-affected lung sound signals, which exploits the complex non-linear neural architecture in contrary to the linear neural perceptron model used by convolutional neural networks. The proposed framework consists of three stages: (a) pre-processing, (b) triplet TFR extraction, and (c) classification using the proposed MHSONN architecture. Upon experimental evaluation, the proposed work outperforms the existing noteworthy research works by achieving the highest performance rates of 99.81%, 99.85%, and 99.73% for accuracy, sensitivity, and specificity, respectively. The onboard implementation of the proposed framework on a Raspberry Pi-4 microcontroller also exhibits its viability for developing a point-of-care COPD detection system in real-world clinical scenarios.
Respiratory disorders have become the third largest cause of death worldwide, which can be assessed by one of the two key diagnostic modalities: breathing patterns (BPs) or the airflow signals, and respiratory sounds (RSs). In recent years, few studies have been conducted on finding correlation between these two modalities which indicate the structural flaws of lungs under disease condition. In this letter, we propose 'RS-2-BP': a unified deep learning framework for deriving the electrical impedance tomography-based airflow signals from respiratory sounds using a hybrid neural network architecture, namely ReSTL, that comprises cascaded standard and residual shrinkage convolution blocks, followed by feature refined transformer encoders and long-short term memory (LSTM) units. The proposed framework is extensively evaluated using the publicly available BRACETS dataset. Experimental results suggest that our ReSTL can accurately derive the BPs from RSs with an average mean absolute error of 0.024 +/- 0.011, 0.436 +/- 0.120, 0.020 +/- 0.011, 0.134 +/- 0.0680.024 +/- 0.011, 0.436 +/- 0.120, 0.020 +/- 0.011, 0.134 +/- 0.068 , and 0.031 +/- 0.0190.031 +/- 0.019 , respectively for five different tasks. Furthermore, these derived BPs can be used for extracting different respiratory vitals, identifying disease conditions efficiently, and retrieving salient breathing cycle information from the RSs.
Interstitial lung disease (ILD) is a collection of pulmonary adventitious conditions that induce scarring of the lung parenchyma, fibrosis, and inflammation. ILD encompasses over 200 chronic respiratory diseases that gradually damage the lung tissues and make it difficult to acquire adequate oxygen in the lungs. Therefore, it is essential to identify and diagnose diseases early to prevent their progression. ILDs are often characterized by abnormal respiratory sounds (RSs) such as crackles and squawks as a result of anatomical faults in the respiratory pathway produced by the disease. In this paper, for the first time, we propose a novel sinc convolution-based residual convolutional deep learning architecture, namely the ILDNet, for categorizing the ILD-affected RSs. The proposed framework comprises two major stages: (a) preprocessing of the input RS and (b) classification of the RSs using the proposed ILDNet. The proposed framework is extensively evaluated using the RSs from the publicly available BRACETS and KAUH datasets, and the experimental results show that our proposed ILDNet framework achieves an accuracy, sensitivity, and specificity of 81.25%, 78.85%, and 83.33%. These results also pave the way for future research on the potential use of RSs to identify reliable biomarkers for early-stage ILD identification.
Pulmonary disorders (PDs) are one of the substantial hazards to human life, which can be diagnosed by a variety of clinical modalities, including peak flowmeter and spirometry measurements, chest auscultation-based respiratory sound (RS) measurements, etc. Analyzing the acoustic RS measurements is one of the inexpensive yet essential diagnostic methods for identifying PDs as these RSs are correlated with structural flaws of the lungs that occur due to PDs. Additionally, the development of the digital stethoscope facilitates the continuous measurement of acoustic RSs of any individual, which can be exploited to identify a variety of PDs. In this article, we have proposed a triple time-frequency feature set driven triple-scale self-operational neural network (TS2ONN) architecture, namely Pulmo-TS2ONN, to classify a wide spectrum of PDs using the RSs. The proposed pulmo-TS2ONN comprises three major stages: preprocessing, triplet time-frequency feature set (TTFFS) extraction, and finally classification of seven class PDs by using TS2ONN architecture which utilizes the improved nonlinear neural backbone of self-operational neural network (SONN) in place of the linear neural architecture used in conventional deep learning (DL) networks. Upon experimental evaluation, the proposed framework outperforms the existing noteworthy research works by achieving the highest performance rates of 98.88%, 98.27%, and 99.84% for accuracy, sensitivity, and specificity, respectively. Lastly, the proposed framework is implemented on a quad-core ARM-A7-based Raspberry Pi-4 microcontroller, allowing the possibility of translating the research into real clinical situations for RS-based PD screening.
Radio-frequency (RF) fingerprint identification leverages the inherent unique transmitter hardware impairments to authenticate an emitter through an analysis of the received signal at the receiver. LoRa devices have gained widespread popularity in various Internet of Things (IoT) applications, primarily owing to their cost-effectiveness and impressive long-range communication capabilities. Recently, few studies have concentrated on the precise fingerprint identification of LoRa devices, however, these are focused on open-set authentication and security issues such as rogue device detection. In this paper, we propose a Self Operational Neural Network (SONN) learning framework, namely the LoRaSONN, for fingerprint identification of 60 distinct LoRa devices. The proposed framework exploits the spectrogram time-frequency representations (TFRs) obtained from the IQ samples of the preamble part of the received signals from the LoRa devices. The proposed framework achieves an overall accuracy of 97% over more than 8500 testing samples, across 60 devices from the same and different manufacturers evaluated on the recent publicly available LoRaRFFI dataset.
Chronic obstructive pulmonary disease (COPD) is a major public health concern across the world. Since it is an incurable disease, early detection and accurate diagnosis are very crucial for preventing the progression of the disease. Lung sounds provide reliable and accurate prognoses for identifying respiratory diseases. Recently, Altan et al. recorded 12-channel real-time lung sound dataset, namely RespiratoryDatabase@TR, for five different severity levels of COPD at Antakya State Hospital Turkey, and proposed deep learning frameworks for two-class COPD classification and five-class classification using a deep belief network (DBN) classifier and extreme learning machine (ELM) classifier respectively. A classification accuracy of 95.84% and 94.31% were achieved for two-class and five-class respectively. In this paper, we have proposed a melspectrogram snippet representation learning framework for both two-class and five-class COPD classification. The proposed framework consists of the following stages: preprocessing, melspectrogram snippet representation generation from lung sound and fine tuning of a pretrained YAMNet. Experimental analysis on the RespiratoryDatabase@TR dataset demonstrates that the proposed framework achieves accuracies of 99.25% and 96.14% for binary and multi-class COPD severity classification respectively, which is superior to the only existing methods proposed by Altan et al. for severity analysis of COPD using lung sounds.
Respiratory diseases are the world's third leading cause of mortality. Early detection is critical in dealing with respiratory diseases, as it improves the effectiveness of intervention, including treatment and reducing the spread. The main aim of this article is to propose a novel lightweight inception network to classify a wide spectrum of respiratory diseases using lung sound signals. The proposed framework consists of three stages: 1) preprocessing; 2) mel spectrogram extraction and conversion into a three-channel image; and 3) classification of the mel spectrogram images into different pathological classes using the proposed lightweight inception network, namely, respiratory disease lightweight inception network (RDLINet). Utilizing the proposed architecture, we have achieved a high classification accuracy of 96.6%, 99.6%, and 94.0% for seven-class classification, six-class classification, and healthy versus asthma classification. To the best of our knowledge, this is the first work on seven-class respiratory disease classification using lung sounds. Whereas, our proposed network outperforms all the existing published works for six-class and binary classifications. The suggested framework makes use of deep-learning methods and offers a standardized evaluation with strong categorization capabilities. In order to distinguish between a wide range of respiratory diseases, our study is a pioneering one that focuses exclusively on lung sounds. The proposed framework can be translated into real-time clinical application, which will facilitate the prospect of automated respiratory health screening using lung sounds.
Asthma is one of the most prevalent respiratory disorders, which can be identified by different modalities such as speech, wheezing of lung sounds (LSs), spirometric measures, etc. In this paper, we propose AsthmaSCELNet, a lightweight supervised contrastive embedding learning framework, to classify asthmatic LSs by providing adequate classification margin across the embeddings of healthy and asthma LS, in contrast to vanilla supervised learning. Our proposed framework consists of three steps: pre-processing, melspectrogram extraction, and classification. The AsthmaSCELNet consists of two stages: embedding learning using a lightweight embedding extraction backbone module that extracts compact embedding from the melspectrogram, and classification by the learnt embeddings using multi-layer perceptrons. The proposed framework achieves an accuracy, sensitivity, and specificity of 98.54%, 98.27%, and 98.73% respectively, that outperforms existing methods based on LSs and other modalities.
Asthma is one of the most severe chronic respiratory diseases which can be diagnosed using several modalities, such as lung function test or spirometric measures, peak flow meter-based measures, sputum eosinophils, pathological speech, and wheezing events of the lung auscultation sound, etc. Lung sound examinations are more accurate for diagnosing respiratory problems since these are associated with respiratory abnormalities occurred due to pulmonary disorders. In this paper, we propose a time-frequency domain self-operational neural network (SONN) based framework, namely, AsTFSONN, to efficiently categorize asthmatic lung sound signals, which uses the SONN-based heterogeneous neural model by incorporating an additional non-linearity into the neural network architecture, unlike the vanilla convolutional neural model that uses homogeneous perceptions which resemble the fundamental linear neuron model. The proposed framework comprises three major stages: pre-processing of the input lung sounds, mel-spectrogram time-frequency representation (TFR) extraction, and finally, classification using AsTFSONN based on the mel-spectrogram images. The proposed framework supersedes the notable prior works of asthma classification based on lung sounds and other diagnostic modalities by achieving the highest accuracy, specificity, sensitivity, and ICBHI-score of 98.50%, 98.80%, 98.11%, and 98.46%, respectively, using lung sounds as the input diagnostic modality, as evaluated on publicly available chest wall lung sound dataset.
Chronic obstructive pulmonary disease (COPD) is one of the most severe respiratory diseases and can be diagnosed by several clinical modalities such as spirometric measures, lung function tests, parametric response mapping, wheezing events of lung sounds (LSs), etc. Since LSs are related to the respiratory irregularities caused by pulmonary illnesses, examining them is more effective for identifying respiratory issues. In this letter, we propose a visibility graph (VG)-based adjacency matrix representation of LS in conjunction with a residual deep neural network (ResNet) for accurate detection of COPD, namely, the VGAResNet. The proposed framework comprises four stages: preprocessing, visibility graph creation, adjacency matrix (AdjM) generation, and lastly, classification of these AdjMs using the ResNet architecture. The proposed framework is extensively evaluated using the publicly available LS database and outperforms the existing noteworthy research works by achieving the highest performance rates of 95.13%, 96.33%, and 94.37% for accuracy, sensitivity, and specificity, respectively.
Lung sound is a non-invasive diagnostic tool for assessing several respiratory disorders, such as chronic obstructive pulmonary disorder (COPD). Due to the severe implications of COPD, it is essential to distinguish different severity levels of COPD. In this paper, we have utilized 12 channel lung sound data from RespiratoryDatabase@TR to distinguish two extreme severity levels of COPD namely, COPD-0 (lesser risk) and COPD-4 (very severe level). In this study, we have proposed an efficient COPD severity classification framework using variational mode decomposition (VMD); which involves four major stages: (a) preprocessing, (b) signal decomposition using VMD and feature extraction from the decomposed modes, (c) feature selection and ranking, (d) classification using machine learning (ML) classifiers. Employing this proposed technique, we have achieved high classification measures of 97.43%, 100%,93.75% for accuracy, specificity, and sensitivity respectively, which is superior to the only existing work presented by Altan et al.
Noninvasive monitoring of respiratory activity is an emerging research area in biomedical health monitoring. This article describes a neural network-based model, intelligent Photoplethysmography derived Respiration signal Extraction, and Tracking ( ${i}$ -PRExT). Here, an ensemble empirical mode decomposition (EEMD) is used to select the appropriate intrinsic mode functions (IMFs) through filtering in the respiration band and reconstruct by a linear weighted sum to obtain the photoplethysmography derived respiration (PDR) signal. The weight factors are derived by a multilayer perceptron neural network (MLPNN) fed with respiratory induced amplitude variation (RIAV) features extracted by a deep autoencoder (DAE). The tracking of respiration rate (RR) is done by an adaptive filter-based predictor. ${i}$ -PRExT was tested and validated with BIDMC data set under PhysioNet and 30 volunteers’ data collected under resting condition. The PDRs achieved over 90% correlation and low error (NRMSE~0.2) with reference respiration signal, while RRs have almost 100% correlation even under motion artifact (MA) corrupted photoplethysmography (PPG). The PDR shows improved performance, while RR tracking outperforms the published research on respiration signal extraction based on PPG.