Cuffless blood pressure (BP) estimation has emerged as an alternative and effective technique to mitigate the limitations of conventional sphygmomanometers for prolonged BP monitoring. Cuffless BP can be estimated from cardiovascular measurements, including photoplethysmogram and electrocardiogram signals. Several machine learning-based BP estimation methods are available in the literature. However, the effectiveness of higher order spectral features, such as bispectrum, bicoherence, and trispectrum, for BP estimation has never been explored. This letter proposes efficient ensemble learning-based approaches for cuffless BP estimation utilizing the higher order spectrum of the cardio signals. The extracted higher order spectral features are incorporated in ensemble learning-based extra trees and categorical boosting models. These methods incorporate multiple weak learners to produce the desired estimates. The novel features capture the nonlinear interactions and phase coupling between different frequency components. The proposed techniques are validated using various international standards for cuffless BP estimation tasks. The experimental results demonstrate that the proposed methods outperform state-of-the-art techniques. Furthermore, the proposed machine learning models are executed on the Xilinx PYNQ-Z2 board to verify the hardware compatibility.
The decimation of high-frequency (HF) features during exhaustive low-dose computed tomography(LDCT) denoising introduces structural deformation. This paper addresses the aforementioned issue by introducing a GAN framework that provides novel adversarial training via discriminators in the wavelet and spatial domains. The wavelet-domain discriminator forces the generator to gain knowledge of the HF features via HF wavelet details (LH, HL, HH) and minimizes structural distortion. The generator network uses a spatial-domain discriminator to preserve local and global pixel correlations without altering the low-frequency (LF) features. Furthermore, we develop a generator network using a novel stationary wavelet-based residual block (SWTRB), which adaptively integrates spatial and frequency-domain information. In addition, we propose a wavelet-domain objective function on HF components, further improving the diagnostic quality of CT images. The experimental results demonstrate that the proposed method outperforms several state-of-the-art techniques on publicly available datasets, including "2016 NIH-AAPM-Mayo Clinic LDCT " and "Low-dose CT image and projection."
Hypertension or high blood pressure is a significant global health issue. Having high blood pressure is a big risk for conditions like coronary heart disease, including ischemic and hemorrhagic stroke. In general, the measurement of blood pressure is performed using a sphygmomanometer. However, this technique has several limitations in continuous and long-term monitoring due to bulky electronic devices with pneumatic systems (pump, valve, battery) to inflate and deflate the cuff. Cuffless blood pressure estimation has recently emerged as a good alternative to overcome these limitations. This paper proposes a machine learning-based approach using wavelet-based time-frequency features and adaptive boosting regression for cuffless blood pressure estimation from photoplethysmogram signals. The efficacy of the proposed approach is evaluated using various parameters concerning different state-of-the-art approaches. The proposed approach is found to perform better than various state-of-the-art methods. Furthermore, the proposed approach is implemented on the Xilinx PYNQ-Z2 board to validate the hardware compatibility.
Biomedical measurements are generally contaminated with substantial noise from various sources, including thermal noise, interference from other physiological signals, environment, or electrode movements. Complete restoration of biosignals from noisy measurements using linear filtering techniques is not feasible owing to the spectral overlapping problem. This paper proposes an arithmetic–geometric mean inequality-based robust denoising method. The proposed method incorporates a novel convexified cost function using the concept of majorization-minimization. A two-step algorithm is derived using the forward–backward splitting technique. An optimality condition is derived to set the hyperparameters of the new algorithm. The proposed convex optimization-based method effectively denoises cardiovascular signals, including both electrocardiogram and photoplethysmography. Furthermore, the efficacy of the proposed approach is verified over different datasets. The quantitative and qualitative results obtained using the proposed method demonstrate the superiority of the proposed method in biosignal denoising concerning state-of-the-art techniques.
The reconstruction of biomedical signals from noisy measurements has been an indispensable research topic. A majority of biosignals exhibit typical piecewise characteristics. The recovery of these piecewise biomedical signals embedded in noise through conventional nonlinear filtering schemes fails due to the lack of proper balance between strict sparsity and smoothness-inducing property of regularizers at large noise levels. This work proposes a nonlinear convex optimization-based filtering approach, which incorporates a Moreau envelope-based regularizer using the majorized version of the total variation function. The source signals are restored by exploiting their piecewise characteristics through a majorized cost function. The majorized functions provide some relaxation in solving non-convex functions. The relaxation of the stringent sparsifying penalty provides the balance between the smoothness property and the piecewise features of biosignals. The optimality criterion for the proposed method is analyzed in this work. Furthermore, we evaluate the new method using a standard IoT platform. The recovery performance of this method is found to be superior to various state-of-the-art techniques for piecewise synthetic and real-world physiological signals corrupted by additive noise.
Ambulatory electroencephalography (EEG) is a comparatively recent technology and a gold standard for the diagnosis of brain activity via prolonged EEG measurements. In general, motion artifact is the major concern during the ambulatory EEG signal acquisition. Removal of motion artifacts and other noise components from these EEG measurements has been a challenging task for researchers. In this paper, we propose a robust approach by employing an optimized Laplacian of Gaussian (LoG) filtering technique and empirical wavelet transform (EWT) to suppress the motion artifacts from ambulatory EEG measurements. The artifacts are suppressed via multi-resolution filtering of desired frequency components using optimized LoG filtering. Furthermore, the denoising performance of the proposed method is enhanced by optimizing the filter order. The efficacy of the proposed approach is evaluated using signal-to-noise ratio and correlation coefficient based metrics. The proposed method seems to have superior performance as compared to state-of-the-art techniques.
Schizophrenia is a long-term brain complication that impairs speech, behavior, mood, and cognitive function. A psychiatrist’s manual patient examination is highly subjective and labor-intensive. An automated classification tool is developed using 16-channel electroencephalogram measurements on the PYNQ-Z2 platform to diagnose schizophrenia patients. The linear discriminate analysis (LDA)-based feature extraction technique is incorporated in the extreme gradient boosting (XGBoost) to achieve promising classification accuracy. The LDA handles the multicollinearity (correlation between features) in the measurements. The XGBoost provides a more direct route to the minimum error with faster convergence and has built-in support for handling missing values, making it robust toward real-world data. A comprehensive study with state-of-the-art classification techniques is performed. The efficacy of the proposed approach is evaluated concerning accuracy, sensitivity, precision, specificity, and F1 score. The proposed method is found to outperform various state-of-the-art techniques. Furthermore, we implement our algorithms on the PYNQ-Z2 board, substantiating that the proposed method is hardware-compatible.
An ambulatory electrocardiogram (ECG) is a comparatively recent technology and a gold standard for diagnosing heart activities. Motion artifacts and baseline wander in ambulatory ECG measurements may hinder the detection of vital ST segments due to varying electrical isolines. It is challenging to completely suppress the motion artifact and baseline wander from ambulatory ECG measurements. This work proposes a novel attention-based deep recurrent neural network (DRNN) using bidirectional long short-term memory (Bi-LSTM) and total variation denoising (TVD) for motion artifact and baseline wander suppression. The attention mechanism, along with the Bi-LSTM, enhances the relevant features of the model by allocating a specific attention score, which improves the denoising performance. The proposed method is implemented on a Xilinx PYNQ-Z2 platform. The efficacy of the proposed technique is evaluated using the parameters, i.e., increment in signal-to-noise ratio (SNR), increment in correlation coefficient concerning ground truth signal, root mean square error (RMSE), maximum absolute distance (MAD), and cosine similarity index (CosSim). The proposed denoising network is validated using MIT-BIH arrhythmia and MIT-BIH NSTDB database. Compared to state-of-the-art techniques, the proposed method seems superior in suppressing motion artifacts and baseline wander from corrupted ambulatory ECG measurements.
Developing an automatic speech recognition (ASR) system for children’s speech is extremely challenging due to the unavailability of data from the child domain for the majority of the languages. Consequently, in such zero-resource scenarios, we are forced to develop an ASR system using adults’ speech for transcribing data from child speakers. However, the acoustic mismatch due to differences in formant frequencies and speaking rate between the two groups of speakers results in poor recognition rates as reported in earlier works. To reduce the said mismatch, an out-of-domain data augmentation approach based on formant and time-scale modification is proposed in this work. For that purpose, formant frequencies of adults’ speech data are up-scaled using warping of linear predictive coding coefficients. Next, the speaking rate of adults’ speech data is decreased through time-scale modification. Due to simultaneous altering of formant frequencies and duration of adults’ speech and then pooling the modified data into training, the ill effects of the acoustic mismatch due to the aforementioned factors get reduced. This, in turn, enhances the recognition performance significantly. Additional improvement in recognition rate is obtained by combining the recently reported voice-conversion-based data augmentation technique with the proposed approach. As demonstrated by the experimental evaluations presented in this paper, compared to an adult data trained ASR system, a relative reduction of 37.6 % in word error rate is achieved through data augmentation. Furthermore, the proposed approach yields large reductions in word error rates even under noisy test conditions.
Developing an automatic speech recognition (ASR) system for children’s speech is extremely challenging due to the unavailability of data from the child domain for the majority of the languages. Consequently, in such zero-resource scenarios, we are forced to develop an ASR system using adults’ speech for transcribing data from child speakers. However, differences in formant frequencies and speaking-rate between the two groups of speakers degrade recognition performance. To reduce the said mismatch, out-of-domain data augmentation approaches based on formant and duration modification are proposed in this work. For that purpose, formant frequencies of adults’ speech training data are up-scaled using warping of linear predictive coding coefficients. Next, the speaking-rate of adults’ data is also increased through time-scale modification. Due to simultaneous altering of formant frequencies and duration of adults’ speech and then pooling the modified data into training, the acoustic mismatch due to the aforementioned factors gets reduced. This, in turn, enhances the recognition performance significantly. Additional improvement is obtained by combining the recently reported voice-conversion-based data augmentation technique with the proposed ones. On combining the proposed and voice-conversion-based data augmentation techniques, a relative reduction of nearly 32.3% in word error rate over the baseline is obtained.
User applications such as voice-based web search, online learning, and video gaming require an effective speech recognition module to take user commands. Nowadays, even children are frequently using such tools, especially for online learning and gaming. This has increased the demand for developing a noise-robust automatic speech recognition (ASR) system that can effectively transcribe children’s data under varied ambient conditions. However, automatic recognition of children’s speech is extremely challenging due to the insufficiency of data from child speakers in the majority of the languages across the world. Consequently, in this zero-resource condition, we are forced to decode children’s speech on systems trained using adults’ data. However, the acoustic mismatch between adults’ and children’s speech, such as differences in pitch, formant frequencies, and speaking-rates, leads to highly degraded recognition performance. To enhance the recognition rate under zero-resource conditions, we have explored the role of formant and duration-modification-based out-of-domain data augmentation in this paper. For that purpose, the formant frequencies of the adults’ speech data are upscaled using warping of linear predictive coding coefficients. On pooling original and formant modified adults’ speech data into training, the mismatch in formant locations is reduced leading to better recognition performance. Further improvement in recognition rate can be achieved by simultaneously modifying the duration as well as the formant frequencies of the training data. This case of out-of-domain data augmentation has also been studied in this work and found to yield added gains. In addition to data augmentation, a noise- and pitch-robust front-end acoustic feature extraction approach exploiting higher-order spectral analysis (simple and cross-bispectrum) is also proposed in this paper. The proposed features are noise-robust due to the inherent immunity of the bispectrum towards additive noises. An added advantage of bispectrum is reduced pitch sensitivity as demonstrated in this work. This, in turn, helps alleviate the aforementioned pitch-induced acoustic mismatch. The experimental evaluations presented in this paper demonstrate that the use of proposed acoustic features, as well as the out-of-domain data augmentation techniques, are highly suited for zero-resource children’s speech recognition tasks under clean and noisy conditions.
A simplistic and an efficient approach toward color image contrast enhancement is proposed in this paper. The proposed approach effectively combines linear and principal component analysis (PCA) based fusion methods for contrast enhancement. In the proposed technique, we first obtain three different enhanced images using three pre-existing established techniques. Next, two of the enhanced images are combined using the PCA-based fusion method. The image resulting from first level of fusion is then combined with the third image using linear fusion technique to secure the final enhanced image. Since two different fusion techniques are employed, therefore, the proposed method is referred to as hybrid image fusion technique. Again, as we are employing three different images to secure the final enhanced image, there are 3 possible ways in which hybrid image fusion can be implemented. All the possible combinations of hybrid image fusion are compared with the existing techniques qualitatively as well as quantitatively in this paper. The proposed approach is observed to be better than the contemporary image enhancement methods.
A simple and effective approach for color image contrast enhancement is proposed in this paper. The proposed approach effectively combines linear and discrete wavelet transform (DWT) based fusion methods for contrast enhancement. In the proposed approach, we first derive three different enhanced images using three dominant existing techniques. Next, two of the enhanced images are combined using the linear fusion method. The image resulting from linear fusion is then combined with the third image through DWT-based fusion technique to obtain the final enhanced image. Since two-levels of fusion using two different techniques are employed, therefore, the proposed approach is referred to as two-level hybrid image fusion technique. Since three different images are employed to obtain the final enhanced image, there are 3 possible ways in which two-level hybrid image fusion can be performed. All the possible combinations of two-level hybrid image fusion are compared with the existing techniques qualitatively as well as quantitatively in this paper. The proposed approach is noted to be better than the existing image enhancement methods.