Speech enhancement is crucial both for human and machine listening applications. Over the last decade, the use of deep learning for speech enhancement has resulted in tremendous improvement over the classical signal processing and machine learning methods. However, training a deep neural network is not only time-consuming; it also requires extensive computational resources and a large training dataset. Transfer learning, i.e. using a pretrained network for a new task, comes to the rescue by reducing the amount of training time, computational resources, and the required dataset, but the network still needs to be fine-tuned for the new task. This paper presents a novel method of speech denoising and dereverberation (SD&D) on an end-to-end frozen binaural anechoic speech separation network. The frozen network requires neither any architectural change nor any fine-tuning for the new task, as is usually required for transfer learning. The interaural cues of a source placed inside noisy and echoic surroundings are given as input to this pretrained network to extract the target speech from noise and reverberation. Although the pretrained model used in this paper has never seen noisy reverberant conditions during its training, it performs satisfactorily for zero-shot testing (ZST) under these conditions. It is because the pretrained model used here has been trained on the direct-path interaural cues of an active source and so it can recognize them even in the presence of echoes and noise. ZST on the same dataset on which the pretrained network was trained (homo-corpus) for the unseen class of interference, has shown considerable improvement over the weighted prediction error (WPE) algorithm in terms of four objective speech quality and intelligibility metrics. Also, the proposed model offers similar performance provided by a deep learning SD&D algorithm for this dataset under varying conditions of noise and reverberations. Similarly, ZST on a different dataset has provided an improvement in intelligibility and almost equivalent quality as provided by the WPE algorithm.
This paper presents a novel sound event detection (SED) system for rare events occurring in an open environment. Wavelet multiresolution analysis (MRA) is used to decompose the input audio clip of 30 seconds into five levels. Wavelet denoising is then applied on the third and fifth levels of MRA to filter out the background. Significant transitions, which may represent the onset of a rare event, are then estimated in these two levels by combining the peak-finding algorithm with the K-medoids clustering algorithm. The small portions of one-second duration, called ‘chunks’ are cropped from the input audio signal corresponding to the estimated locations of the significant transitions. Features from these chunks are extracted by the wavelet scattering network (WSN) and are given as input to a support vector machine (SVM) classifier, which classifies them. The proposed SED framework produces an error rate comparable to the SED systems based on convolutional neural network (CNN) architecture. Also, the proposed algorithm is computationally efficient and lightweight as compared to deep learning models, as it has no learnable parameter. It requires only a single epoch of training, which is 5, 10, 200, and 600 times lesser than the models based on CNNs and deep neural networks (DNNs), CNN with long short-term memory (LSTM) network, convolutional recurrent neural network (CRNN), and CNN respectively. The proposed model neither requires concatenation with previous frames for anomaly detection nor any additional training data creation needed for other comparative deep learning models. It needs to check almost 360 times fewer chunks for the presence of rare events than the other baseline systems used for comparison in this paper. All these characteristics make the proposed system suitable for real-time applications on resource-limited devices.
Reverberation results in reduced intelligibility for both normal and hearing-impaired listeners. This paper presents a novel psychoacoustic approach of dereverberation of a single speech source by recycling a pre-trained binaural anechoic speech separation neural network. As training the deep neural network (DNN) is a lengthy and computationally expensive process, the advantage of using a pre-trained separation network for dereverberation is that the network does not need to be retrained, saving both time and computational resources. The interaural cues of a reverberant source are given to this pretrained neural network to discriminate between the direct path signal and the reverberant speech. The results show an average improvement of 1.3% in signal intelligibility, 0.83 dB in SRMR (signal to reverberation energy ratio) and 0.16 points in perceptual evaluation of speech quality (PESQ) over other state-of-the-art signal processing dereverberation algorithms and 14% in intelligibility and 0.35 points in quality over orthogonal matching pursuit with spectral subtraction (OSS), a machine learning based dereverberation algorithm.
The outcome of source separation (SS) algorithms founded on spatial location cues, degrades in echoic conditions, due to corruption of these cues, that otherwise act as discriminative features for such systems. One of the solutions, for improving the performance of these systems, is to dereverberate the speech mixtures, ahead of the separation process. In this paper, we explore various dereverberation algorithms for preprocessing the reverberant speech mixture signal, before it can be given as an input to the model-based expectation-maximization source separation and localization (MESSL); a SS system based on location cues, working in varying echoic conditions. We then find the most optimum dereverberation algorithm, which can provide significant improvement in quality and intelligibility of the output speech signals from MESSL. It is found that the objective metrics advocate the use of the "weighted prediction error (WPE)" algorithm, providing an improvement of 3% in short term objective intelligibility (STOI) and 3.4 dB in signal to distortion ratio (SDR), while the subjective metrics favor the use of the "precedence effect (PE)" algorithm, which provides an improvement of 6% in average intelligibility score and 1% in average quality score, over the stand-alone MESSL system.
Deep learning models do not perform well if they are not trained for the acoustic conditions under which they have to operate. Due to this reason, the neural network based speech separation models specifically designed for anechoic conditions do not perform well in reverberant conditions. As training a deep neural network is a lengthy and computationally expensive process, training it every time for any change in the acoustic conditions is almost always impossible. This paper presents a comparative study of few of the state-of-the-art dereverberation algorithms and suggests the best among them, which enables a U-Net based speech separation model, trained in anechoic conditions, to work under different reverberant conditions for online and offline applications. The results show that dereverberating the audio mixtures before they enter the anechoic U-Net based speech separation network by spectral subtraction (SS) dereverberation algorithm improves the signal to distortion ratio (SDR) by almost 0.84 dB for online applications over the anechoic U-Net based speech separation model if it is directly exposed to reverberations. For offline applications, there is an average improvement of 2 dB in SDR and 4% in short term objective intelligibility (STOI), when the mixtures are dereverberated by the cascaded system containing the weighted prediction error (WPE) and the spectral subtraction (SS) dereverberation models. (C) 2021 Elsevier Ltd. All rights reserved.
In computer networks, vertices represent hosts or servers, and edges represent as the connecting medium between them. In localization, some special vertices (resolving sets) are selected to locate the position of all vertices in a computer network. If an arbitrary vertex stopped working and selected vertices still remain the resolving set, then the chosen set is called as the fault-tolerant resolving set. The least number of vertices in such resolving sets is called the fault-tolerant metric dimension of the network. Because of the variety of applications of the metric dimension in different areas of sciences, many generalizations were proposed, and fault tolerant is one of them. In this paper, we computed the fault-tolerant metric dimension of triangular snake, ladder, Mobius ladder, and hexagonal ladder networks. It is important to observe that, in all these classes of networks, the fault-tolerant metric dimension and metric dimension differ by 1.
Objective: To determine the association of carcinoembryonic antigen (CEA) levels in colorectal cancer with regard to the site and the stage of the tumor. Methodology: This cross-sectional study was conducted in Surgical unit 3, colorectal ward, Civil Hospital, Karachi, from January 2010 to June 2015. CEA level was monitored in all patient admitted to the unit. Normal pre-operative serum levels for CEA in of patients with colorectal carcinoma were considered <2.5 ng/mL for nonsmokers and <5.0 ng/mL for smokers. Any patient above this level was labeled as CEA positive case. Results: Out of 310 biopsy-proven patients, only 228 (73.6%) were found to have elevated CEA level at the time of diagnosis. The majority (72%) of the patient had adenocarcinoma and 83.4% were CEA positive. CEA positive cases with sigmoid colon was found significantly higher (86.8%) followed by rectum (79.5%) and descending colon (68.4%) (p<0.001). CEA positive cases were significantly higher in Dukes stage D and C, 88.5% and 86.02%, respectively (p<0.001). Conclusion: Patients with positive CEA levels were found to be associated more in the sigmoid colon and rectal carcinoma. While the DUKE D stage was associated with maximum positive cases of CEA levels.
This paper considers the problem of multiple human target tracking in a sequence of video data. A solution is proposed which is able to deal with the challenges of a varying number of targets, interactions, and when every target gives rise to multiple measurements. The developed novel algorithm comprises variational Bayesian clustering combined with a social force model, integrated within a particle filter with an enhanced prediction step. It performs measurement-to-target association by automatically detecting the measurement relevance. The performance of the developed algorithm is evaluated over several sequences from publicly available data sets: AV16.3, CAVIAR, and PETS2006, which demonstrates that the proposed algorithm successfully initializes and tracks a variable number of targets in the presence of complex occlusions. A comparison with state-of-the-art techniques due to Khan et. al., Laet et. al., and Czyz et. al. shows improved tracking performance.
The analysis of human sperm as part of infertility investigations or assisted conception treatments is a labor intensive process reliant upon the skill of the observer and as such prone to human error. Therefore, there is a need to develop automated systems that can adequately assess the concentration, motility and morphology of live sperm. This paper presents an algorithm for analyzing the morphology of motile sperm. Techniques for eliminating the background, segmentation of the cells and template matching techniques are used to analyze the morphology in two stages: first stage eliminates the immotile cells and at the second stage the morphology of the motile cells is analyzed. Results are presented with real sperm samples recorded in the andrology lab at the University of Sheffield. The performance of the proposed algorithm is analyzed in terms of accuracy and complexity. The proposed algorithm demonstrates high accuracy under variable conditions.
Objectives: To assess the role of incentive spirometry in trauma patients managed with tube thoracostomy in preventing postoperative pulmonary complications. Trauma injury accounts for 30% of all life years lost in the U.S.4 Chest trauma constitutes the major part of trauma. The majority of chest trauma requires careful surveillance and no surgical intervention. Tube thoracostomy may be required in the treatment of chest trauma. Incentive spirometer as a mechanical device helps in the lung expansion and encourages the residual collection either fluid or air to come out of the pleural space and drain it out through the chest tube. Materials and Methods: The study was conducted on patients coming with chest trauma to accident and emergency department of Civil Hospital Karachi, from January 2013 till July 2014. A total of 100 patients with chest trauma admitted through A&E department were enrolled in this research after taking written consent for tube thoracostomy and agreed to be the part of this research protocol. After assessment and consent the patients underwent tube thoracostomy under local anesthesia. The patients were divided into two groups by envelope technique, group A (n=50), who were advised to use incentive spirometer post procedure and the other was group B (n=50) who were not advised the use of incentive spirometer. Both the groups were then managed on same protocol of antibiotics and pain killers and were observed for the recovery in terms of removal of chest tube.
This paper presents a particle filter for multiple target tracking. The main contribution of this work is in the proposed likelihood function accounting for the interactions between the objects. The filter likelihood function is calculated by combining a social force model for human behaviour with image features such as colour and motion. The added social force model contributes to coping with occlusions between the objects. The performance of the developed algorithm is validated on real video data. The results demonstrate the algorithm accuracy during complex interactions between the objects.
Source separation algorithms that utilize only audio data can perform poorly if multiple sources or reverberation are present. In this paper we therefore propose a video-aided model-based source separation algorithm for a two-channel reverberant recording in which the sources are assumed static. By exploiting cues from video, we first localize individual speech sources in the enclosure and then estimate their directions. The interaural spatial cues, the interaural phase difference and the interaural level difference, as well as the mixing vectors are probabilistically modeled. The models make use of the source direction information and are evaluated at discrete time-frequency points. The model parameters are refined with the well-known expectation-maximization (EM) algorithm. The algorithm outputs time-frequency masks that are used to reconstruct the individual sources. Simulation results show that by utilizing the visual modality the proposed algorithm can produce better time-frequency masks thereby giving improved source estimates. We provide experimental results to test the proposed algorithm in different scenarios and provide comparisons with both other audio-only and audio-visual algorithms and achieve improved performance both on synthetic and real data. We also include dereverberation based pre-processing in our algorithm in order to suppress the late reverberant components from the observed stereo mixture and further enhance the overall output of the algorithm. This advantage makes our algorithm a suitable candidate for use in under-determined highly reverberant settings where the performance of other audio-only and audio-visual methods is limited.
This paper proposes an improved data association technique for dealing with occlusions in tracking multiple people in indoor environments. The developed technique can mitigate complex inter-target occlusions by maintaining the identity of targets during their close physical interactions. It can cope with the origin uncertainty of the multiple measurements and performs measurement to target association by automatically detecting the measurement relevance. The measurements are clustered by using the variational Bayesian method. An improved joint probabilistic data association filter (JPDAF) is proposed to associate measurements to targets with the aid of clustering process and extracting image features. A particle filter is used to track the multiple targets by exploiting the data association information. Both qualitative and quantitative evaluations are presented on real data sets which demonstrate that the proposed algorithm successfully tracks targets while solving complex occlusions.
A novel two stage data association technique for multi-target tracking is proposed which assigns multiple measurements to a target to mitigate information loss. At the first stage a variational Bayesian (VB) clustering technique is used which groups the measurements automatically into a determined number of clusters. In the second stage a belief propagation (BP) based cluster to target association method is proposed to assign multiple clusters to a target. This is achieved by exploiting the inter-cluster dependency information. The proposed technique is suitable to accommodate non-rigid targets such as humans. Both location and features of clusters are used to re-identify the targets when they emerge from occlusions. The proposed technique is compared with state of the art method due to Laet et al. and evaluations are presented on a real data set.
ABSTRACT... Objective: The aim of this study was to determine head-dipping exploratory test parameter as a measure of strongmodulating effect on brain and behavior. Design: It was an observational animal study. Setting: University of Karachi. Period: Jan 2004 toJuly 2006. Material & methods: In this present study, drugs used reserpine, nux- vomica; anacardium and chlorpromazine were widerange of pharmacological actions. We evaluate the effectiveness of these drugs as agents with modulating effect on brain and behavioraccessed by head dipping parameter. In this study, 25 mice were included belonging to both sexes. The study animals were divided intofive groups of five animals each. Four groups were given drugs and one group was kept as control. Mice (20-35g) of either sex were usedin this study. One group was kept as control for drugs. Mice were kept under room temperature. Tap-water was allowed ad-Libitum.30minutes after giving drugs, animals were observed for 10 minutes with two minutes of interval. Tablet crushed in 10ml of water, 1cc wasgiven. Screening method used was head dipping. Results: Strychnos Nux-Vomica when used in a dose of 0.07mg has strong action oncholinergic system, CNS activity and frequent head dipping (39.8±28.8) was observed. Rauwolfia serpentine is an active alkaloidparticularly present in reserpine (62.2±43.4) no significant head dipping effect was observed. Anacardium (37.2±28.6) &Chlorpromazine (39.4±32.4), show decrease effects. Keeping in view, the medicinal importance of these herbs, our present study wasdesigned to screen these drugs for CNS activity on albino mice.
An improvement is proposed in the audio-visual approach to solve the problem of source separation of physically moving speakers by exploiting multiple video cameras, a circular microphone array and robust spatial beamforming. The challenge of separating moving sources is that the mixing filters are time varying; as such the unmixing filters should also be time varying but these are difficult to determine from only audio measurements. Therefore the visual modality is utilized to track the direction of each speaker to the microphone array by using a Markov chain Monte Carlo particle filter (MCMC-PF). The proposed direction of arrival (DOA) tracker improves the computational complexity with respect to a previously employed 3-D multi-speaker position tracker. The DOA information is used in a robust least squares frequency invariant data independent (RLSFIDI) beamformer to separate the audio sources. Experimental results show that the proposed technique efficiently tracks the DOA with improved computational complexity and enhanced source separation.
In this paper a new combination of the model of the interaural spatial cues and a model that utilizes spatial properties of the sources is proposed to enhance speech separation in reverberant environments. The algorithm exploits the knowledge of the locations of the speech sources estimated through vision. The interaural phase difference, the interaural level difference and the contribution of each source to all mixture channels are each modeled as Gaussian distributions in the time-frequency domain and evaluated at individual time-frequency points. An expectation-maximization (EM) algorithm is employed to refine the estimates of the parameters of the models. The algorithm outputs enhanced time-frequency masks that are used to reconstruct individual speech sources. Experimental results confirm that the combined video-assisted method is promising to separate sources in real reverberant rooms.