针对低信噪比下的声音事件检测问题,提出基于能量压缩和灰度增强的多频带能量分布图的声音事件检测方法.将声音数据的gammatone频谱转成能量谱,对不同频带的能量进行不同比例的能量压缩,计算其多频带能量分布图,并对其进行灰度增强;对调整后的多频带能量分布图进行8×8的分块,对每一子块进行奇异值分解,提取主要数值作为声音事件的特征;利用随机森林分类器对特征建模与检测.实验结果表明,在低信噪比环境下,该方法具有良好的检测效果.
An improved noise reduction algorithm based on feedforward denoising neural network (DnCNN) is proposed for the noise removal problem of noisy seismic data. The previous DnCNN originally used for noise reduction of seismic data had the problem of large network depth and thus reduced training efficiency. The improved DnCNN algorithm was first proposed for the noise reduction of natural data sets, and this paper applies the algorithm to the noise reduction of seismic data after adjusting the relevant parameters. The analysis and comparison of the experimental results show that the DUDnCNN algorithm can remove noise with high efficiency, and the algorithm has certain feasibility and significance for further research in seismic data noise reduction.
Few shot learning aims to recognize novel categories with only few labeled data in each class. We can utilize it to solve the problem of insufficient samples during training. Recently, many methods based on meta-learning have been proposed in few shot learning and have achieved excellent results. However, unlike the human visual attention mechanism, these methods are weak in filtering critical regions automatically. The main reason is meta-learning usually treats images as black boxes. Therefore, inspired by the human visual attention mechanism, we introduce the salient region into the few shot learning and propose the SRFS-Net. In addition, considering the introduction of the salient region, we also modify the embedding function to improve the feature extraction capabilities of the network. Finally, the experimental results in miniImagenet dataset show that our model performs better in 5-way 1-shot than few shot learning models in recent years.
Most current multi modalities medical image registration approaches are concerned about registering one modality image to another. However, in the real world, medical image registration may be involved in multiple modes, not just two specific modalities. To this end, we propose a multi-contrast modalities medical image registration modal (Star-Reg net). It uses a single generator and discriminator for all contrasts of registrations amount several modalities. Furthermore, the proposed approach is trained in an unsupervised way, which alleviates the requirement of manual annotation data. The experiment on the IXI dataset demonstrates the Star-Reg net effectiveness in multi-contrast modalities medical image registration.
Due to existence of different environments and noises, the existing method is difficult to ensure the recognition accuracy of animal sound in low Signal-to-noise (SNR) conditions. To address these problems, we propose a double feature, which consists of projection feature and Local binary pattern variance (LBPV) feature, combined with Random forest (RF) for animal sound recognition. In feature extraction, an operation of projecting is made on spectrogram to generate the projection feature. Meanwhile, LBPV feature is generated by means of accumulating the corresponding variances of all pixels for every Uniform local binary pattern (ULBP) in the spectrogram. Short-time spectral estimation algorithm is used to enhance sound signals in severe mismatched noise conditions. In the experiments, we classify 40 kinds of common animal sounds under different SNRs with rain noise, traffic noise, and wind noise. As the experimental results show, the proposed framework consisting of shorttime spectrum estimation, double feature, and RF, can recognize a wide range of animal sounds and still remains a recognition rate over 80% even under 0dB SNR.
Environmental sound classification (ESC) is an important but challenging issue. In this paper, we propose a new deep convolutional neural network, which uses concatenated spectrogram as input features, for ESC task. This concatenated spectrogram feature we adopt can increase the richness of features compared with single spectrogram. It is generated by concatenating two regular spectrograms, the Log-Mel spectrogram and the Log-Gammatone spectrogram. The network we propose uses convolutional blocks to extract and derive high-level feature images from concatenated spectrogram, and each block is composed of three convolutional layers and a pooling layer. In order to keep depth of the network and reduce numbers of parameters, we use filter with a small receptive field in each convolutional layer. Besides, we use the average pooling to keep more information. Our method was tested on ESC-50 and UrbanSound8K and achieved classification accuracy of 83.8% and 80.3%, respectively. The experimental results show that the proposed method is suitable for ESC task.
As to the problem of sound event detection in low Signal-Noise-Ratio (SNR) noise environments, a method is proposed based on discrete cosine transform coefficients extracted from multi-band power distribution image. First, by using gammatone spectrogram analysis, sound signal is transformed into multi-band power distribution image. Next, 8x8 size blocking and discrete cosine transform are applied to analyze the multi-band power distribution image. Based on the main Zigzag coefficients which are scanned from the discrete cosine transform coefficients, features of sound event are constructed. Finally, features are modeled and detected through random forests classifier. The results show that the proposed method achieves a better detection performance in low SNR comparing to other methods.
论文针对各种背景声音中低信噪比声音事件的检测问题,提出把背景声音与声音事件混合,形成带噪声样本来训练分类器.在预处理阶段,使用基于经验模态分解与2-6级固有模态函数的投票方法,对背景声音与声音事件端点进行预测并估算信噪比.接着使用子带能量分布方法,提取声音数据的特征.最后,论文将背景声音与声音事件样本库中所有声音样本按照估算的信噪比相混合,生成混合声音特征训练多随机森林,用于低信噪比声音事件的检测.实验证实,所提出的方法可以用于各种声场景下低信噪比声音事件的检测,并能在信噪比为-5dB的情况下保持67.1%的平均检测率.
As a technology of context analysis, the detective method of polyphonic sound event detection has a widespread prospect of application. In this paper, a detective method of polyphonic sound event were proposed to resolve the challenge in IEEE DCASE2017 task 3 based on full convolutional DenseNet. Relevant results illustrated that the method is higher in F -score and lower in ER than the baseline method proposed by IEEE DCASE2017 based on DNNs, and also get higher performance -7.4% higher in F -score and 3% lower in ER than the best method in IEEE DCASE2017 challenge based on CRNN.
In this paper,we consider the influence of complex background environments on the automatic recognition of animal sounds with low signal-to-noise ratios(SNRs).We propose a method for identifying low-SNR animal sounds in various background environments.In this method,the sound signal is decomposed by a Bark scale wavelet packet,and the decomposition coefficient is used to generate a spectrogram of the reconstructed signal,which is projected onto a spectrogram to generate a Bark spectral projection(BSP)feature.Random forests(RF)are then used to identify animal sounds with low SNRs.We classified 40 common animal sounds with different SNRs in noise environments such as flowing water,highway,wind,and loud speech.The experimental results show that by combining the proposed meth-ods of short-time spectrum estimation,BSP,and RF in various background environments with different SNRs,the mean identification rate for animal noises can reach 80.5%.In addition,a recognition rate above 60%can be maintained even at –10 dB.
Impulse noise corruption in digital images frequently occurs because of errors generated by noisy sensors or communication channels, such as faulty memory locations in devices, malfunctioning pixels within a camera, or bit errors in transmission. Although recently developed big data streaming enhances the viability of video communication, visual distortions in images caused by impulse noise corruption can negatively affect video communication applications. In addition, ${{sparsity}}$, ${{density}}$, and ${{multimodality}}$ in large volumes of noisy images have often been ignored in recent studies, whereas these issues have become important because of the increasing viability of video communication services. To effectively eliminate the visual effects generated by the impulse noise from the corrupted images, this study proposes a novel model that uses a devised cost function involving semisupervised learning based on a large amount of corrupted image data with a few labeled training samples. The proposed model qualitatively and quantitatively outperforms the existing state-of-the-art image reconstruction models in terms of the denoising effect.
Impulse noise corruption in digital images frequently occurs because of errors generated in noisy sensors or communication channels, such as faulty memory locations in devices, malfunctioning pixels within the camera, and bit errors in transmission. Although the recently developed big data streaming enhances the viability of video communication, visual distortions in images that are caused by impulse noise corruption can negatively affect the viability of video communication applications. This paper develops a novel model that uses a devised cost function through semisupervised learning on a vast amount of corrupted image data with sparse labeled training samples to effectively remove the visual effects of impulse noise from these corrupted images. The experiments demonstrated that the proposed model significantly outperformed the existing state-of-the-art image reconstruction models when tested on a large image data set. To the best of our knowledge, this study is the first to specifically address the impulse noise removal problem for such large volumes of image data corrupted by high-density impulse noise.
Intelligent vehicles use advanced driver assistance systems (ADASs) to mitigate driving risks. There is increasing demand for an ADAS framework that can increase driving safety by detecting dangerous driving behavior from driver, vehicle, and lane attributes. However, because dangerous driving behavior in real-world driving scenarios can be caused by any or a combination of driver, vehicle, and lane attributes, the detection of dangerous driving behavior using conventional approaches that focus on only one type of attribute may not be sufficient to improve driving safety in realistic situations. To facilitate driving safety improvements, the concept of dangerous driving intensity (DDI) is introduced in this paper, and the objective of dangerous driving behavior detection is converted into DDI estimation based on the three attribute types. To this end, we propose a framework, wherein fuzzy sets are optimized using particle swarm optimization for modeling driver, vehicle, and lane attributes and then used to accurately estimate the DDI. The mean opinion scores of experienced drivers are employed to label DDI for a fair comparison with the results of our framework. The experimental results demonstrate that the driver, vehicle, and lane attributes defined in this paper provide useful cues for DDI analysis; furthermore, the results obtained using the framework are in favorable agreement with those obtained in the perception study. The proposed framework can greatly increase driving safety in intelligent vehicles, where most of the driving risk is within the control of the driver.
In order to rapidly determine the content of raw juice in blending pear juice by near-infrared spectroscopy (NIR), experiments using the same soluble solids content of fresh pear juice and juice powder were conducted. Four common swarm intelligence optimization algorithms, including Genetic Algorithm (GA), Particle Swarm Optimization (PSO), Glowworm Swarm Optimization (GSO) and Firefly Algorithm (FA), were combined with PLS to select wavelength variables. The results showed that the four kinds of models could remove most of the wavelength variables, and the FA-PLS model achieved the optimal performance, which simplified the model and improved the accuracy of prediction. Then, the successive projections algorithm (SPA) was used to select wavelength variables after Firefly Algorithm (FA). The results indicated the generalization ability were as follow: FA-PLS>PLS> FA-SPA-PLS>SPA-PLS. The root mean square errors of prediction (RMSEP) was 0. 029 1, 0. 033 3, 0. 033 9, 0. 137 0, respectively, and the corresponding wavelength variables number were 367, 765, 20, 18. The wavelength variables of SPA-PLS model were the least, but RMSEP was much higher than the other three models. Considering the prediction precision and the number of wavelength variables, the FA-SPA-PLS model was validly improved with less wavelength variables and higher prediction accuracy. This study provides a convenient way for rapid identification of blending fruit juice using NIR.
In recent years, food adulteration of various kinds has become a severe problem in food safety detection. In order to get rid of the limitations of traditional qualitative identification of new food adulteration, Fourier transform near-infrared spectroscopy (FT-NIR) was used to collect the spectrum ranging from 12400 to 4000 cm(-1). The pure ganoderma lucidum spore oil adulterated with peanut oil, corn oil, coix seed oil, and hogwash oil were investigated in this study, where the ganoderma lucid urn spore oil adulterated with hogwash oil was taken as the new category of food adulteration. Then, Multiple Relevance Vector Machine (RVM) classifiers were constructed with calibration samples of the first 4 categories. The prediction samples and ganoderma lucidum spore oil adulterated with hogwash oil were discriminated by the 4 kinds of classifier. In addition, the discriminated results were further verified with new clustering algorithm. Results showed that the discriminant accuracy of the first four categories was close to 93.75% with RVM classifier, but the ganoderma lucidum spore oil adulterated with hogwash oil was mistaken for pure ganoderma lucidum spore oil because of the limitations of model. So a new clustering algorithm based on local density and distance decision graph was applied to verify that. It was found that the cluster centers were 1 when the samples only contained pure ganoderma lucidum spore oil, however, the cluster centers were 2 when the samples mixed with pure ganoderma lucidum spore and adulterated with hogwash oil. The results demonstrated the FT-NIR in combination with RVM classifier and new clustering algorithm could be used for the identification of the adulterant in the pure ganoderma lucidum spore oil and qualitatively identify new category of food adulteration, providing a new method to solve the problem of food diversified adulteration.
A sound event recognition method based on optimized Orthogonal Matching Pursuit (OMP) is proposed for decreasing the influence of sound event recognition on various environments. Firstly, OMP is used for sparse decomposition and reconstruction of sound signal to decrease the influence of noise and reserve the main body of sound signal, where Particle Swarm Optimization (PSO) is adopted to accelerate the best atom searching in the process of sparse decomposition. Then, an optimized composited feature of Mel-Frequency Cepstral Coefficients (MFCCs), time-frequency OMP feature, and PITCH feature is extracted from reconstructed signal. Finally, Random Forests (RF) classifier is employed to recognize 40 classes of sound events in different environments and Signal-to-Noise Rates (SNRs). The experiment result shows that the proposed method can effectively recognize sound events in various environments.
针对生态自然环境中噪声对声音识别产生干扰的问题,提出利用混合优化的匹配追踪(MP)进行生态声音识别的方法.首先,使用萤火虫算法(GSO)和粒子群算法(PSO)对匹配追踪算法进行混合优化,加快匹配追踪有限次稀疏分解的速度并重构声音信号,保留高相关成分,滤除低相关噪声;其次,根据所选最优原子的时频信息结合MFCCs提取复合抗噪特征;最后,结合支持向量机(SVM)对40种生态声音在不同背景噪声与信噪比的情境下进行分类与识别.实验表明,优化后的匹配追踪算法去噪性能优于谱减法和小波去噪法.与常用的MFCCs方法相比,本方法对生态声音在不同信噪比下的识别性能有不同程度的改善,并且具有较好抗噪性.
随着Android系统迅速发展,Android应用软件广泛应用于人们日常生活与工作中,由此Android 智能手机中存储了用户许多敏感数据.而Android应用软件漏洞的存在,造成用户敏感数据被恶意窃取的矛盾日渐突出.为了更好地保护用户隐私信息,论文提出一种Android应用软件漏洞检测方法:首先,对不同漏洞进行归类整理分析;然后,根据不同类别漏洞特征采用相应的检测方式;最后,基于所提出的分析方法,实现了Android软件漏洞检测原型工具,并用该工具检测了886个Android应用软件样本,检测结果表明论文所提的方法简便有效.
The application of wavelength variable selection before partial least squares (PLS) regression to rapidly discriminate the adulteration of apple juice by Fourier transform near-infrared (FT-NIR) was investigated in this study. Successive projections algorithm (SPA) combined with four swarm intelligence optimization algorithms, including genetic algorithm (GA), particle swarm optimization (PSO), group search optimizer (GSO), and firefly algorithm (FA), was applied to extract effective wavelength variables. The results demonstrated that the variable number of SPA-PSO-PLS models was validly improved with a wavelength variable of four. The accuracy of model was satisfactory with the coefficients of determination of prediction (R 2 p = 0.9986) and good root mean square errors of prediction (RMSEP = 0.0628). The results suggested that SPA combined with swarm intelligence optimization algorithms for wavelength variable selection could rapidly and efficiently discriminate the adulteration of apple juice.
The paper proposes a robust ecological environmental sounds identification system by using optimized matching pursuit algorithm which is optimized by Glowworm Swarm Optimization(GSO)to improve the performance of sound recognition in real environmental noisy conditions. It uses the Matching Pursuit(MP) to decompose the sound signal sparsely, and reconstructs its inner structure to reduce the influence of the noise. GSO is employed to speed up the searching for the best atom in each process of decomposition. Different feature sets are extracted. As the performance of popular Mel-Frequency Cepstral Coefficients(MFCC)degrades due to sensitivity to noise, MP based time-frequency features and Pitch are adopted to supplant the MFCCs feature. Through the SVM classifier, 56 subclasses of 4 classes of ecological envi-ronmental sounds are tested for the comparison experiments in different environments under different SNRs. The experi-mental results show that this approach outperforms traditional methods of MFCCs and SVM, as the average identification accuracy and robustness for ecological environmental sounds are improved to a different degree, especially under the con-ditions of SNRs lower than 30 dB.