The relevant model based on convolutional neural networks (CNNs) has been proven to be an effective solution in speech enhancement algorithms. However, there needs to be more research on CNNs based on microphone arrays, especially in exploring the correlation between networks associated with different microphones. In this paper, we proposed a CNN-based feature integration network for speech enhancement in microphone arrays. The input of CNN is composed of short-time Fourier transform (STFT) from different microphones. CNN includes the encoding layer, decoding layer, and skip structure. In addition, the designed feature integration layer enables information exchange between different microphones, and the designed feature fusion layer integrates additional information. The experiment proved the superiority of the designed structure.
Currently, a significant portion of acoustic scene categorization (ASC) research is centered around utilizing Convolutional Neural Network (CNN) models. This preference is primarily due to CNN's ability to effectively extract time-frequency information from audio recordings of scenes by employing spectrum data as input. The expression of many dimensions can be achieved by utilizing 2D spectrum characteristics. Nevertheless, the diverse interpretations of the same object's existence in different positions on the spectrum map can be attributed to the discrepancies between spectrum properties and picture qualities. The lack of distinction between different aspects of input information in ASC-based CNN networks may result in a decline in system performance. Considering this, a feature pyramid segmentation (FPS) approach based on CNN is proposed. The proposed approach involves utilizing spectrum features as the input for the model. These features are split based on a preset scale, and each segment- level feature is then fed into the CNN network for learning. The SoftMax classifier will receive the output of all feature scales, and these high-level features will be fused and fed to it to categorize different scenarios. The experiment provides evidence to support the efficacy of the FPS strategy and its potential to enhance the performance of the ASC system.
How to effectively extract features with high representation ability has always been a research topic and a challenge for classification tasks. Most of the existing methods mainly solve the problem by using deep convolutional neural networks as feature extractors. Although a series of excellent network structures have been successful in the field of Chinese ink-wash painting classification, but most of them adopted the methods of only simple augmentation of the network structures and direct fusion of different scale features, which limit the network to further extract semantically rich and scale-invariant feature information, thus hindering the improvement of classification performance. In this paper, a novel model based on multi-level attention and multi-scale feature fusion is proposed. The model extracts three types of feature maps from the low-level, middle-level and high-level layers of the pretrained deep neural network firstly. Then, the low-level and middle-level feature maps are processed by the spatial attention module, nevertheless the high-level feature maps are processed by the scale invariance module to increase the scale-invariance properties. Moreover, the conditional random field module is adopted to fuse the optimized three-scale feature maps, and the channel attention module is followed to refine the features. Finally, the multi-level deep supervision strategy is utilized to optimize the model for better performance. To verify the effectiveness of the model, extensive experimental results on the Chinese ink-wash painting dataset created in this work show that the classification performance of the model is better than other mainstream research methods.
Different artists have their unique painting styles, which can be hardly recognized by ordinary people without professional knowledge. How to intelligently analyze such artistic styles via underlying features remains to be a challenging research problem. In this paper, we propose a novel multi-task feature fusion architecture (MTFFNet), for cognitive classification of traditional Chinese paintings. Specifically, by taking the full advantage of the pre-trained DenseNet as backbone, MTFFNet benefits from the fusion of two different types of feature information: semantic and brush stroke features. These features are learned from the RGB images and auxiliary gray-level co-occurrence matrix (GLCM) in an end-to-end manner, to enhance the discriminative power of the features for the first time. Through abundant experiments, our results demonstrate that our proposed model MTFFNet achieves significantly better classification performance than many state-of-the-art approaches. In this paper, an end-to-end multi-task feature fusion method for Chinese painting classification is proposed. We come up with a new model named MTFFNet, composed of two branches, in which one branch is top-level RGB feature learning and the other branch is low-level brush stroke feature learning. The semantic feature learning branch takes the original image of traditional Chinese painting as input, extracting the color and semantic information of the image, while the brush feature learning branch takes the GLCM feature map as input, extracting the texture and edge information of the image. Multi-kernel learning SVM (supporting vector machine) is selected as the final classifier. Evaluated by experiments, this method improves the accuracy of Chinese painting classification and enhances the generalization ability. By adopting the end-to-end multi-task feature fusion method, MTFFNet could extract more semantic features and texture information in the image. When compared with state-of-the-art classification method for Chinese painting, the proposed method achieves much higher accuracy on our proposed datasets, without lowering speed or efficiency. The proposed method provides an effective solution for cognitive classification of Chinese ink painting, where the accuracy and efficiency of the approach have been fully validated.
Talent cultivation is the primary task of universities. Local general undergraduate colleges and universities should adhere to the basic guidelines of systematization, practicality and integration, continuously explore the concept of "student-centered" talent cultivation, and build a three-dimensional practical teaching system from three aspects: strengthening the planning and design of the three-dimensional practical teaching system; building an internal and external practical teaching platform; and improving the evaluation and guarantee system of practical teaching quality. The system of practical teaching quality evaluation and guarantee is improved. In order to improve the cultivation ability of applied talents in all aspects.
Automatic speech emotion recognition is a challenging task due to the gap between acoustic features and human emotions, which rely strongly on the discriminative acoustic features extracted for a given recognition task. We propose a novel deep neural architecture to extract the informative feature representations from the heterogeneous acoustic feature groups which may contain redundant and unrelated information leading to low emotion recognition performance in this work. After obtaining the informative features, a fusion network is trained to jointly learn the discriminative acoustic feature representation and a Support Vector Machine (SVM) is used as the final classifier for recognition task. Experimental results on the IEMOCAP dataset demonstrate that the proposed architecture improved the recognition performance, achieving accuracy of 64% compared to existing state-of-the-art approaches.
3D point cloud has gained significant attention in recent years. However, raw point clouds captured by 3D sensors are unavoidably contaminated with noise resulting in detrimental efforts on the practical applications. Although many widely used point cloud filters such as normal-based bilateral filter, can produce results as expected, they require a higher running time. Therefore, inspired by guided image filter, this paper takes the position information of the point into account to derive the linear model with respect to guidance point cloud and filtered point cloud. Experimental results show that the proposed algorithm, which can successfully remove the undesirable noise while offering better performance in feature-preserving, is significantly superior to several state-of-the-art methods, particularly in terms of efficiency.
Movie highlights are composed of video segments that induce a steady increase of the audience’s excitement. Automatic movie highlights’ extraction plays an important role in content analysis, ranking, indexing, and trailer production. To address this challenging problem, previous work suggested a direct mapping from low-level features to high-level perceptual categories. However, they only considered the highlight as intense scenes, like fighting, shooting, and explosions. Many hidden highlights are ignored because their low-level features’ values are too low. Driven by cognitive psychology analysis, combined top-down and bottom-up processing is utilized to derive the proposed two-way excitement model. Under the criteria of global sensitivity and local abnormality, middle-level features are extracted in excitement modeling to bridge the gap between the feature space and the high-level perceptual space. To validate the proposed approach, a group of well-known movies covering several typical types is employed. Quantitative assessment using the determined excitement levels has indicated that the proposed method produces promising results in movie highlights’ extraction, even if the response in the low-level audio-visual feature space is low.
3D point clouds have become increasingly popular in recent year due to the rapid development of low-cost 3D sensors. One of the most interesting challenges is to filter point cloud, which undoubtedly becomes a crucial part of the point cloud processing pipeline. Based on normal information, this paper proposes a simple but effective point cloud filter framework. In this framework, a kd-tree structure is constructed for representing point cloud to search neighborhood and estimate normal for each point at first. Then, iteratively performing the processing that a bilateral filter is applied to the normal field obtained from the previous iteration, using the same normal field as the guidance; afterward, adjusting point positions is performed depending on the filtered normals. Experimental results indicate the effectiveness of our algorithms.
The introduction of inexpensive 3D data acquisition devices has promisingly facilitated the wide availability and popularity of 3D point cloud, which attracts more attention on the effective extraction of novel 3D point cloud descriptors for accurate and efficient of 3D computer vision tasks. However, how to de- velop discriminative and robust feature descriptors from various point clouds remains a challenging task. This paper comprehensively investigates the exist- ing approaches for extracting 3D point cloud descriptors which are categorized into three major classes: local-based descriptor, global-based descriptor and hybrid-based descriptor. Furthermore, experiments are carried out to present a thorough evaluation of performance of several state-of-the-art 3D point cloud descriptors used widely in practice in terms of descriptiveness, robustness and efficiency.
Different from the western paintings, Chinese ink-wash paintings (IWPs) have own distinctive art styles. Furthermore, Chinese IWPs can be divided into two classes, Gongbi (traditional Chinese realistic painting) and Xieyi (freehand style). The extraction of Chinese IWP features with good classification results is challenging because of similar content. This paper presents a novel framework by combining a discrete cosine transformation (DCT) and convolutional neural networks (CNNs). In this framework, a CNN automatically extracts Chinese IWP features from a small subset of the DCT coefficients of an image instead of raw pixels commonly because of its good performance. We evaluate the proposed framework on a dataset including 1400 Chinese IWPs. Experimental results show that the proposed framework achieves competitive classification performance compared to existing benchmark methods.
This paper proposes a new approach to detect fire from a video stream. It takes full advantage of the motion feature and color information of fire. Firstly, motion detection using Gaussian Mixture Model-based background subtraction is applied to extract moving objects from a video stream. Then, multi-color-based detection combining the RGB, HSI and YUV color space is employed to obtain possible fire regions. Finally, the results of the above two steps are combined to identify the accurate fire areas. The experimental results obtained by applying this method on different fire videos show that the proposed method can achieve better effectiveness, adaptability and robustness.
In recent years, 3D point cloud has gained increasing attention as a new representation for objects. However, the raw point cloud is often noisy and contains outliers. Therefore, it is crucial to remove the noise and outliers from the point cloud while preserving the features, in particular, its fine details. This paper makes an attempt to present a comprehensive analysis of the state-of-the-art methods for filtering point cloud. The existing methods are categorized into seven classes, which concentrate on their common and obvious traits. An experimental evaluation is also performed to demonstrate robustness, effectiveness and computational efficiency of several methods used widely in practice.
This paper conducts revisions of the existing elastic-plastic contact models for a single asperity with the fractal surface topographies generated from W-M functions. The expressions of the critical parameters and the contact load-area relationship are deduced in the terms of fractal parameters, asperity's base diameter and material properties. Numerical analysis suggest that the contact models' critical parameters derived strictly in a convincing and easier way in this study are scale-dependent on the asperity's base size and determined jointly by fractal parameters, asperity's base diameter and material properties. The transformation condition for an asperity transfers from elastic to plastic contact and the contact load-area relationship are consistent with those in classical contact mechanics but contrary to the existing fractal contact models.
Based on the fractal theory, this paper quantitatively studies the influence of machining parameters on fractal characterization of typical metallic surface topography. Seven kinds of materials were selected including Q235, 45#, OCr18Ni9, HT200, T2, HPb59-1 and YL12. The surface topography was obtained under different machining parameters and the corresponding fractal dimension D and characteristic roughness Ra* were calculated using root mean square method. The conventional parameter roughness Ra was used for comparison to evaluate the influence of turning speed and feed rate and tries to reveal the feasibility of fractal parameters to characterize the machinability of materials and eventually establish the relationship between fractal parameters and material properties.
Objective: The technique of detecting moving regions has been playing an important role in computer vision and intelligent surveillance. Gaussian mixture models provide an advanced modeling approach for us. Although this method is very effective, it is not robust when there are some lighting changes and shadows in the scenes. This paper proposes a mixture model based on gradient images. Method: We firstly calculate gradient images of a video stream using the Scharr operator. We then mix RGB and gradient, and use a morphological approach to remove noise and connect moving regions. To further reduce false detection, we make an AND operation between two modeling results and result in final moving regions. Result: Finally, we use three video streams for analysis and comparison. Experiments show that this method has effectively avoided false detection regions resulting from lighting change and shadow, and improves the accuracy of detection. Conclusion: The approach demonstrates its promising characteristics and is more applicable in real-time detection.
We introduce a novel method for the consolidation of unorganized point clouds with noise, outliers, non-uniformities as well as sharp features. This method is feature preserving, in the sense that given an initial estimation of normal, it is able to recover the sharp features contained in the original geometric data which are usually contaminated during the acquisition. The key ingredient of our approach is a weighting term from normal space as an effective complement to the recently proposed consolidation techniques. Moreover, a normal mollification step is employed during the consolidation to get normal information respecting sharp features besides the position of each point. Experiments on both synthetic and real-world scanned models validate the ability of our approach in producing denoised, evenly distributed and feature preserving point clouds, which are preferred by most surface reconstruction methods.