
In the originally published chapter 16 the first- and last name order of one of the authors was incorrect. The author’s name has been corrected as “Ghembaza, Moulay Ibrahim El-Khalil”.
In an older version of this paper, there was an orthographical error in the title. This has been corrected.
Language is an effective way to express human emotions, but emotions are difficult to describe and judge with computers, so it is an important task to analyze their emotions through speech. We summarized the current situation of speech emotion recognition from five aspects: the development of speech emotion recognition, emotion description model, emotion speech database, feature extraction, and emotion recognition algorithm. By summarizing and analyzing these five aspects, we can predict the future development trend of speech emotion recognition and improve recognition accuracy combined with other information for joint analysis. In this paper, we summarized the commonly used emotion database, compared the popular attention mechanism in speech emotion recognition, and finally, proposed the prospect of speech emotion recognition.
This paper proposes a method of crowd counting. We use ResNeSt-50 as the backbone network of YOLOv3. After the backbone network, we add SPP (Spatial Pyramid Potential) and PANet (Path Aggregation Network) to enhance the receptive field of convolutional neural network and improve the accuracy of stream of people or crowd counting in real application scenarios. In the application scenario of high-density crowd counting, an improved VGG network is used to design a deep network to capture high-level semantic information. At the same time, a shallow network is constructed to detect the head blob of people far away from the camera. The deep network and the shallow network are combined to detect high-density crowd. Finally, through the effective fusion of the above two network models, the accuracy and applicability of the algorithm are further improved. It can improve the detection accuracy in the case of small number of people and occlusion, and effectively reduce the estimation error in the scene with high density crowd.
Because pedestrians are always in the active state, each target is at a different distance from the camera, resulting in a certain difference in the size of similar targets in the figure. Therefore, an infrared pedestrian detection algorithm is proposed in the paper based on Yolov4 algorithm. Aiming at the problems of low recognition rate and high background influence in infrared image downlink human small target detection, the network structure of YOLOv4 is optimized. Compared with YOLOv4 and YOLOv3, the mean Average Precision is improved by 0.53
The electrocardiogram reflects the temporal changes in the body’s cardiac potential; This is also an important technology for diagnosing cardiovascular disease, so the classification of electrocardiogram has gradually become the focus of many scholars. This paper designs an ECG classification algorithm based on convolutional neural network, which aims to automatically classify ECG using artificial intelligence algorithm. The algorithm has the characteristics of less parameters and more layers, and the classification speed is faster and has very strong real-time performance. The experimental results show that the accuracy of the algorithm reaches 98.1
Channel estimation is very important and challenging for underwater acoustic (UWA) systems that use orthogonal frequency division multiplexing (OFDM) technique. The conventional methods are not appropriate to the severe frequency selective fading channel; also current channel estimation algorithms do not use the sparse characteristics of the UWA channel efficiently. In this paper, a novel algorithm about channel estimation is addressed that applies channel sparse features. First, Least Square (LS) algorithm is used to get the pilot channel valuation. Secondly, the sparse degree of channel is estimated through DFT for noise reduction processing. Then the autocorrelation matrix of the channel is obtained approximately. Finally a preliminary calculation of the error threshold is acquired, and the high quality data subcarrier channel impulse can be estimated by the pilot symbols through using the orthogonal matching pursuit (OMP) algorithm. In terms of simulating, we study the performance of the system through bit error rate (BER) and the constellation diagram, and it is indicated that the new method has excellent performance at less computation. So it is very important that the method can be implemented in OFDM systems.
The sphere detection algorithm is a low SNR algorithm in the MIMO-OFDM system with low complexity, but it still has a certain complexity. It is challenging to choose its initial radius and determine whether a point is in the ball. The improved algorithm is to find the CL algorithm's initial radius by the particle swarm algorithm's optimization ability. Then the convergence factor is given to speed up the shrinkage speed of the CL algorithm's radius. The improved algorithm is compared with the CL algorithm. The simulation results clear that when the SNR ratio is below 16 dB, the enhanced algorithm significantly affects algorithm complexity. The improvement of the algorithm is proved to be effective and reliable.
Steepest descent algorithm (SD) itself can find a better convergence direction, but its own convergence speed is relatively slow, resulting in multiple iterations to approach the true solution. In order to speed up its convergence, this paper proposes the SD-NM algorithm first, whose principle is to improve the approximate solution obtained after each execution of the steepest descent algorithm. In order to further reduce the computational complexity and improve the detection performance of the SD-NM algorithm., SD-NM-RC detection algorithm and local optimal LO-SD-NM-RC detection algorithm are proposed in this paper. The principle is to divide the constellation into three regions using the idea of regional division. The estimated value that fall into the reliable area are directly extracted for judgment, and the estimated value are no longer involved in subsequent iterations. The estimated value that fall into the unreliable region or the normal iteration region are not processed during the iteration. After the iteration, the constellation points around the estimated value that fall into the unreliable region are traversed. Simulations show that when the number of iterations of the SD-NM algorithm is consistent with that of SD algorithm, the detection performance of SD-NM algorithm is one order of magnitude higher than that of the SD algorithm. The bit error rate (BER) of the LO-SD-NM-RC algorithm is lower than the BER of the minimum mean square error (MMSE) algorithm when the signal-to-noise ratio is greater than 4dB and the number of iterations is 2. When the number of iterations is 2, the BER of the SD-NM-RC algorithm is close to that of the SD-NM algorithm but the computational complexity of the SD-NM-RC algorithm is only 79% of that of the SD-NM algorithm, the computational complexity of the SD-RC algorithm will be further reduced as the number of iterations increases.
Data-intensive applications have achieved great success in the field of machine learning. How to ensure that the machine can still learn correctly in the absence of labeled samples is the next challenging problem to be solved. This paper first introduces the problem definition of few-shot learning. Secondly, the existing small few-shot learning methods based on meta-learning are comprehensively summarized. Specifically, they are divided into three categories: metric-based learning methods, optimization-based learning methods and model-based learning methods. We conducted a series of comparisons among various methods in each category to show the advantages and disadvantages of each method. Finally, the limitations of existing methods are analyzed, and the future development direction of few-shot learning research is prospected.
Adaptability of the partition method of fractal image compression to gray level textures directly influences the total number of partition blocks and image decoding effects. Hence, it is of critical significance to find a partition method which can accurately reflect image gray level distribution and visual threshold linkage relations in order to speed up encoding and enhance de-coding quality. Therefore, in this paper, the SNAM (Square of Non-symmetry and Anti-packing Model) partition method is optimized by thresholds. The optimized method is employed to improve fractal encoding. On the basis of the organic relations between local image textures, human vision threads and encoding efficiency as well as decoding quality, a self-adaption sub-blocks partition method based on a square non-symmetry, anti-packing model and human vision system is proposed. With such method, partitioned image sub-blocks can accurately reflect gray level distribution of images, while the number of partitioned image sub-blocks is reduced. In this way, the calculation and matching times are reduced in encoding. Encoding time is reduced in addition to improvement of restored image quality. Compared with the basic fractal encoding method, the speed is increased by over 30 times.
This paper introduced four common methods to improve the voltage balance of supercapacitor. Because of the advantages of high efficiency and high speed, the flyover capacitor method is selected to solve the problem of different charging and discharging speed of supercapacitor caused by different parameters of supercapacitor. In order to further improve the efficiency of the supercapacitor charging and discharging control system, a fuzzy control algorithm is designed based on the balance strategy of the flyover capacitor to further improve the efficiency of the supercapacitor charging and discharging control system. Finally, the validity of the design is verified by PSIM simulation software.
In response to the current spam flooding problem, this paper uses Python language machine learning and natural language processing technology to study the identification classification of spam messages. The Jieba algorithm is used to distinguish the Chinese word, and the TF-IDF algorithm is used to conduct feature extraction. On the basis of the analysis of the classifier algorithm, the experimental data is finalized. The results show that the classification effect of the polynomial plain Bayes classifier is optimal, and the identification of garbage text is best optimized.
The essence of image classification task is to extract high-level semantic content features of images. The traditional data enhancement methods based on convolutional neural network (CNN) are translation, rotation, clipping, noise adding, etc. These methods have not changed the content and style of image data. This paper proposes a fast style migration data enhancement method, which can quickly apply the style art of one image to another image without changing the high-level semantic content characteristics of the image. Through the experimental comparison, it is found that the method of fast style migration data enhancement proposed here can further improve the accuracy of the model compared with the traditional data.
The acceleration of the population ageing and the frequent accidents have caused the rapid growth of the number of the elderly and the disabled in our country. The Traditional Electric Wheelchair can no longer meet the diversified needs of the elderly and the disabled in the new era, therefore, it is necessary to research and develop the intelligent wheelchair according to this demand. In this paper, an intelligent wheelchair based on medical health detection is designed. The system is mainly divided into three parts: lower computer, Web end and mobile end. It is mainly composed of STM32 main control chip, 4g communication module, GPS module, sensor module, alarm module and motor driving module, with the function of detecting human health indicators, after experimental testing, the system can effectively remind obstacle avoidance and detect health indexes such as blood pressure, heart rate, body temperature, PM2.5, temperature and humidity.
As a large forest region in Northeast China, forest fire prevention has always been an important matter concerned by Heilongjiang Province. The emergence of fire, the damage to the ecology and environment is inestimable. Therefore, it is particularly important to use the intelligent image recognition technology to monitor the forest fires in northeast China in real time to ensure that the forest areas in northeast China are not damaged by fire. In fact, in the field of forest fire prevention, many scholars at home and abroad have done a lot of forest fire research, mainly in the field of forest fire monitoring, a lot of research and practical application. The forest fire detection and recognition system based on machine vision can effectively reduce the impact of forest fire, reduce the loss, and improve the real-time and accuracy of forest fire recognition.
Recognition of desert plants has been a difficult activity for both human and computers due to similarities between these plants. In this paper, we propose an approach for recognizing desert plants by images of the bark. This approach depends on deep learning techniques for image recognition. The recognition process depends on texture of the bark. Therefore, we use Prewitt edge detection and Hough transform to detect the bark from original image. Further, we build a bark dataset for desert plants; this dataset consists of 1660 bark images for five species of desert plants. Each species in the dataset has 332 images. These species are Palm Dates, Mimosa Scabrella, Sidr, Lemon and Pomegranate. Convolutional Neural Network (CNN) is a deep learning technique that used in image classification tasks. Therefore, we test CNN on our dataset, and it gives an accuracy of 99.8
This article focuses on the face recognition model in real life scenarios, because the possible occlusion affects the recognition effect of the model, resulting in a decline in the accuracy of the model. An improved WGAN network is proposed to repair occluded facial images. The generator in the improved WGAN network is composed of an encoder-decoder network, and a jump connection is used to connect the bottom layer with the high-level feature information to generate missing facial images. The low-level feature information is connected with the deep-level feature information, and the network's ability to extract features and generate pictures is enhanced at the same time. The paper also uses a global discriminator and a local discriminator, taking all the restored pictures as input to measure the overall authenticity, and taking the restored part of the pictures as input to judge whether the content structure is reasonable. After comparison and analysis of experiments, the improved face image has a complete structure and clear content, which is helpful for face recognition with partial occlusion.
Advancements of data mining and machine learning have paved the road for establishing an efficient attack prediction paradigm to protect large scaled networks. In this study, computer network intrusions had been eliminated using smart machine learning algorithm to eliminate network intrusion. Referring a big dataset named KDD computer intrusion dataset which includes large number of connections that diagnosed with several types of attacks; the model is established for predicting the type of attack by learning through this data. Feed forward neural network model is outperformed over the other proposed clustering models in attack prediction accuracy.
The aim of this research is to use the sentiment analysis techniques to deal with large dataset corpus, which has been collected, to detect and classify anti-Islamic online contents. Anti-Islamic websites have spread a lot in the last decade causing a lot of hate toward the Muslims communities; there have been many websites that attack Islam and Muslims and insult the Messenger, blessings and peace be upon him. We have gathered our proper dataset from different sources into a large corpus, and we have produced two datasets (balanced and non-balanced) for the English language. The framework of our proposed methodology has been described. Two approaches are used in this framework, the first one is based on supervised Machine Learning (ML) approach using Support Vector Machines (SVM) model as classifier and Term Frequency-Inverse Document Frequency (TF-IDF) as feature extraction; the second one is a hybrid approach combining lexicon-based dictionary and TF-IDF as feature extraction with SVM algorithm. We conducted different experiments and we compared the obtained results. We first use TF-IDF on word level, and then we have improved the model using tri-gram level. The experimental results show that the ML approach is the best approach for both datasets that produces high accuracy of 97