
With the development of cities and the prevalence of networks, interpersonal relationships have become increasingly distant. When people crave communication, they hope to find someone to confide in. With the rapid advancement of deep learning and big data technologies, an enabling environment has been established for the development of intelligent chatbot systems. By effectively combining cutting-edge technologies with human-centered design principles, chatbots hold the potential to revolutionize our lives and alleviate feelings of loneliness. A multi-topic chat companion robot based on a state machine has been proposed, which can engage in fluent dialogue with humans and meet different functional requirements. It can chat with users about movies, music, and other related topics, and recommend movies and music that may interest them to alleviate their loneliness and provide companionship. The interaction platform of the companion robot is realized through the QQ communication platform, with two chat modes: Conversation mode and recommendation mode. First, the KdConv open-source corpus was selected, and Python was used to crawl information on movies and music from Douban and QQ Music to establish and pre-process the dataset. Then, the dialogue function was implemented using generative language models and retrieval systems, while the recommendation function was achieved using user profiling and collaborative filtering. Finally, a state machine algorithm was used to achieve real-time switching between the two chat modes of the companion robot. In conclusion, test participants gave high ratings for the accuracy of the companion robot's responses and the satisfaction with its content recommendations. Compared to traditional large-scale integrated models, this robot employs a state-machine framework to achieve diverse functions through seamless state transitions, thereby enhancing computational speed and precision. Additionally, the robot can recommend movies and music, providing companionship and alleviating loneliness for users, which is of great significance in modern society where interpersonal relationships are increasingly alienated.
As a result of the introduction of new infectious illnesses, key infection prevention measures were implemented. Now, a new coronavirus (SARS-CoV-2) epidemic has expanded swiftly, causing the coronavirus illness 2019 (COVID-19). Many microorganisms spread illness via hospital surfaces due to environmental pollution. This virus has been associated to close contact between persons in tight situations such as houses, hospitals, assisted living, and residential institutions. Aside from health care settings, public buildings, faith-based community centers, marketplaces, transportation, and corporate environments are prone to COVID-19 transmission. Physical contact with the sanitizer device may cause for spread Covid virus. That’s why we have proposed an automatic fogger mechanism-based hand sanitizer that may be able to reduce covid risk. Disinfectant fog will flow when an object will pass through the machine. This project will save cost, time, and wastage along with Covid spreading risk. This project is about designing a good healthcare system. In recent years, sophisticated automation has influenced the health industry. Health care in poor nations is costly. So, the project is an attempt to tackle this issue.
YouTube videos on sustainable fashion enable the public to gain basic knowledge about this concept. In this paper, we analyse user comments on YouTube videos that contain sustainable fashion content. The paper’s main objective is to help content creators and business managers effectively understand the perspectives of viewers, thus improving video quality and developing business. We analysed a dataset of 17,357 comments collected from 15 sustainable fashion YouTube videos. First, we use Latent Dirichlet Allocation (LDA), a topic modelling technique, to discover the abstract topics. In addition, we use two approaches to rank these topics: ranking based on proportion and Rank-1 method. Second, we apply sentiment analysis to identify the user’s emotional tone in the comments. As a result, 14 topics were identified. The most common positive and negative scores are 1 and −1, respectively. In total, there are 28.42% positive comments, 22.35% negative comments and 49.23% neutral comments.
Secret image sharing (SIS) is a significant research topic of image information hiding, which divides the image into multiple shares and distributes them to multiple parties for management and preservation. In order to reconstruct the original image, a subset with predetermined number of shares is needed. And just because it is not necessary to use all of the shares to make a reconstruction, SIS creates a high which breaks the limitations of traditional image protection methods, but at the same time, it causes a reduce of safety. Recently, new technologies, such as deep learning and blockchain, have been applied into SIS to improve its and . This paper gives an overall review of SIS, discusses four important approaches for SIS, and makes a comparison analysis among them from the perspectives of pixel expansion, tamper resistance, etc. At the end, this paper indicates the possible research directions of SIS in the future.
As a result of the introduction of new infectious illnesses, key infection prevention measures were implemented. Now, a new coronavirus (SARS-CoV-2) epidemic has expanded swiftly, causing the coronavirus illness 2019 (COVID-19). Many microorganisms spread illness via hospital surfaces due to environmental pollution. This virus has been associated to close contact between persons in tight situations such as houses, hospitals, assisted living, and residential institutions. Aside from health care settings, public buildings, faith-based community centers, marketplaces, transportation, and corporate environments are prone to COVID-19 transmission. Physical contact with the sanitizer device may cause for spread Covid virus. That’s why we have proposed an automatic fogger mechanism-based hand sanitizer that may be able to reduce covid risk. Disinfectant fog will flow when an object will pass through the machine. This project will save cost, time, and wastage along with Covid spreading risk. This project is about designing a good healthcare system. In recent years, sophisticated automation has influenced the health industry. Health care in poor nations is costly. So, the project is an attempt to tackle this issue.
Person re-identification (ReID) is a sub-problem under image retrieval. It is a technology that uses computer vision to identify a specific pedestrian in a collection of pictures or videos. The pedestrian image under cross-device is taken from a monitored pedestrian image. At present, most ReID methods deal with the matching between visible and visible images, but with the continuous improvement of security monitoring system, more and more infrared cameras are used to monitor at night or in dim light. Due to the image differences between infrared camera and RGB camera, there is a huge visual difference between cross-modality images, so the traditional ReID method is difficult to apply in this scene. In view of this situation, studying the pedestrian matching between visible and infrared modalities is particularly crucial. Visible-infrared person re-identification (VI-ReID) was first proposed in 2017, and then attracted more and more attention, and many advanced methods emerged.
A groundbreaking method is introduced to leverage machine learning algorithms to revolutionize the prediction of success rates for science fiction films. In the captivating world of the film industry, extensive research and accurate forecasting are vital to anticipating a movie’s triumph prior to its debut. Our study aims to harness the power of available data to estimate a film’s early success rate. With the vast resources offered by the internet, we can access a plethora of movie-related information, including actors, directors, critic reviews, user reviews, ratings, writers, budgets, genres, Facebook likes, YouTube views for movie trailers, and Twitter followers. The first few weeks of a film’s release are crucial in determining its fate, and online reviews and film evaluations profoundly impact its opening-week earnings. Hence, our research employs advanced supervised machine learning techniques to predict a film’s triumph. The Internet Movie Database (IMDb) is a comprehensive data repository for nearly all movies. A robust predictive classification approach is developed by employing various machine learning algorithms, such as fine, medium, coarse, cosine, cubic, and weighted KNN. To determine the best model, the performance of each feature was evaluated based on composite metrics. Moreover, the significant influences of social media platforms were recognized including Twitter, Instagram, and Facebook on shaping individuals’ opinions. A hybrid success rating prediction model is obtained by integrating the proposed prediction models with sentiment analysis from available platforms. The findings of this study demonstrate that the chosen algorithms offer more precise estimations, faster execution times, and higher accuracy rates when compared to previous research. By integrating the features of existing prediction models and social media sentiment analysis models, our proposed approach provides a remarkably accurate prediction of a movie’s success. This breakthrough can help movie producers and marketers anticipate a film’s triumph before its release, allowing them to tailor their promotional activities accordingly. Furthermore, the adopted research lays the foundation for developing even more accurate prediction models, considering the ever-increasing significance of social media platforms in shaping individuals’ opinions. In conclusion, this study showcases the immense potential of machine learning algorithms in predicting the success rate of science fiction films, opening new avenues for the film industry.
The issue of finding available parking spaces and mitigating congestion during parking is a persistent challenge for numerous car owners in urban areas. In this paper, we propose a novel method based on the A-star algorithm to calculate the optimal parking path to address this issue. We integrate a road impedance function into the conventional A-star algorithm to compute path duration and adopt a fusion function composed of path length and duration as the weight matrix for the A-star algorithm to achieve optimal path planning. Furthermore, we conduct simulations using parking lot modeling to validate the effectiveness of our approach, which can provide car drivers with a reliable optimal parking navigation route, reduce their parking costs, and enhance their parking experience.
Campus network provides a critical stage to student service and campus administration, which assumes a paramount part in the strategy of 'Rejuvenating the Country through Science and Education' and 'Revitalizing China through Talented Persons'.However, with the rapid development and continuous expansion of campus network, network security needs to be an essential issue that could not be overlooked in campus network construction.In order to ensure the normal operation of various functions of the campus network, the security risk level of the campus network is supposed to be controlled within a reasonable range at any moment.Through literature research, theory analysis and other methods, this paper systematically combs the research on campus network security at home and abroad, analyzing and researching the campus network security issues from a theoretical perspective.A series of efficient solutions accordingly were also put forward.
In recent years, deep learning algorithms have been popular in recognizing targets in synthetic aperture radar (SAR) images. However, due to the problem of overfitting, the performance of these models tends to worsen when just a small number of training data are available. In order to solve the problems of overfitting and an unsatisfied performance of the network model in the small sample remote sensing image target recognition, in this paper, we uses a deep residual network to autonomously acquire image features and proposes the Deep Feature Bayesian Classifier model (RBnet) for SAR image target recognition. In the RBnet, a Bayesian classifier is used to improve the effect of SAR image target recognition and improve the accuracy when the training data is limited. The experimental results on MSTAR dataset show that the RBnet can fully exploit effective information in limited samples and recognize the target of the SAR images more accurately. Compared with other state-of-the-art methods, our method offers significant recognition accuracy improvements under limited training data. Noted that the RBnet is moderately difficult to implement and has the value of popularization and application in engineering application scenarios in the field of small-sample remote sensing target recognition and recognition.
Big data is a comprehensive result of the development of the Internet of Things and information systems. Computer vision requires a lot of data as the basis for research. Because skeleton data can adapt well to dynamic environment and complex background, it is used in action recognition tasks. In recent years, skeleton-based action recognition has received more and more attention in the field of computer vision. Therefore, the keypoints of human skeletons are essential for describing the pose estimation of human and predicting the action recognition of the human. This paper proposes a skeleton point extraction method combined with object detection, which can focus on the extraction of skeleton keypoints. After a large number of experiments, our model can be combined with object detection for skeleton points extraction, and the detection efficiency is improved.
Recent advances in OCR show that end-to-end (E2E) training pipelines including detection and identification can achieve the best results. However, many existing methods usually focus on case insensitive English characters. In this paper, we apply an E2E approach, the multiplex multilingual mask TextSpotter, which performs script recognition at the word level and uses different recognition headers to process different scripts while maintaining uniform loss, thus optimizing script recognition and multiple recognition headers simultaneously. Experiments show that this method is superior to the single-head model with similar number of parameters in end-to-end identification tasks.
Microphone array-based sound source localization (SSL) is widely used in a variety of occasions such as video conferencing, robotic hearing, speech enhancement, speech recognition and so on. The traditional SSL methods cannot achieve satisfactory performance in adverse noisy and reverberant environments. In order to improve localization performance, a novel SSL algorithm using convolutional residual network (CRN) is proposed in this paper. The spatial features including time difference of arrivals (TDOAs) between microphone pairs and steered response power-phase transform (SRP-PHAT) spatial spectrum are extracted in each Gammatone sub-band. The spatial features of different sub-bands with a frame are combine into a feature matrix as the input of CRN. The proposed algorithm employ CRN to fuse the spatial features. Since the CRN introduces the residual structure on the basis of the convolutional network, it reduce the difficulty of training procedure and accelerate the convergence of the model. A CRN model is learned from the training data in various reverberation and noise environments to establish the mapping regularity between the input feature and the sound azimuth. Through simulation verification, compared with the methods using traditional deep neural network, the proposed algorithm can achieve a better localization performance in SSL task, and provide better generalization capacity to untrained noise and reverberation.
In this paper, we propose a intrusion detection algorithm based on auto-encoder and three-way decisions (AE-3WD) for industrial control networks, aiming at the security problem of industrial control network.The ideology of deep learning is similar to the idea of intrusion detection.Deep learning is a kind of intelligent algorithm and has the ability of automatically learning.It uses self-learning to enhance the experience and dynamic classification capabilities.We use deep learning to improve the intrusion detection rate and reduce the false alarm rate through learning, a denoising AutoEncoder and three-way decisions intrusion detection method AE-3WD is proposed to improve intrusion detection accuracy.In the processing, deep learning AutoEncoder is used to extract the features of high-dimensional data by combining the coefficient penalty and reconstruction loss function of the encode layer during the training mode.A multi-feature space can be constructed by multiple feature extractions from AutoEncoder, and then a decision for intrusion behavior or normal behavior is made by three-way decisions.NSL-KDD data sets are used to the experiments.The experiment results prove that our proposed method can extract meaningful features and effectively improve the performance of intrusion detection.
Accurate electricity forecasting is the key basis for guiding the power sector to arrange operation plans and guaranteeing the profitability of electric power companies.However, with the increasing demand of enterprises and departments for data security, the phenomenon of "Isolated Data Island" becomes more and more serious, resulting in the accuracy loss of the traditional electricity prediction model.Federated learning, as an emerging artificial intelligence technology, is designed to ensure data privacy while carrying out efficient machine learning, which provides a new way to solve the problem of "Isolated Data Island" in terms of electricity forecasting.Nonetheless, due to the popularity of smart meters, the collected electricity data presents the characteristics of uneven distribution and huge data volume, so it is difficult to apply the electric quantity prediction model generated only by federated learning in practice.To solve this problem, a clustering federated learning method (C-FL) is proposed to protect data privacy while improving the accuracy of power prediction.Firstly, C-FL uses K-means algorithm to cluster power data locally in power enterprises, and then builds accurate power forecasting models for each class of power data combined with other local clients through federated learning.A large number of experimental results show that the clustering federated learning method proposed in this paper is superior to the existing federated learning models in terms of the accuracy of electric power forecasting.
Anomaly detection in images has attracted a lot of attention in the field of computer vision.It aims at identifying images that deviate from the norm and segmenting the defect within images.However, anomalous samples are difficult to collect comprehensively, and labeled data is costly to obtain in many practical scenarios.We proposes a simple framework for unsupervised anomaly detection.Specifically, the proposed method directly employs CNN pre-trained on ImageNet to extract deep features from normal images and reduce dimensionality based on Principal Components Analysis (PCA), then build the distribution of normal features via the multivariate Gaussian (MVG), and determine whether the test image is an abnormal image according to Mahalanobis distance.We further investigate which features are most effective in detecting anomalies.Extensive experiments on the MVTec anomaly detection dataset show that the proposed method achieves 98.6% AUROC in image-level anomaly detection and outperforms previous methods by a large margin.
Plateau forest plays an important role in the high-altitude ecosystem, and contributes to the global carbon cycle.Plateau forest monitoring request in-suit data from field investigation.With recent development of the remote sensing technic, large-scale satellite data become available for surface monitoring.Due to the various information contained in the remote sensing data, obtain accurate plateau forest segmentation from the remote sensing imagery still remain challenges.Recent developed deep learning (DL) models such as deep convolutional neural network (CNN) has been widely used in image processing tasks, and shows possibility for remote sensing segmentation.However, due to the unique characteristics and growing environment of the plateau forest, generate feature with high robustness needs to design structures with high robustness.Aiming at the problem that the existing deep learning segmentation methods are difficult to generate the accurate boundary of the plateau forest within the satellite imagery, we propose a method of using boundary feature maps for collaborative learning.There are three improvements in this article.First, design a multi input model for plateau forest segmentation, including the boundary feature map as an additional input label to increase the amount of information at the input.Second, we apply a strong boundary search algorithm to obtain boundary value, and propose a boundary value loss function.Third, improve the Unet segmentation network and combine dense block to improve the feature reuse ability and reduces the image information loss of the model during training.We then demonstrate the utility of our method by detecting plateau forest regions from ZY-3 satellite regarding to Sanjiangyuan nature reserve.The experimental results show that the proposed method can utilize multiple feature information comprehensively which is beneficial to extracting information from boundary, and the detection accuracy is generally higher than several state-of-art algorithms.As a result of this investigation, the study will contribute in several ways to our understanding of DL for region detection and will provide a basis for further researches.
In the medical field, the classification and analysis of blood samples has always been arduous work. In the previous work of this task, manual classification maneuvers have been used, which are time consuming and laborious. The conventional blood image classification research is mainly focused on the microscopic cell image classification, while the macroscopic reagent processing blood coagulation image classification research is still blank. These blood samples processed with reagents often show some inherent shape characteristics, such as coagulation, attachment, discretization and so on. The shape characteristics of these blood samples also make it possible for us to recognize their classification through computer vision algorithms. Blood sample classification focuses on the texture and shape of the picture. HOG feature is a kind of feature descriptor used for object detection in computer vision and image processing. It can better extract the outline and texture features of the image by calculating and counting the histogram of oriented gradient of the local region of the image. Because the medical machines that need to identify and classify blood samples often lack strong calculation power, the current popular machine-learning classification algorithms cannot play a good role in these machines. In addition, the characteristics of blood samples produced by different types of reagents and processing methods are different, and it is difficult to obtain real samples, so the amount of data that can be used for training is small. Combining the above conditions and the experimental comparison of a variety of classification algorithms, we find that the lightweight SVM model has a better performance on this problem, and the combination of HOG and SVM has also been widely used in other research. The experiment demonstrated that the classification algorithm based on SVM and HOG can give a good result of both performance and accuracy in the classification of blood samples problem.
With the speedy development of communication Internet and the widespread use of social multimedia, so many creators have published posts on social multimedia platforms that fake news detection has already been a challenging task. Although some works use deep learning methods to capture visual and textual information of posts, most existing methods cannot explicitly model the binary relations among image regions or text tokens to mine the global relation information in a modality deeply such as image or text. Moreover, they cannot fully exploit the supplementary cross-modal information, including image and text relations, to supplement and enrich each modality. In order to address these problems, in this paper, we propose an innovative end-to-end Cross-modal Relation-aware Networks (CRAN), which exploits jointly models the visual and textual information with their corresponding relations in a unified framework. (1) To capture the global structural relations in a modality, we design a global relation-aware network to explicitly model the relation-aware semantics of the fragment features in the target modality from a global scope perspective. (2) To effectively fuse cross-modal information, we propose a cross-modal co-attention network module for multi-modal information fusion, which utilizes the intra-modality relationships and inter-modality relationship jointly among image regions and textual words to replenish and heighten each other. Extensive experiments on two public real-world datasets demonstrate the superior performance of CRAN compared with other state-of-the-art baseline algorithms.
With the improvement of people's security awareness, numerous monitoring equipment has been put into use, resulting in the explosive growth of surveillance video data. Key frame extraction technology is a paramount technology for improving video storage efficiency and enhancing the accuracy of video retrieval. It can extract key frame sets that can express video content from massive videos. However, the existing key frame extraction algorithms of surveillance video still have deficiencies, such as the destruction of image information integrity and the inability to extract key frames accurately. To this end, this paper proposes a key frame extraction algorithm of surveillance video based on quaternion Fourier saliency detection. Firstly, the algorithm used colors, and intensity features to perform quaternion Fourier transform on surveillance video sequences. Next, the phase spectrum of the quaternion Fourier transformed image was obtained, and he image visual saliency map was obtained according to the quaternion Fourier phase spectrum. Then, the image visual saliency map of two adjacent frames is used to characterize the change of target motion state. Finally, the frames that can accurately express the motion state of the target are selected as key frames. The experimental results show that the method proposed in this paper can accurately capture the changes of the local motion state of the target while maintaining the integrity of the image information.