As consumer electronics evolve towards greater intelligence, their automation and complexity also increase, making it difficult for users to diagnose faults when they occur. To address the problem where users, relying solely on their own knowledge, struggle to diagnose faults in consumer electronics promptly and accurately, we propose a multimodal knowledge graph-based text generation method. Our method begins by using deep learning models like the Residual Network (ResNet) and Bidirectional Encoder Representations from Transformers (BERT) to extract features from user-provided fault information, which can include images, text, audio, and even olfactory data. These multimodal features are then combined to form a comprehensive representation. The fused features are fed into a graph convolutional network (GCN) for fault inference, identifying potential fault nodes in the electronics. These fault nodes are subsequently fed into a pre-constructed knowledge graph to determine the final diagnosis. Finally, this information is processed through the Bias-term Fine-tuning (BitFit) enhanced Chinese Pre-trained Transformer (CPT) model, which generates the final fault diagnosis text for the user. The experimental results show that our proposed method achieves a 4.4% improvement over baseline methods, reaching a fault diagnosis accuracy of 98.4%. Our approach effectively leverages multimodal fault information, addressing the challenges users face in diagnosing faults through the integration of graph convolutional network and knowledge graph technologies.
This paper introduces a solution to address the intricacy of the model employed in the deep learning-based diagnosis of musculoskeletal abnormalities and the limitations observed in the performance of a single deep learning network model. The proposed approach involves the integration of an improved EfficientNet-B2 model with MobileNetV2, resulting in the creation of FusionNet. First, EfficientNet-B2 is combined with coordinate attention (CA) to obtain CA-EfficientNet-B2. Furthermore, aiming to minimize the model parameter count, we further enhanced the mobile inverted residual bottleneck convolution module (MBConv) employed for feature extraction in EfficientNet-B2, resulting in the development of CA-MBC-EfficientNet-B2. Next, the features extracted from CA-MBC-EfficientNet-B2 and MobileNetV2 are fused. Finally, the final diagnosis of musculoskeletal abnormalities was performed by using fully connected layers. The experimental results demonstrate that, first, compared to EfficientNet-B2, CA-MBC-EfficientNet-B2 not only significantly improves the diagnostic performance of musculoskeletal abnormalities, it also reduces the parameter count and storage space by 17%. Moreover, as compared to other models, FusionNet demonstrates remarkable performance in the area of anomaly diagnosis, particularly on the elbow dataset, achieving a precision of 92.93%, an AUC of 93.89% and an accuracy of 87.10%.
In the field of the Internet of Things, image acquisition equipment is the very important equipment, which will generate lots of invalid data during real-time monitoring. Analyzing the data collected directly from the terminal by edge calculation, we can remove invalid frames and improve the accuracy of system detection. SSD algorithm has a relatively light and fast detection speed. However, SSD algorithm do not take full advantage of both shallow and deep information of data. So a multiscale feature fusion attention mechanism structure based on SSD algorithm has been proposed in this paper, which combines multiscale feature fusion and attention mechanism. The adjacent feature layers for each detection layer are fused to improve the feature information expression ability. Then, the attention mechanism is added to increase the attention of the feature map channels. The results of the experiments show that the detection accuracy of the optimized model is improved, and the reliability of edge calculation has been improved.
Aiming at the problems of short duration, low intensity, and difficult detection of micro-expressions (MEs), the global and local features of ME video frames are extracted by combining spatial feature extraction and temporal feature extraction. Based on traditional convolution neural network (CNN) and long short-term memory (LSTM), a recognition method combining global identification attention network (GIA), block identification attention network (BIA) and bi-directional long short-term memory (Bi-LSTM) is proposed. In the BIA, the ME video frame will be cropped, and the training will be carried out by cropping into 24 identification blocks (IBs), 10 IBs and uncropped IBs. To alleviate the overfitting problem in training, we first extract the basic features of the pre-processed sequence through the transfer learning layer, and then extract the global and local spatial features of the output data through the GIA layer and the BIA layer, respectively. In the BIA layer, the input data will be cropped into local feature vectors with attention weights to extract the local features of the ME frames; in the GIA layer, the global features of the ME frames will be extracted. Finally, after fusing the global and local feature vectors, the ME time-series information is extracted by Bi-LSTM. The experimental results show that using IBs can significantly improve the model's ability to extract subtle facial features, and the model works best when 10 IBs are used.
针对皮肤镜采集的光学黑色素瘤图像,由于其背景信息复杂,干扰信息过多,导致检测精度较低,容易出现误检、漏检等问题,提出一种重参数化大核卷积的光学黑色素瘤图像检测算法.首先,在主干部分设计一种融合大核卷积与C3的新模块C3_RepLK,以增大模型的感受野,提取更多的有效信息.其次,引入感受野模块RFB,融合不同尺度的特征信息,减少错检.颈部网络中采用混合密集稀疏卷积GSConv和轻量化上采样算子CARAFE,使得网络能够捕捉到丰富的上下文信息,抑制漏检.最后,在算法中融入二阶通道注意力模块SOCA,加强特征之间的关联性,关注更有用的特征.实验表明,所提检测算法较原YOLOv5算法,所有类别平均精度从85.0%提升至89.4%,证明所提出的算法对于检测黑色素瘤的有效性.
In order to realize the remote control of the meeting documents in progress, the traditional method uses infrared remote control or 2.4 GHz wireless remote control. However, the shortcomings of carrying and storing the remote control, the infrared itself cannot pass through obstacles or the remote control of the device from a large angle, the 2.4 GHz cost is slightly higher, etc., this article introduces the use of PyTorch model and YOLO network gesture control to facilitate this practical problem. The plan proposes to use the PyTorch model to establish a neural network, train to achieve the purpose of classifying gestures, and use the YOLO network to cooperate with the corresponding control algorithm to achieve the purpose of controlling conference documents. The experimental results show that the proposed scheme is feasible and complete to achieve the required functions.
Recently, Siamese trackers have attracted extensive attention because of their simplicity and low computational cost. However, for most Siamese trackers, only a frame of the video sequence is used as the template, and the template is not updated in inference process, which makes the tracking success rate inferior to the trackers that can update the template online. In the current study, we introduce an enhanced visual attention Siamese network (ESA-Siam). The method is based on a deep convolutional neural network, which integrates channel attention and spatial self-attention to improve the discriminative ability of the tracker for positive and negative samples. Channel attention reflects different targets according to the response value of different channels to achieve better target representation. Spatial self-attention captures the correlation between two arbitrary positions to help locate the target. At the same time, a template search attention module is designed to implicitly update the template features online, which can effectively improve the success rate of the tracker when the target is interfered by the background. The proposed ESA-Siam tracker shows superior performance compared with 18 existing state-of-the-art trackers on five benchmark datasets including OTB50, OTB100, VOT2016, VOT2018, and LaSOT.
Garbage classification is a social issue related to people’s livelihood and sustainable development, so letting service robots autonomously perform intelligent garbage classification has important research significance. Aiming at the problems of complex systems with data source and cloud service center data transmission delay and untimely response, at the same time, in order to realize the perception, storage, and analysis of massive multisource heterogeneous data, a garbage detection and classification method based on visual scene understanding is proposed. This method uses knowledge graphs to store and model items in the scene in the form of images, videos, texts, and other multimodal forms. The ESA attention mechanism is added to the backbone network part of the YOLOv5 network, aiming to improve the feature extraction ability of the network, combining with the built multimodal knowledge graph to form the YOLOv5-Attention-KG model, and deploying it to the service robot to perform real-time perception on the items in the scene. Finally, collaborative training is carried out on the cloud server side and deployed to the edge device side to reason and analyze the data in real time. The test results show that, compared with the original YOLOv5 model, the detection and classification accuracy of the proposed model is higher, and the real-time performance can also meet the actual use requirements. The model proposed in this paper can realize the intelligent decision-making of garbage classification for big data in the scene in a complex system and has certain conditions for promotion and landing.
The Internet has become one of the important channels for users to obtain information and knowledge. It is crucial to work out how to acquire personalized requirement of users accurately and effectively from huge amount of network document resources. Group recommendation is an information system for group participation in common activities that meets the common interests of all members in the group. This paper proposes a group recommendation system for network document resource exploration using the knowledge graph and LSTM in edge computing, which can solve the problem of information overload and resource trek effectively. An extensive system test has been carried out in the field of big data application in packaging industry. The experimental results show that the proposed system recommends network document resource more accurately and further improves recommendation quality using the knowledge graph and LSTM in edge computing. Therefore, it can meet the user’s personalized resource need more effectively.
Haze-fog, which is an atmospheric aerosol caused by natural or man-made factors, seriously affects the physical and mental health of human beings. PM2.5 (a particulate matter whose diameter is smaller than or equal to 2.5 microns) is the chief culprit causing aerosol. To forecast the condition of PM2.5, this paper adopts the related the meteorological data and air pollutes data to predict the concentration of PM2.5. Since the meteorological data and air pollutes data are typical time series data, it is reasonable to adopt a machine learning method called Single Hidden-Layer Long Short-Term Memory Neural Network (SSHL-LSTMNN) containing memory capability to implement the prediction. However, the number of neurons in the hidden layer is difficult to decide unless manual testing is operated. In order to decide the best structure of the neural network and improve the accuracy of prediction, this paper employs a self-organizing algorithm, which uses Information Processing Capability (IPC) to adjust the number of the hidden neurons automatically during a learning phase. In a word, to predict PM2.5 concentration accurately, this paper proposes the SSHL-LSTMNN to predict PM2.5 concentration. In the experiment, not only the hourly precise prediction but also the daily longer-term prediction is taken into account. At last, the experimental results reflect that SSHL-LSTMNN performs the best.
The air quality in urban areas seriously affects the physical and mental health of human beings. And PM2.5 (a particulate matter whose diameter is smaller than or equal to 2.5 microns) is the chief culprit causing haze-fog. Since the meteorological data and air pollutes data are typical time series data, it’s reasonable to adopt a single hidden-layer LSTMNN (Long Short-Term Memory Neural Network) containing memory capability to implement the prediction. As for deciding the best structure of the neural network, this paper employs a self-organizing algorithm, which uses Information Processing Capability (IPC) to adjust the number of the hidden neurons automatically during a learning phase. In a word, to predict PM2.5 concentration accurately, this paper proposes a Self-organizing Single Hidden-Layer Long Short-Term Memory Neural Network (SSHL-LSTMNN) to predict PM2.5 concentration. In the experiment, not only the hourly precise prediction but also the daily longer-term prediction is taken into account. At last, the experimental results reflect that SSHL-LSTMNN performs the best.
This paper proposes a new object classification method based on an improved bacterial foraging optimisation algorithm. Firstly, a dynamic step size is used instead of the fixed step size of the chemotaxis. Secondly, the fixed elimination-dispersal probability is replaced by the dynamic probability. Features are extracted to distinguish the objects, such as pedestrians, cars and pets. Ultimately, all the objects are classified using the improved bacterial foraging optimisation algorithm. The experimental results prove that the effectiveness of the object classification method proposed in this paper is better than that of other algorithms.
This paper proposes a novel algorithm to solve the challenging problem of classifying error-diffused halftone images. We firstly design the class feature matrices, after extracting the image patches according to their statistics characteristics, to classify the error-diffused halftone images. Then, the spectral regression kernel discriminant analysis is used for feature dimension reduction. The error-diffused halftone images are finally classified using an idea similar to the nearest centroids classifier. As demonstrated by the experimental results, our method is fast and can achieve a high classification accuracy rate with an added benefit of robustness in tackling noise.