
Cardiac magnetic resonance imaging (CMR) is crucial for the pathological segmentation of the central muscle in the diagnosis of myocardial infarction (MI) patients. The automatic segmentation technology of medical images has achieved significant success with the development of deep learning technology. However, due to the low contrast of the target area edges, irregular lesion areas, and insufficient medical image data, automatic segmentation of myocardial pathology still has great challenges. In this article, we propose an improved RU-Net model called a Multimodal RU- -Net. Used to segment edema and scar areas in multimodal cardiac CMR data. In this network, we use RU-Net as the basic model, embed our proposed multimodal image feature extraction module (MFF) in the encoding path to enhance the extraction of complementary information between different modal images, and add attention modules in the skip connection and encoding path to enhance attention to the regions of interest in the image. The experiment shows that both modules mentioned above effectively improve segmentation accuracy. In addition, we have adopted methods such as data augmentation, deep supervision, and combination loss to further improve segmentation accuracy. We evaluated multimodal RU-Net on the MyoPS2020 challenge dataset and achieved a Dice score of 64.4% in scar segmentation and 70.7% in edema and scar segmentation. We achieved almost equivalent performance to the most advanced single stage segmentation methods on the MyoPS 2020 ranking, and our proposed method outperformed it in the standard deviation of dice scores, The test results are more stable. This indicates that our proposed method is meaningful for automatic segmentation of myocardial pathology.
In the autonomous driving domain, the fisheye camera is often used as an important sensor of environment perception. In order to realize the simulation and verification of advanced intelligent driving assistance system (ADAS), the simulation of fisheye camera is becoming more and more important. The traditional camera calibration has many problems, such as long process, poor accuracy and complex process. The checkerboard was used as the external parameters calibrating board, which was imported into VTD scene simulation software, and the external parameter calibration of the fisheye camera was realized through the coordinate system transformation. The tool LensDistortion is used to generate intrinsic parameter distortion. The results show that the fisheye camera calibration method can be well applied to validate the function of ADAS domain controller.
Aiming at the fatigue failure problem caused by the cycling influence of electrical stress and thermal stress of IGBT during working state, this paper presents a remaining useful life (RUL) prediction method based on Snake Optimizer (SO) to optimize the stacked bidirectional long short-term memory (Bi-LSTM) neural network model. Firstly, the peak voltage of collector-emitter in the IGBT aging data provided by NASA data center is selected as the characteristic parameter, and the wavelet transformation is used to preprocess the data. Secondly, the Bi-LSTM neural network model is constructed and the attention mechanism is introduced to calculate the allocation weight of the time series, which enhances the expression ability of the nonlinear features of the hidden layer, and adopts SO to optimize the BiALSTM neural network training hyperparameters and construct the SO-Bi-ALSTM neural network prediction model. Finally, the mean absolute error (MAE), mean absolute percentage error (MAPE) and root mean squared error (RMSE) are selected as the evaluation criteria, and the prediction accuracy of extreme learning machine (ELM), Transformer, LSTM and SO-Bi-ALSTM models is compared and analyzed. The results show that the MAE predicted by the SO-Bi-ALSTM model is 0.0173, the MAPE is 0.0921%, and the RMSE is 0.0160, which indicates the prediction performance of SO-Bi-ALSTM is better than the other methods.
As the digital wave sweeps the world, digital twin emerge as the times require, which digitally establishes a multi-dimensional dynamic virtual model of a physical entity. At present, there is a broad space for intelligent development in the medical field. Digital twin medicine has become an important development direction of intelligent medicine. The application of digital twin technology in the field of personalized medicine and digital medical system will promote the fundamental change of traditional electronic health and push it into a new era of personalized medicine and digital medical system. This paper investigates the research status of digital twin technology in the medical field and conducts quantitative statistical analysis. The results show that digital twin has a certain basis in medical development and can play a pivotal role in formulating highly personalized treatment plans in the future. In addition, the main concepts of digital twin and some applications in other fields are summarized, and some application examples and visions of digital twin in personalized medicine and digital medical systems are introduced. Finally, other issues and challenges of digital twin in medicine are briefly discussed.
With the development of computer vision technology and smart agriculture, deep learning techniques have been widely applied to crop pest identification tasks. However, existing studies do not consider the problem of large differences in pests across multiple growth stages, leading to unsatisfactory performance in pest identification in practical applications. This article proposes a simple framework for multi-stage prediction of pests, which can effectively predict the growth stage of pests and improve pest classification performance. The framework consists of a classification branch and a stage prediction branch. The stage prediction branch predicts the growth stage of pests based on feature similarity using K-Means, and guides the classification branch to classify images from different stages into different categories to avoid interfering with the performance of the classifier. In addition, to update the entire network parameters, we propose a multi-stage cross-entropy loss that optimizes feature extractors and classifiers by fusing image labels and stage prediction outputs. Experimental results show that the proposed multi-stage prediction framework for pest identification can accurately classify pest stages and improve pest classification accuracy. In addition, our work provides research ideas for pest stage prediction and identification, which is expected to help achieve more efficient pest control.
With the development and application of the Internet of Things, introducing a variety of power distribution intelligent sensing devices into the traditional power network to collect real-time information has become the mainstream in the industry. Such network is called the power distribution intelligent sensing Internet of Things. Due to the narrow bandwidth of most low-power wireless gateways, narrow bandwidth IoT communication protocol is often used. At the same time, the access environment of the IoT presents the characteristics of open network and weak computing ability of the devices. Therefore, we design a special trusted access scheme based on dynamic token for the communication protocol adapted to the IoT of power distribution intelligent sensing equipment. The trusted access scheme we proposed includes the network framework, access flow, protocol parameter exchange flow, data message format and error handling scheme. Then, we analyzed the scheme and compared it with other access schemes in the simulation environment. The results show that the scheme has good performance and can adapt to the terminal access scenario of power distribution equipment.
Occlusion effects have been a serious issue in industrial grasping for poorly textured workpiece bit position estimation. In this paper, we present an improved LineMod-2D algorithm called I-Line2D, which combines the matching mechanism of the voting statistics of the Generalized Hough Transform to achieve the bit-pose estimation of partially occluded objects. The method first calculates the similarity score of each feature point by LineMod-2D’s unique fast similarity calculation system, then indexes the high scoring feature points using their gradient directions, votes on the center position by finding the reference table established in advance, and gives different weights to the votes according to the different scores, and gives higher voting weight to the high scoring points. The experimental results show that among the data collected by ourselves, the method proposed in this paper can achieve 88%, 87% and 86.3% matching results in the actual occlusion data under the blurring of nonlinear illumination, noise ratio of 0.1, and filter kernel of 5, and also achieves a better matching result in the simulated occlusion data with rotation and scale change. And it outperforms the traditional LineMod-2D method in tests with different occlusion degrees.
Human pose forecasting predicts future human postures based on a given sequence of human postures. The fusion of spatio-temporal joint dependencies has a significant impact on model output, but research in this area is relatively scarce. This paper introduces a spatiotemporal graph convolutional neural network that effectively fuses multi-channel information with learnable multi-channel topology. We have designed a multi-branch module to capture temporal features from diverse spatial domains. Moreover, we initiatively adopted DST to optimize long-term forecasting. Our network has been evaluated on 3DPW and AMASS datasets commonly used in motion prediction. The results of the experiments showed that our methods were up to the current state of the art.
The Computing First Network is a new type of information infrastructure that integrates various resources such as cloud, network, and edge, and performs unified scheduling of computing power at the network level according to business needs. When the scale of the computing power network is large, it may include multiple autonomous domains, and it is necessary to dispatch computing power equipment in other domains to meet the needs of the domain. In order to schedule computing power equipment more efficiently from a global perspective, multidomain scheduling of computing power is necessary. However, the problem of multi-domain scheduling of computing power equipment is more complicated. It is not easy to take care of multi-dimensional requirements simply by using protocols for scheduling. The existing machine learning models have poor generalization and poor performance when processing topologies that have not been seen in training. Simultaneously applied in domains with different topologies. In response to the above problems, this paper proposes a multi-domain computing power network path planning model based on the combination of Graph Neural Networks (GNN) and Deep Q Network (DQN). The model can handle unseen topologies well during the training process, and generate multiple alternative scheduling paths based on computing power requests and topology structures, and then use the decision-making ability of DQN to evaluate these paths to choose the most appropriate scheduling The path satisfies the computing power request. The experimental results show that, compared with the simple DQN model and Dijkstra algorithm, when the model proposed in this paper performs computing power scheduling, the task completion time is shorter and the network load is more balanced.
Nowadays, convolutional neural networks (CNNs) are widely adopted in medical image analysis. Nevertheless, the inherent local nature of the convolution operator leads to limitations in capturing long remote dependencies. Fortunately, the self-attentive mechanism in Transformers can effectively model long remote dependencies. Therefore, we construct a hybrid network (CTW-Net) for medical image segmentation that combines the benefits of CNN and Transformer more effectively. We first construct a pure convolutional multiscale transformer (PCM Transformer) to multiscale the features obtained from the hierarchical encoder and decoder. It utilizes Hadamard product for second-order spatial interactions to fuse features at multiple scales, enabling the network to concentrate on more significant regions of lesion characteristics. Besides, we developed an efficient maximum pooling residual block (EMPR) based on convolution and maximum pooling. The EMPR block enables the complementation of locally important features in the downsampling process, which raises the performance of lesion segmentation. Experimental outcomes reveal that our network has achieved excellent segmentation performance in breast ultrasound and colon polyp lesion segmentation.
At present our country road traffic accident rate high all year round, motor vehicle safety detection becomes more and more important. Due to the limitation of manpower and material resources, the traditional vehicle detection technology can not meet the current requirements. Based on the above deficiencies, this paper designed and implemented an artificial intelligence vehicle safety detection platform, self-built vehicle detection image data set, and used a series of methods, such as YOLOv4, SAR, multidimensional convolution, attention mechanism, to automatically complete the detection of external hardware and internal information of motor vehicles. The experimental results show that our method has high speed and high precision, which can not only be used for motor vehicle detection, but also can be extended to real-time road safety detection in the future. It has high innovation, advanced nature and broad engineering application prospect.
The Vehicle Routing Problem (VRP) is a well-known NP-hard combinatorial optimization problem, which has wide spread applications in real world. This paper investigates a novel VPR, called multi-depot split delivery vehicle routing problem with customers’ multi-requirement (MSDVRP-CMR) according to an actual process of logistics distribution. It owns multiple depots, customers, and vehicles, and customers’ multiple goods requirements can be split delivered by vehicles. A mixed-integer linear programming formulation is presented. Then, an enhanced K-means and Adaptive Large Neighborhood Search (KALNS) algorithm is proposed for addressing this problem. First, K-means, C-W, and Greedy operations are designed for better initial solutions. Furthermore, ALNS global search with various destroy and repair operators are developed, and 3-OPT local search are proposed to further enhance its search ability. Extensive simulation is conducted and its results show that the augmented KALNS is effective and well outperforms ACO, ALNS and Gourbi.
The speech enhancement task aims to improve the perceptual quality of speech signals. Most of the existing models cannot guarantee the perceptual quality of the enhanced speech due to the use of Ll or L2 loss. In this paper, we designed a new perceptual loss based on lip movement for the speech enhancement model. There are two motivations for using lip-based perceptual loss: one is that the content of spoken language has a strong correlation with lip movements; the other is that studies have shown that multimodal speech enhancement models based on audio-clip information perform better than audio-only models by using lip information as auxiliary information for speech enhancement task. In addition, to obtain perceptual features and calculate the perceptual loss on the audio-only dataset, we designed a network that generates lip movements from speech signals guided by the Wav2Lip model for this task. Experiment results show that using lip movement-based perceptual loss can effectively improve speech enhancement performance.
The paper established a new load prediction model according to the data of regional population migration scale index, and further introduced temperature for correction. Then solved the load prediction model by CNN-BiLSTM-Attention algorithm. Finally, through the analysis of cases compared with the traditional prediction models, the accuracy and validity of this model were verified. And the accuracy of load prediction could be effectively improved by taking into account the migration scale index and temperature through the ablation test. The model provided a more accurate method for load prediction in regions of mass movement of population during special holidays.
scSPRITE is an emerging method to capture multi-way chromatin interaction and generate high-resolution, genome-wide three dimensional (3D) genome organization in single cells. However, the data obtained from scSPRITE often contain missing values and exhibit sparsity, making it difficult to directly observe local structures of the 3D genome at the level of single cells or even small cell populations. Additionally, comprehending the chromatin conformation composition entails a large number of cell populations by scSPRITE with huge cost. Therefore, it is valuable to impute these missing values and extract higher-order structures using computational methods. Here, we propose SpriteHyper2vec, a single-cell imputation algorithm based on hypergraph, specifically designed for scSPRITE datasets, which effectively capture local and global topological structures by generalizing Node2vec method. When compared with original method, SpriteHyper2vec significantly improves the similarity between small cell populations and the benchmark data. Furthermore, the completion capability of SpriteHyper2vec enables the identification of A/B compartment-like structures even in small cell populations. In summary, SpriteHyper2vec effectively enhances the visualization and comparison of 3D organization in small cell populations based on scSPRITE data.
Depression, a pervasive psychiatric disorder characterized by concealment, dependence on expert judgment, and a notable rate of misdiagnosis, poses a substantial burden on society. To enhance the diagnosis and treatment of depression, this study puts forth a proposition of employing knowledge-enhanced pre-training technology leveraging large language models. By integrating domain knowledge and depression knowledge graph directives, the pre-trained model undergoes optimization. Expert involvement in depression diagnosis and treatment fosters a guided learning process facilitated by expert feedback. Through the application of dialogue therapy, the efficacy of treatment is augmented. This technical approach aims to ameliorate the societal burden by improving the diagnosis and treatment of depressed individuals.
This paper presents a new approach for multi-agent communication via emergent language in a navigation task. The task involves a tourist and a guide who communicate through an emerged language grounded in the environment to help the tourist reach a target place. The paper proposes a collaborative multi-agent reinforcement learning framework that enables the agents to generate and understand emergent language, with the goal of solving the task. We evaluate the results on a simulated navigation game with 3,000 scenarios and multi-turn dialogues. Results show that the proposed approach achieves competitive performance in both language understanding and task success rate. The paper also provides explanations of the emerged language by comparing the language patterns with the environment structures.
As one of the main border gateway protocol (BGP) security schemes at present, the research on machine learning-based BGP anomaly detection technology faces severe data imbalance due to the lack of BGP anomaly data. To solve this problem, this paper proposes an unbalanced BGP data generation method based on undersampling technology and DoppelGANger, called US-DPGAN. Firstly, random undersampling is performed on majority-class samples to reduce the imbalanced ratio of the dataset and reduce the bias impact of majority-class samples on the data generated by DoppelGANger. Then, we use the trained DoppelGANger model to generate anomaly samples that match the original data distribution and add these samples to the original dataset to create a balanced BGP dataset. Through qualitative and quantitative experiments, we demonstrate that the data generated by the proposed US-DPGAN method better fits the distribution of the original data compared with the original DoppelGANger method. In addition, the anomaly detection model constructed using the US-DPGAN generated data outperforms some other methods in terms of detection accuracy and F1-Score.
Partial Label Learning (PLL) aims to train a multi-class classifier to deal with the problem that each training instance is associated with multiple candidate labels, but only one of them is true. For PLL, the averaging-based disambiguation strategy treats all candidate labels equally, however, this strategy may lead to the model prediction of candidate labels being stronger than the model prediction of valid labels. In this paper, decision trees are retrofitted by considering the inherent properties of PLL, and a method named PL-RF (label confidence based ensemble partial label learning) is proposed that can overcome the defects of averaging-based disambiguation strategies and obtain better classification performance by ensembleing multiple partial label decision trees, then determining the final predicted label by voting on each tree. Extensive experiments over both real-world and synthetic partial label data sets empirically justify the effectiveness of PL-RF, which achieves more competitive performance than other state-of-the-art PLL methods.
The paper presents a method to obtain Pythagorean fuzzy information (PFI) based on a bidirectional long short-term memory (BiLSTM) neural network. The method first uses the Word2Vec word embedding model for vectorized text pre-processing. Then, to solve the problems of one-way propagation and gradient explosion in traditional recurrent neural networks, BiLSTM is used to extract the global features of the comment information. Finally, the PFI is constructed based on the Softmax classifier and Pythagorean fuzzy set definition. Experimental results show that the proposed method can accurately capture online reviews’ satisfaction, dissatisfaction, and hesitation sentiment tendencies. Compared with other methods, such as RNN, CNN, LSTM, and GRU, the proposed method has higher accuracy, precision, and F1 values.