
With the rapid development of deep learning algorithms,the computational complexity and functional diversity are increasing rapidly.However,the gap between high computational density and insufficient memory bandwidth under the traditional von Neumann architecture is getting worse.Analyzing the algorithmic characteristics of convolutional neural network(CNN),it is found that the access characteristics of convolution(CONV) and fully connected(FC) operations are very different.Based on this feature,a dual-mode reronfigurable distributed memory architecture for CNN accelerator is designed.It can be configured in Bank mode or first input first output(FIFO) mode to accommodate the access needs of different operations.At the same time,a programmable memory control unit is designed,which can effectively control the dual-mode configurable distributed memory architecture by using customized special accessing instructions and reduce the data accessing delay.The proposed architecture is verified and tested by parallel implementation of some CNN algorithms.The experimental results show that the peak bandwidth can reach 13.44 GB·s -1 at an operating frequency of 120 MHz.This work can achieve 1.40,1.12,2.80 and 4.70 times the peak bandwidth compared with the existing work.
Subjective scales have different kinds of applicability in diverse fields.This study intends to implement a quantitative approach to determine the applicability of subjective scales in manual as-sembly work and evaluate the cognitive load of assembly workers.A multi-scale research paradigm based on subjective evaluation method is proposed.Three typical task stages are extracted from the process of assembly work.The National Aeronautics and Space Administration Task Load Index(NASA-TLX)scale,PAAS scale and Workload Profile Index Ratings(WP)scale are selected for the design of 3×3 multi-factor mixed experiment.The power spectrum density(PSD)characteris-tics of electroencephalogram(EEG)are utilized to identify the difficulty levels of the three task sta-ges.The relevant indicators of scale applicability are assessed.The results show that in terms of sensitivity,NASA-TLX scale reaches the highest sensitivity(F =999.137,P =0<0.05).In terms of validity,NASA-TLX scale possesses the best concurrent validity(P =0.0255<0.05).In terms of diagnosticity,NASA-TLX scale based on 6 dimensions takes on the best diagnostic performance.In terms of subject acceptability,WP scale performs the worst.According to the analytic hierarchy process(AHP)model,the applicability scores of NASA-TLX scale,PAAS scale and WP scale are determined as 3,2.55 and 1.6714,respectively.Therefore,NASA-TLX scale is regarded as the most suitable subjective evaluation questionnaire for assembly workers,which is also an effective quantitative evaluation method for the cognitive load of assembly workers.
As the research of knowledge graph(KG)is deepened and widely used,knowledge graph com-pletion(KGC)has attracted more and more attentions from researchers,especially in scenarios of in-telligent search,social networks and deep question and answer(Q&A).Current research mainly fo-cuses on the completion of static knowledge graphs,and the temporal information in temporal knowl-edge graphs(TKGs)is ignored.However,the temporal information is definitely very helpful for the completion.Note that existing researches on temporal knowledge graph completion are difficult to process temporal information and to integrate entities,relations and time well.In this work,a rotation and scaling(RotatS)model is proposed,which learns rotation and scaling transformations from head entity embedding to tail entity embedding in 3D spaces to capture the information of time and rela-tions in the temporal knowledge graph.The performance of the proposed RotatS model have been evaluated by comparison with several baselines under similar experimental conditions and space com-plexity on four typical knowl good graph completion datasets publicly available online.The study shows that RotatS can achieve good results in terms of prediction accuracy.
With the increasing demand of computational power in artificial intelligence(AI)algorithms,dedicated accelerators have become a necessity.However,the complexity of hardware architectures,vast design search space,and complex tasks of accelerators have posed significant challenges.Tra-ditional search methods can become prohibitively slow if the search space continues to be expanded.A design space exploration(DSE)method is proposed based on transfer learning,which reduces the time for repeated training and uses multi-task models for different tasks on the same processor.The proposed method accurately predicts the latency and energy consumption associated with neural net-work accelerator design parameters,enabling faster identification of optimal outcomes compared with traditional methods.And compared with other DSE methods by using multilayer perceptron(MLP),the required training time is shorter.Comparative experiments with other methods demonstrate that the proposed method improves the efficiency of DSE without compromising the accuracy of the re-sults.
Cataract is the leading cause of visual impairment globally.The scarcity and uneven distribution of ophthalmologists seriously hinder early visual impairment grading for cataract patients in the clin-ic.In this study,a deep learning-based automated grading system of visual impairment in cataract patients is proposed using a multi-scale efficient channel attention convolutional neural network(MECA_CNN).First,the efficient channel attention mechanism is applied in the MECA_CNN to extract multi-scale features of fundus images,which can effectively focus on lesion-related regions.Then,the asymmetric convolutional modules are embedded in the residual unit to reduce the infor-mation loss of fine-grained features in fundus images.In addition,the asymmetric loss function is applied to address the problem of a higher false-negative rate and weak generalization ability caused by the imbalanced dataset.A total of 7 299 fundus images derived from two clinical centers are em-ployed to develop and evaluate the MECA_CNN for identifying mild visual impairment caused by cataract(MVICC),moderate to severe visual impairment caused by cataract(MSVICC),and nor-mal sample.The experimental results demonstrate that the MECA_CNN provides clinically meaning-ful performance for visual impairment grading in the internal test dataset:MVICC(accuracy,sensi-tivity,and specificity;91.3%,89.9%,and 92%),MSVICC(93.2%,78.5%,and 96.7%),and normal sample(98.1%,98.0%,and 98.1%).The comparable performance in the external test dataset is achieved,further verifying the effectiveness and generalizability of the MECA_CNN model.This study provides a deep learning-based practical system for the automated grading of visu-al impairment in cataract patients,facilitating the formulation of treatment strategies in a timely man-ner and improving patients'vision prognosis.
To enhance the efficiency of warehouse order management,this study investigates a dual-com-mand operation mode in the Flying-V non-traditional warehouse layout.Three dual-command opera-tion strategies are designed,and a dual-command operation path optimization model is established with the shortest path as the optimization goal.Furthermore,a genetic algorithm based on a dynamic decoding strategy is proposed.Simulation results demonstrate that the Flying-V layout warehouse management and access cooperation operation can reduce the operation time by an average of 25%-35%compared with the single access operation path,and by an average of 13%-23%compared with the'deposit first and then pick'operation path.These findings provide evidence for the effec-tiveness of the optimization model and algorithm.
In the field of target recognition based on the temporal-spatial information fusion,evidence the-ory has received extensive attention.To achieve accurate and efficient target recognition by the evi-dence theory,an adaptive temporal-spatial information fusion model is proposed.Firstly,an adaptive evaluation correction mechanism is constructed by the evidence distance and Deng entropy,which realizes the credibility discrimination and adaptive correction of the spatial evidence.Secondly,the credibility decay operator is introduced to obtain the dynamic credibility of temporal evidence.Finally,the sequential combination of temporal-spatial evidences is achieved by Shafer's discount criterion and Dempster's combination rule.The simulation results show that the proposed method not only considers the dynamic and sequential characteristics of the temporal-spatial evidences com-bination,but also has a strong conflict information processing capability,which provides a new refer-ence for the field of temporal-spatial information fusion.
Dynamic path planning is crucial for mobile robots to navigate successfully in unstructured envi-ronments.To achieve globally optimal path and real-time dynamic obstacle avoidance during the movement,a dynamic path planning algorithm incorporating improved IB-RRT∗ and deep reinforce-ment learning(DRL)is proposed.Firstly,an improved IB-RRT∗ algorithm is proposed for global path planning by combining double elliptic subset sampling and probabilistic central circle target bi-as.Then,to tackle the slow response to dynamic obstacles and inadequate obstacle avoidance of tra-ditional local path planning algorithms,deep reinforcement learning is utilized to predict the move-ment trend of dynamic obstacles,leading to a dynamic fusion path planning.Finally,the simulation and experiment results demonstrate that the proposed improved IB-RRT∗ algorithm has higher con-vergence speed and search efficiency compared with traditional Bi-RRT∗,Informed-RRT∗,and IB-RRT∗algorithms.Furthermore,the proposed fusion algorithm can effectively perform real-time obsta-cle avoidance and navigation tasks for mobile robots in unstructured environments.
Hyperparameter optimization is considered as one of the most challenges in deep learning and dominates the precision of model in a certain.Recent proposals tried to solve this issue through the particle swarm optimization(PSO),but its native defect may result in the local optima trapped and convergence difficulty.In this paper,the genetic operations are introduced to the PSO,which makes the best hyperparameter combination scheme for specific network architecture be located easier.Spe-cifically,to prevent the troubles caused by the different data types and value scopes,a mixed coding method is used to ensure the effectiveness of particles.Moreover,the crossover and mutation opera-tions are added to the process of particles updating,to increase the diversity of particles and avoid local optima in searching.Verified with three benchmark datasets,MNIST,Fashion-MNIST,and CIFAR10,it is demonstrated that the proposed scheme can achieve accuracies of 99.58%,93.39%,and 78.96%,respectively,improving the accuracy by about 0.1%,0.5%,and 2%,respectively,compared with that of the PSO.
Compared with RGB videos and images,human bone data is less vulnerable to external factors and has stronger robustness.Therefore,behavior recognition methods based on skeletons are widely studied.Because graph convolution network(GCN)can deal with the irregular topology data of hu-man skeletons very well,more and more researchers apply GCN to human behavior recognition.Tra-ditional graph convolution methods only consider the joints with physical connectivity or the same type when building the behavior recognition model based on human skeletons structure,which cannot capture higher-order information better.To solve this problem,Motif-GCN is used in this paper to ex-tract spatial features.The relationship between the joints with natural connection in the human body is encoded by the first Motif-GCN,and the possible relationship between the unconnected joints in the human skeleton is encoded by the second Motif-GCN.In this way,the relationship between non-physical joints can be strengthened.Then a two stream framework combining joint and bone informa-tion is used to capture more action information.Finally,experiments are conducted on two subdata-sets X-Sub and X-View of NTU-RGB +D,and the accuracy shown in Top-1 classification results is 89.5%and 95.4%respectively.The experimental results are 1.0%and 0.3%higher than those of the 2S-AGCN model respectively.The superiority of this method is also proved by the experimental results.
Aiming at the problems of low accuracy,long time consumption,and failure to obtain quantita-tive fault identification results of existing automatic fault identification technic,a fault recognition method based on clustering linear regression is proposed.Firstly,Hough transform is used to detect the line segment of the enhanced image obtained by the coherence cube algorithm.Secondly,the endpoint of the line segment detected by Hough transform is taken as the key point,and the adaptive clustering linear regression algorithm is used to cluster the key points adaptively according to the lin-ear relationship between them.Finally,a fault is generated from each category of key points based on least squares curve fitting method to realize fault identification.To verify the feasibility and pro-gressiveness of the proposed method,it is compared with the traditional method and the latest meth-od on the actual seismic data through experiments,and the effectiveness of the proposed method is verified by the experimental results on the actual seismic data.
The superconducting rapid single flux quantum(RSFQ)integrated circuit is a promising solu-tion for overcoming speed and power bottlenecks in high-performance computing systems in the post-Moore era.This paper presents an architecture designed to improve the speed and power limitations of high-performance computing systems using superconducting technology.Since superconducting microprocessors,which operate at cryogenic temperatures,require support from semiconductor cir-cuits,the proposed design utilizes the von Neumann architecture with a superconducting RSFQ mi-croprocessor,cryogenic semiconductor memory,a room temperature field programmable gate array(FPGA)controller,and a host computer for input/output.Additionally,the paper introduces two key circuit designs:a start/stop controllable superconducting clock generator and an asynchronous communication interface between the RSFQ and semiconductor chips used to implement the control system.Experimental results demonstrate that the proposed design is feasible and effective,provi-ding valuable insights for future superconducting computer systems.
Influenced by its training corpus,the performance of different machine translation systems va-ries greatly.Aiming at achieving higher quality translations,system combination methods combine the translation results of multiple systems through statistical combination or neural network combina-tion.This paper proposes a new multi-system translation combination method based on the Transform-er architecture,which uses a multi-encoder to encode source sentences and the translation results of each system in order to realize encoder combination and decoder combination.The experimental veri-fication on the Chinese-English translation task shows that this method has 1.2-2.35 more bilingual evaluation understudy(BLEU)points compared with the best single system results,0.71-3.12 more BLEU points compared with the statistical combination method,and 0.14-0.62 more BLEU points compared with the state-of-the-art neural network combination method.The experimental re-sults demonstrate the effectiveness of the proposed system combination method based on Transformer.
This paper examines the prediction of film ratings. Firstly, in the data feature engineering, fea-ture construction is performed based on the original features of the film dataset. Secondly, the clus-tering algorithm is utilized to remove singular film samples, and feature selections are carried out. When solving the problem that film samples of the target domain are unlabelled, it is impossible to train a model and address the inconsistency in the feature dimension for film samples from the source domain. Therefore, the domain adaptive transfer learning model combined with dimensionality re-duction algorithms is adopted in this paper. At the same time, in order to reduce the prediction error of models, the stacking ensemble learning model for regression is also used. Finally, through com-parative experiments, the effectiveness of the proposed method is verified, which proves to be better predicting film ratings in the target domain.
The robotic drilling always generates the axial vibration along the drill bit and the torsional vi-bration around the drill bit, which will adversely affect the drilling precision. A vibration control mechanism fixed between the end-effector and the robot is proposed, which can suppress the axial and torsional vibrations based on the principle of vibro-impact ( VI) damping. The energy dissipa-tion of the system by vibro-impact damping is analyzed. Then, the influence of the structure parame-ters on the vibration attenuation effect is studied, and a semi-active vibration control method of vari-able collision clearance is presented. The simulation results show that the control method has effec-tive vibration control performance.
Aiming at the problem that ensemble empirical mode decomposition ( EEMD ) method can not completely neutralize the added noise in the decomposition process, which leads to poor reconstruc-tion of decomposition results and low accuracy of traffic flow prediction, a traffic flow prediction model based on modified ensemble empirical mode decomposition ( MEEMD) , double-layer bidirec-tional long-short term memory ( DBiLSTM) and attention mechanism is proposed. Firstly, the intrin-sic mode functions( IMFs) and residual components( Res) are obtained by using MEEMD algorithm to decompose the original traffic data and separate the noise in the data. Secondly, the IMFs and Res are put into the DBiLSTM network for training. Finally, the attention mechanism is used to en-hance the extraction of data features, then the obtained results are reconstructed and added. The ex-perimental results show that in different scenarios, the MEEMD-DBiLSTM-attention ( MEEMD-DBA ) model can reduce the data reconstruction error effectively and improve the accuracy of the short-term traffic flow prediction.
The failure rate of crankpin bearing bush of diesel engine under complex working conditions such as high temperature, dynamic load and variable speed is high. After serious wear, it is easy to deteriorate the stress state of connecting rod body and connecting rod bolt, resulting in serious acci-dents such as connecting rod fracture and body damage. Based on the mixed lubrication characteris- tics of connecting rod big endbearing shell of diesel engine under high explosion pressure impact load, an improved mixed lubrication mechanism model is established, which considers the influence of viscoelastic micro deformation of bearing bush material, integrates the full film lubrication model and dry friction model, couples dynamic equation of connecting rod. Then the actual lubrication state of big end bearing shell is simulated numerically. Further, the correctness of the theoretical re-search results is verified by fault simulation experiments. The results show that the high-frequency impact signal with fixed angle domain characteristics will be generated after the serious wear of bear- ing bush and the deterioration of lubrication state. The fault feature capture and alarm can be real- ized through the condition monitoring system, which can be applied to the fault monitoring of con-necting rod bearing bush of diesel engine in the future.
In order to improve the elbow passing performance and different diameter adaptability of pipe-line robot, a supported crawler pipeline robot is designed, which adopts screw nut mechanism and hinge four-bar mechanism to adapt to the complex environment such as variable diameter pipeline and elbow. The steering characteristics passing through the elbow are studied, the kinematic of pipe-line robot bending steering is established, the geometric constraint ( GC ) and steering constraint ( SC) in the elbow are analyzed, and the steering experiment is conducted. The results show that the robot can pass through the elbow by the SC model. The SC model can reduce the motor current and energy consumption when the robot passes through the elbow.
Image matching refers to the process of matching two or more images obtained at different time, different sensors or different conditions through a large number of feature points in the image. At present, image matching is widely used in target recognition and tracking, indoor positioning and navigation. Local features missing, however, often occurs in color images taken in dark light, mak-ing the extracted feature points greatly reduced in number, so as to affect image matching and even fail the target recognition. An unsharp masking ( USM) based denoising model is established and a local adaptive enhancement algorithm is proposed to achieve feature point compensation by strengthe-ning local features of the dark image in order to increase amount of image information effectively. Fast library for approximate nearest neighbors ( FLANN) and random sample consensus ( RANSAC) are image matching algorithms. Experimental results show that the number of effective feature points obtained by the proposed algorithm from images in dark light environment is increased, and the ac-curacy of image matching can be improved obviously.
To improve the inference efficiency of convolutional neural networks ( CNN) , the existing neu-ral networks mainly adopt heuristic and dynamic programming algorithms to realize parallel schedu-ling among operators. Heuristic scheduling algorithms can generate local optima easily, while the dynamic programming algorithm has a long convergence time for complex structural models. This pa-per mainly studies the parallel scheduling between operators and proposes an inter-operator parallel-ism schedule ( IOPS) scheduling algorithm that guarantees the minimum similar execution delay. Firstly, a graph partitioning algorithm based on the largest block is designed to split the neural net-work model into multiple subgraphs. Then, the operators that meet the conditions is replaced ac-cording to the defined operator replacement rules. Finally, the optimal scheduling method based on backtracking is used to schedule the computational graph. Network models such as Inception-v3, ResNet-50, and RandWire are selected for testing. The experimental results show that the algorithm designed in this paper can achieve a 1 . 6 × speedup compared with the existing sequential execution methods.