The extraction of buildings in aerial remote sensing applications is an important and challenging task. Most existing methods extract buildings based on local area attention, ignoring the loss of accuracy due to the global structure of the building. However, global structural features of buildings with strong coupling relationships in complex scenes are difficult to extract, such as the edges and bodies of buildings, leading to discontinuous results. Therefore, multiscale decoupled body and edge supervision network (MDBES-Net), which can consider both edge optimization and inner consistency, is proposed to solve these problems. MDBES-net consists of the body-mask-edge consistency constraint base network (BMECC), decoupling the body and edge aware module (DBEA), and the channel decoupled attention module (CDA). First, body-mask-edge consistency constraint supervision is established by body and edge labels to jointly improve the segmentation effect in the BMECC base network. Second, in the mutiscale DBEA module, building features are warped by a learnable flow field to make body parts more consistent and edges more detailed. Finally, the CDA module performs adaptive calibration of the recoupled feature map channel response to minimize external background noise interference. Experiments on the open Massachusetts building dataset, WHU Building Dataset show that the proposed MDBES-Net can accurately extract buildings in complex scenarios, enabling complete building segmentation with refined boundaries and improved internal consistency.
A two- stage dynamic multi-object positioning and grasping method is proposed to solve the problem of fast and accurate grasping of various types of dynamic objects on a factory assembly line. In the first stage, the proposed multiscale context-aware single- branch fusion semantic segmentation network is used to obtain the mask area of the target object: first, the feature extraction network adopts a single- branch structure, which reduces the number of network parameters while ensuring the extraction of rich spatial information and high- level semantic information; subsequently, the feature fusion network improves the expression ability of spatial data and semantic information through the bilateral guided feature fusion module; finally, the feature enhancement network is designed, and the feature assisted convergence module is embedded in the shallow and deep networks to accelerate the convergence speed of the network. In the second stage, a quick pose estimation strategy based on contour point detection is applied to predict the optimum posture of the grasping point in the mask region. The test results on the self- built dataset and the pipeline platform grab experiments demonstrate that the proposed method can detect and predict the position and posture of the object grab points in real time and accurately complete the object grab. Furthermore, its segmentation accuracy, prediction time, and grab success rate are better than the comparison method.
针对多样性目标在非结构化环境中的抓取位姿难以估计的问题,提出一种基于上下文聚合策略的轻量级编/解码抓取位姿检测网络.首先,以编/解码网络架构为基础,利用深度可分离卷积层与混洗单元构建目标特征深度分离-融合提取块,减少编码网络参数量,增强网络对抓取区域特征的提取能力;其次,利用双线性插值法和深度可分离卷积层建立深度分离-重构块,在恢复高层特征丢失信息的同时,有效减少解码网络的参数量;最后,针对可抓取区域像素点与目标物体全貌之间的非一致性问题,基于交叉熵辅助损失和自注意力机制,提出一种抓取区域上下文聚合策略,引导网络增强可抓取目标区域特征的表征能力,抑制非抓取像素点的冗余特征.实验结果表明,所提网络在Cornell数据集的图像拆分与对象拆分子集上抓取检测准确率分别可达97.8%与93.8%,单张图像检测速度可达64.93张/秒;在Jacquard数据集上抓取检测准确率可达95.1%,单张图像检测速度可达60.6张/秒.与对比网络相比,所提网络不仅计算量与参数量较小,而且抓取检测的准确率与速度均有明显提升,在真实场景下对9种物体的抓取检测验证中,抓取成功率达到93.3%.
针对非结构化场景中存在的多工件堆叠遮挡等问题,提出了基于多尺度特征注意Yolact网络的堆叠工件识别定位算法;所提算法首先在Yolact网络的掩码模板生成分支中加入多尺度融合与特征注意机制,提升网络预测堆叠工件掩码的质量,并设计了基于膨胀编码的目标检测模块,增强网络对不同尺度堆叠工件的适应能力,构建了多尺度特征注意Yolact网络;其次,利用构建的多尺度特征注意Yolact网络预测堆叠工件的掩码与边界框,并对堆叠工件掩码进行最小外接矩形生成,根据掩码边界框与掩码的最小外接矩形确定目标工件的抓取点与旋转角度;最后,基于堆叠工件识别定位算法研发了视觉机器人工件分拣系统;实验结果表明,所提模型在边界框回归、掩码预测两项任务上的识别精度均有提升,机器人工件分拣系统进行堆叠工件分拣作业的成功率达到97.5%.
针对复杂实际场景中模糊、污损、扭曲、倾斜等车牌图像关键信息缺失以及新能源车牌背景与字符对比度低难以识别的问题,提出了一种编解码结构的车牌图像超分辨率网络.首先,构建一种基于编解码结构的车牌重构生成器网络,利用编码器对车牌图像的纹理、字符等特征进行提取,解码器对车牌特征进行重构;然后,设计一种基于语义监督的判别器网络,在网络损失中引入了对抗损失与CTC(connectionist temporal classification)损失,增强生成器网络对车牌图像语义特征的表征能力;最后,基于VGG16网络提取车牌顶角点特征,利用坐标变换方法对车牌图像进行矫正,进一步提高重构清晰度与识别准确率.采用所提网络在自建XAUAT-Parking数据集和公开CCPD数据集上进行超分辨率重构与识别实验,结果表明:所提网络在CCPD数据集上的平均峰值信噪比可达25.5 dB,结构相似性(SSIM)可达0.989;在XAUAT-Parking数据集上峰值信噪比可达26.6 dB,结构相似性可达0.997.研究结果表明,该网络有较好的车牌图像超分辨率重建效果,而且对车牌关键信息缺失问题具有较强的鲁棒性.
In this paper, we propose a novel crack detection algorithm based on feature enhanced whole nested network to resolve the issue of inaccurate crack segmentation caused by complex background and changeable texture of concrete cracks in natural scenes. First, based on the holistically-nested network (a deep learning edge detection network), the multi-scale supervision mechanism was adopted to integrate the prediction results of concrete cracks of different scales to enhance the expression ability of the network to the linear topology of concrete cracks. Then, we used a convolution-deconvolution feature fusion module to effectively integrate the deconvolution deep semantic features and convolution shallow detail features of concrete cracks. The deep semantic features can reduce the interference of complex backgrounds and improve the feature response of the fuzzy crack area. The shallow features can improve the expression ability of crack details and the quality of crack features. Finally, we proposed a hybrid void convolution boundary thinning module that used residual network and void convolution group to refine the fracture boundary and improve the accuracy of fracture segmentation. Using the Bridge_Crack_ Image_Data dataset and Crack Forest Dataset, the accuracy of the proposed algorithm was 92. 1% and 91. 6% and the F-1-score was 80. 2% and 91. 1%, respectively. The experimental results show that the proposed algorithm obtains stable and accurate segmentation results in complex natural environments and attains strong generalizations.
Aimed at the challenge of low accuracy of building segmentation caused by poor continuity of remote-sensing-image regions and blurred boundaries, a remote sensing building semantics segmentation algorithm based on multiscale regional consistent attention supervision is proposed. First, based on the Unet encoder–decoder architecture, the proposed algorithm constructs the region attention network (ReA-Net), which employs a multiscale receptive field-guidance model to simultaneously focus on regional features and edge details of remote sensing image objects. Second, the self-attention mechanism is employed to establish the correlation representation of regional-level features of remote sensing images, and multiscale regional attention features of remote sensing images are obtained through weighted regional-level correlation mapping. Finally, to address the lack of spatial correlation constraints on the prediction of remote sensing images segmentation, a loss function with multiscale neighborhood consistency supervision is suggested to constrain the consistency of pixel label assignment related to a local region. Experimental results on WHU building dataset showed that intersection over union (IOU) reached 91.6%, precision reached 95.61%, recall reached 95.68% recall, and F1-score reached 95.64%; On the Massachusetts building dataset, IOU reached 74.77% and precision reached 83.93%, recall reached 87.53%, and F1-score reached 85.69%. Therefore, the proposed algorithm not only has a good segmentation effect but also has a strong robustness for remote sensing building image segmentation.
针对遥感图像建筑物易受背景中道路、树木、阴影干扰而导致分割边界不清晰的问题,提出了一种融合分形几何特征的Resnet网络.所提模型基于编码-解码框架,以Resnet网络为主干网络,在编码阶段中引入融合分形先验的空洞空间金字塔池化模块(FD-ASPP),利用分形维数捕获遥感图像的分形特征,增强了Resnet网络的几何特征描述能力.解码阶段提出一种深度可分离卷积注意力融合机制(DSCAF),有效融合高层次特征和低层次特征,获取更加丰富的遥感图像语义信息和位置细节信息.在WHU遥感图像数据集上的实验表明,精确率达到0.9448,召回率达到0.9462,F1分数达到0.9455,平均交并比mIoU达到0.9415.所提模型与FCN、Segnet、Deeplab V3、U-net、SETR和AlignSeg等现有建筑物遥感语义分割模型相比,具有更好的分割精度,有效克服了道路、树木、阴影等因素的干扰,得到了较清晰的建筑物边界.
In order to improve the accuracy and applicability of person re-identification(Re-ID),a Re-ID method based on vector attention mechanism GoogLeNet is proposed.Firstly,three groups of images(anchor,positive and negative) are input into the GoogLeNet-GMP network to obtain segmented feature vectors.Then,spatial pyramid pooling(SPP) is used to aggregate the features from different pyramid levels,and attention mechanism is introduced.By integrating the multi-scale pooling regions which represent the visual information of the target,the distinguishable features on multiple semantic levels are obtained.At the same time,the mixed form of two different loss functions is taken as the final loss function.Experiments on Market-15012 and Duke-MTMC3 data set show that the proposed method performs better in Rank-1 and mAP indicators than other excellent methods.
针对现有的PCB缺陷检测存在检测精度低、速率慢等问题,提出一种用于PCB缺陷检测的增强上下文信息Yolov4_tiny算法;该算法首先通过Transformer编码单元对特征提取网络深层特征冗余的问题进行优化,增强网络捕获不同尺度局部特征信息的能力;然后利用浅层特征增强PCB缺陷小目标上下文信息,提升FPN网络对小目标缺陷的表征能力;最后引入注意力机制对特征提取网络输出的有效特征层加权,强化目标特征表征能力;实验结果表明,该算法对于整体缺陷的平均检测精度的均值(mAP)达到98.70%,较Yolov4_tiny提升了 3.12%,实现了 PCB缺陷精准定位和识别,满足工业检测的实际需求.
Automatic crack detection on concrete surfaces has become increasingly important for the health diagnosis of concrete structures to prevent possible malfunctions or accidents. In this paper, a concrete crack segmentation network based on convolution–deconvolution feature fusion with holistically nested networks is proposed. The proposed network adopts an encoder–decoder structure and uses VGG-16 as the basic feature extraction network. First, considering the problem that the VGG-16 network can extract redundant features in the encoding stage, based on the channel attention mechanism, the channel spatial correlation and global information are used to emphasize crack features to remove redundant features. Second, through the convolution–deconvolution feature fusion module, the deep semantic information of the deconvolution is effectively fused with the shallow features of convolution, which effectively improves the semantic crack feature information extracted at each stage of the VGG-16 network. Finally, based on a multiscale supervised learning mechanism, holistically nested networks are used to fuse the prediction results from different scales, which enhances the network's ability to express linear topological structures and improves the accuracy of crack segmentation. Through a large number of experiments on the Bridge_Crack_Image_Data dataset and CFD dataset, we demonstrate that compared with other deep networks, the proposed network not only achieves better segmentation results for cracks of different widths but is also more robust.
Mahjong is a popular tabletop game in China, Japan, and other Asian countries. As an incomplete information game, it is more complicated than other complete information games, such as Go, chess, and shogi. With the rapid development of service robot, there have been several human-robot interactive playing systems for complete information games so far, but mahjong is not the case. In this paper, a human-robot interactive mahjong playing system (HRMPS) was developed. HRMPS consists of five modules, including central host, robot players, visual kit, mahjong conveyor, and interactive software. The central host serves the communication between four robot players, which grab a tile on mahjong conveyor and recognize its face via visual kit, and make an action decision by themselves or by human opponents with the help of the interactive software. In HRMPS, to visually recognize a total of 27 different mahjong faces in uncontrolled conditions, a deep convolutional neural network was adopted to achieve an accuracy of 99.71% with a running time of 29ms. The experimental results tell that HRMPS is applicable in human-robot interactive mahjong game.
The disadvantages of convolutional neural networks are that they cannot run on mobile devices and embedded devices due to their large memory requirements and large amounts of computation. With a small loss of accuracy, the lightweight neural network greatly reduces the number of parameters and the amount of computation compared with the ordinary convolution neural network. In order to use robot arms accurately grasp the chess pieces and place on the right place of the Chinese chess board, the chess pieces should be recognized and classified quickly in real time. In this paper, the MobileNet is embedded into the Xavis platform, which is combined with the robot arms to realize the recognition of Chinese chess. Firstly, convolutional neural network methods are introduced. Then network structure of MobileNet is analyzed, and the pattern recognition method is shown. Xavis platform with MobileNet is further investigated. And the dataset of Chinese Chess is collected. At last, experiments on the self-made Chinese chess data set show that the MobileNet has good chess classification ability. It lays a solid foundation for the grasping work of the robot arms. Next, we will further improve the method of Chinese Chess recognition with Xavis Platform and compare the proposed schedule with state-of-art methods.
高分辨率图像具有特征尺度差异较大的特点,针对其造成的细粒度特征难以捕获、多尺度特征融合不佳问题,提出一种共享核空洞卷积与注意力引导(Kernel-Sharing Dilated Convolutions and Attention-guided FPN,KDA-FPN)的复杂场景文本检测方法;提出最小交集(Intersection Over Minimum,IOM)后处理策略,改善因文本长宽比变化较大特性导致的掩膜重叠现象,提升检测效果.首先,模型以Resnet50为主干网络采用FPN结构捕获多尺度特征;然后,利用空洞卷积扩大特征感受野,提高特征信息的多尺度捕获能力,深层次挖掘文本细粒度特征,并通过共享核手段减少模型参数量,降低计算成本;同时,采用上下文注意模块(Context Attention Module,CxAM)捕捉多感受野间的语义信息关系,通过内容注意模块(Content Attention Module,CnAM)精确定位目标位置信息,增强多尺度融合能力,提升特征图质量;最后,将同一文本区域预测的候选框按大小排列,提出将面积最大的框与相邻文本框之间区域的交集面积占较小框面积的比值作为候选框筛选指标,抑制检测结果的掩模重叠现象,实现文本的精准检测.采用ICDAR2013、ICDAR2015、Total-Text数据集进行对比实验,实验结果表明,本文模型对于水平场景文本检测的精度和召回率分别为95.3和90.4;对于倾斜文本检测的精度和召回率分别为87.1和84.2;对于任意形状文本检测的精度和召回率分别为69.6和57.3.提出的算法有效克服了图像分辨率、文本形状与长度等因素的影响,提高了检测精度,得到了更为精准的文本边界.
The spatial interaction of chromosomes is regarded as an important issue affecting the regulation of gene expression, and the high-throughput chromosome conformation capture (Hi-C) technology has become the primary tool to explore the temporal and spatial interactions of chromosomes in three-dimensional genomics. With the continuous accumulation of Hi-C samples and the increasing complexity of pipelines, the bioinformatic analysis of Hi-C data has been considered an opportunity and a challenge for understanding the spatial regulation mechanism of gene expression. In this paper, the current status and development outline of bioinformatic methods for Hi-C data are introduced, including data normalization, multi-level structure analysis, data visualization and 3D modeling, especially of multi-level structure at A/B compartments, topological associated domains (TADs) and chromain looping levels. Based on this, we provide the outlook of future hotspots and trends in this area. Hopefully our insight will be beneficial for the exploration of gene expression regulation from the traditional linear model to the 3D mode.
The flexibility and adaptability of traditional robots are pretty poor.This paper focuses on the key techniques of visual robots and puts forward a complete set of recognition,location and grasping algorithms.To solve the problem of target recognition and tracking,an improved method of dynamic target recognition and tracking based on Camshift is designed.In order to improve the success rate of grabbing target,the monocular visual ranging and visual navigation theory are used to measure the target center's spatial position.Then,a robot arm model is established using modified D-H parameters.The forward and inverse kinematics are used to get the control scheme of mechanical arm.The final test shows that the recognition rate of the algorithm is 100%,the relative error between the measured and actual coordinates is less than 5%,and the overall assembly success rate is 96.7%.The results prove that the present key techniques of visual robot can satisfy the application requirements.
Transmembrane region (TR) is a conserved region of transmembrane (TM) subunit in envelope (env) glycoprotein of retrovirus. Evidences have shown that TR is responsible for anchoring the env glycoprotein on the lipid bilayer and substitution of the TR for a covalently linked lipid anchor abrogates fusion. However, universal software could not achieve sufficient accuracy as TM in env also has several motifs such as signal peptide, fusion peptide and immunosuppressive domain composed largely of hydrophobic residues. In this paper, a support vector machine-based (SVM) model is proposed to identify TRs in retroviruses. Firstly, physicochemical and evolutionary information properties were extracted as original features. And then, the feature importance was analyzed by minimum Redundancy Maximum Relevance (mRMR) feature selection criterion. Our model achieved an Sn of 0.955, Sp of 0.998, ACC of 0.995, MCC of 0.954 using 10-fold cross-validation on the training dataset. These results suggest that the proposed model can be used to predict TRs in non-annotation retroviruses and 11917, 3344, 2, 289 and 6 new putative TRs were found in HERV, HIV, HTLV, SIV, MLV, respectively.
通过对工业革命发展类人比较分析,得出了前四次工业革命的每一次都是以机器(广义机器)衍生出类人某种重要器官肌能的机器机能为标志,使得各种机器不断转型升级和广泛应用,从而形成工业X.0的发展规律。采用工业发展的这一规律,推论出类人认知学习能力的机器学习机能将引发第五次工业革命,诞生各种学习机能机器和广泛应用,即为工业5.0,并研究定义了工业5.0机器机能的定义。工业发展的这一规律为未来工业发展重点乃至科学研究方向提供了战略性的理论依据。依据工业5.0机器机能的定义,我们研制成功可体现工业5.0特征的一种群机器人智动化作业系统模型,为工业5.0智动化系统的关键技术研究奠定了试验基础和模型示范。 A comparison of human-like attributes with machines in previous four industrial revolutions was performed in this study. Each industrial revolution was found to be symbolized by a machine that acquired an important humanoid organ function. The development of industry revolutions will provide strategic theory support for future industrial development and research direction. It was then applied to explore the higher level of function emerged in the next industrial revolution. We found that the fifth industrial revolution will be triggered by learning function, and all types of machines with learning function will continuously emerge, leading to Industry 5.0. On the basis of the definition of Industry 5.0, we developed a flexible intelligent swarm robotic system, which helps to research key technology of intelligentized automation system.
Integrase catalytic domain (ICD) is an essential part in the retrovirus for integration reaction, which enables its newly synthesized DNA to be incorporated into the DNA of infected cells. Owing to the crucial role of ICD for the retroviral replication and the absence of an equivalent of integrase in host cells, it is comprehensible that ICD is a promising drug target for therapeutic intervention. However, annotated ICDs in UniProtKB database have still been insufficient for a good understanding of their statistical characteristics so far. Accordingly, it is of great importance to put forward a computational ICD model in this work to annotate these domains in the retroviruses. The proposed model then discovered 11,660 new putative ICDs after scanning sequences without ICD annotations. Subsequently in order to provide much confidence in ICD prediction, it was tested under different cross-validation methods, compared with other database search tools, and verified on independent datasets. Furthermore, an evolutionary analysis performed on the annotated ICDs of retroviruses revealed a tight connection between ICD and retroviral classification. All the datasets involved in this paper and the application software tool of this model can be available for free download at https://sourceforge.net/projects/icdtool/files/?source=navbar.
Human endogenous retroviruses (HERVs) encode active retroviral proteins, which may be involved in the progression of cancer and other diseases. Matrix protein (MA), in group-specific antigen genes (gag) of retroviruses, is associated with the virus envelope glycoproteins in most mammalian retroviruses and may be involved in virus particle assembly, transport and budding. However, the amount of annotated MAs in ERVs is still at a low level so far. No computational method to predict the exact start and end coordinates of MAs in gags has been proposed yet. In this paper, a computational method to identify MAs in ERVs is proposed. A divide and conquer technique was designed and applied to the conventional prediction model to acquire better results when dealing with gene sequences with various lengths. Initiation sites and termination sites were predicted separately and then combined according to their intervals. Three different algorithms were applied and compared: weighted support vector machine (WSVM), weighted extreme learning machine (WELM) and random forest (RF). G - mean (geometric mean of sensitivity and specificity) values of initiation sites and termination sites under 5-fold cross validation generated by random forest models are 0.9869 and 0.9755 respectively, highest among the algorithms applied. Our prediction models combine RF & WSVM algorithms to achieve the best prediction results. 98.4% of all the collected ERV sequences with complete MAs (125 in total) could be predicted exactly correct by the models. 94,671 HERV sequences from 118 families were scanned by the model, 104 new putative MAs were predicted in human chromosomes. Distributions of the putative MAs and optimizations of model parameters were also analyzed. The usage of our predicting method was also expanded to other retroviruses and satisfying results were acquired.