Cross-chain mechanism functions as typical approaches for information interaction between diverse blockchains tackling the problem of information silos in the big data era. Most of the existing cross-chain mechanisms are targeted at virtual currency blockchains in the financial sector. With more and more engineering documents manufactured by the development of modern smart farming, the need for engineering document management and cross-chaining between various blockchains has become increasingly urgent. This paper proposes a novel attainable cross-chain mechanism for agricultural engineering document management blockchains concerning the unique structure and operation principals of the specific domain. The methodology sufficiently integrated the characteristics of the agricultural engineering document management with the notary scheme, constructed by government supervision nodes with high credibility. Meanwhile, the authentication technology and cryptographic algorithms are internally fused, solving the authentication problem of the document cross-chain and protecting the cross-chain information respectively, which ensures the integrity and security of the file attribute information, alongside file ontology data in the cross-chain process. Adequate security proof and experiments illustrate that the developed mechanism can guarantee the feasibility of the mechanism, authenticity of the cross-chain parties, and the integrality and reliability of the document information, thus catering to the requirements of the cross-chain performance of blockchain in the field of agricultural engineering document management.
This article proposes a mask refinement method for chromosome instance segmentation. The proposed method exploits the knowledge representation capability of Neural Knowledge DNA (NK-DNA) to capture the semantics of the chromosome's shape, texture, and key points, and then it uses the captured knowledge to improve the accuracy and smoothness of the masks. We validate the method's effectiveness on our latest high-resolution chromosome image dataset. The experimental results show that our proposed method's mask average precision (MaskAP) is 3.66% higher than Mask R-CNN and outperforms advanced Cascade Mask R-CNN by 1.35%.
In recent years, there have been significant advancements in deep learning-based lightweight image super-resolution (SR) reconstruction techniques. However, in practical applications, there are still challenges to be addressed. Many existing lightweight SR reconstruction algorithms simplify the model by reducing the number of model parameters or changing the combination of convolutions, which results in significantly decreased model performance and stability, leading to poor reconstruction results. To address this issue, we proposed a Cosine Self-Attention mechanism and deepens the network depth to improve the Swin Transformer, enhanced the model’s performance and stability. Experimental results show that the proposed approach achieves stronger and more stable reconstruction performance with lower parameter count than existing lightweight SR models, outperforming them in reconstruction results and model complexity. The PSNR/SSIM of our proposed method for ×2 and ×4 SR reconstruction on the Set5, Set14, BSD100, Urban100, and Manga109 datasets exceed those of most existing lightweight SR models.
Effective segment extraction has broad application prospects in the field of video surveillance. In action recognition or object recognition, a large number of repetitive background segments will reduce the efficiency of computation. This causes the algorithm to consume more time. Effective segments can greatly save computing and storage resources for both target detection and storage. Common algorithms are to process the whole surveillance video, such as frame difference method and optical flow method. However, false detection can occur when the picture is disturbed, such as snowflakes. Will result in low accuracy. The effective segment detection method adopted in this paper. First convert the color space. The surveillance video is then segmented. Features were extracted and repeated segments filtering was carried out to reduce the interference of bad images. Finally, effective segments are extracted. Experimental results show that the accuracy of the effective segment detection method adopted in this paper is increased by more than 50%, and the F-metric value is increased by more than 24%.
Real-time detection and alert systems for distracted driving are pivotal areas of research. With Edge AI, it is feasible to process data in real-time without relying on Internet connectivity, thereby safeguarding user privacy. However, the computational and storage constraints of edge devices can hamper the deployment of deep learning models, such as VGG-16. In response, we introduce a framework tailored for distracted driving detection using an ensemble model. This model taps into the potential of lightweight pre-trained networks like Inception, Xception, DenseNet, and MobileNet. Our objective is to amalgamate these models to achieve high accuracy, low latency, and fewer parameter set. Experiments on the State Farm dataset reveal a remarkable 99% accuracy rate, underlining its viability for devices with limited resources, such as the Nvidia Jetson Nano. These results underscore the efficacy and feasibility of our framework in the realm of edge computing.
Video summarization can help people retrieve videos quickly. Existing unsupervised video summarization models make insufficient use of video frame uniqueness in the self-attention mechanism and have redundancy in the computation of uniqueness. In this paper, we rethink the relationship between attentive uniqueness and diversity of video frames, improve the calculation method of video frame uniqueness by masking adjacent frame information in the model framework based on a self-attention mechanism, and simplify the concentrated attention mechanism with it. In addition, we propose a new loss function called Deviation loss, which uses the uniqueness of video frames to calculate and helps train video summarization models to obtain diverse representations. The experimental results show that the proposed method outperforms other state-of-the-art unsupervised summarization approaches on the SumMe dataset, and also demonstrates competitiveness on the TVSum dataset.
目的 针对ASPP(atrous spatial pyramid pooling)在空洞率变大时空洞(atrous)卷积效果会变差的情况,以及图像分类经典模型ResNet(residual neural network)并不能有效地适用于细粒度图像分割任务的问题,提出一种基于改进ASPP和极化自注意力的自底向上全景分割方法.方法 重新设计ASPP模块,将小空洞率卷积的输出与原始输入进行拼接(concat),将得到的结果作为新的输入传递给大空洞率卷积,然后将不同空洞率卷积的输出结果拼接,并将得到的结果与ASPP中的其他模块进行最后拼接,从而改善ASPP中因空洞率变大导致的空洞卷积效果变差的问题,达到既获得足够感受野的同时又能编码多尺度信息的目的;在主干网络的输出后引入改进的极化自注意力模块,实现对图像像素级的自我注意强化,使其得到的特征能直接适用于细粒度像素分割任务.结果 本文在Cityscapes数据集的验证集上进行测试,与复现的基线网络Panoptic-DeepLab(58.26%)相比,改进ASPP模块后分割精度PQ(panoptic quality)(58.61%)提高了 0.35%,运行时间从103 ms增加到124 ms,运行速度没有明显变化;通过进一步引入极化自注意力,PQ指标(58.86%)提高了 0.25%,运行时间增加到187 ms;通过对该注意力模块进一步改进,PQ指标(59.36%)在58.86%基础上又提高了 0.50%,运行时间增加到192 ms,速度略有下降,但实时性仍好于大多数方法.结论 本文采用改进ASPP和极化自注意力模块,能够更有效地提取适合细粒度像素分割的特征,且在保证足够感受野的同时能编码多尺度信息,从而提升全景分割性能.
With the national Belt and Road strategy in place, meteorological operational systems will face domestic and international dangers, and it is very urgent and important to protect the safety of meteorological equipment and resources. The existing meteorological operational system uses a role-based access control model for permission management, whose permissions do not change during one authorized access, and cannot sense the threat of user behavior to the system in real time, much less adjust user permissions in time, thus failing to cope with the security needs of the weather business system. In this paper, we propose a Meteorological Operational System-oriented Usage Control (MS-UCON) model to address this problem. The model refines the UCON model according to industry characteristics and business needs, and adds a periodic trust assessment module and a collection of trust value assessment factors to monitor user behavior in real time, perform periodic trust assessment on user behavior, sense user risky behavior in time, dynamically update user trust value attributes, and dynamically adjust permissions. Finally, the security of the MS-UCON model is proved based on state machine theory. The results show that the MS-UCON model can solve the problem that user privileges cannot be dynamically adjusted in meteorological operational systems, and also meet the security requirements of meteorological operational systems.
Personnel behavior understanding under complex scenarios is a challenging task for computer vision. This paper proposes a novel Compact model, which we refer to as CGARPN that incorporates with Global Association relevance and Adaptive Routing Pose estimation Network. Our framework firstly introduces CGAN backbone to facilitate the feature representation by compressing the kernel parameter space compared with typical algorithms, effectively lowering the calculation capacity and consumption. The framework integrates the Global Association information between keypoints, and learns the correlation between high-dimensional feature parameters. ARPN introduced by our structure is established to sufficiently excavate the resembling properties of outcome concealed in the network, adaptively achieving remarkable performance by selecting compatible paths for optimization. Meanwhile, Parametric Content Similarity NMS (PCSNMS) is developed where detailed information on proposal boxes is associated. Comparative experiments (datasets on FLIC, MPII, etc.) with CNN-based counterparts have empirically demonstrated the effectiveness and competitiveness of the model in perspective of accuracy, memory consumption, and computation perplexity. Our model contributes to an efficient and feasible framework of human behavior apprehension.
Despite the great performance of language identification (LID), there is a lack of large-scale singing LID databases to support the research of singing language identification (SLID). This paper proposed a over 3200 hours dataset used for singing language identification, called Slingua. As the baseline, we explore two self-supervised learning (SSL) models, WavLM and Wav2vec2, as the feature extractors for both SLID and universal singing speech language identification (ULID), compared with the traditional handcraft feature. Moreover, by training with speech language corpus, we compare the performance difference of the universal singing speech language identification. The final results show that the SSL-based features exhibit more robust generalization, especially for low-resource and open-set scenarios. The database can be downloaded following this repository: https://github.com/Doctor-Do/Slingua.
In the era of deep learning, various segmentation tasks have been studied. As a new segmentation task, panoptic segmentation has been proposed and studied by researchers recently. This article summarizes the basic ideas of the panoptic segmentation method based on deep learning and classifies the current image panoptic segmentation into four categories: top-down, bottom-up methods, single-path methods and other methods. In some methods, they are further divided into several small categories. The characteristics and limitations of each method are analyzed, and the segmentation effects are compared. In addition, video panoptic segmentation and LiDAR data panoptic segmentation are also involved. Finally, the possible future research directions are prospected.
Pedestrian reidentification is a popular research topic in the field of computer vision in recent years, and is a technique that uses computer vision techniques to determine whether a specific pedestrian is present in an image or video. After research and experiment, we found that YOLOv3-based pedestrian reidentification in practice has the problem of low accuracy rate of recognizing pedestrian pictures taken at night and cannot recognize multiple pedestrians at one time. In this paper, we improve the above problems by introducing a picture enhancement module to improve the brightness and defogging of night pictures before recognition, and improve the practice of averaging the distance values of multiple results for the same pedestrian to enable multiple targets pedestrian recognition. The experimental results demonstrate that the average accuracy rate of recognizing pedestrian pictures taken at night has improved from 6.85% to 80%, while the average accuracy rate of multiple targets pedestrian recognition has reached 85.9%, which is competent for multiple targets pedestrian recognition tasks at night.
Web application brings us convenience but also has some potential security problems. SQL injection attacks topped the list of Top 10 Network Security Problems released by OWASP, and the detection technology of SQL injection attacks has been one of the hotspots of network security research. In this paper, we propose a SQL injection detection method that does not rely on background rule base by using a natural language processing model and deep learning framework on the basis of comprehensive domestic and international research. The method can improve the accuracy and reduce the false alarm rate while allowing the machine to automatically learn the language model features of SQL injection attacks, greatly reducing human intervention and providing some defense against 0day attacks that never occur.
In view of the disadvantages of the existing popular RFID (Radio Frequency Identification) book inventory in the industry, such as expensive equipment, low accuracy, and high business threshold, this paper proposes a deep neural network-based inventory framework. By loading different deep neural network algorithm modules, the proposed framework is able to complete target detection of 1D/2D barcodes in photos or video streams. Faster R-CNN is used for 2D barcode and SSD-ResNet for 1D barcode. Then the target barcode and the coordinates of the barcode are intercepted, the books are sorted by the barcode coordinates, and the values are obtained through the barcode recognition to realize the book inventory. Through inventory test on 1D/2D barcode book, the accuracy and recall rate of the framework reached above 99%, with precision rate close to 100%. Compared with RFID inventory, the proposed deep neural network based inventory method is more accurate and precise, and increased the processing speed by around 35%.
This paper describes the ByteDance speaker diarization system for the fourth track of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). The VoxSRC-21 provides both the dev set and test set of VoxConverse for use in validation and a standalone test set for evaluation. We first collect the duration and signal-to-noise ratio (SNR) of all audio and find that the distribution of the VoxConverse's test set and the VoxSRC-21's test set is more closer. Our system consists of voice active detection (VAD), speaker embedding extraction, spectral clustering followed by a re-clustering step based on agglomerative hierarchical clustering (AHC) and overlapped speech detection and handling. Finally, we integrate systems with different time scales using DOVER-Lap. Our best system achieves 5.15\% of the diarization error rate (DER) on evaluation set, ranking the second at the diarization track of the challenge.
Object detection is the basic and key step of remote sensing image analysis. In optical remote sensing images, object detection faced many challenges such as multi-scale and small objects, appearance ambiguity and complicated background. To address these problems, a new method of object detection based on convolutional neural networks (CNN) and hybrid restricted boltzmann machine (HRBM) is proposed. Firstly, the detail-semantic feature fusion network (D-SFN) is designed to extract fusion features from low-level and high-level CNNs, which can make the target representation more distinguishable, especially for small objects. Secondly, context information is incorporated to further boost feature discrimination, which also improves the detection accuracy. Experiments on NWPU datasets show that the proposed method can significantly improve the accuracy of object detection and has certain robustness.