The utilization of QR codes in commodity anti-counterfeiting is a prevalent phenomenon. Their traceability via smartphones represents an effective authentication method. However, the captured codes are susceptible to misjudgment due to the influence of blur. To address the above issue, this paper proposes an innovative and efficient blind deblurring optimization method. The method first uses the multi-channel pulse enhancement technique to denoise and the spectral double-feature prior to detect blur. Then, the minimum brightness difference prior and the edge gradient prior are defined, and the code is partitioned into several sub-regions based on graphical guidance. Priors are integrated as constraints in both the maximum a posteriori estimation framework and the maximum likelihood estimation framework to estimate the blur kernel for each sub-region. After applying blur kernels to the corresponding sub-regions, a deblurring operation is performed. Finally, the deblurred sub-regions are stitched together to reconstruct a clear code. Experimental results demonstrate that the proposed method significantly improves the deblurring effect and reduces computation time, providing a novel technical approach to enhance the quality and recognition accuracy of code. The proposed method, as presented in this paper, can be generalized and applied to the deblurring of QR codes in a variety of contexts, including identity verification, product identification, medicine traceability, and food safety.
In order to simplify the complexity and reduce the cost of the microphone array, this paper proposes a dual-microphone based sound localization and speech enhancement algorithm. Based on the time delay estimation of the signal received by the dual microphones, this paper combines energy difference estimation and controllable beam response power to realize the 3D coordinate calculation of the acoustic source and dual-microphone sound localization. Based on the azimuth angle of the acoustic source and the analysis of the independent quantity of the speech signal, the separation of the speaker signal of the acoustic source is realized. On this basis, post-wiener filtering is used to amplify and suppress the voice signal of the speaker, which can help to achieve speech enhancement. Experimental results show that the dual-microphone sound localization algorithm proposed in this paper can accurately identify the sound location, and the speech enhancement algorithm is more robust and adaptable than the original algorithm.
Roller bearings are some of the most critical and widely used components in rotating machinery. Appearance defect inspection plays a key role in bearing quality control. However, in real industries, bearing defects are usually extremely subtle and have a low probability of occurrence. This leads to distribution discrepancies between the number of positive and negative samples, which makes intelligent data-driven inspection methods difficult to develop and deploy. This paper presents a small data-driven convolution neural network (SDD-CNN) for roller subtle defect inspection via an ensemble method for small data preprocessing. First, label dilation (LD) is applied to solve the problem of an imbalance in class distribution. Second, a semi-supervised data augmentation (SSDA) method is proposed to extend the dataset in a more efficient and controlled way. In this method, a coarse CNN model is trained to generate ground truth class activation and guide the random cropping of images. Third, four variants of the CNN model, namely, SqueezeNet v1.1, Inception v3, VGG-16, and ResNet-18, are introduced and employed to inspect and classify the surface defects of rollers. Finally, a rich set of experiments and assessments is conducted, indicating that these SDD-CNN models, particularly the SDD-Inception v3 model, perform exceedingly well in the roller defect classification task with a top-1 accuracy reaching 99.56%. In addition, the convergence time and classification accuracy for an SDD-CNN model achieve significant improvement compared to that for the original CNN. Overall, using an SDD-CNN architecture, this paper provides a clear path toward a higher precision and efficiency for roller defect inspection in smart manufacturing.
Microservice architecture is a promising architectural style. It decomposes monolithic software into a set of loosely coupled containerized microservices and associates them into multiple microservice chains to serve service requests. The new architecture creates flexibility for service provisioning but also introduces increased energy consumption and low service performance. Efficient resource allocation is critical. Unfortunately, existing solutions are designed at a coarse level for virtual machine (VM)-based clouds and not optimized for such chain-oriented service provisioning. In this paper, we study the resource allocation optimization problem for service request routing and microservice instance placement, so as to jointly reduce both resource usage and chains’ end-to-end response time for saving energy and guaranteeing Quality of Service (QoS). We design detailed workload models for microservices and chains and formulate the optimization problem as a bi-criteria optimization problem. To address it, a three-stage scheme is proposed to search and optimize the trade-off decisions, route service requests into instances and deploy instances to servers in a balanced manner. Through numerical evaluations, we show that while assuring the same QoS, our scheme performs significantly better than and faster than benchmarking algorithms on reducing energy consumption and balancing load.
In feature-based image matching, implementing a fast and ultra-robust feature matching technique is a challenging task. To solve the problems that the traditional feature matching algorithm suffers from, such as long running time and low registration accuracy, an algorithm called feedback unilateral grid-based clustering (FUGC) is presented which is able to improve computation efficiency, accuracy and robustness of feature-based image matching while applying it to remote sensing image registration. First, the image is divided by using unilateral grids and then fast coarse screening of the initial matching feature points through local grid clustering is performed to eliminate a great deal of mismatches in milliseconds. To ensure that true matches are not erroneously screened, a local linear transformation is designed to take feedback verification further, thereby performing fine screening between true matching points deleted erroneously and undeleted false positives in and around this area. This strategy can not only extract high-accuracy matching from coarse baseline matching with low accuracy, but also preserves the true matching points to the greatest extent. The experimental results demonstrate the strong robustness of the FUGC algorithm on various real-world remote sensing images. The FUGC algorithm outperforms current state-of-the-art methods and meets the real-time requirement.
首先将一幅100像元×100像元的高光谱图像划分为以5像元×5像元构成的区域,作为大像元的图像,大像元中25个像元特征值的均值作为一个大像元的特征值.每组图像有175幅图像,实验中将他们分成四段,发现第三、四段的信息熵较小(其数值仅为二或三位数的数量级),仅用第一、二段的数据进行分类.选用图像信息熵≥1 500的图幅,进行该类图像信息熵均值的计算.农田、山体、居民地、水体4类不同地物的信息熵的均值和4种地物中某个地物的属性,利用其信息熵值与4种不同地物信息熵的均值比较,差值最小者的属性,即为待定地物的属性.
Many deep learning models, such as convolutional neural network (CNN) and recurrent neural network (RNN), have been successfully applied to extracting deep features for hyperspectral tasks. Hyperspectral image classification allows distinguishing the characterization of land covers by utilizing their abundant information. Motivated by the attention mechanism of the human visual system, in this study, we propose a spectral-spatial attention network for hyperspectral image classification. In our method, RNN with attention can learn inner spectral correlations within a continuous spectrum, while CNN with attention is designed to focus on saliency features and spatial relevance between neighboring pixels in the spatial dimension. Experimental results demonstrate that our method can fully utilize the spectral and spatial information to obtain competitive performance.
引用信息论中信息熵可以区分不同信息源包含不同信息量的思想,解决图像分类的问题.在图像分类中为了节省计算时间,将每幅100×100像元图像化算为以5×5像元构成的区域,作为大像元的图像,大像元中25个像元特征值的均值作为一个大像元的特征值,本文的图像分类是在大像元图像上进行的.首先计算出每幅大像元图像的信息熵,按照已知信息将图像分为5类,在每一类别中以图像总数的1/4的原则确定该类图像的取样数目,由每一类别中每幅大像元图幅的信息熵,计算各类别大像元图像取样的信息熵均值.在这个基础上,选择3类图像作为一个组合,计算待检验图像的信息熵Hl,分别与3个取样信息熵的均值HCPa、HCPb、HCPc之间的绝对值 Δ1、Δ2、Δ3,其中 Δi的最小者所属的类别,便是待检验图像的类别.通过实验证明本文提出的大像元信息熵,用于图像分类是可行的、有潜力的.
Quickly establishing reliable correspondence between two feature sets is a challenging task for feature matching. However, the key to successful feature matching is not only matching robustness but also the precision and real-time performance. It is difficult to achieve both efficiency and efficacy using the current algorithms. In this paper, we propose unilateral grid-based clustering (UGC), which creates a unilateral grid of an image's features and meanshift clustering constraints of the other image correspondence features. UGC removes a large number of mismatches using clustering center statistical analysis of the match feature points in a grid region. For low texture, blur and wide-baselines feature matching of images, UGC provides a real-time, ultra-robust correspondence system. Extensive experiments on image data sets demonstrate the higher precision and real-time performance of UGC, which outperforms current state-of-the-art methods, including conditions such as low contrast and high exposure.
随着人工智能的新一轮崛起,嵌入式技术有了更为广阔的发展空间.传统教授方法下培养的学生个性化缺失,与当前社会发展的需求明显不相适应.针对武汉大学电子信息类学科理工结合的特点,该文通过优化课程结构,扩充和丰富嵌入式技术教学内容,自制教学仪器设备,开展多样化教学等改革举措,探索了在综合性大学背景下,培养掌握嵌入式技术的电子信息大类个性化人才的方法.实践结果表明,该方法教学效果显著,取得了较好的成果.
SLAM is the current research hot spot in robot area and is considered to be the key of achieving robot's full autonomous movement. Traditional RGB-D SLAM algorithms compute camera's pose via SIFT descriptor. On account of the complexity of extracting SIFT descriptor, siftGPU is applied to accelerate this procedure, which makes it unsuitable for embedded equipment. Besides, traditional algorithm is inefficient in loop closure detection and is bad at instantaneity. Therefore, a new approach combined ORB feature and visual dictionary is proposed. In the front end of the algorithm, ORB feature of the adjacent images is firstly extracted and then the nearest neighbor and the second nearest neighbor is found by k-Nearest Neighbor(kNN)algorithm. Once found, ratio test and cross test is adopted to remove outliers. Secondly, a modified PROSAC-PnP algorithm is used to calculate the high-accuracy estimation of camera's pose. In the back end, a loop closure detection algorithm based on visual dictionary is performed to reduce accumulated error of the robot's move-ment, which adds new constraint to pose graph. Finally, generalized graph optimization tool is used to perform global pose optimization and global consistent camera pose and point cloud is got. The test and comparison on the standard dataset shows that this algorithm can obtain better robustness.
We put forward a method for image classification based on data gravitation.The quality of the data particles we use is the feature of the images(such as the fractal dimension of the image).For each kind of training data particle set,we use the mean of several image characteristics mi in the set to be the image characteristic,and the number of images wi to be its weight.So the quality of the i-thtraining data particle set is wi mi,and the data particle for inspection is atomic data particle with the mass of 1.Assuming that the characteristic of the image data particle for inspection j is tm,the distance between the i-th training data particle set and the data particles for inspection j of the image is |mi-tmj |.Assuming that there are three different catagories of images,we choose the part of images from all kinds in order to compose three kinds of training data particle set,then work out the characteristic mean of data particle set of each kind and characteristic value of a data particle for inspection,which can be used to calculate the gravitation between data particle set of each kind and characteristic value of data particle for inspection.The kind with largest gravitation is the one of data particle for inspection.It is proved by the experimental result that the method for image classification based on data gravitation has a certain advantage.
For the crowd safety and social stability, it's important to take crowd density estimation on public areas, such as scenic spots.Due to illumination changes, different camera height and angles and the pedestrian occlusion, it's difficult to make an accurate estimation with the existing methods.A method based on ensemble learning with support vector regression is proposed to estimate the crowd density.First, the scene is divided to multiple levels patches using the head width as a reference, then the first layer support vector regression mode is used to make coarse prediction on three feature descriptors extracted on the image patches.The prediction results are used as a new feature and to make fine prediction with the second layer support vector regression mode.At last, the crowd density is estimated with the sum of all patches prediction results and the classification standard which is set according to the characteristics of the scene.Experimental results show that the method proposed can achieve the classification accuracy above 85% on many scenes of the scenic spots, and it's an effective and robust crowd density estimation algorithm.
In this article a new method based on MRF to classify image texture texton has been put forward.The constraint relationship between the center pixel feature value and the neighbor pixels feature value in MRF can reflect the features of image texture texton as well as different MRF parameters.Standard deviation based on the MRF parameter of the same category is the smallest.So we can use this property to classify image texture.By comparing the different experimental scheme and different classification method,we can come to the conclusion that the method of image texture element classification proposed in this paper has certain advantages,and it is a good methold of image classification.
Crowd motion estimation is an important part of crowd action analysis.Crowd motion Analysis in special places is a necessary action for maintaining the safety and social stability in public place and there is a research difficulty in the field of intelligent video monitoring.Existing approaches for crowd motion estimation based on traditional cameras have the limitation of small field-of-view and more blind spots.This paper proposes a crowd motion estimation approach based on the feature point optical flow employing the advantages of large field-of-view and no blind spot of fisheye cameras.Firstly,the original images are preprocessed using the method of background difference based on Gaussian Mixture Model with area feedback,and the region of interest (ROI) is obtained by circle fitting.Secondly,a feature point extraction method based on non-uniform sampling of edge density is presented to describe the moving crowd for improving the real-time performance as the same time as ensuring the accuracy of describing the crowd.And then the optical flow field is calculated using the method by Lucas & Kanade.Finally,a perspective weight model of the fisheye camera is developed to weighting the compute the motion vector and the motion direction and speed of the crowd in fisheye camera images in order to solve the issues of the size differences of the crowd in long and short distances and the distortion of fisheye images in this paper.The experimental results show that the proposed approach is effective and feasible for estimating the motion speed and orientation of the crowd in dense crowd.In addition,the proposed method provides an important research basis for crowd behavior analysis.
For color correction in the stitching of outdoor scene,there may be grain noise and block artefacts in result images because of the large difference of brightness between images.To solve this problem,a method of color correction using gradient region segmentation was proposed.The gradient and color information were combined to divide each image into two regions to match the areas with similar gradient level and color character.To avoid the mutual influence between regions,simple statistical characteristics correction method was used for regions with simple color,while modified multidimensional probability density function transfer was used to deal with regions with complex color.The approach was tested on a number of different scenes and compared with multidimensional probability density function transfer method.The stitching results after correction and quantita-tive indexes were given to evaluate two methods.The experimental results show that the method works well to reduce the color distortion and keeps the structure of the original image,thus it obtains better performance of color correction.
In order to overcome the difficulty of traditional multi-camera video stitching method based on CPU or GPU in satisfying both the real-time performance and visual effect,this paper proposes a CUDA-based real-time seamless HD video stitching method.It solves the visual troubles caused by seams in stitching the moving objects by combining the static seam masks of Graphcut pre-treatment and the blending algorithm of image spatial domain,meanwhile puts the emphasis on studying the optimisation strategy of implementation of stitching procedures including perspective transform and image blending in CUDA.Experimental results demonstrate that under the condition of obtaining extra wide filed-of-view video by real-time stitching with four 1080 HD web cameras,the method achieves higher speedup ratio compared with CPU-based algorithm,and satisfies the real-time requirement on GPUs with different computing capability and architecture,and possesses better visual quality as well.
提出带有确定度的关联度的图像模糊分类新方法.该方法在求得每幅图像相对各个图像类别的关联度基础上,求得一幅图像相对各个类别的确定度,将一幅图像的确定度作为相应图像关联度的权,以带权的关联度为最大准则,确定待识别图像类别的属性.识别图像类别的第二个准则是,以带有确定度的关联度为特征,采用与聚类中心距离最小准则,确定待识别图像的类别.通过由三种不同类别图像组成的多种组合的试验,试验中满足两个准则之一的实验结果表明,该方法的结果具有一定的优势.
针对现有智能交通系统仅仅通过车牌信息获取车辆信息存在不准确的情况,提出一种基于联合层特征的卷积神经网络(Multi-CNN)进行车标识别。该方法将通过卷积神经网络中不同层提取的特征联合起来,一起作为全连接层的输入,训练获得分类器。通过理论分析和实验表明,与传统的卷积神经网络训练获得的分类器相比,MultiCNN方法能够减少训练所需计算量,同时将车标识别准确率提升至98.7%。
To improve the real-time performance of high-definition video stitching,a stitching method of GPU-based real-time multi-channel high-definition YUV video is proposed.Firstly,the perspective model of YUV422 image stitching is deduced,and then parallel optimization of the stitching steps including perspective warping and image blending are realized and performed on GPU by Compute Unified Device Architecture(CUDA)technology.The experimental results show that,on the condition of implementing four-channel 1080p video stitching,the method in this paper achieves 20%~40% improvement in real-time performance over RGB color model-based stitching method on different GPU,and video frame rate of 33 frame per second on GTX 780.