PurposePresently, 6 Degree of Freedom (6DOF) visual pose measurement methods enjoy popularity in the industrial sector. However, challenges persist in accurately measuring the visual pose of blank and rough metal casts. Therefore, this paper introduces a 6DOF pose measurement method utilizing stereo vision, and aims to the 6DOF pose measurement of blank and rough metal casts.Design/methodology/approachThis paper studies the 6DOF pose measurement of metal casts from three aspects: sample enhancement of industrial objects, optimization of detector and attention mechanism. Virtual reality technology is used for sample enhancement of metal casts, which solves the problem of large-scale sample sampling in industrial application. The method also includes a novel deep learning detector that uses multiple key points on the object surface as regression objects to detect industrial objects with rotation characteristics. By introducing a mixed paths attention module, the detection accuracy of the detector and the convergence speed of the training are improved.FindingsThe experimental results show that the proposed method has a better detection effect for metal casts with smaller size scaling and rotation characteristics.Originality/valueA method for 6DOF pose measurement of industrial objects is proposed, which realizes the pose measurement and grasping of metal blanks and rough machined casts by industrial robots.
The purpose of this thesis is to propose a 6DoF grasping position measurement strategy for robots based on monocular vision in response to the problems of high cost and poor stability of 6DoF grasping position measurement of target workpieces when industrial robots are grasping metal workpieces under vision system. The strategy adopts a simple structure, low-cost, easy-to-deploy monocular vision system to collect industrial sample data, and through the introduction of virtual reality technology and generative adversarial network, to realize the purpose of data enhancement for industrial metal small sample objects with surface defects and interference from the external environment, to solve the problem of insufficient data due to the difficulty of acquiring data in complex environments, and to achieve the problem of insufficient data due to the later application of monocular vision system as well. The purpose of continuously updating and expanding the sample data can be achieved by applying the monocular vision system in the later stage. At the same time, combining the deep learning-based multi-keypoint object detection technology and 2D-3D affine transformation, the method realizes the accurate measurement of the 6DoF position of the target workpiece. This method can measure the 6DoF position of industrial metal parts only by 2D plane images, and is a new method to estimate the 6DoF pose of a given object from a single RGB image using monocular vision, which possesses the ability to detect the pose of target objects with low-cost, stable, and strong antiinterference, reduces the cost of using robotic vision guidance technology in the industrial field, and provides a new way of thinking for monocular vision to be adopted by the industrial field. Provide a new idea for monocular vision to be adopted by the industrial field. Experiments show that the method has good detection accuracy for industrial small sample objects.
For the problems of insufficient number of samples and lack of sample diversity faced by industrial scenes in robot vision applications, this paper proposes an improved data enhancement strategy for industrial small-sample objects. The method combines a stochastic algorithm with a deep learning-based image generation technique to generate a large number of realistic real-life style target images from a small number of template images, thus effectively improving the recognition ability of target objects in complex backgrounds. In addition, virtual reality technology is introduced to generate virtual artifact images with similar styles to those captured by actual cameras using a virtual engine, further enriching the diversity and coverage of the dataset. This technique also employs an advanced algorithm that incorporates a convolutional network and a self-attention module in a cyclic adversarial generative network (CycleGan) to achieve efficient generation and style migration of industrial object images. This data enhancement strategy enriches the training data of the robot vision system and significantly improves the robot's stable detection ability and data quality of targets in complex environments.
In the face of the promotion of robot positioning and grasping applications in the industrial field, accurate identification and positioning of target objects in complex industrial environments have become a research highlights. To solve the problem of visual positioning interference caused by local reflections on industrial metal workpieces and obtain high-precision pose measurement results, this paper proposes a binocular vision 6DOF pose measurement strategy. Firstly, in view of the difficulties in manually collecting images and the high cost of collection, a data augmentation method based on generative adversarial networks is applied to achieve data augmentation and expand the the metal object image datasets. Then, in response to the problem of deformation and missing target objects in the images generated by the image data augmentation model, this paper proposes a method of changing the model architecture to improve the model's feature extraction ability for target objects and enrich the detailed features of the output images. Finally, aiming at the problem that the reflection of metal workpieces leads to incomplete information collection of the target object by the visual system and affects the pose measurement, this paper proposes a deep learning-based multi-keypoint detection method to provide effective data input for binocular vision 6DOF pose measurement. Combined with the binocular vision triangulation method to obtain the three-dimensional information of the target object, the pose measurement accuracy of the target object's 6DOF is improved. Experimental results show that this strategy outputs pose measurement data with high accuracy, providing data support for binocular vision-guided robot positioning and grasping.
To address the challenges related to poor positioning accuracy and high usage cost of 6DOF visual measurement systems in industrial settings, this paper presents a monocular vision-based robot vision guidance approach. The goal is to address the issues of expensive 6DOF pose measurement and limited measurement robustness when robots need to manipulate metal objects in industrial environments. The proposed approach enables precise and robust measurement of the 6DOF pose of the target workpiece. The approach integrates two main algorithms: a virtual reality-based image data enhancement algorithm and a 6DOF pose measurement algorithm that combines a multi-keypoint detection model and the Efficient Perspective-n-Points (EPnP) algorithm. The image data enhancement algorithm enhances the data of small-sample industrial objects using image enhancement techniques. This improves the robustness of the detection model by mitigating the challenges of high-cost image acquisition and long acquisition time associated with industrial objects. On the other hand, the 6DOF pose measurement algorithm performs the pose measurement of the target workpiece using a single image, enabling cost-effective 6DOF pose measurement by utilizing only a monocular camera. Experimental results demonstrate that the proposed method achieves measurement errors of 4.21% in the X direction, 2.94% in the Y direction, and 0.39% in the Z direction of the target workpiece. These results highlight the effectiveness of the proposed approach in achieving accurate and reliable pose measurement.
Deep learning models rely heavily on large amounts of data for training. This paper proposes a virtual sample generation technique (VSG) aimed at addressing the problem of few-shot object detection in the industrial field. The technique first obtains a new sample dataset by performing operations such as image rotation, cropping, and splicing on the original samples, and then generates a virtual sample dataset through CycleGAN to expand the sample dataset of industrial components, thus improving the performance of deep learning object detection networks. Finally, experiments are conducted and compared on YOLOv3 and YOLOv5. The mean average precision (mAP) obtained by this method is 2.3% and 2.4% higher than the two benchmark networks, respectively. The experimental results show that the proposed virtual sample generation technique for few-shot object detection in the industrial field is feasible, effective, and practical value.
Although strides in target detection, challenges persist in remote sensing image target detection due to complex backgrounds, small targets, and interference from background and other target information within the horizontal bounding box. To enhance remote sensing image target detection, this study introduces a YOLOv7-based attention mechanism for rotation target detection. The network model's loss function is first enhanced to extract rotating target angle information, ensuring accurate detection of the rotation angle. Subsequently, the ELAN module undergoes optimization to augment the model's feature extraction capability in remote sensing images by integrating an improved attention mechanism module and structural re-parameterization concept. Experimental testing on a self-built dataset demonstrates an overall detection accuracy of 92.7% for the proposed method, surpassing other mainstream rotational target detection algorithms in mAP. Results affirm that the proposed approach effectively enhances the rotation detection performance of remote sensing images.
The local stereo matching algorithm based on traditional census transform has some problems, such as over-reliance on the center pixel, fixed transform window and poor matching effect in weak texture region. An improved census transform stereo matching algorithm combined with adaptive window is proposed. Firstly, according to the mean square error of pixel gray level in the initial transformation window, the size of the transformation window is adjusted. Secondly, the census transform is improved and combined with the improved AD algorithm to form the initial matching cost; Then, the cost aggregation adopts the cross domain algorithm; Finally, the winner-take-all algorithm and multi-step optimization are used to obtain disparity map. Experimental results show that the proposed algorithm has good anti-interference performance against noise. By evaluating the standard images on Middlebury dataset, the average error of this method is reduced by 7.05 % compared with the traditional census transform stereo matching algorithm.
To solve the problem that factors such as the similar color of polyps and background in colon polyp images and the different sizes of polyps affect the segmentation accuracy, an improved SegFormer, U-SegFormer, is proposed for colon polyp image segmentation. Firstly, in the feature fusion phase of the decoder, features at multiple scales from the encoder output are fused in a cascade fashion and the feature representation is enhanced using the Unified Attention Fusion Module; Then, training the network in combination with transfer learning, using a loss function combining Dice Loss and Focal Loss to mitigate the effect of positive and negative sample imbalance on model training. Finally, the U-SegFormer was compared with other segmentation models, and the experimental results showed that the polyp segmentation method based on the U-SegFormer model was superior to the current mainstream segmentation methods and had a certain potential for clinical application.
Due to the long-term relatively fixed state of industrial installations, the diversity of industrial data is low, with high repetitiveness and uneven distribution, which makes industrial few-shot object detection still challenging. Therefore, from the perspective of handling the dataset, this paper proposes a virtual sample generation method based on virtual reality. Firstly, by using a virtual engine, the target workpiece model is created to directly simulate the states of the workpieces captured by the visual system in various environments. Secondly, random algorithms are added and improved to simulate workpieces with different materials and degrees of wear, distributed in random positions with different poses and types in the visual system, thereby generating virtual samples containing workpieces. Finally, virtual samples are introduced for experimental testing in the target detection network. The proposed method is tested in both YOLOv5 and YOLOv7 and compared with datasets without virtual samples, resulting in an improvement of 3.06% and 2.23% in mAP values, respectively. The experimental results demonstrate that the proposed method effectively improves the issue of industrial small-sample object detection.
Less effective information is obtained by the object detection network, due to the small size of the detection object in the entire image, the complex background, and the dense object in unmanned aerial vehicle (UAV) images. In response to the difficulties encountered, a small object detection method in UAV images is proposed as an improved YOLOv5-based algorithm in this paper. First, the space-to-depth(SPD) conv module is introduced into the basic feature extraction network, to improve significant loss of image information during downsampling. Then, various attention mechanisms are added, to intensify the acquisition of regions of interest in UAV images. Finally, the multiscale detection module is improved, to enhance the network's ability to detect small objects in UAV images. By conducting experiments on the VisDrone-DET2019 dataset, the test results of the established model show. The improved algorithm achieved a Mean Average Precision (mAP) of 41.8%, which is 7.8% better than the baseline network. In addition, the detection performance is better than most current mainstream target detection algorithms and is of some practical value.
The small and dense objects in unmanned aerial vehicle (UAV) images deteriorate the detection accuracy of neural networks. This article proposed an improved YOLOv5-based algorithm for small and dense object detection in UAV images. To enhance the capability of acquiring the feature information and the receptive field of the network in the backbone feature extraction network, we proposed an enhanced feature extraction (EFE) module, while incorporating the advantages of different pooling methods, and introduced the receptive field block (RFB) module, which realized fusion of different features. Meanwhile, we improved the multi-scale detection module, while enhancing the detection capability of the network for small objects in UAV images. Experiments were done on the VisDrone-DET2019 dataset. The improved algorithm achieved 39.4% mean average precision (mAP), which was 5.5% better than the benchmark network. The experimental results showed that the YOLOv5 algorithm proposed in this article was feasible and effective.