Robotic bin-picking of disordered, randomly stacked workpieces remains challenging because reliable grasping depends on an accurate estimate of object pose, yet many established solutions require high-precision 3D sensing, detailed object models, or large annotated datasets that raise the cost and effort of deployment on a new production line. This work presents a complete binocular vision framework that estimates workpiece pose by image matching and executes vision-guided grasping on a 6-DOF manipulator. A pose-annotated multi-view template library is constructed automatically through robot-driven image acquisition and compressed by a coarse-to-fine clustering scheme, and object pose is estimated by discriminative template matching with rigid refinement. To characterize the geometric reliability of the matched poses, an offline cross-modal analysis relates the 2D templates to a 3D reference model of the object and measures their agreement through region and contour reprojection metrics. Grasp configurations are then generated under orientation and collision constraints and corrected online by closed-loop visual feedback. Experiments on two representative workpieces show template-matching accuracy of 89–90% against classical and learned similarity measures, and grasp success between 81 and 87% across single-object and mixed scenes, outperforming the GraspNet baseline under the tested conditions. The framework offers an accurate and deployment-friendly route to robotic bin-picking.
Displacement is an important indicator for structural performance evaluation and its measurement accuracy directly affects the assessment of structural safety and service life. Although machine vision-based displacement measurement methods have exceptional advantages of longdistance shooting, non-contact operation, as well as time and labor efficiency, the surface features loss for monitoring target caused by the exposure to complex environmental conditions, such as illumination variations, fog, and surface occlusion, poses challenges for measurement accuracy in practical applications. In order to minimize the adverse effect of features loss, this paper proposes a novel deep learning-based displacement identification and quantification method with gradient correlation matching techniques. First, the end-to-end YOLOv8n algorithm is utilized to automatically extract calibration targets in complex scenarios, and exact position detection and reliable track could be realized even for mobile target under varying feature loss levels. Then, a gradient correlation matching technique based on machine vision is employed to analyze the correlation between image gradient vectors and corresponding pixels in overlapping regions. Finally, the actual displacement of structure is obtained through coordinate calculations. In laboratory experiments, the RMSE of displacement measured by this method does not exceed 0.03 mm and the results of outdoor experiment show high accuracy even the target exposed to illumination variations, heavy fog, and high occlusion rate; In-site bridge test further verifies its practicality and flexibility in engineering application.
Transferring the style from source domain to target domain for learning target models is a widely used strategy in domain adaptive segmentation. Although diffusion-based image translation has enabled flexible style transfer, it is often difficult to maintain the original structure of the image realistically during the reverse diffusion, which provides very little control over the generated image. To tackle this issue, we present the diffusion-based approach toward domain adaptive segmentation of general retinal image, which conditions diffusion models with carefully crafted input noise artifacts as explicit guidance at the inference step. Concretely, the input cross-domain image and the segmentation map of source domain are merged by summing the output of two encoders. Then, the encoder-decoder framework is adopted to iteratively refine the segmentation map by using a diffusion model. In order to enhance the general understanding of target domain distribution, we also establish the frequency-adaptive conditions for each sampling step. Moreover, this paper takes the channel-wise information and coarse semantic mask with noise of target image as guidance in the denoising process, which is different from existing approaches that input Gaussian noise and further establishes controllable conditions at the inference step. Extensive experiments on domain adaptation (DA)-based retinal image segmentation demonstrate the superiority of our approach over some state-of-the-art methods.
In high-definition imaging, solving the problem of out-of-focus blurring is one of the core challenges. For example, focus drift in telephoto lenses for distant targets affects subsequent analysis, necessitating a quantitative assessment of blur. Existing methods have drawbacks: subjective assessment lacks accuracy; reference-based objective methods rely on original, clear images, often unavailable in practice; feature-based no-reference methods have limitations and may misjudge complex images. For instance, EMBM is less sensitive to weak edges. Thus, a time-frequency domain no-reference assessment algorithm is proposed, with core innovations: first, a multi-scale feature extraction model integrating time-domain and frequency-domain features to comprehensively capture edge information across dimensions; second, a principal component analysis feature optimization module for dimensionality reduction and redundancy removal, enhancing key feature representation; finally, a dynamic weight allocation mechanism that specifically increases weak edge feature weights, solving EMBM weak edge neglect. Tests on the TID2013 dataset show that its SROCC index is 1.39% and 2.44% higher than that of EMBM, respectively.
Fisheye cameras, with their ultra-wide field-of-view (FOV) characteristics, are highly valuable for applications such as intelligent surveillance requiring large-scale human monitoring. However, human pose estimation in top-down fisheye images faces two critical challenges: first, the radial distribution of human targets caused by wide-angle imaging demands multi-directional annotated data due to the lack of rotational invariance in traditional algorithms; second, the non-linear distortion of fisheye cameras dynamically varies with spatial positions, creating strong correlations between geometric deformation and imaging locations that existing methods fail to model effectively. To address these limitations, we propose FishPoseNet, a distortion-adaptive pose estimation framework based on YOLOv8Pose. Our framework includes three core innovations: (i) a Geometric Anchor Alignment Module (GAAM) that standardizes human orientations via affine transformations, enabling full-scene directional coverage with single-direction annotated samples; (ii) a position-sensitive Dynamic Distortion Coefficient Estimator (DDCE) establishing continuous mapping from pixel coordinates to distortion coefficients; (iii) an enhanced Distortion-Aware YOLOPose (DA-YOLOPose) network that leverages distortion coefficients to guide feature fusion, improving adaptability to nonlinear deformations. Experimental results on a self-built top-down fisheye surveillance dataset demonstrate significant improvements, offering an innovative solution for high-precision human pose estimation in complex distortion scenarios.
Fighting detection helps maintain order and safeguard the lives of people in public. Using video surveillance to detect fighting in public places is a common and effective solution. Conventional detection methods focus on analyzing the angles between joints, striking, kicking actions, etc., of individuals engaged in fighting. These methods predominantly capture the action features of individuals, without considering the variation patterns of the motion trajectories during fighting. Moreover, this research faces a significant challenge: many non-fighting behaviors such as hugging, running and dancing may visually present similar action features to fighting. This similarity makes it difficult to accurately differentiate between fighting and normal behavior in practical applications. In contrast, during fighting, individuals exhibit intense motions and keypoints’ motion trajectories lack obvious regularity, which provides higher discriminability compared to other behaviors. Based on this, we proposed an effective and practical detection method that analyzes the motion trajectory patterns of different individuals’ keypoints over a period of time to identify the distinct patterns between fighting and normal behaviors. The steps are as follows: firstly, YOLO-Pose is employed to estimate the human pose, obtaining bounding box and individuals’ keypoints, in which centroid of the bounding box is considered one of the keypoints; secondly, DeepSORT is utilized to track individuals, acquiring the motion trajectories of keypoints; thirdly, using variance to identify the individual with abnormal speed variations in keypoints’ motion; finally, the entropy of motion directions of abnormal individuals’ is used to detect fighting behavior. The algorithm achieved 91.1
Due to the influence of both internal and external factors, fixed installed surveillance cameras often suffer from shakiness. In dynamic and complex scenes, frequent discontinuous depth variations and large foreground moving objects lead to multi-plane motion. This can lead video stabilization algorithms to misjudge local plane motion as global camera shakiness, resulting in stabilization failure or degraded performance. To address this problem, we propose a video stabilization algorithm based on the MeshFlow motion model. First, we propose a shakiness detection method and rules, which enables the stabilization algorithm to process only shaky frames, thus improving computational efficiency. Then, during motion estimation, we divide each frame into multiple mesh and construct a sparse motion field using motion vectors from mesh vertices to extract the camera’s shakiness trajectory. Finally, we apply Kalman filtering for trajectory smoothing, and use motion compensation to generate stabilized video. Experimental results show that the stabilized video achieves a PSNR improvement of 30
The nondestructive judgment of watermelon ripeness is a key issue in everyday life. Whether watermelons are ripe or not affects the interests of watermelon farmers and consumers. So far, acoustic, electrical, machine vision and other methods have been applied to the nondestructive judgment of agricultural product ripeness, but the equipment used in the lab, such as spectrometers and laser Doppler vibrometers, is expensive, difficult to carry, and has limited practical value in daily life. Based on practical experience, the ripeness of watermelons can be judged by observing specific physical characteristics, such as the rind's glossiness, texture, fuzz on the fruit stalk and navel. In this study, images of the watermelon's texture, fuzz and navel were captured using a smartphone. We utilized ResN et with an embedded Coordinate Attention mechanism and Swin-Transformer networks to independently analyze the ripeness of each feature. Then we combined the output prediction scores using post-fusion technology, with the highest weighted prediction score judging the final ripeness assessment. Our method achieved a 90.2% accuracy rate on the test set. The experiments demonstrate that a deep learning model utilizing multi-feature fusion of image data can reliably and accurately identify watermelon ripeness, even under varying image conditions.
The accuracy of defect detection in various locations is directly affected by the influence of precise segmentation of radial tire images. As the need for high-quality tires has grown, so has interest in the steel belt region. This region contains a diversity of textures and is impacted by shadows, and only a rough segmentation of this region has been achieved in the available literature, but this is interfered with by other areas, resulting in unsatisfactory defect detection in this region. As a result, we present a fine segmentation method for boundary mutation that takes advantage of the projection property of the local narrow window. First, the histogram adhesion change detection algorithm is designed; next, the narrow sliding window’s histogram peak width change algorithm is designed; finally, a wide sliding window is used for vertical projection, and the absolute value of the histogram change rate at the boundary is designed to maximize the algorithm. To calculate the outside boundary of the steel belt, the difference between the three sets of coordinates is kept within a specific error range. The detection was performed on 1,000 self-constructed data sets, with an accuracy rate of 94
In robotic tasks involving object grasping, pose estimation using accurate depth information combined with RGB images and neural network models has become a standard solution. Industrial-grade structured light cameras can capture precise depth information of targets and generally exhibit strong adaptability and stability. However, they are often unsuitable for deployment on robots due to their large size, high power consumption, high cost, and poor real-time performance. Although consumer-grade structured light cameras are cost-effective and offer real-time capabilities, they suffer from large measurement errors and depth loss due to reflective materials, especially in dynamic scenes where movement occurs, which severely impacts downstream grasping tasks. To address these issues, we propose the DiffDRNet model, designed to mitigate the limitations of consumer-grade structured light cameras through neural network-based depth completion. DiffDRNet reformulates this task as a conditional denoising diffusion process, using RGB images and incomplete depth maps as guiding conditions to “denoise” random depth distributions into high-precision depth maps. The network leverages Swin Transformers to extract multi-scale features from RGB images and incomplete depth maps with varying degrees of noise. These features are then fused using Content-Guided Attention to produce a feature map with channel specificity and interactive information. By executing the diffusion process in latent space through dedicated encoder and decoder designs, DiffDRNet achieves high-resolution, efficient depth completion. Experimental results on the NYU-Depth-V2 dataset demonstrate that DiffDRNet achieves high-quality depth completion, showcasing the potential of diffusion models for depth completion tasks.
During the production process of all-steel radial tires, defects are inevitable, resulting in nonconformities. By using digital image processing technology to analyze tire X-ray images, nonconforming tire products can be automatically detected. The cord slack defect occurs in the sidewall area, and the characteristic of this defect is that the up-and-down trend of a single or multiple cords is abnormal. Currently, no literature or algorithm for detecting this defect has been found. According to the characteristic of the abnormal trend of the cords in the slack defect, this paper designs a defect detection algorithm. The algorithm preprocesses the tire X-ray image, including region cropping, adaptive binarization, and skeletonization. Then, the relative pixel height of all pixels of each cord relative to the starting point of the cord is calculated, and the maximum and minimum slopes of each cord are calculated through the height to reflect the up-and-down trend of the cord. Finally, the presence of cord slack defect is determined by the change in slope. The experimental results show that the accuracy rate of this algorithm is 97.7%, and the detection takes 0.4 seconds. It can effectively detect the slack cord defect and meet the real-time requirements.
Weakly supervised semantic segmentation (WSSS) mainly adopts class activation map (CAM) to recognize different categories with generated pseudo masks of only image-level labels. Recently, many advanced works focus on learning the semantic correlation to refine the conventional CAM, which only identify sparse and discriminative semantic regions, severely weakening the further learning ability of spatial features. To copy with the above problem, a spatial correlation-guided learning framework is proposed to exploit the spatial and semantic correlation between adjacent pixels for weakly supervised fine-grained semantic segmentation (WSFGSS). From the spatial perspective, self-supervised multi-view clustering (SMC) is designed to fully mine spatial correlation by clustering of multiple views including scale, angle and position. Moreover, a hybrid self-supervised (HS) loss function is used to further promote the optimized speed and accuracy of three spatial representations. From the semantic perspective, the affinity matrix is applied to describe the semantic similarity between different pixels by building a weighted graph, and combine above robust pseudo label into a probability transition matrix. Therefore, the initial CAM is gradually corrected during the iterative optimization by the random walk algorithm. Finally, the refined CAM is utilized as supervision information for training the standard segmentation network effectively. Sufficient experimental results on BSDS500, PASCAL VOC 2012 and MS COCO datasets show that the proposed SMC method obtains more accurate pseudo labels than the recent unsupervised segmentation models. Meantime, with these pseudo labels, the proposed fine-grained framework achieves the state-of-the-art performance for WSFGSS.
Optical coherence tomography angiography (OCTA) is a new non-invasive imaging technology that provides detailed visual information on retinal biomarkers, such as the retinal vessel (RV) and the foveal avascular zone (FAZ). Ophthalmologists use these biomarkers to detect various retinal diseases, including diabetic retinopathy (DR) and hypertensive retinopathy (HR). However, only limited study is available on the parallel segmentation of RV and FAZ, due to multi-scale vessel complexity, inhomogeneous image quality, and non-perfusion, leading to erroneous segmentation. In this paper, we proposed a new adaptive segmented deep clustering (ASDC) approach that reduces features and boosts clustering performance by combining a deep encoder–decoder network with K-means clustering. This approach involves segmenting the image into RV and FAZ parts using separate encoder–decoder models and then employing K-means clustering on each part separated by the encoder–decoder models to obtain the final refined segmentation. To deal with the inefficiency of the encoder–decoder network during the down-sampling phase, we used separate encoding and decoding for each task instead of combining them into a single task. In summary, our method can segment RV and FAZ in parallel by reducing computational complexity, obtaining more accurate interpretable results, and providing an adaptive approach for a wide range of OCTA biomarkers. Our approach achieved 96% accuracy and can adapt to other biomarkers, unlike current segmentation methods that rely on complex networks for a single biomarker.
In the process of stitching images with equal depth of field, there are certain differences in brightness and colour between the stitched images, leading to the problem of large differences in brightness and colour at the seams of the stitched images. This affects people's visual effect, which in turn affects some related algorithmic effects on the basis of stitching big picture, so solving the problem of colour inconsistency in equal depth of field image stitching is an important research topic. Some existing methods for processing colour correction of a stitched image include global colour correction methods and local colour correction methods, while the local colour correction methods includes simple local correction methods and colour correction methods that fully considers local features of the image. However, the above methods suffer from inefficiency or poor image correction. In this regard, under the assumption that the colour features around the seams of a stitched large image are similar, a method for eliminating colour inconsistencies in stitched images using local statistical features in the region near the stitch line is proposed. The colour inconsistency in the stitched image is effectively solved by firstly calculating the difference between the colour statistical features in the region near the seams of the stitched images, then using the difference to correct the region near the two sides of the stitch, and finally extending it to the whole image for colour correction. The experimental results show that the method in this paper is simple and effective, which can keep the structure of the original image intact and improve the colour correction effect of the stitched image with better visual effect.
An effective semantic segmentation approach should contain fine object boundaries and continuous regions. However, recent mask-based segmentation cannot extract boundary features well on the coarse prediction, which causes obvious problems of blurry edges. Although several segmentation methods embed boundary detection branch to calculate contours directly, this type of architecture will lead to the increased computational complexity and miss the edge detailed information such as inter-class distinction. In order to obtain fine boundaries, we present a lightweight boundary refinement module with point supervision named BRPS to improve the boundary quality for the segmentation result generated by various existing segmentation models. Firstly, a direction field is learned to complete the initial feature rectification, which is defined as pointing away from the nearest object boundary to each pixel, where the weighted Euclidean and Cosine distance function is used as the loss between predicted boundary pixels and Ground-truth labels. Then, point-based supervised learning is performed at uncertain and certain locations including random distribution and key feature points based on a new point convolutional operation to output final crisp object boundaries. Finally, we verify that our BRPS module can effectively reduce the prediction errors for segmentation results generated from various state-of-the-art models such as DeepLabv3 and HRNet on the Pascal VOC2012, NYUD v2 datasets, Cityscapes and BDD100K datasets.
In the manufacturing process of piston, most of the piston cavities are printed with different character sequences to describe the specifications of piston. It is labor-consuming and inefficient to read the piston cavity character sequences manually. Although scholars have done a lot of researches in the field of industrial character recognition, there are few researches on piston cavity character recognition. A piston cavity character recognition method based on Faster R-CNN and priori knowledge library of character sequences is presented. First, we design a ring light source and an imaging device for the piston cavity based on the characters of the piston cavity protruding upward and the texture of the metal being easily reflective. Second, we use the character images in piston cavity collected by the imaging device to make character dataset. Third, according to the noisy background of the piston cavity image, Gaussian filtering and morphological operations were used to obtain a clean background image of the piston cavity. Fourth, use the Faster R-CNN training dataset to get the character recognition model, and then use the character type and position information detected by the recognition model to form a character sequence. Fifth, according to the highly similar characteristics of the character sequences of the same piston model, a character sequences prior knowledge library is constructed to correct the recognition results of Faster R-CNN. The experimental results show that the accuracy rate of the character sequences detected by the character recognition model is 95.5%. And when combined with the prior library of character sequences, the accuracy rate of the character sequences is 99%.
At present, the main idea of CNN-based unsupervised image segmentation is clustering a single image in the framework of CNNs. However, the single image clustering is very difficult to obtain enough supervision information for network learning. For solving this problem, we propose a Self-supervised Multi-view Clustering (SMC) structure for unsupervised image segmentation to mine additional supervised information. Based on the observation that the predicted pixel-level labels and the input images have the same spatial features, the multi-view images acquired by data augmentation are clustered to obtain the multi-view results and the proposed SMC uses the differences among these results to learn self-supervised information. Moreover, a Hybrid Self-supervised (HS) loss is proposed to make full use of the self-supervised information for further improving the prediction accuracy and the convergence speed. Extensive experiments in BSD500 and PASCAL VOC 2012 datasets demonstrate the superiority of our proposed approach.
Autonomous driving needs to obtain various traffic information promptly to make decisions, and the accurate positioning of traffic light in real time is one of the key points to realize autonomous driving. Up to now, there are many interesting and practical researches related to the positioning. However, most of the scholars are doing researches of daylight traffic light location. At night, Neon lights, street lights, car lights and other factors may have a serious impact on the accurate positioning of traffic lights. How to remove the influence of these factors is a big challenge in autonomous driving. We propose a practical approach to address this challenge. The motivation of this approach is as follows: Firstly, GPS information and electronic map are used to estimate the distance and direction of the current position relative to the traffic light; secondly, binocular vision is applied to get the depth, orientation and other information of each area in the image, excluding the interference area outside the scope of the traffic light; thirdly, use some methods to positioning the traffic light, including brightness value division, geometric features analysis, circular degree detection, and then the status of the traffic light can be recognized through the HSV color space. The experimental result shows that the positioning accuracy of this method is more than 95%, the average of processing time for each image is 301ms, which can meet the requirements of accuracy and real-time.