Accurate positioning of robot end effectors is essential for many industrial applications. Conventional robot calibration methods primarily identify deviations in kinematic parameters. At the same time, complex nonkinematic errors, such as those caused by structural deformation and varying external payloads, are often neglected, resulting in reduced positioning accuracy. To address this issue, this article proposes a two-step calibration method to improve robot positioning accuracy under external payloads. The method is developed based on the modified Denavit-Hartenberg (MDH) model and incorporates a data-driven error estimation strategy. In the first step, a convolutional neural network (CNN) predicts the positioning error of the robot end-effector under no-load conditions. In the second step, a dual-branch convolutional neural network with an attention mechanism (DB-CNN-AM) estimates the additional positioning error induced by varying payloads. The predicted errors are combined and compensated within the MDH-based model to obtain corrected target positions. Experiments on a TB6-R10 industrial robot demonstrate the effectiveness of the proposed method. When the external payload is less than 7 kg, the maximum positioning error is reduced from 4.197 mm to 0.240 mm, and the average error decreases from 1.792 mm to 0.096 mm, corresponding to a reduction of 94.64%.
High-precision hand-eye calibration for robotic chromatic confocal sensors (CCSs) is often hindered by heteroscedastic noise caused by varying inclination angles. At large detection angles, the significant attenuation of reflected energy leads to deteriorated signal-to-noise ratios, causing traditional equal-weight methods to suffer from severe rotational bias. This paper proposes an adaptive calibration framework that integrates intensity-based confidence weighting with Variable Projection (VarPro) optimization. A multi-pose concentric scanning strategy is designed to construct a comprehensive geometric constraint space. By establishing a physical coupling model between backscattered intensity and measurement uncertainty, an adaptive weighting function is derived to suppress the negative impact of low-SNR boundary data. Furthermore, the VarPro algorithm is utilized to eliminate sphere center parameters analytically, resolving the numerical coupling between translation components and the artifact centers. Experimental results demonstrate that the proposed method effectively mitigates edgeinduced artifacts, reducing the RMSE from 0.153 mm to 0.082 mm and lowering the across-pose standard deviation to 0.026 mm.
The pursuit of high-precision manufacturing necessitates advancements in coordinate metrology, where scanning probes on Coordinate Measuring Machines are pivotal. However, the accuracy of dynamic measurements is critically influenced by data segment selection and the setting of the deflection threshold, aspects that have not been systematically optimized. This study introduces a novel dynamic measurement methodology to address these challenges. First, an in-depth analysis of the probe's motion dynamics during the approaching and retraction stages reveals distinct characteristics. This finding provides the theoretical foundation for selecting the retraction stage data. Second, accounting for the nonlinear elastic behavior of the probe's parallel leaf-springs, a linearized deflection-error model is established, enabling the determination of an optimal deflection threshold. Experimental validation on a self-developed CMM platform demonstrates that the proposed strategy achieves a measurement error of 0.9 mu m for a gauge block, representing an 86.4% improvement over the comparative method, which verifies the effectiveness of the proposed method.
Shape from focus is widely used for microscopic three-dimensional morphology reconstruction, but its performance is often limited by noise-sensitive discrete peak localization, weak-texture regions, boundary over-smoothing, and the computational burden of large-scale image sequences. To address these issues, this paper proposes a Gaussian process regression-based SFF method with uncertainty-guided filtering. The discrete focus responses along the optical axis are treated as samples from an underlying continuous focus function, and Gaussian process regression is used for Bayesian interpolation and sub-frame focal-position estimation. To improve efficiency, shared kernel hyperparameters are learned from representative pixels, and the kernel matrix decomposition is reused for full-image inference. In the refinement stage, the posterior uncertainty is converted into a confidence weight to adaptively fuse the initial GPR depth map with the structure-guided filtering result. Experiments on synthetic surfaces show that the proposed method achieves RMSE values of 0.36021, 4.39122, and 3.19075 on the Cone, Cosine, and Sinusoidal datasets, respectively. On the HCI14 public synthetic benchmark, the proposed method obtains best or near-best RMSE and CORR. results in most scene-wise comparisons, with an average running time of 1.99 s. Practical measurement performance is further evaluated using calibrated cross-groove and step blocks. The proposed method shows height-difference deviations of −0.031 μm and −0.014 μm from the certified values, with RMSE values of 0.060 μm and 0.048 μm, respectively. These results indicate that the proposed framework reduces peak-localization fluctuation and improves height-difference measurement accuracy for weak-texture shallow grooves and sharp step discontinuities.
To address the challenges in microscopic three-dimensional morphology measurement, including noise-sensitive peak localization, the trade-off between boundary preservation and noise suppression, and the high computational cost of large-scale image sequences, this paper proposes a Gaussian process regression-based shape from focus morphology recovery method combined with uncertainty-guided filtering. First, the discrete focus responses of pixels along the optical axis are modeled as continuous stochastic processes, and Gaussian process regression is used to probabilistically model the focus curves. The predicted mean is then used for sub-pixel focal position estimation, which reduces the effects of false peaks and peak-position fluctuations in conventional discrete peak searching. To reduce the computational burden of pixel-wise probabilistic inference, kernel hyperparameters are learned from a representative pixel subset, and the kernel matrix decomposition results are reused for fast full-image inference. Furthermore, in the depth refinement stage, an uncertainty-driven structure-guided fusion mechanism is introduced. The mean intensity image is used as guidance information, while the uncertainty of the Gaussian process is used to characterize pixel-level uncertainty. The filtering weights are adaptively adjusted to preserve continuous estimates in high-confidence regions and enhance structural constraints in low-confidence and boundary regions. Experiments on synthetic surfaces, the HCI14 public dataset, and standard samples demonstrate that the proposed method achieves a favorable balance among reconstruction accuracy, structural consistency, and computational efficiency. Experiments on standard cross-groove and step blocks yield root mean square errors of and , respectively, verifying the robustness and practical measurement capability of the proposed method in weak-texture, shallow-structure, and depth-discontinuity scenarios.
Industrial robots suffer from limited absolute positioning accuracy, which restricts their applications in precision manufacturing. To address this limitation, this paper proposes a high-order joint-dependent kinematic error modeling and compensation method that explicitly accounts for joint flexibility induced by the robot’s self-weight and external payloads. Unlike conventional calibration models that assume constant kinematic errors, the proposed approach incorporates flexibility-related parameters into a high-order joint-dependent error formulation, enabling more accurate representation of configuration- and load-dependent positioning errors. In addition, a hybrid sampling strategy is developed to optimize the selection of measurement configurations for parameter identification. Joint-related geometric error parameters and flexibility parameters associated with self-weight and external payloads are identified, and the resulting model is applied for positioning error compensation. Experimental results demonstrate that the proposed method significantly improves the robot’s absolute positioning accuracy. Specifically, the maximum positioning error is reduced from 4.197 mm to 0.115 mm, while the average positioning error decreases from 1.405 mm to 0.043 mm. Furthermore, comparative experiments under different external payloads show that the proposed method consistently achieves the lowest root mean square error (RMSE) among several existing error models, demonstrating superior generalization capability.
Kinematic modeling and parameter identification are essential for achieving high-precision robot calibration. A widely used strategy involves utilizing the end-effector position error for parameter identification. However, the strong coupling between length and angular parameters often impedes calibration accuracy. In addition, substantial differences in their scales further exacerbate this issue. To overcome these limitations, following the variable projection method, this paper reformulates the conventional Modified Denavit–Hartenberg (MDH) model into a separable nonlinear structure. This allows independent identification of the two parameter types. Non-geometric errors such as joint compliance and backlash are also explicitly taken into account. The backlash errors are separated from the angular positions of each joint by modeling their bidirectional positioning errors with Chebyshev polynomials. This method enables the establishment of a comprehensive positioning error model to mitigate the influence of backlash errors. Based on the variable projection method, an improved variable projection with modified Gram–Schmidt (IVPMGS) identification method is proposed, which also eliminates redundant parameters that hinder identification robustness. Simulations indicate that the proposed method achieves faster convergence and higher identification accuracy. Compensation experiments demonstrate that the average absolute positioning error is reduced from 0.1804 mm to 0.0917 mm compared with the traditional MDH model, corresponding to a 49.17% improvement in positioning accuracy. These findings confirm the accuracy and effectiveness of the proposed approach.
Uncalibrated photometric stereo aims to recover surface normals from images captured under varying and unknown illumination. This problem remains challenging due to the ambiguity among geometry, reflectance, and lighting, especially in regions affected by non-Lambertian effects and local shape discontinuities. In this paper, we propose a quality-aware two-stage framework for normal estimation in uncalibrated photometric stereo. The method first extracts per-image features with a visual encoder and aggregates them through a quality-aware fusion module that jointly considers local pixel-wise reliability and global image-level reliability across illumination observations. An initial normal map is then predicted from the fused representation. To further recover fine geometric details, a refinement stage is introduced by combining the initial normal estimate with explicit multi-illumination statistical cues, including per-pixel mean, standard deviation, and intensity range. The refinement module learns a residual correction to progressively improve the normal prediction. Unlike conventional black-box regression methods that treat all input images equally, the proposed framework adaptively weights more informative observations and enhances reconstruction robustness in challenging regions. The whole network is trained end-to-end with supervision on surface normals. Experimental results show that the proposed method achieves accurate and spatially consistent normal estimation while preserving fine-scale details.
Addressing the complexity of composite error coupling modeling and compensation for coordinate measuring machines (CMM), this paper proposes a collaborative optimization method for error element modeling and compensation. Traditional composite error models typically separate and integrate errors using function approximation approaches, which result in limited prediction accuracy under varying temperature conditions. As a result, a deep pyramid convolutional neural network (DPCNN) model is constructed. It achieves a nonlinear mapping from position and temperature parameters to composite errors. The complexity and low accuracy issues of composite error modeling are resolved. To address the limitations of conventional coupling effect evaluation methods, an improved sensitivity analysis method is employed to quantify error coupling effects. Geometric errors are classified based on first-order sensitivity. It avoids issues arising from small differences between total and first-order indices that hinder the evaluation of coupling effects. The improved method enables a clearer analysis of coupling effects while reducing computational complexity and cost. To mitigate the influence of error coupling and enhance compensation efficiency and accuracy, an error proportion compensation approach is proposed. The compensation ratio is calculated using the error distribution characteristics output by the DPCNN, thereby enabling targeted adjustment of key error components. Experimental results show that this strategy enhances compensation accuracy while reducing the number of compensation terms. Compared with traditional methods, the compensation accuracy is improved by 48.42%. This study demonstrates the practical impact of precise error modeling on compensation strategies and provides a systematic solution for multi-source coupled error analysis.
Accurate positioning of a robot end effector is essential for precision operations. Conventional calibration methods typically identify kinematic parameters from end-effector spatial errors, but they overlook parameter coupling and thus offer limited accuracy improvements. To address this limitation, this article proposes a kinematic comprehensive error model based on error separation technology. The method separates the pose error caused by bidirectional angular positioning deviations of joint rotation axes from the spatial error of the end effector. Subsequently, it identifies rigid-flexible coupling kinematic parameters under the influence of link self-gravity. To enhance parameter identification, a pose error measurement method is developed using a specially designed experimental tool, and the iterative one-by-one forward floating search (IOOFFS) algorithm is proposed as the pose selection strategy. Based on the identified parameters, joint angle compensation is used to realize kinematic error compensation. Experimental validation on the TB6-R10 robot demonstrates that the proposed method can significantly enhance the position accuracy of the robot end effector. Specifically, the maximum absolute positioning error is reduced from 2.995 to 0.223 mm, and the average positioning error is reduced from 1.738 to 0.094 mm, corresponding to an overall reduction of 94.61%.
Density-based Spatial Clustering of Applications with Noise (DBSCAN) is a widely used clustering algorithm based on density measures; however, it performs poorly on multi-density datasets. Additionally, it necessitates two key input parameters: the radius (Eps) and the minimum number of points (MinPts), both of which need to be experimentally determined for accuracy. To address these issues, this paper proposes an adaptive clustering method based on an improved DBSCAN, which determines parameter pair based on the data distribution. The method does not require pre-set parameters and can effectively cluster multi-density datasets. In order to assess the efficacy of the proposed method, experiments were carried out using synthetic datasets and the results were compared with those of the conventional DBSCAN algorithm. The experimental results demonstrate that the proposed method can successfully cluster multi-density datasets without the need for manual parameter adjustment.
Solar photovoltaic (PV) cells are the primary elements of the PV power generation process, and their quality directly influences the overall efficiency and reliability of the power generation. Visual inspection of PV electroluminescence (EL) images in the factory is a classical method for defect detection, but it is a time-consuming and labor-intensive process. Therefore, an improved YOLOv8 model YOLOv8-DGN was proposed for EL images. In this paper, we introduced depthwise separable convolution (DWConv) and GhostConv into YOLOv8n to reduce the number of parameters and computational complexity. To improve the model’s detection performance on small-size defects, the Normalized Gaussian Wasserstein distance (NWD) was employed to replace the original loss function of YOLOv8. The experimental results showed that the proposed model YOLOv8-DGN was superior to the baseline model YOLOv8, with a mAP50 of 91.68
Six-degrees-of-freedom (6-DoF) object pose accurate estimation is an important visual task. However, due to complex application scenarios such as complicated target shapes and mutual occlusion, etc., traditional feature-matching methods result in low-pose estimation accuracy. Deep learning methods use multi-source information about the target and employ dense fusion prediction networks to estimate the target pose, providing high accuracy and robustness. However, the accuracy of target pose estimation depends on the input data accuracy of the prediction network. Therefore, we proposed an improved instance segmentation network SOLOv2 to identify and segment targets in complex scenarios based on the RGB-D data of targets, extracting multi-source information of the target such as color, depth, and mask. Then, we used an improved dense fusion pose estimation network SOLO-Dense to fuse the extracted multi-source information to achieve object 6-DoF pose estimation. Experiments showed that the average precision (AP) of the improved segmentation network increased by 5.1% compared with the original network. The results of our pose estimation method showed an excellent level of evaluation of the LineMOD and YCB-Video datasets. This method accurately segments targets and estimates their pose in cluttered scenarios, demonstrating the effectiveness of the proposed approach.
Solar photovoltaic (PV) cells are inevitably subject to defects during the production process, affecting their power generation efficiency and life. Electroluminescence (EL) imaging is the mainstream non-destructive method for PV cell defect detection. Aiming at PV cell EL images, an unsupervised defect detection method was proposed. Specifically, an unsupervised convolutional autoencoder (CAE), the scale structure perception convolutional autoencoder (SSP_CAE), was constructed, whose Squeeze-and-Excitation Attention (SE Attention) and skip connections avoid the blurring of image structure information and the loss of pixel-level details in the encoding and decoding process. Furthermore, to balance the global and local information of the image, a scale perception loss function called SP_SSIM was proposed for model training. The defect segmentation was achieved by using Otsu thresholding method to binarize the obtained Mean Absolute Error (MSE) residual heat image in the testing stage of the model. Finally, the experiments were performed on the test dataset and the experimental results showed that the proposed SSP_CAE can effectively detect PV cell defects. The experimentally obtained defect detection performance metrics Precision, Recall, IoU, F1-score and AUROC values were 0.739, 0.886, 0.723, 0.764 and 0.841, respectively. Compared with other classical methods, the proposed SSP_CAE had a better comprehensive performance for defect detection.
The resolution of underwater raw images is typically low, making them inadequate for practical applications. Therefore, high-resolution reconstruction of these images is imperative. However, most deep learning-based super-resolution (SR) algorithms tend to produce images that are excessively smoothed and lack fine details. To address underwater image SR, we introduce an underwater image super-resolution convolutional neural network based on multiscale dense residual blocks. First, utilizing two convolutional kernels of distinct scales, low-level features of the low-resolution original underwater images are extracted. Deepening the network's architecture and enhancing model perception are achieved through the concatenation of multilevel residual dense blocks and dilated convolution blocks. Each dense block incorporates local skip connections to fully harness feature extraction capabilities across all convolutional layers, facilitating the fusion of local features. Subsequently, global residual connection blocks amalgamate low-level features from diverse pathways, concurrently enriching fine-grained information and elevating training efficiency. Finally, within the reconstruction block, a sub-pixel convolution layer is introduced to replace deconvolution layers, thereby circumventing issues of manual information redundancy and significantly elevating the quality of image reconstruction. Experimental results show that the proposed method achieves excellent SR reconstruction of underwater images.
Grasp detection is a significant research direction in the field of robotics. Traditional analysis methods typically require prior knowledge of the object parameters, limiting grasp detection to structured environments and resulting in suboptimal performance. In recent years, the generative convolutional neural network (GCNN) has gained increasing attention, but they suffer from issues such as insufficient feature extraction capabilities and redundant noise. Therefore, we proposed an improved method for the GCNN, aimed at enabling fast and accurate grasp detection. First, a two-dimensional (2D) Gaussian kernel was introduced to re-encode grasp quality to address the issue of false positives in grasp rectangular metrics, emphasizing high-quality grasp poses near the central point. Additionally, to address the insufficient feature extraction capabilities of the shallow network, a receptive field module was added at the neck to enhance the network's ability to extract distinctive features. Furthermore, the rich feature information in the decoding phase often contains redundant noise. To address this, we introduced a global-local feature fusion module to suppress noise and enhance features, enabling the model to focus more on target information. Finally, relevant evaluation experiments were conducted on public grasping datasets, including Cornell, Jacquard, and GraspNet-1 Billion, as well as in real-world robotic grasping scenarios. All results showed that the proposed method performs excellently in both prediction accuracy and inference speed and is practically feasible for robotic grasping.
In industrial visual measurement, converting point clouds into depth maps is a widely adopted technique to enhance data processing efficiency and structural representation. However, the process is plagued by voids and structural distortions arising from non-uniform sampling, occlusions, and projection ambiguities. To address these issues, we propose an efficient method for generating orthographic dense depth maps. The method's novelty lies in three key contributions: a visibility-prioritized preprocessing framework to suppress depth distortion, a robust depth fusion strategy to resolve projection ambiguities, and a composite inpainting algorithm to effectively restore void regions. Extensive experiments validate our method's state-of-the-art (SOTA) performance. For the task of generating orthographic depth maps, our framework improves the Chamfer Distance by up to 14.38% compared to the commercial platform VisionMaster. For the critical sub-task of depth completion, our sep_repair algorithm demonstrates superior robustness over the recent SOTA deep learning method, long-short range recurrent updating (LRRU) network. In the most challenging 'Severe missing' scenarios-where the deep learning model's performance degrades sharply-our method achieves a 23.87% reduction in root mean square error while completing the task in seconds. Furthermore, our entire framework achieves this SOTA-level performance efficiently on a standard CPU, highlighting its practical applicability for edge devices in smart manufacturing without the need for training data or GPU acceleration.