To overcome the limitations of manual measurement in prismatic truss wireframe modeling, including low efficiency and high subjectivity, this study presents a vision-based modeling method. A measurement system that integrates stereo imaging, coded targets, and visual measurement algorithms is established. Through multi-view image acquisition and three-dimensional reconstruction, the system obtains the spatial coordinates of coded points on truss members. An iterative rigid transformation method with adaptive composite weighting is introduced to achieve robust registration and coordinate unification of multi-view point sets. The same weighting mechanism is further applied to plane and line fitting to determine the intersections on each side. On this basis, a parameterized geometric optimization is employed to correct the vertical-web intersections and ensure global consistency. An affine transformation is then used to align the diagonal-web intersections with the corrected main structure, resulting in a standardized wireframe model. Simulation results demonstrate that the main geometric parameters of the reconstructed wireframe are recovered with relative errors below 0.12%. Real-world experiments on different truss structures further validate the feasibility of the proposed method for close-range engineering measurement and standardized wireframe reconstruction.
To address the subjectivity and inconsistency of manual centerline extraction for single-side truss members, this study presents a vision-based measurement method that integrates deep stereo matching with geometric constraints. A single pair of rectified stereo images serves as the input, from which a dense point cloud is reconstructed using FoundationStereo, a zero-shot stereo matching model. The region of interest within the point cloud is then projected orthogonally to obtain a high-contrast binary image for subsequent analysis. Based on this image, line detection is performed, and a triple geometric-constraint strategy is applied to automatically classify redundant line segments and select a unique representative main line for each truss member. Building upon this, a robust endpoint extraction and uniform perpendicular sampling scheme is introduced to enable accurate estimation of the centerline for each member through least-squares regression. Experimental results demonstrate that the proposed method achieves stable and precise performance in both simulated and real-world truss scenes with varying complexity. Its overall accuracy approaches that of the coded-target measurement method and demonstrates good applicability in practical truss measurement. The proposed approach provides a reliable solution for centerline extraction and geometric measurement of single-side truss members.
Few-shot Semantic Segmentation (FSS) aims to segment novel classes using only a limited number of annotated samples, where prototype learning approach extracts prototypes from the support images to guide the segmen tation of query images. However, when applied to query image segmentation, prototypes derived from limited support images cannot cover all the variations of the target class. This paper proposes a dual-branch few-shot segmentation network with prototype batch updating to enhance prototype stability. The dual-branch architec ture consists of Support Image Prototype (SIP) generated via masked average pooling of support features, and Language-Image Prototype (LIP) derived from Contrastive Language-Image Pre-training (CLIP). Building upon this architecture, a Global Prototype Pool (GPP) is designed to accumulate prototypes across batches, while in tegrating SIP through a momentum updating mechanism. To mitigate the distribution discrepancy between SIP and LIP, a Prototype Alignment Module (PAM) employs a contrastive loss to enhance prototype-level consistency. Furthermore, an adaptive dual-prototype fusion mechanism leverages learnable weights to dynamically balance SIP and LIP, which enhances adaptability under complex scenes and ambiguous semantic conditions. The net work demonstrates consistent performance improvements on the SBD, Pascal VOC 2012, and Cityscapes datasets, with average mIoU gains of 1.84%, 11.16%, and 0.77% over the baseline, respectively. Moreover, these results suggest that the proposed framework remains effective in FSS settings with limited training samples.
Monocular depth estimation is a pivotal task in computer vision, aimed to predict a dense depth map from a single image. Existing methods can achieve satisfactory performance through carefully designed network architectures. However, they often ignore the semantic information of scene structures, leading to poor performance in the semantic boundary areas of the predicted depth maps. To address this issue, we propose SSRDepth, a monocular depth estimation method that leverages semantic segmentation to improve metric depth prediction. Specifically, we decompose metric depth into segmentation scale and semantic-aware relative depth, which are predicted separately through Segmentation Proposal Scale (SPS) module and an Adaptive Relative Depth (ARD) module. The SPS module predicts the segmentation scale under the constraints of semantic segmentation and scale alignment, endowing the model with semantic awareness and maintaining sharp semantic boundaries. Meanwhile, the ARD module adaptively predicts semantic-aware relative depth within a normalized depth space through a hybrid regression approach. Furthermore, to effectively capture structural details in semantic segments, we propose the Segmentation Query Guided Aggregation (SQGA) module, which utilizes a novel feature interaction approach to enhance the aggregation of multi-scale features, guided by masks produced through a Segmentation Mask Generator (SMG). Extensive experiments show that our method achieves competitive results on the indoor, outdoor, and unseen datasets.
Rail-mounted gantry (RMG) cranes are critical handling equipment. The problem of traveling gantry gnawing on rail affects equipment safety and operational efficiency, and also increases maintenance costs, presenting a long-standing technical challenge in the crane industry. This paper proposes a method for establishing a rigid-flexible coupling dynamic model to investigate the rail gnawing phenomenon of RMG cranes. Focusing on a 37 m span RMG crane, a rigid-flexible coupling dynamic model is developed using ADMAS software and ANSYS software. The influence of three factors of different gantry dynamic stiffness, different installation methods of the wheels, and different wheel tread forms on the rail gnawing is tested on the rigid-flexible coupling dynamic model. The results show that: 1) when the dynamic stiffness of the gantry in the direction of the trolley is reduced to less than 0.60 Hz, the rail gnawing is significantly increased, while when the dynamic stiffness is increased to around 0.65 Hz, the rail gnawing is significantly alleviated; 2) the overall inclination angle of the traveling gantry wheels is between 0.1 deg and 0.15 deg, which can effectively reduce rail gnawing; 3) the use of specific tread shapes for traveling gantry wheels, such as curved or M-shaped shapes, can significantly alleviates the phenomenon of rail gnawing. Therefore, adopting the method proposed in this paper during the RMG crane design stage can prevent the rail gnawing phenomenon.
Bolt connections play a crucial role in the manufacturing of crane steel structures; however, the potential issue of bolt loosening cannot be overlooked. Thus, regular inspection of bolt connections is essential. To automate the detection of bolt loosening, this paper proposes a batch detection method for bolt loosening angles based on machine vision. Firstly, YOLOv8 is employed for bolt target detection, and an improved lightweight segmentation network is used to achieve precise segmentation of the bolt areas. Subsequently, edge line detection and clustering of the segmented images are performed to accurately obtain all edge lines and corner points of the bolts. The centroids and reference points of the bolts are determined based on the corner points, completing the perspective correction. Finally, the corrected bolt corner points are compared with the reference corner points to measure the loosening angles of the bolts. Experimental results indicate that this method achieves high accuracy and sensitivity in bolt loosening measurements. Under different shooting angles, the maximum allowable loosening threshold is 2.62°, with a maximum relative error of only 6.6
Accurate visual measurement depends on precise camera calibration. For cameras with a large field of view (FOV), combined small targets (CST) are commonly used to construct a large calibration object, balancing accuracy and flexibility. However, calibration accuracy is significantly affected when the calibration object is defocused. To overcome this challenge, this paper proposes a CST-based calibration method incorporating defocus deblurring. An image restoration method based on fast defocus estimation is introduced to efficiently restore defocus blur. The method estimates defocus blur through dual-scale re-blurring and region-level transductive inference, and then performs deconvolution accordingly. Building upon this, a novel calibration strategy based on defocus estimation and CST is developed. Multiple small targets (STs) are placed within the camera FOV, and images are captured by adjusting the relative pose between the camera and CST. To enhance feature extraction accuracy, deblurring is applied to defocused ST regions. Extracted features from each ST are then integrated using a global nonlinear optimization algorithm, achieving high-precision calibration. Experimental results demonstrate that the proposed method effectively mitigates the impact of CST defocus on calibration precision, with good stability and computational efficiency. This study provides reliable technical support for calibrating cameras with a large FOV in non-ideal imaging environments and holds significant application potential.
Circular coded targets are widely used in visual measurement, but their identification can be hindered by non-uniform illumination. Additionally, localization methods that rely on the central circular contour are prone to projection errors. This study introduces an improved method for target identification and localization to address these challenges. To enhance identification rates under uneven lighting conditions, homomorphic filtering is applied during image preprocessing, with filter parameters optimized using the artificial hummingbird algorithm. For more precise target center localization, ellipse parameters for both the central circular contour and the inner contour of the coded band are estimated using a spatial median consensus fitting method. These parameters are then employed to achieve sub-pixel localization of the target center through concentric circle projection error compensation. Experimental results demonstrate that the proposed method achieves high identification rates and localization accuracy under non-uniform illumination, offering good practical performance.
In this work, an integrated monitoring system was applied to the shape and strain monitoring of the base boom of a crawler crane. Two industrial cameras and 16 fiber Bragg grating (FBG) strain sensors were installed on the base boom. A simplified model of the main boom was established, the deflection and strain formulae were deduced, and a finite element simulation of the main boom was achieved. Loading tests of the crawler crane were done, and the images taken by the two cameras and the Bragg wavelengths of 16 FBG sensors under various loads and elevation angles were recorded. The marker point lines on the base boom displayed good linear relations, meaning that the bending strains of the base boom in the monitoring region were approximately zero. The 16 FBG sensors displayed their wavelength shifts (WLSs) with the loads and elevation angles in real time, and the change relations were consistent with the curves given by the theoretical model and the finite element simulations. The values of compression rigidity and bending rigidity of the base boom were obtained, by which the strain distribution along the base boom can be estimated.
Point cloud registration is usually divided into two processes: coarse registration and fine registration. Coarse registration provides initial values for fine registration. Generally speaking, the more accurate the result of coarse registration is, the better the effect of fine registration will be. In order to further improve the accuracy of coarse registration, a point cloud coarse registration framework based on local feature description is proposed. First, a key point detection method based on shape index is proposed, the local features of key points are used to match points, the point correspondences are filtered using the rigid transformation geometric consistency criterion, and the initial rigid transformation matrix and inliers can be get by using RANSAC The rigid transformation matrix is re-estimated using the inliers, and the inliers are recalculated by the new rigid transformation matrix, the above process is iterated until the error is less than the threshold or the number of iteration is reached. A more accurate rigid transformation matrix is get again through a re-estimation method based on truncation ratio, and the registration verification is completed by utilizing the compatibility of the matrix norm and the vector norm. Finally, the proposed method is verified by common datasets. The experimental results show that the proposed method can effectively improve the accuracy and stability of coarse registration.
Road crack detection is crucial for maintaining the aesthetics and safety of roads. The varying morphology of cracks often results in insufficient road crack samples, limiting the effectiveness of existing detection methods in few-sample scenarios. Further, when visual samples are insufficient, employing textual information to extract visual information from images is a cutting-edge technology. In this paper, we propose a Dual Prototype Network (DPNet) for few-shot crack detection. Firstly, we introduce an Improved Pixel Weight (IPW) data enhancement to strengthen the foreground and edges of cropped samples, improving learning efficiency in the case of insufficient samples. Next, we design a dual prototype prediction method. Specifically, we employ domain related text input to generate a Language-Image Prototype (LIP) with general domain knowledge through Contrastive Language-Image Pre-training (CLIP). Then, we generate a Support Prototype (SuP) with specialized domain knowledge from crack dataset images. The final prediction is obtained by linearly combining the predictions of the two prototypes. Additionally, we design an Embedding Attention Module (EAM), which leverages the characteristics of the embedding dimension to simultaneously satisfy both spatial and channel attention mechanisms in the transformer structure. Finally, our DPNet achieves superior performance on the FCrack-i and MixCrack few sample datasets, with an average mIoU improvement of 8.52% and 1.44% compared to the baseline. Moreover, we demonstrate the zero-shot capability of DPNet on CFD crack dataset.
Pavement cracks significantly affect road safety and longevity, making accurate crack segmentation essential for effective maintenance. Although deep learning methods have demonstrated excellent performance in this task, their large network architectures limit their applicability on resource-constrained devices. To address this challenge, this paper proposes a lightweight, fully convolutional neural network model, enhanced with spatial information. First, the backbone network structure is optimized to improve the efficiency of spatial information utilization. Second, by incorporating adaptive feature reassembly and wavelet transforms, the up-sampling and down-sampling processes are refined, enhancing the model capacity to capture spatial information. Lastly, a dynamic combined loss function is employed during training to further improve model attention on crack edge details. To validate the model performance, we trained and tested it on the Crack500 dataset and applied the trained model directly to the AsphaltCrack300 dataset. Experimental results indicate that the proposed model achieved an MIoU of 80.37% and an F1-score of 78.22% on the Crack500 dataset, representing increases of 3.08% and 5.62%, respectively, compared to EfficientNet. On the AsphaltCrack300 dataset, the model exhibited strong robustness, significantly outperforming other mainstream models. Additionally, its lightweight design provides clear advantages, making it well suited for realworld applications with limited computational resources.
Abstract In the process of automatic grabbing of bridge segment beams, it is crucial to accurately locate and align the corner points of the crane's boom with the beam's lifting holes. This requires the utilization of image processing techniques to precisely detect and locate the corner points of the crane's boom. Existing feature matching methods face challenges such as low detection accuracy and unsuitability for this specific scenario. This article proposes a novel approach for corner point localization by using the intersection points of lines, facilitating the matching of feature points between left and right images. The method consists of three steps: first, a grayscale difference map is constructed by utilizing the R and G channels of the RGB color space. This enhances the bimodal characteristics of the grayscale histograms between the foreground and background, which facilitates the subsequent binarization process. Additionally, opening and closing operations are employed to remove small artifacts from the Canny edge detection results, effectively reducing noise. Second, an adaptive thresholding method based on the mean and variance of Hough transform voting scores is proposed. This method filters out interference lines from the clustering results by selecting appropriate voting scores. Furthermore, an improved centroid calculation method is introduced, which utilizes weighted formulas based on different proportions of voting scores. These weighted formulas replace the original clustering centroids as the basis for line fitting. Finally, the corner coordinates of the crane's boom are computed based on the line fitting results, and the recognition accuracy is compared under different lighting conditions. Experimental results demonstrate that the proposed algorithm exhibits smaller detection errors and higher robustness compared to other corner detection algorithms, particularly when there are numerous interference edge points in the edge detection results. The computed corner coordinates achieve pixel‐level accuracy. The algorithm performs optimally under strong supplementary lighting conditions, with an average detection error percentage of 97.1% within 0–2 pixels and a recognition accuracy of 98.6%. The recognition success rates under different lighting conditions are all above 92.9%. This method is superior to traditional corner detection methods, meets the requirements for automatic grabbing of the boom, and holds practical engineering value. It provides a basis for addressing the accuracy and robustness challenges of crane algorithms influenced by multiple environmental factors.
In the field of visual measurement, camera calibration is the first step and an important step in determining the accuracy of the measurement. Achieving rapid and accurate camera calibration has been a significant focus of research among scholars. Therefore, a high-precision camera calibration framework based on ellipse eccentricity compensation is proposed. In the first step, the Canny algorithm is used to detect ellipse edges. In the second step, a high-precision fitting of the ellipse is accomplished by a weighted fitting method. In the third step, the initial calibration parameters are calculated by using the coordinates of the elliptical center. Subsequently, the eccentricity deviation of the ellipse is compensated based on the initial calibration parameters. The updated coordinates are then used to recalculate the calibration parameters. This iterative process is repeated using an improved swarm intelligence optimization algorithm until the error is below the threshold or the number of iterations is reached. Simulations and experiments are used to verify the proposed method. The results show that the proposed method has high accuracy and stability, and can be widely used in engineering.
Transverse vibration measurement of beam structures based on computer vision has been widely used, but most visual measurement methods require the installation of artificial markers at the measured positions. This makes the measurement of vibrations in some key position of beam structures difficult due to their inaccessibility. Therefore, the development of target-less approaches is important for some engineering applications. In this article, a method for measuring the displacement and vibration of beam structures visually is presented based on subpixel edge point tracking. Combining the principle of coordinate space transformation and the 2-D normal distribution characteristics of false edge parameters, false edges are eliminated to determine the position of structural edges accurately. This allows the measurement of beam deflection and point tracking of the beam structure's vibration. The effectiveness of the proposed method is verified through a cantilever beam vibration experiment and the measured data are compared using a traditional piezoelectric sensor (eddy current displacement sensor and accelerometer) and digital image correlation (DIC) method. Compared with the traditional piezoelectric sensor methods, the maximum difference in the measured displacement was less than 2.4%. Compared with the DIC method, the time-series correlation of the proposed method with the reference measurements was 0.35% higher, and the main vibration frequency error was 0. The experiment proves that the proposed method can be used as an effective visual measurement method for target-less tracking of the transverse vibration of beam structures.
Pose estimation plays a vital role in numerous application fields, such as photogrammetry, machine vision, robotics, et al. Although various intelligent algorithms can be used to solve the prediction of the estimated pose, it is difficult to acquire the high-accurate pose when the observed coordinates are contaminated by gross errors, especially when the number of 3D-2D correspondences is small. To address this problem, a weighted least squares solution with data snooping (DSWLS) for pose estimation based on the generalized errors-in-variables (EIV) model is proposed. We first utilize the collinearity equation to present the pose estimation as a generalized EIV model. Then, the Euler-Lagrange method is used to the weighted least squares (WLS) solution of the generalized EIV model. Finally, data snooping is introduced into the generalized EIV model to eliminate the impact of gross errors on the estimated pose parameters. Two types of test statistics for data snooping are constructed based on the least squares theory with known and unknown variance components. The simulated and empirical experimented results demonstrate that the proposed method can reduce the influence of gross errors effectively, ensuring reliable pose estimation even with limited data.
In order to achieve safe and economical design results for crane bridge structures, an optimization method with safety assessment constraints is proposed. In this approach, the limit state method is used to test the structure design. Safety assessment is conducted based on the fuzzy analytic network process, allowing the safety score of designs to be quantified. Aiming to minimize the self-weight of the bridge, the dimensions of critical sections are regarded as design variables and constraints related to the process size, strength, stiffness, stability, and safety score are imposed, utilizing the artificial hummingbird algorithm for the structure optimal design. After validating through an engineering example, the results demonstrated that under the settled safety constraints, the proposed optimization method successfully reduced the structural self-weight by 19.709%, while maintaining a safety score proximate to the original design. In comparison to optimization without safety assessment constraints, this method resulted in a 10.441% increase in weight, but its safety was significantly improved by 21.740%. This validates the effectiveness and practicality of incorporating safety assessment into structure optimization, thereby ensuring a balanced trade-off between safety and material economy in the design of crane bridge structures.
Camera calibration is a crucial step in binocular measurement, and the accuracy of camera calibration largely determines the measurement accuracy of binocular vision. However, the calibration accuracy of existing calibration methods is difficult to satisfy the requirements of variable calibration environments and engineering applications. Therefore, based on the principle of Zhang's calibration method, a calibration method is proposed by combining bundle adjustment and diagonal constraints of the calibration target. Firstly, the improved Canny edge extraction algorithm is used to obtain the sub-pixel center of mass of ellipse (CME). Then, the Zhang's calibration method is used to obtain the initial values of the calibration parameters. The camera calibration parameters are optimized by using the bundle adjustment. Finally, diagonal constraints are used to further optimize the extrinsic camera position parameters. The experimental results show that compared to the other excellent methods, the proposed method can significantly improve calibration accuracy and has certain engineering application value.
Semantic segmentation is one of the directions in image research. It aims to obtain the contours of objects of interest, facilitating subsequent engineering tasks such as measurement and feature selection. However, existing segmentation methods still lack precision in class edge, particularly in multi-class mixed region. To this end, we present the Feature Enhancement Network (FE-Net), a novel approach that leverages edge label and pixel-wise weights to enhance segmentation performance in complex backgrounds. Firstly, we propose a Smart Edge Head (SE-Head) to process shallow-level information from the backbone network. It is combined with the FCN-Head and SepASPP-Head, located at deeper layers, to form a transitional structure where the loss weights gradually transition from edge labels to semantic labels and a mixed loss is also designed to support this structure. Additionally, we propose a pixel-wise weight evaluation method, a pixel-wise weight block, and a feature enhancement loss to improve training effectiveness in multi-class regions. FE-Net achieves significant performance improvements over baselines on publicly datasets Pascal VOC2012, SBD, and ATR, with best mIoU enhancements of 15.19%, 1.42% and 3.51%, respectively. Furthermore, experiments conducted on Pole&Hole match dataset from our laboratory environment demonstrate the superior effectiveness of FE-Net in segmenting defined key pixels.
The hexapod wall climbing robots have the advantages of traversing complex wall surfaces. To traverse complex environments autonomously, it must possess the capability to select gait parameters and paths appropriate for the wall surface. Path planning and gait optimization is a fundamental issue in the aspect of stable, energy efficient robot navigation in complex environments with static and dynamic obstacles. Traditional statistical models have been developed to get the optimal path and gait parameters but the result obtained was very poor. Metaheuristic algorithms are gaining importance in robotic gait planning. In this paper, we proposed robust two stage gait planning approach for predicting collision-free, distance-minimal, smooth navigation path and ensuring stable, energy efficient gait patterns for robots using hybrid metaheuristic algorithms. In the first stage, optimal climbing path for robot is predicted using Tri-objective Grey Wolf Path Optimization (TGWPO) based on obstacle and target detection. In the second stage, the gait parameters adaptive to the constructed climbing path are optimized using Adaptive multi-objective Particle swarm optimization (AMPSO). The hexapod wall climbing robot is designed with STM32F103 as core controller modeled with optimal path planner (using TDWPO) and gait optimizer module (using AMPSO). STM32F103 controller commands and controls the robot to climb on wall with optimized gait parameters according to the optimal path. We analyzed the efficacy of the proposed two stage gait planning approach using TDWPO-AMPSO for hexapod wall climbing robots with existing gait planning approaches in terms of path length, climbing time, gait stability, obstacle avoidance, and energy efficiency. The result analysis showed that the suggested gait planning approach is efficient over conventional strategies for climbing robots.