The foveal mechanism of the human retina produces sharp central vision with a blurred periphery. By generating foveated images that more closely replicate this mechanism, it is possible to significantly enhance immersion and deliver a more natural visual experience in virtual environments and related applications. In this paper, we proposed a two-stage foveated image reconstruction method to simulate a biologically plausible foveated retina. In the first stage, a monocular humanoid field of view (FOV) model is designed based on the mapping relationship between the human retina and the camera sensor, enabling the capture of images with a human-like FOV using a commonly available uniform-pixel camera. During the second stage, a non-uniform pixel sampling approach is presented to approximate the spatial distribution of photoreceptors across the retina, combined with biharmonic spline-based interpolation for natural-looking resampling. Qualitative and quantitative results demonstrate that the foveated images generated by the proposed method exhibit better consistency with human visual perception than existing representative methods.
In traditional autofocus algorithms, the determination of the optimal focus position relies on identifying the lens location that yields the maximum image sharpness. However, during the initial search stage, the difficulty in selecting an appropriate step size often causes the algorithm to be trapped in local optima, and under complex lighting and scene conditions, it further suffers from time-consuming iterative computations, leading to prolonged processing time. Meanwhile, current deep learning-based autofocus methods still face challenges in achieving high focusing accuracy. To address these challenges, we propose a hybrid high-precision autofocus algorithm that combines an improved deep learning network with a variable-step hill-climbing strategy. Specifically, the convolutional block attention module (CBAM) and efficient channel attention (ECA) mechanisms are integrated into ShuffleNetV2 to enhance multi-level feature extraction. The classification layer is replaced with a three-layer fully connected structure to directly regress the defocus distance. The predicted value is then used to guide the variable-step local search, thus achieving final accurate localization. Experimental results demonstrate that the proposed method reduces the average focusing error by 47% to 91% and decreases processing time by 63% to 88%. In addition, robustness is improved, with a 26% to 81% reduction in the standard deviation of the average focusing errors. Together, these improvements offer an efficient and reliable autofocus solution for complex imaging scenarios.
During the process of enabling rollers and pavers with intelligence, the accurate measurement of the braking distance between these two machines impacts on the timely compaction of asphalt mixtures and the construction safety. Although unmanned rollers are equipped with obstacle detectors to detect the braking distance, these detectors can only mechanically detect obstacles without the ability to identify them. To solve this problem, a novel method based on object detection and inverse perspective mapping (IPM) is proposed to measure the braking distance between unmanned rollers and pavers. First, YOLOv12 is employed to identify the paver. Second, IPM is applied to the object detection results to generate the inverse perspective image. After that, the pixel coordinates of the object detection box are extracted from the inverse perspective image. Finally, combined with the proposed distance measurement model and the distance scale factor, the braking distance can be accurately measured. The experimental results show that the proposed distance measurement method can accurately and rapidly measure the braking distance. The relative error is less than 1.54%, and the processing time is 76.9ms. The proposed method can fully meet the requirements of asphalt road construction.
Unmanned rollers are typically equipped with satellite-based positioning systems for positional monitoring. However, satellite-based positioning systems may result in unmanned rollers driving out of the specified compaction areas during asphalt road construction, which affects the compaction quality and has potential safety hazards. Additionally, satellite-based positioning systems may encounter signal interference and cannot locate unmanned rollers. To solve this problem, a lateral positioning method for unmanned rollers is proposed to realize the positioning of unmanned rollers relative to asphalt road. First, we captured images from different perspectives and developed a dataset for asphalt road construction. Second, a method for boundary extraction of asphalt road is proposed to accurately locate pixels of asphalt road boundary. Subsequently, the lateral distances are measured by the designed lateral positioning methods. Finally, field validation experiments are conducted to evaluate the effectiveness of the proposed lateral positioning method. The results indicate that the method excels in extracting the asphalt road boundary. Furthermore, the proposed lateral positioning method shows excellent performance, with a mean relative error of 3.40% and a frequency of 6.25 Hz. The proposed lateral positioning method meets the performance requirements for lateral positioning in both accuracy and real-time in asphalt road construction for unmanned rollers.
The positioning for unmanned rollers is typically achieved by global positioning systems (GPS), but GPS easily suffers from signal blockage. In contrast, visual simultaneous localization and mapping (SLAM) provides a robust solution because it does not rely on satellite signals. However, dynamic features in construction environment have a negative influence on the positioning accuracy of unmanned rollers. To solve this problem, a highefficiency semantic segmentation model is constructed to eliminate dynamic features. Additionally, a static feature compensation mechanism is proposed to increase the number of effective features available for SLAM. To verify the positioning accuracy of the proposed SLAM, a specialized image dataset for roller construction environments with dynamic features is developed. Field validation experiments demonstrate that the proposed SLAM has a high positioning accuracy, with a mean relative error of 5.33 %. Compared with state-of-the-art SLAM methods, the proposed SLAM effectively reduces the positioning errors of unmanned rollers in dynamic construction environments, with the reduction of mean relative error at least 40.65 %. The proposed SLAM has been applied to monitor compaction counts of construction areas and has achieved commendable results.
To address the challenge of decreased accuracy in sand and gravel particle size detection under varying shooting distances—caused by changes in image characteristics—this paper constructs a multi-distance aggregate image dataset containing 4,800 images across four particle size ranges. Based on the classical image classification network EfficientNet, we propose an improved MC-EfficientNet model. Structurally, MogaNet convolutions are introduced to replace the original convolutional operators, enhancing feature extraction capability. In addition, the Convolutional Block Attention Module (CBAM) replaces the squeeze-and-excitation (SE) module in MBConv blocks, improving the model’s sensitivity to local edge and texture features. The proposed method achieves accurate particle size classification under multiple distances using training data from a single shooting distance, thus significantly reducing data acquisition requirements. Experiments conducted under various camera distance scenarios demonstrate that MC-EfficientNet improves particle size classification accuracy by 0.32% compared to the original EfficientNet. Moreover, it outperforms mainstream image classification networks such as ResNet, MobileNet, DenseNet, RegNet, and EfficientNetV2 in terms of both accuracy and robustness under different conditions, validating its effectiveness and adaptability for intelligent particle size detection of sand and gravel aggregates under complex imaging environments.
Rotor blades in gas turbines and aircraft engines are subject to harsh conditions, making them prone to damage. Blade tip timing (BTT) is widely used to measure blade vibrations. However, its undersampling property causes spectrum aliasing, complicating frequency identification when monitoring the blade condition. Traditional anti-aliasing methods apply priors or require complex frequency identification steps. Therefore, this paper investigated the aliasing pattern for single-sensor BTT sampling under varying speeds. A simple frequency identification method based on aliasing pattern is proposed. The correctness of theoretical derivation and the feasibility of the proposed method have been verified by simulations and experiments.
Rotor blades are critical components in aeroengines, and monitoring their condition through blade vibration is essential to ensure engine safety. Blade tip timing (BTT) is a non-contact blade vibration measurement method for blade condition monitoring; however, traditional BTT relies on the use of an once-per-revolution (OPR) sensor and multiple BTT sensors, which limits its practical application. Existing OPR-free methods often involve complex computations, and the traditional abnormal blade detection technique based on blade displacement requires quantitative comparison. To address these limitations, this paper proposes an OPR-free BTT approach for abnormal blade detection using a single BTT sensor. The first part of the proposed approach is the blade spacing change with speed smoothing (BSC-SS) method, which measures the spacing change between adjacent blades. The speed smoothing method provides a simple way to eliminate speed calculation errors present in traditional BTT. The second part is the BSC-based abnormal blade detection (BABD) strategy, which detects the abnormal blade through qualitative analysis of the time-frequency spectrum. Simulated and experimental results demonstrate that the BSC obtained by BSC-SS exhibits less disturbance compared to blade displacement and displacement differences. Furthermore, the BABD strategy proves effective as an intuitive approach for abnormal blade detection.
In the fields of multi-focus image fusion and shape from focus, accurate registration of multi-focus images is a crucial prerequisite. Due to the difficulty of feature detection in defocused regions and the limitations of global registration methods, traditional multi-focus image registration methods have low accuracy. To solve this problem, a novel multi-focus image registration method based on optical flow tracking and Delaunay triangulation is proposed. The innovation includes two aspects. The first is that optical flow tracking is utilized to extract and match the non-salient features of multi-focus images. It greatly increases the number of matching features in the defocused regions of multi-focus images. The second is that Delaunay triangulation is adopted for local registration. It makes the matching features strictly aligned. The results of the experiments show that the proposed method is superior to the traditional methods in terms of image registration accuracy and image fusion quality.
In practical engineering scenarios, machines are seldom in a faulty operating state, so it is difficult to get enough available sample data to train the fault diagnosis model, leading to the problem of the small and unbalanced number of rotating machinery fault samples and low fault diagnosis accuracy. To solve this problem, this paper introduces a novel approach to machinery fault diagnosis. This approach involves the integration of a Convolutional Attention Residual Network (CBAM-ResNet) with a Graph Convolutional Neural Network (GCN). Firstly, to comprehensively exploit time-domain information from one-dimensional vibration signals, this study utilize Gram Angular Field (GAF) coding to transform traits of vibration signals into two-dimensional image characteristics. The resultant two-dimensional image is then expanded by applying the Wasserstein Distance Gradient Penalty Generation Adversarial Network (WGAN-GP) to produce a representative sample image. Secondly, the image is input to CBAM-ResNet to perform focused feature extraction and construct the feature matrix. Lastly, the adjacency matrix is derived through Graph Generation Layer (GGL); subsequently, the feature matrix and adjacency matrix are utilized as inputs for the GCN. After deep feature extraction, fault feature classification is executed via Softmax. Performance tests were conducted using the Case Western Reserve University bearing dataset and the planetary gearbox dataset. The method demonstrated remarkable results, achieving an accuracy of over 99% on the unbalanced dataset and surpassing 98% in 0dB noise compared to various other models. This illustrates the effectiveness and feasibility of the proposed method.
Machine vision is crucial for detecting surface defects on cylindrical objects.While correcting the perspective distortion of images of cylindrical objects is feasible,shooting conditions pose challenges.Cylin-drical objects are often observed in tilted positions.To address distortion caused by the perspective projection of inclined cylindrical object surfaces,a method for image correction based on cylindrical surface pose estima-tion was proposed.This method first extracted the side edges of cylindrical images and then estimated the pose relationship between the cylindrical surface and the camera by using the cylinder radius and the system's imaging parameters.An iterative algorithm with variable step size was employed to precisely calculate the cy-lindrical surface pose.Subsequently,a correspondence relationship between the subdivision mesh points of cylindrical surface and the original image pixel coordinates was established.The surface was unfolded into a plane,and a correspondence relationship between the unfolded subdivision mesh points and the corrected im-age pixel coordinates was established.This created a mapping relationship between the corrected image pixel coordinates and the original image pixel coordinates,allowing for resampling of the original image to obtain the corrected image.Experimental results demonstrate that the average distance error of cylindrical surface pose estimation is 0.3 mm,and the average angle error is 0.60°.The average distance standard deviation of adjacent corner points of a chessboard-patterned cylindrical object surface decreases from 12.2 pixels pre-cor-rection to 0.8 pixels after correction.The corrected image effectively identifies text on the cylindrical sur-face,with a measurement error of defects on the cylindrical surface not exceeding 0.1 mm.The corrected image eliminates the inclined projection distortion and"near large,far small"perspective deformation of the cylindrical surface,validating the effectiveness of the proposed method.
Cracks are one of the main road surface diseases,and timely and effective crack detection and evaluation are crucial for road maintenance.To achieve fast and accurate semantic segmentation of road crack images,a road crack detection method based on the DeepLabv3+model is proposed.To reduce the number of model parameters and improve inference speed,MobileNetv3 is used as the model's backbone feature extraction network,and Ghost convolution is used instead of ordinary convolution in the atrous spatial pyramid pooling module to make the model lightweight.To avoid degrading model accuracy by replacing the backbone network,the following measures are adopted.First,a strip pooling module is used in the atrous spatial pyramid pooling module to effectively capture the contextual information of crack structures while avoiding interference from irrelevant regional noise.Second,a lightweight channel attention mechanism,the effective channel attention(ECA)module,is introduced to enhance the feature expression ability,and a shallow feature fusion structure is designed to enrich the image's detailed information,optimizing the model's crack recognition effect.Finally,a mixed loss function is proposed to address the issue of low detection accuracy caused by imbalanced categories in the crack dataset,and transfer learning training is used to improve the model's generalization ability.The experimental results show that the proposed road crack detection model's parameters are only 14.53 MB,which is 93.04%less than the original model parameters,and the average frame rate reaches 47.18,meeting the requirements of real-time detection.In terms of accuracy,the intersection to union ratio and F1 value of this model's crack detection results are 57.21%and 72.76%,respectively,which are superior to classic DeepLabv3+,PSPNet,and U-Net models,as well as advanced FPBHN,ACNet,and other models.The proposed method can significantly reduce the number of model parameters while maintaining road crack detection accuracy and meeting real-time requirements,thus laying the foundation for online detection of road cracks based on semantic segmentation.
Objective Shape from focus is a passive three-dimensional reconstruction technology that restores three-dimensional topography from multi-focused image sequences of target objects.To improve the reconstruction accuracy of this technology in practical applications,the existing methods mostly remove image jitter noise,improve focus measure operator and evaluation window,and optimize data interpolation or fitting algorithms.Although these methods can improve the accuracy of shape from focus,the influence of imaging parameters on reconstruction accuracy is not considered,and the accuracy of shape from focus should be further improved.We explore the influence of imaging parameters on the accuracy of shape from focus of large-depth objects and then clarify the improvement measures of the imaging system when the reconstructive accuracy of shape from focus does not meet the requirements in practical applications.Finally,our study helps select imaging parameters in the application of shape from focus technology to obtain better reconstruction accuracy. Methods Based on constructing the evaluation index of 3D reconstruction accuracy of shape from focus,we firstly analyze the influence degree of focal length,F-number,pixel size,and other parameters in the imaging system on the accuracy of shape from focus by the equal-level orthogonal experiment of a single index.Meanwhile,the primary and secondary orders of the influence of these imaging parameters on the accuracy of shape from focus are determined.Then,the influence of main and sub-main imaging parameters on the 3D reconstruction accuracy is analyzed emphatically by experiments,and the relationship between the optimal imaging parameters and the sampling interval of multi-focus images is revealed.Finally,considering that the change of imaging parameters affects the restoration accuracy of shape from focus by changing the depth of field of the system,it is necessary to explore the influence of imaging parameters on the restoration accuracy of shape from focus of large-depth objects via the depth of field.The experiments help establish the empirical formula between the sampling interval of multi-focus images and the optimal depth of field,providing a theoretical basis for setting imaging parameters of the system. Results and Discussions According to the orthogonal experiment results(Table 3),focal length and F-number are the main and sub-main parameters affecting the accuracy of shape from focus,the influence of pixel size is less than focal length and F-number,and the influence of blank column is the least,which means that there are no important parameters that have not been analyzed.In practical applications,adjusting the focal length and F-number can be realized by adjusting the zoom lens with variable apertures,and meanwhile adjusting the pixel size usually requires replacing the camera,which is costly and usually not considered.Thus,the pixel size is regarded as a non-main influencing parameter.Analyzing the influence of main and sub-main parameters on the accuracy of shape from focus shows that there is the best focal length(Table 4)and the best F-number(Table 5)for the highest reconstruction accuracy under a given multi-focus image sampling interval,and with the decreasing sampling interval,the best focal length increases(Fig.3)and the best F-number reduces(Fig.4).Considering that the change of imaging parameters affects the accuracy of shape from focus by changing the depth of field of the system,we establish an empirical formula between the sampling interval of multi-focus images and the optimal depth of field.The fitting accuracy of the empirical formula is 97.28%(Table 6),and the verification accuracy is 94.76%(Table 7),which can be adopted to calculate the optimal depth of field.The optimal depth of field can significantly improve the accuracy of shape from focus(Table 9),which provides a new way for improving the accuracy of shape from focus of large-depth objects. Conclusions The primary and secondary orders of the influence of imaging parameters on the accuracy of shape from the focus of large-depth objects are discovered,including focal length,F-number,and pixel size.The influence of main and sub-main imaging parameters,focal length,and F-number is analyzed emphatically.It is known that the root mean square error of object reconstruction results decreases first and then increases with the rising focal length or F-number in a given multi-focus image sampling interval,and there is an optimal focal length and F-number that leads to the highest reconstruction accuracy.With the decreasing sampling interval,the optimal focal length increases and the optimal F-number reduces.We consider that the change of imaging parameters affects the accuracy of shape from focus by changing the depth of field of the system.The experiments indicate that the empirical formula between the optimal depth of field and the sampling interval of multi-focused images is obtained.The accuracy of the empirical formula obtained by the verified data is 94.76%,which can be employed to calculate the optimal depth of field.Our experiments show that adjusting the focal length and F-number of the imaging system according to the optimal depth of field can significantly improve the 3D reconstruction accuracy of large-depth objects.
Aiming at the problems of color distortion and detail loss in underwater images due to water scattering and absorption,a generative adversarial network model integrating multi-scale information and attention mechanism was proposed to enhance underwater images.Firstly,to fully exploit and enhance both local and global information of the image,local encoders and global encoders were employed to ex-tract local and global features respectively,which were then fused to achieve complementarity.Next,a multi-scale hybrid convolution was designed to capture multi-scale information,increasing the network's adaptability to features at different scales.Subsequently,attention mechanisms were utilized to enhance the accuracy of feature extraction,emphasizing the focus on high-value features.Finally,by iteratively ap-plying multi-scale hybrid convolution and attention mechanisms to refine features,the enhanced image was gradually up-sampled.Compared with the six classical and state-of-the-art methods,the proposed model not only achieved the best visual perception in subjective evaluations but also outperformed the six compar-ative methods on the entire test set in terms of four objective evaluation metrics peak signal-to-noise ratio(PSNR),structural similarity(SSIM),underwater image quality measurement(UIQM),and natural im-age quality evaluation(NIQE)with average scores of 22.499,0.789,2.911,and 4.175,respectively.The improvements over the best scores among the comparative methods are 0.353,0.002,0.025,and 0.307,respectively.These results indicate that the proposed model not only corrects image color distor-tion but also performs well in restoring image details,increasing image contrast,and enhancing clarity.Therefore,it shows promising prospects for practical applications in underwater image enhancement.
In the modern industrial environment, the accumulation and discharge of static electricity pose potential threats to safe production. High-risk areas urgently require efficient and automated detection technologies for electrostatic discharge (ESD) behaviors. However, research on the automatic detection of workers’ ESD behaviors remains scarce. This paper proposes a video recognition method based on deep learning, utilizing RESNET50 for key point detection and an GRU network for long-term sequence behavior classification. By constructing a high-quality video dataset and systematically training and testing the model, we validate the effectiveness of this approach in detecting workers’ ESD behaviors. Experimental results demonstrate that the proposed method can efficiently and accurately identify various ESD behaviors, providing a reliable technical solution for industrial safety management regarding human ESD.
Research on systems that imitate the gaze function of human eyes is valuable for the development of humanoid eye intelligent perception. However, the existing systems have some limitations, including the redundancy of servo motors, a lack of camera position adjustment components, and the absence of interest-point-driven binocular cooperative motion-control strategies. In response to these challenges, a novel biomimetic binocular cooperative perception system (BBCPS) was designed and its control was realized. Inspired by the gaze mechanism of human eyes, we designed a simple and flexible biomimetic binocular cooperative perception device (BBCPD). Based on a dynamic analysis, the BBCPD was assembled according to the principle of symmetrical distribution around the center. This enhances braking performance and reduces operating energy consumption, as evidenced by the simulation results. Moreover, we crafted an initial position calibration technique that allows for the calibration and adjustment of the camera pose and servo motor zero-position, to ensure that the state of the BBCPD matches the subsequent control method. Following this, a control method for the BBCPS was developed, combining interest point detection with a motion-control strategy. Specifically, we propose a binocular interest-point extraction method based on frequency-tuned and template-matching algorithms for perceiving interest points. To move an interest point to a principal point, we present a binocular cooperative motion-control strategy. The rotation angles of servo motors were calculated based on the pixel difference between the principal point and the interest point, and PID-controlled servo motors were driven in parallel. Finally, real experiments validated the control performance of the BBCPS, demonstrating that the gaze error was less than three pixels.
Existing camera calibration methods using a single image have exhibited some limitations. These limitations include relying on large datasets, using inconveniently prepared calibration objects instead of commonly used planar patterns such as checkerboards, and requiring further improvement in accuracy. To address these issues, a high-quality and convenient camera calibration method is proposed, which only requires a single image of the commonly used planar checkerboard pattern. In the proposed method, a nonlinear objective function is derived by leveraging the linear distribution characteristics exhibited among corners. An algorithm based on enumeration theory is designed to minimize this function. It calibrates the first two radial distortion coefficients and principal points. The focal length and extrinsic parameters are linearly calibrated from the constraints provided by the linear projection model and the unit orthogonality of the rotation matrix. Additionally, a guideline is explored through theoretical analysis and numerical simulation to ensure calibration quality. The quality of the proposed method is evaluated by both simulated and real experiments, demonstrating its comparability with the well-known multi-image-based method and its superiority over advanced single-image-based methods.
To overcome the contradiction between flame retardancy and mechanical properties of rigid polyurethane foams (RPUFs), a reaction-type Schiff base DOPO@chitosan-vanillin Schiff polymer (DOPO@CS-V) is synthesized from chitosan (CS), vanillin, and DOPO. DOPO@CS-V is linked with RPUFs matrix while KH-550 functionalized molybdenum tailings (MTs) and dimethyl methyl phosphate (DMMP) are mixed into RPUFs. The results show that appropriate DOPO@CS-V (4 wt%) and MTs (6 wt%) enhanced the flame retardancy, the p-HRR decreased from 237.5 kW center dot m(-2) to 92.1 kW center dot m(-2), and passed the V-0 rating in the vertical burning test (UL 94). In addition, appropriate MTs increased the compressive strength from 0.15 MPa to 0.25 MPa. This work promotes the application of biomass materials and solid waste in high-performance buildings and industrial materials.
In recent years, fire accidents caused by smoking occur frequently, so it is particularly important to detect cigarettes in non-smoking places. Aiming at the problems that cigarette targets are small, their features are not obvious, and the detection accuracy of existing target detection methods is low, a cigarette small target detection model based on improved YOLOv8 is proposed. Firstly, a small target detection layer is added to the YOLOv8 model, so that the model retains more small target feature information, pays more attention to shallow features, and improves the model’s ability to detect small targets. Finally, the boundary loss function is modified to Wise-IoU v3, and the loss function is optimized through the dynamic non-monotonic mechanism and the gradient gain allocation strategy to optimize the weighted processing of small targets dynamically, so as to improve the model’s detection ability of small targets. The experimental results show that compared to the original model, XW-Improved YOLOv8’s mAP@ 50 increased by 10.60%, and F1 score increased by 13.44%, which can be used for cigarette detection in no-smoking occasions.
This paper presents an object displacement measurement system based on a one-dimensional position-sensitive detector (PSD) and a rotating laser. The system uses the plane formed by the rotating laser as a reference. A one-dimensional PSD mounted on the object captures the laser signal, outputting two signals containing displacement information. These signals are amplified and sent to a microcontroller for parallel ADC conversion and data filtering. Finally, the data is transmitted to a computer via USART communication for further processing to obtain displacement information. Experimental results show that the displacement measurement system offers good measurement accuracy and can be applied in fields such as industrial measurement.