Citrus fruit instance segmentation is an important prerequisite for citrus fruit detection, recognition, and yield estimation in complex environments of citrus orchards. In order to address the challenges of leaf shading, fruit overlapping, and light variations in citrus fruit detection, identification, and yield estimation, this study proposes the YOLO11n-LMPS based on the improved You Only Look Once-11n-Segmentation (YOLO11n-Seg) for citrus instance segmentation and yield estimation. The YOLO11n-LMPS proposes a Parallel Multi-scale Channel Enhancement (PMCE) module for enhancing different sizes of fruit feature extraction and suppressing background noise. An improved Lightweight Separated and Enhancement Attention Module (LSEAM) is used in the Cross Stage Partial with Pyramid Squeeze Attention (C2PSA) to enhance attention and masking robustness and reduce computational complexity. The Multi-Scale Dilated Attention (MSDA) is introduced to optimise detail and global semantic information capture and extend the sensory field. The experimental results in Section 3.2 show that the mAPmask 50 of the YOLO11n-LMPS reaches 91.6%, and the mAPmask 50∶95 reaches 71.1%. The mAPbox 50 of the YOLO11n-LMPS reaches 92.1%, and the mAPbox 50∶95 reaches 74.1%. mAPmask 50, mAPmask 50∶95, mAPbox 50and mAPbox 50∶95 are higher than other mainstream models. Under the complex orchard environment, YOLO11n-LMPS improves citrus fruit detection and segmentation accuracy in complex orchard scenarios, offering a reliable solution for yield estimation.
To enable effective mechanical harvesting of Citrus Reticulata ‘Ehime 38’ (Guodongcheng), a dexterous five-finger robotic gripper was previously developed. However, the lack of quantitative biomechanical data for this cultivar posed a major challenge in designing appropriate grasping and control strategies. This study aimed to systematically analyze the mechanical properties of Guodongcheng through uniaxial compression, puncture, and creep experiments. The viscoelastic response was characterized by the Burgers model. Results showed that the fruit exhibited higher rupture strength and stiffness in the axial direction than in the radial direction, while rupture displacement remained relatively stable. The creep deformation was consistently lower in the axial direction, confirming anisotropic mechanical behavior. Among the statistical comparison of five classical viscoelastic models, the Burgers model provided the best characterization and was extended to a dynamic form capable of predicting the coupled time–force–displacement response. These findings filled a critical knowledge gap in the biomechanical characterization of ‘Ehime 38’ and provided a theoretical basis for optimizing grasping and control strategies of a five-fingered robotic gripper in the mechanized harvesting of the citrus, including determining optimal grasping positions, force ranges, grasp rate, etc.
Compared with typical close-range photogrammetry scenarios, port machinery is larger in size and requires higher measurement accuracy. Traditional measurement methods necessitate photographing from a farther distance to ensure the entire object fits within the frame. However, this approach leads to excessive measurement size per pixel due to the distance, significantly degrading measurement accuracy and failing to meet high-precision requirements. To address this issue, this paper proposes a distributed vision measurement system based on pre-calibration. This system can achieve coordinate system alignment across multiple measurement stations using total station pre-calibration point sets, enabling the integration of data from multiple binocular vision measurement systems to conduct close-range measurements of large components, thereby significantly improving measurement accuracy. To optimize measurement results, the paper introduces redundant observational data within the pre-calibration point set and proposes a selection method based on an optical distortion model that quantifies and weights the confidence of each point. By assigning different weights to various observational data, this method enhances the reliability and accuracy of the measurement data. Measurement experiments simulating real port conditions demonstrate that the proposed method can significantly improve the accuracy of vision measurement in large component measurement scenarios, providing a new theoretical tool for the application of this method in ports.
Slope entropy (SlE) is a recently proposed effective metric for quantifying the complexity of nonlinear signals. However, its representation of fine-grained features and dynamic information remains incomplete, primarily due to rigid hard-threshold slope symbolic segmentation and inadequate characterization of the state transition information in symbol patterns. To address these issues, transition fuzzy slope entropy (TFuSlE) is proposed, which for the first time integrates fuzzy sign partitioning of the slopes derived from two consecutive data samples with the dynamic transition probabilities of the resulting symbol mode sequences into the SlE framework. This integration enables TFuSlE to capture more detailed dynamic information from nonlinear signals, thereby yielding more accurate and comprehensive entropy estimates. Through experiments on simulated data, the optimal parameter configuration for TFuSlE is first determined, and the superiority of its multiple performance metrics in accurately characterizing signal complexity is validated. Finally, TFuSlE is evaluated on two real-world datasets related to mechanical components. The experimental results demonstrate that the health-state features extracted from vibration signals using TFuSlE exhibit well-separated visual distributions and achieve superior classification accuracy and noise robustness compared to other entropy methods.
Self-powered sensor based on nanogenerator provides a potential strategy for addressing the power concern of intelligent sensor. However, the low output of piezo-triboelectric coupled nanogenerator (PTCNG) severely hinders its applications in self-powered sensor. To manage this issue, a novel PTCNG with synergistic regulation of microcrystalline phases and patterned structures was developed, which is composed of the multi-wall carbon nanotube (MWCNT)-polyvinylidene difluoride (PVDF) membrane prepared by solution electrospinning and the polycaprolactone (PCL) membrane with patterned structures manufactured by melt-electrospinning direct writing. To establish the synergistic effects on the PTCNG output, four different assembled modes of PTCNG were constructed. In comparison to the PTCNG assembled with PVDF/PCL and MWCNT-PVDF/PCL with irregular structures, the obtained PTCNG demonstrates a significantly enhanced output performance with four times and 5.2 times higher. Specifically, the PTCNG exhibits a stable open-circuit voltage (260 V) and contact response time (40 ms) and recovery response time (10 ms) after 75 min of continuous operations at a 4 Hz frequency (over 18 000 loading-unloading cycles). Additionally, this PTCNG can charge a 3.3µF capacitor to a voltage of around 29 V within 120 s, power 104 series-connected LEDs, etc., suggesting a promising potential in self-powered sensor.
Tool wear state monitoring under varying operating conditions is important for machining quality and production reliability. However, changes in cutting parameters can shift monitoring-signal distributions and reduce the generalization ability of data-driven models. This paper proposes a cross-condition tool wear state monitoring method based on multi-source sensor signal fusion and supervised transfer learning. X-axis vibration, Z-axis vibration, and spindle current signals are organized as multi-channel time-series inputs. A deep model integrating a multi-scale convolutional neural network, bidirectional long short-term memory, and an attention mechanism is developed to extract discriminative wear-related features. Source-domain pretraining, target-domain warm-up fine-tuning, and source-target joint fine-tuning are organized as a progressive supervised transfer procedure to improve target-condition adaptation. Experiments are conducted on a custom multi-condition dataset using an hp0 + hp1 → hp2 transfer task. Under the unified XZI input configuration, the proposed method outperforms CNN-LSTM, DANN, and CORAL. Input ablation results show that X, XZ, and XZI achieve accuracies of 0.6000, 0.7647, and 0.8588, respectively. In repeated random-seed experiments, the method obtains an Accuracy of 0.7929 ± 0.0499, a Macro-F1 of 0.7292 ± 0.0706, and a Cohen’s Kappa of 0.6542 ± 0.0840. The results demonstrate the effectiveness of multi-source sensor fusion and supervised target-condition adaptation for cross-condition tool wear monitoring.
To address the challenge of picking citrus fruits that grow in clusters and are easily obscured by foliage, a flexible end-effector was designed. It is capable of isolating individual fruits within dense canopies, and its design was inspired by the nonlinear deformation of petal contours during petunia blossoming. A twisting-based picking strategy was adopted to replace conventional shearing methods. A model was developed to describe the relationship between citrus geometry, picking posture, and the required driving force. Finally, the relationship between the driving force and the peak torque required for successful citrus picking was derived, providing a theoretical basis for the optimal application of gripping force. Field experiments showed that when the picking angle is less than 30 degrees, the success rate for citrus fruits with diameters of 75-95 mm was 100%. In the range of 30 degrees-45 degrees, the success rate was 100% for fruit diameters less than 80 mm, and only 64.3% for diameters greater than 80 mm. When the picking angle exceeded 45 degrees, the success rate is 29%. The average picking time per fruit was 3.4 s. Since a complete visual recognition system was not available in this experiment, timing began only after the end-effector had stabilized its grip. Timing ended when the citrus fruit stalk was observed to be snapped off. When the picking angle was controlled within 45 degrees, the flexible end-effector achieved the highest picking success rate. These results demonstrated the feasibility of the "flexible end-effector + twist picking" strategy for citrus picking.
In recirculating aquaculture systems, fish detection is an essential component for maintaining effective farming operations. The availability of high-quality fish datasets is limited because of the richness of fish species, and the annotation of large-scale data, which is used to train models, is often labor-intensive and time-consuming. The presence of different fish species across batches introduces further challenges for consistent detection performance. This work introduces a few-shot learning approach for fish detection, utilizing a customized dataset as novel classes and the Fish4Knowledge dataset for base classes, thereby establishing a framework that enhances adaptability in data-scarce scenarios. Within the model architecture, multi-scale feature extraction is enhanced through an attention mechanism, which is integrated as a dedicated module to strengthen representation learning, thus enhancing the model’s capability to differentiate visually similar fish species. Two distinct customized fish datasets are employed to evaluate the robustness of the proposed method. Experimental results show that the proposed model performs competitively against TFA, Meta-RCNN, and VFA. In the base-training phase, it achieves a mAP of 0.775, slightly surpassing VFA, while in the 1-shot, 5-shot, and 10-shot fine-tuning settings, it obtains mAP values of 0.152, 0.247, and 0.265, respectively. A similar trend is observed on a subset of black fish, with mAP scores of 0.169, 0.253, and 0.286 in the corresponding few-shot settings. These results indicate that the proposed approach can maintain relatively stable detection accuracy and adaptability across different fish batches, offering a practical solution for fish detection tasks in aquaculture when annotated data is scarce. To further demonstrate the efficacy and practical utility of the proposed methodology, a case study in fish farming confirms that the enhanced model achieves consistent and precise detection across diverse fish species, even when trained with limited annotated data.
Industrial defect detection is a critical process for ensuring product quality and production safety. Yet accurate detection remains challenging because defects are often small, multi-scale, morphologically diverse, and embedded in cluttered backgrounds. Existing YOLO-based detectors are limited by homogeneous backbone representations, rigid convolutional geometry, and insufficient coupling between feature extraction and feature fusion. To address these issues, we propose SPF-YOLO, an integrated framework that harmonizes multi-level feature perception with adaptive structural fusion built upon YOLOv11 for high-precision defect detection. Specifically, a Dual-Branch Attentive Refinement (DBAR) module is introduced into the backbone to learn complementary texture-sensitive local cues and channel-wise semantic cues through structurally complementary parallel branches with attention-guided fusion. In addition, a Scale-Adaptive Kernel Perception (SAKP) module combines depthwise separable and dynamic convolutions to adapt its receptive field to defect morphology and scale, improving localization of small and irregular targets. By jointly enhancing feature diversity and geometric adaptability at the source, SPF-YOLO provides higher-quality representations for downstream multi-scale fusion. Experiments on NEU-DET, GC10-DET, and a private welding-defect dataset show higher five-run mean performance than YOLOv11 on all three datasets, although the margin on GC10-DET is small relative to the reported run-to-run variation. SPF-YOLO achieves 83.5 ± 1.1% mAP50 and 48.2 ± 1.3% mAP75 on NEU-DET, 66.2 ± 1.3% mAP50 and 36.0 ± 1.2% mAP75 on GC10-DET, and 95.0 ± 0.8% mAP50-95 on the private dataset.
In the context of vision measurement systems, ports represent a typical harsh measurement environment. The presence of docking ships, operational port cranes, and heavy-duty vehicles often subjects vision measurement equipment to frequent vibrations. These vibrations can cause relative movement between the lens groups, leading to changes in the principal distance. Current vision measurement equipment lacks the capability for rapid on-site principal distance error compensation in port settings, resulting in swift and severe degradation of measurement accuracy. This limitation significantly hinders the application of this advanced non-contact measurement technology in ports. To address this issue, this paper proposes a principal distance compensation method based on the symmetrical constraints of port crane structures. Utilizing the principle of approximate threshold scanning, this method achieves efficient and precise compensation of principal distance errors. Furthermore, the paper introduces a quantitative weighting method for the confidence level of constraint points, further enhancing the accuracy of the compensation values. The experiments demonstrate that the method proposed in this paper can significantly improve the accuracy performance of visual measurement systems under conditions of slight variations in focal length. It can be effectively applied in mechanical structure scenarios with geometric constraints.
The method is applicable for solving the obstacle avoidance workspace of a snake-like robot working on high-voltage transmission cables, based on an improved Monte Carlo method, to address the issues of uneven distribution of scattered points, difficulty in extracting point cloud boundaries, and insufficient accuracy in traditional Monte Carlo methods. The proposed method first generates a seed workspace for the snake-like robot using traditional Monte Carlo method and then envelops the seed workspace with a cube and divides it into several smaller cubes that contain points in the workspace equally. Next, Gaussian distribution probability density function is used to extend and sample the seed workspace of the robot, generating the workspace of the snake-like robot. Finally, the α - shape algorithm is used to extract the point cloud boundaries of the snake-like robot workspace and calculate its volume, accurately determining the workspace. Simulation experiments comparing the reconstructed surface obtained from the α - shape algorithm with the point cloud of the snake-like robot workspace show high accuracy.
The efficiency of helical locomotion in snake-like robots along high-voltage transmission lines is often hindered by low motion efficiency, high joint signal noise, and challenges in traversing obstacles. This study aims to address these issues by proposing a gait generation method that leverages a standardized Central Pattern Generator (CPG). We modify the traditional Hopf-CPG model by incorporating constraint functions and a frequency-tuning mechanism to regulate the oscillator, which allows for the generation of asymmetric waveform signals for deflection joints and facilitates rapid convergence. The method begins by determining initial and obstacle-crossing state parameters, such as deflection angles and helical radii of the snake-like robot, using the backbone curve method and the Frenet–Serret framework. Subsequently, a CPG neural network is constructed based on Hopf oscillators, with a limit cycle convergent speed adjustment factor and amplitude bias signals to establish a fully connected matrix model for calculating multi-joint output signals. Simulation analysis using Simulink–CoppeliaSim evaluates the robot’s obstacle-crossing ability and the optimization of deflection joint signal noise. The results indicate a 55.70% increase in the robot’s average speed during cable traversal, a 57.53% reduction in deflection joint noise disturbance, and successful crossing of the vibration damper. This gait generation method significantly enhances locomotion efficiency and noise suppression in snake-like robots, offering substantial advantages over traditional approaches.
To simplify the soft gripper structure while improving its gripping range, load capacity, and radial load-carrying capacity, Single-axis Deformation (SD) soft grippers were designed, inspired by the blooming behavior of petunias. Four SD finger structures were designed by analyzing the anatomical structure of petunia petals. The stress concentration resulting from the deformation of the fingers with different structures was analyzed using numerical methods under the same deformation load conditions. To verify the gripping ability of the SD soft gripper, we conducted experiments on opening and closing the SD soft gripper, as well as gripping tests on objects of various shapes and sizes, and also tested its radial load-carrying capacity. The experimental results showed that the opening and closing range of the SD soft gripper is 18-144 mm, the maximum gripping load is 39.4 N, and the radial bending angle is only 2.76 degrees when the radial load is 24 N.
The application of citrus-picking robotic hands in orchard environments is constrained by the diversity in fruit size and shape, as well as the need to control fruit damage during harvesting. To address this issue, this study proposes a passively compliant citrus-picking robotic hand and experimentally evaluates its performance. The robotic hand employs a spring-assisted grasping mechanism, optimizing the gripping force range and adjusting spring parameters to achieve passive, compliant encapsulation of citrus fruits of varying sizes while preventing damage. Furthermore, to accommodate citrus fruits with varying ellipticity, the robotic hand incorporates a floating linkage mechanism, enabling each finger to move independently under the control of a single stepper motor, thereby enhancing adaptability to morphological variations. Experimental results indicate that the robotic hand can reliably grasp citrus fruits of various sizes and ellipticities, and complete the harvesting process by rotating four times without applying tensile force, with a damage rate of only 2.6%. The proposed passively compliant robotic hand features a simple structure and strong adaptability, offering a reference for enhancing the applicability of citrus-picking robots in complex orchard environments. Future research will focus on further optimizing the robotic hand’s structure, improving harvesting efficiency, and exploring its adaptability in various operational environments.
The high noise in the automotive body surface image makes it difficult to extract defects. Moreover, a single feature cannot describe the complex automotive body surface defects leading to low classification accuracy. This paper proposes a highly robust method for classifying body surface defects. Firstly, an edge detection method that integrates the wavelet transform and the mathematical morphology is applied to detect defects. Subsequently, the geometric features of detected defects are combined with scale-invariant feature transform features to be the classification basis. Finally, the classification accomplishes through a support vector machine(SVM) with the parameters optimized via the grey wolf optimizer-SVM. Experimental results show the proposed classification method based on feature fusion achieves an average of 93% accuracy in automotive body surface defects classification and exhibits a 100% classification accuracy for pseudo-defects, which demonstrates the fusion of the wavelet transform and the mathematical morphology for automotive body surface defects detection can effectively reduce the impact of image noise for ensuring the extracted edges are intact.
Research on efficient detection methods for rolling bearings is crucial for enhancing the reliability and safety of mechanical equipment. Statistics indicate that over 30% of failures in rotating machinery are attributed to rolling bearings. This paper proposes the wavelet retention transformation and integrates it seamlessly with a residual neural network, resulting in a novel signal processing-based residual neural network framework (MWRC-ResNet). This approach significantly improves the accuracy and interpretability of fault detection in high-noise environments. The proposed method was experimentally validated using both the Case Western Reserve University dataset and the HIT dataset, and the experimental results show that its accuracy and noise resistance are superior to traditional models and other wavelet-based models. This approach not only improves the accuracy of fault detection but also offers better interpretability, providing an effective solution for rolling bearing fault diagnosis.
To address the issues of low trajectory tracking accuracy and difficulties in tuning control parameters for crawler robots operating in uneven terrains, this paper proposes a trajectory tracking control method. The method is based on improved particle swarm optimization and sliding mode active disturbance rejection control (SPSO-SMADRC). Firstly, considering the influence of disturbances such as terrain undulations and soil inhomogeneity on trajectory deviation, the kinematic and dynamic models of the crawler robot are established. A vector field guidance approach is employed to transform the trajectory tracking task into a heading control problem. The heading angle is adaptively adjusted based on the position deviation and path curvature. A nonlinear extended state observer is introduced to estimate external disturbances. A velocity-based SMADRC controller is designed to dynamically regulate the robot’s linear and angular velocities. This allows real-time correction of the robot’s motion. To overcome the tendency of the standard particle swarm optimization (PSO) algorithm to fall into local optima during controller parameter tuning, a nonlinear dynamic adjustment strategy was adopted. This strategy adaptively adjusts the inertia weight and learning factors, enhancing the algorithm’s global search capability. Comparative experiments were conducted using two types of curved trajectories: U-shaped and V-shaped paths. The experimental results show that, under the proposed SPSO-SMADRC method, the crawler robot achieved maximum position errors of 8.28 cm and 9.26 cm, average position errors of 1.41 cm and 2.94 cm, and maximum heading angle deviations of 0.56 rad and 0.87 rad. The standard deviations of the position errors were 3.19 and 4.28, respectively. Compared with conventional PSO-based SMADRC and standard SMADRC methods, the proposed approach improved the navigation tracking accuracy. In the U-shaped trajectory, the maximum position error was reduced by 19.22% and 38.21%, the average position error by 40.00% and 65.53%, and the heading angle error by 28.21% and 74.66%. In the V-shaped trajectory, the maximum position error was reduced by 17.39% and 38.95%, the average position error by 51.71% and 52.04%, and the heading angle error by 80.58% and 84.49%. These results demonstrate that the proposed SPSO-SMADRC method significantly enhances trajectory tracking performance and system robustness. It provides effective support for high-precision autonomous navigation of crawler robots in complex and unstructured environments.
To achieve automation at the inner corner guard installation station in a steel coil packaging production line and enable automatic docking and installation of the inner corner guard after eye position detection, this paper proposes a binocular vision method based on deep learning for eye position detection of steel coil rolls. The core of the method involves using the Mask R-CNN algorithm within a deep-learning framework to identify the target region and obtain a mask image of the steel coil end face. Subsequently, the binarized image of the steel coil end face was processed using the RGB vector space image segmentation method. The target feature pixel points were then extracted using Sobel edges, and the parameters were fitted by the least-squares method to obtain the deflection angle and the horizontal and vertical coordinates of the center point in the image coordinate system. Through the ellipse parameter extraction experiment, the maximum deviations in the pixel coordinate system for the center point in the u and v directions were 0.49 and 0.47, respectively. The maximum error in the deflection angle was 0.45°. In the steel coil roll eye position detection experiments, the maximum deviations for the pitch angle, deflection angle, and centroid coordinates were 2.17°, 2.24°, 3.53 mm, 4.05 mm, and 4.67 mm, respectively, all of which met the actual installation requirements. The proposed method demonstrates strong operability in practical applications, and the steel coil end face position solving approach significantly enhances work efficiency, reduces labor costs, and ensures adequate detection accuracy.
Visual SLAM relies on the motion information of static feature points in keyframes for both localization and map construction. Dynamic feature points interfere with inter-frame motion pose estimation, thereby affecting the accuracy of map construction and the overall robustness of the visual SLAM system. To address this issue, this paper proposes a method for eliminating feature mismatches between frames in visual SLAM under dynamic scenes. First, a spatial clustering-based RANSAC method is introduced. This method eliminates mismatches by leveraging the distribution of dynamic and static feature points, clustering the points, and separating dynamic from static clusters, retaining only the static clusters to generate a high-quality dataset. Next, the RANSAC method is introduced to fit the geometric model of feature matches, eliminating local mismatches in the high-quality dataset with fewer iterations. The accuracy of the DSSAC-RANSAC method in eliminating feature mismatches between frames is then tested on both indoor and outdoor dynamic datasets, and the robustness of the proposed algorithm is further verified on self-collected outdoor datasets. Experimental results demonstrate that the proposed algorithm reduces the average reprojection error by 58.5% and 49.2%, respectively, when compared to traditional RANSAC and GMS-RANSAC methods. The reprojection error variance is reduced by 65.2% and 63.0%, while the processing time is reduced by 69.4% and 31.5%, respectively. Finally, the proposed algorithm is integrated into the initialization thread of ORB-SLAM2 and the tracking thread of ORB-SLAM3 to validate its effectiveness in eliminating feature mismatches between frames in visual SLAM.