Automated leaf removal of tomato plants in greenhouse environments presents significant visual recognition challenges, including severe occlusion, structural complexity, and variability in stem-petiole morphology. To address the challenge of visual recognition in automated leaf removal of tomato plants within greenhouse environments, this study proposes a two-stage visual recognition strategy that integrates 2D image detection with depth-assisted spatial reasoning. After analyzing greenhouse growth conditions, plant structure, and workspace characteristics, three types of datasets were constructed: stem-petiole target detection, sub-image instance segmentation, and main stem semantic segmentation. A hybrid network model, LeafRemoval-YOLO-K, was developed combining YOLOv8 for keypoint detection and instance segmentation with MMSegmentation employing KNet as the decoder head for semantic segmentation. Using self-collected datasets, YOLOv8 models achieved precision, recall, and F1-score of 85.3%, 88.9%, and 87.1% for keypoint detection, and 92.0%, 99.5%, and 95.6% for instance segmentation on sub-images. The K-Net semantic segmentation model demonstrated robust performance with a precision of 84.14%, recall of 87.09%, F1-score of 85.59%, mean Intersection over Union (IoU) of 73.34%, and overall accuracy of 86.09% in delineating the main stem region. The fused 2D features were leveraged in a cutting point localization algorithm to identify and segment the connection regions between the stem and petiole. Morphological processing refined the main stem and petiole masks, enabling accurate extraction of petiole endpoints and cutting point coordinates. Experimental results demonstrate that the proposed strategy achieves a cutting point localization accuracy of 86.5%, providing an effective and reliable approach for the visual perception module of automated leaf removal in tomato plants. This work lays a theoretical and technical foundation for subsequent robotic manipulation in greenhouse agriculture.
In modern greenhouses, complicated tasks and unstructured environments generate the imperious demand for advanced semantic information about each object at work scenes. A significant problem that mainstream methods intend to resolve is that the refinement and understanding of environmental information cannot efficiently cover the entire task in real time. Therefore, this paper proposes a panoptic semantic mapping method to identify each object that is supposed to be concerned in greenhouses. This method builds grid maps with advanced semantic information based on RGB and depth images. For the agricultural task with tomato as the working object, the categories of various objects in the grid map are divided into four groups: fruits, pedicels, stems and obstacles. This method consists of three steps: semantic segmentation from RGB images with K-Net, reconstruction of point cloud data based on depth images and semantic masks and transformation of the point cloud data into OctoMap. Experimental results show the semantic segmentation algorithm reaches a mean precision of semantic segmentation of 93.83
An essential task in maize seed production is the removal of male tassels from the maternal plants, thus ensuring the genetic purity of parental forms of hybrid maize. Hand detasseling is highly precise but labor-intensive and costly regarding human resources. Mechanical detasseling offers high efficiency in maize detasseling, but the high leaf damage rate significantly impacts plant performance and seed yield. With the passing of the rural labor force, there is an urgent need for higher precision and lower leaf damage rate detasseling devices to meet the market demand for maize detasseling. Therefore, a novel maize detasseling device was designed and developed in this paper to address these issues. First, a down-cut rotary detasseling method is proposed based on the phenotypic characteristics of maize plants and the principles of hand detasseling, effectively minimizing leaf damage near the maize tassel during the detasseling process. Second, the maize detasseling device was designed according to the detasseling method, which is compact in overall structure and consists of an RGB camera, claw-R device, claw-L device, sliding device, rotating device, etc. Eventually, a maize detasseling device was developed to conduct field detasseling experiments and evaluate the device's effectiveness. Experimental results demonstrate that the detasseling success accuracy, first-leaf damage rate, second-leaf damage rate, and third-leaf damage rate are 88.33 %, 81.13 %, 15.09 %, and 3.77 %, respectively. Furthermore, damage to only the first leaf or no damage to any maize leaf was considered correct successful detasseling. Therefore, the overall process final correct successful detasseling rate is 71.67 %, and the average detasseling time for a single maize is 11.85 s. This study provides the technical support for precise maize detasseling in maize detasseling devices.
Automated and accurate counting and ripeness level estimation play key roles in the precision management of cherry tomatoes in greenhouses. This study presents an enhanced framework based on tracking by detection to simultaneously estimate the count and ripeness levels for cherry tomato bunches in greenhouses. Ripeness value is determined by the proportional relationship between the number of ripe and unripe fruits in a tomato bunch. To accurately detect fruit kernels in small sizes and blurred ROIs, a lightweight detector called NanoDet is employed. To reduce ripeness estimation errors caused due to shading, the ripeness value with the highest frequency is selected as the estimate when multiple detections of the same bunch are detected. The ripeness values are categorized into five levels and combined with the counting results. The study's results demonstrate that it achieves 92.79% accuracy in counting and 90.84% accuracy in estimating ripeness levels. The improved algorithm has a processing rate of 15 frames per second. These results demonstrate the potential and value of the enhanced method for counting cherry tomatoes and estimating their ripeness levels.
Picking cherry tomatoes is a time-consuming and labour-intensive task, and robots are an alternative solution to address this issue. The end effector is a key component of the harvesting robot, and it is crucial for achieving automated harvesting of cherry tomatoes. To develop efficient end effectors for picking robots, this study proposes a method to aid end effector design by evaluating and analysing picking patterns of cherry tomatoes. Based on manual picking methods, four potential robot picking patterns are proposed: pressing–breaking combination, pulling, pulling–rotating combination and twisting. A dynamic measurement system based on multi-sensor fusion was developed to measure applied forces and angles during the picking process. Based on the selected picking patterns, two pneumatically controlled picking end effectors, namely, a vacuum end effector and a rotating end effector, were designed. The results of the dynamic measurement experiment and the picking pattern evaluation indicated that the recommended order of picking patterns was twisting, pulling, pulling–rotating combination and pressing–breaking combination in descending order. The picking performance test results of the end effector revealed that for the vacuum end effector, the picking success rate was 66.3 %, whereas the detachment failure was the main reason for picking failure. For the rotating end effector, the picking success rate was 70.1 %, whereas localisation failure and collision were the main reasons for picking failure. This study provides a valuable reference and theoretical analysis basis for the development of cherry tomato picking robots and the design of the end effector in the future.
Developing cherry tomato detection algorithms for selective harvesting robots faces many challenges due to the influence of various environmental factors such as lighting, water mist, overlap, and occlusion. To this end, we present LACTA, a lightweight and accurate cherry tomato detection algorithm specifically designed for harvesting robot operation in complex environments. Our approach enhances the model's generalization ability and robustness by selectively expanding the original dataset using a combination of offline and online data augmentation strategies. To effectively capture the small target features of cherry tomatoes, we construct an adaptive feature extraction network (AFEN) that focuses on extracting pertinent feature information to enhance the identification ability. Additionally, the proposed cross-layer feature fusion network (CFFN) preserves the model's lightweight nature while obtaining richer feature representations. Finally, the integration of efficient decoupled heads (EDH) further enhances the model's detection performance. Experimental results demonstrate the adaptability and robustness of LACTA, achieving precision, recall, and mAP values of 94 %, 92.5 %, and 97.3 %, respectively. Compared to the original dataset, the offline-online combined data augmentation strategy improves precision, recall, and mAP by 1.6 %, 1.7 %, and 1.1 %, respectively. The AFEN + CFFN network structure significantly reduces computational complexity by 28 % and number of parameters by 72 %. With a compact size of only 2.88 M, the LACTA model can be seamlessly deployed into selective harvesting robots for the automated harvesting of cherry tomatoes in greenhouses. The code is available at https://github.com/ruyounuo/LACTA.
The automatic harvesting of tomatoes has been achieved for many years in the laboratory. The new research topic is harvesting the tomato more flexibly and nondestructively at any tomato bunch pose according to the agronomic demands. Although the tomato pose can be predicted by keypoints detection, the poor data quality of commercial RGBD cameras, occlusion between plant organs, various tomato poses, and unstructured working environments pose some challenges to the tomato bunch pose detection. Therefore, our research proposed an improved version of the Tomato Pose Method (TPM), namely TPMv2, which is a two-stage end-to-end multi-task network. This network provides comprehensive information on the tomato bunch, including the positions and poses of the stem, peduncle, and fruits, by predicting the two-dimensional bounding box (2D BBox), threedimensional bounding box (3D BBox), two-dimensional key point (2D Kpt), and three-dimensional key point (3D Kpt). Aiming at the problems of occlusion and poor-quality point cloud, this paper specially designs a key point network (KPN) for tomatoes, where a keypoints processing pipeline was innovatively proposed, improving the accuracy of key point positioning and reducing abnormal prediction effectively. TPMv2 makes it possible to detect tomato bunch pose precisely with an economical camera, avoiding dangerous situations caused by abnormal prediction. The precision of 2D BBox and 3D BBox reached 0.9372 and 0.8700, and the Percentage of correct Keypoints (PCK) of 2D Kpt and 3D Kpt reached 0.8882 and 0.7836. About 78.36 % of 3D Kpts' positioning errors are less than 20 mm, sufficient to describe a correct pose trend based on the 3D Kpt, benefiting the manipulator to plan a more reasonable trajectory for non-destructive harvesting.
Natural rubber is a crucial raw material in modern society. However, the process of latex acquisition has long depended on manual cutting operations. The mechanization and automation of rubber-tapping activities is a promising field. Rubber-tapping operations rely on the horizontal cutting of the leading edge and vertical stripping of the secondary edge. Nevertheless, variations in the impact acceleration of the blade can lead to changes in the continuity of the chip, affecting the stability of the cut. In this study, an inertial measurement unit (IMU) and a robotic arm were combined to achieve the real-time sensing of the blade's posture and position. The accelerations of the blade were measured at 21 interpolated points in the optimized cutting trajectory based on the principle of temporal synchronization. A multiple regression model was used to establish a link between impact acceleration and chip characteristics to evaluate cutting stability. The R-squared value for the regression equation was 0.976, while the correlation analysis for the R-squared and root mean square error (RMSE) values yielded 0.977 and 0.0766 mm, respectively. The correlation coefficient for the Z-axis was the highest among the three axes, at 0.22937. Strict control of blade chatter in the radial direction is necessary to improve the stability of the cut. This study provides theoretical support and operational reference for subsequent work on end-effector improvement and motion control. The optimized robotic system for rubber tapping can contribute to accelerating the mechanization of latex harvesting.
The point cloud-based 3D model of forest helps to understand the growth and distribution pattern of trees, to improve the fine management of forestry resources. This paper describes the process of constructing a fine rubber forest growth model map based on 3D point clouds. Firstly, a multi-scale feature extraction module within the point cloud column is used to enhance the PointPillars learning capability. The Swin Transformer module is employed in the backbone to enrich the contextual semantics and acquire global features with the self-attention mechanism. All of the rubber trees are accurately identified and segmented to facilitate single-trunk localisation and feature extraction. Then, the structural parameters of the trunks calculated by RANSAC and IRTLS cylindrical fitting methods are compared separately. A growth model map of rubber trees is constructed. The experimental results show that the precision and recall of the target detection reach 0.9613 and 0.8754, respectively, better than the original network. The constructed rubber forest information map contains detailed and accurate trunk locations and key structural parameters, which are useful to optimise forestry resource management and guide the enhancement of mechanisation of rubber tapping.
The hole fertilization is an effective means of saving fertilizer and increasing yields by applying the fertilizer needed for crop growth to a certain area below the seed during the seed fertilizing stage. In response to the current issues of low fertilizer utilization and soil non-point source pollution caused by strip fertilization, a hole-fertilizing corn planter was designed, which mainly consists of the hole-fertilizing unit, electro-driven seeding unit, and speed measuring device. A dynamic alignment control system with low cost was developed for intermittent fertilization and seeding as well as for precise alignment of the seed and fertilizer. By analyzing the movement process of the seed and fertilizer, a dynamic alignment algorithm was developed to dynamically adjust the alignment and precisely control the placement. In addition, a novel image-based fertilizer detection method was used to study the fertilizer distribution in the soil and evaluate the operating performance of the prototype in the field. Field trials were conducted to investigate the seeding quality, hole-fertilizing effect, and alignment precision of the seed and fertilizer at different operating speeds. Results indicate that (1) the feed index ( $${I}_{qf}$$ ), miss index ( $${I}_{miss}$$ ), and multiple index ( $${I}_{mult}$$ ) are 91.9%, 6.8%, and 1.3%, respectively. (2) 85.4% of the fertilizer distances are greater than 60 mm in the soil, and the amount of fertilizer applied increases progressively with increasing soil depth. (3) The mean value of the offset distances is 28.1 mm, and 94% of the offset distances are less than 50 mm. Moreover, the offset distance increases as the operating speed rises.
Maize tassel detection is essential for future agronomic management in maize planting and breeding,with application in yield estimation,growth monitoring,intelligent picking,and disease detection.However,detecting maize tassels in the field poses prominent challenges as they are often obscured by widespread occlusions and differ in size and morphological color at different growth stages.This study proposes the SEYOLOX-tiny Model that more accurately and robustly detects maize tassels in the field.Firstly,the data acquisition method ensures the balance between the image quality and image acquisition efficiency and obtains maize tassel images from different periods to enrich the dataset by unmanned aerial vehicle(UAV).Moreover,the robust detection network extends YOLOX by embedding an attention mechanism to realize the extraction of critical features and suppressing the noise caused by adverse factors(e.g.,occlusions and overlaps),which could be more suitable and robust for operation in complex natural environments.Experimental results verify the research hypothesis and show a mean average precision(mAP@0.5)of 95.0%.The mAP@0.5,mAP@0.5-0.95,mAP@0.5-0.95(area=small),and mAP@0.5-0.95(area=medium)average values increased by 1.5,1.8,5.3,and 1.7%,respectively,compared to the original model.The proposed method can effectively meet the precision and robustness requirements of the vision system in maize tassel detection.
The utilization of unmanned aerial vehicles (UAVs) for the precise and convenient detection of litchi fruits, in order to estimate yields and perform statistical analysis, holds significant value in the complex and variable litchi orchard environment. Currently, litchi yield estimation relies predominantly on manual rough counts, which often result in discrepancies between the estimated values and the actual production figures. This study proposes a large-scene and high-density litchi fruit recognition method based on the improved You Only Look Once version 5 (YOLOv5) model. The main objective is to enhance the accuracy and efficiency of yield estimation in natural orchards. First, the PANet in the original YOLOv5 model is replaced with the improved Bi-directional Feature Pyramid Network (BiFPN) to enhance the model's cross-scale feature fusion. Second, the P2 feature layer is fused into the BiFPN to enhance the learning capability of the model for high-resolution features. After that, the Normalized Gaussian Wasserstein Distance (NWD) metric is introduced into the regression loss function to enhance the learning ability of the model for litchi tiny targets. Finally, the Slicing Aided Hyper Inference (SAHI) is used to enhance the detection of tiny targets without increasing the model's parameters or computational memory. The experimental results show that the overall AP value of the improved YOLOv5 model has been effectively increased by 22%, compared to the original YOLOv5 model's AP value of 50.6%. Specifically, the AP(s) value for detecting small targets has increased from 27.8% to 57.3%. The model size is only 3.6% larger than the original YOLOv5 model. Through ablation and comparative experiments, our method has successfully improved accuracy without compromising the model size and inference speed. Therefore, the proposed method in this paper holds practical applicability for detecting litchi fruits in orchards. It can serve as a valuable tool for providing guidance and suggestions for litchi yield estimation and subsequent harvesting processes. In future research, optimization can be continued for the small target detection problem, while it can be extended to study the small target tracking problem in dense scenarios, which is of great significance for litchi yield estimation.
In the tomato planting industry, picking is an important step in the fruit harvesting operation. Manual picking is time-consuming and labor-intensive, and automatic picking by machines is the main developing tendency. One of the main reasons for the low picking success rate of current tomato picking robots is picking failure due to collisions between the end-effector and tomato plants during the picking process. Therefore, this paper proposes a cascade deep learning network algorithm to reduce the collisions, identify the maturity, estimate the 3D poses, and search the collision-free picking strategy for tomatoes. The algorithm consists of three steps: tomato bunch detection; tomato detection and occlusion judgment; and the classification of maturity and poses, each task based on a trained YOLOv5s network. Combining with actual harvesting practices, an economical 3D pose estimation method is proposed, where the 3D pose estimation task is divided into two classification tasks, generating four typical 3D poses. Experimental results show that the YOLOv5-based visual detection and pose classification algorithm, whose input is RGB images, can detect unoccluded tomatoes and classify them for maturity and 3D poses with a detection speed of 20 fps. The detection accuracy of unoccluded tomatoes is 82.4 %, the recall rate is 90.9 %; the average accuracy of maturity classification is 96.9 %, and the average accuracy of 3D pose classification is 89.1 %.
The application of robotic grasping for agricultural products pushes automation in agriculture-related industries. Cucumber, a common vegetable in greenhouses and supermarkets, often needs to be grasped from a cluttered scene. In order to realize efficient grasping in cluttered scenes, a fully automatic cucumber recognition, grasping, and palletizing robot system was constructed in this paper. The system adopted Yolact++ deep learning network to segment cucumber instances. An early fusion method of F-RGBD was proposed, which increases the algorithm's discriminative ability for these appearance-similar cucumbers at different depths, and at different occlusion degrees. The results of the comparative experiment of the F-RGBD dataset and the common RGB dataset on Yolact++ prove the positive effect of the F-RGBD fusion method. Its segmentation masks have higher quality, are more continuous, and are less false positive for prioritizing-grasping prediction. Based on the segmentation result, a 4D grab line prediction method was proposed for cucumber grasping. And the cucumber detection experiment in cluttered scenarios is carried out in the real world. The success rate is 93.67% and the average sorting time is 9.87 s. The effectiveness of the cucumber segmentation and grasping pose acquisition method is verified by experiments.
Fruit detachment is one of the essential tasks of cherry tomato harvesting. The harvesting effects of cherry tomato picking robots are greatly influenced by different picking patterns. In this study, to find feasible robotic picking patterns for cherry tomatoes, four potential robotic picking patterns are proposed. A hand-picking dynamic measurement system is developed to measure the applied force, angle variation, displacement, etc. during picking. A series of trials are conducted in a greenhouse to compare and analyze applied force, angle variation, displacement, picking time, calyx retention rate, damage rate, etc. In addition, the compression test is conducted on the ability of the cherry tomato to resist deformation, and the detachment effect of the fruit at high operating speed is tested with a designed air-suction end-effector and a rotating end-effector. The greenhouse picking trials show that pattern 1 (bending) is the standard manual picking method with better indicator results but needs to identify the pedicel and precisely localize the abscission layer for robotic picking. Pattern 2 (pulling) is a simple and effective picking method with large disturbances. Pattern 4 (twisting) is simpler and more suitable for dense environments compared to pattern 3 (pulling with twisting) but requires a large rotation angle. Besides, the compression test results indicate that cherry tomatoes are more resistant to deformation in the axial direction than in the radial direction. The high-speed picking test shows that increasing the detachment speed of the fruit can reduce the disturbance and rotation angle.
Natural rubber latex is an important energy material. However, the harvest of latex is still manual, and there are few researches on automatic work. This paper proposed a method for robot to harvest from each rubber tree without stopping. The robot toggled the collection cup through the flexible actuator, the cup poured out the latex and rotated to the next collection position. In this way, it not only completed the harvest work, but also prepared for the next collection. Aiming at the problem of latex splashing during rapid toggle, the structural parameters of the collection cup and the flexible actuator were modeled and studied. This method made the rotation speed of the collection cup smoother and avoids splashing. To enable the robot to accurately complete the harvesting work, this paper used two-dimensional Light Detection and Ranging (LiDAR) and ranging sensor to locate the space position of the collection cup. The results show that when the toggle speed is 0.5 m/s and the latex volume is 300 ml, the average shaking height is 3.58 mm. In the field test, the lateral error of positioning is less than 8.86 mm, and the height error is less than 0.72 mm; the average harvest rate is 98.18%. The robot has high efficiency and good stability, and can be applied to the rubber plantation to harvest automatically. (C) 2021 Published by Elsevier B.V.
The maize detasseling process was gradually automated with the continuous promotion of the maize detasseling machine. Nevertheless, some problems are gradually becoming more prominent for it. Most maize detasseling machine has a simple maize tassels detection system. That is, the height of the crop was identified, and not the specific information on the maize tassels, leading to low detasseling precision and a high leaf injury rate during maize detasseling machine working. Aiming at these issues, this paper proposed a novel method for the maize tassel pose estimation based on computer vision and oriented object detection. Specifically, a two-step framework for maize tassel pose estimation is developed. Firstly, the maize plant is captured from above, then the maize tassels and the second leaf's vein are posed in the horizontal plane, and their poses are matched to the oriented bounding box. Second, an oriented bounding box is generated according to the Oriented R-CNN model detection. Since maize tassels in the field differ in size and the morphological color of different growth stages, estimating the maize tassel pose accurately by the angle of the oriented bounding box alone is unfeasible. Therefore, these pixels of the oriented bounding box were put into the Look Twice module to extract the critical information about maize tassels, which can accurately determine the final pose of the maize tassels. Finally, evaluation metrics on the test set indicate the proposed method performed with correct maize tassels pose estimation rate of 88.56% and 29.57 Giga Floating-point Operations (GFLOPs), which indicated that the feasibility of maize tassel poses estimation using the proposed method. This study provides the possibility and foundation for precise maize detasseling in maize detasseling machines.
Non-destructive picking of fresh tomatoes is a delicate agronomical operation, based on comprehensive information about the plant organ, such as the location of stem, peduncle, and fruits. The matching between visual information supply and information demand from the agronomical technic is the key power to promote the picking robot from the laboratory to the field. The three-dimensional pose information, containing the location of each organ of the plant, can meet the demand of agronomical technic. It is the premise of precisely handling the cluster of fruits. In order to realize the fine tomato bunch harvesting operation in a bunch, this paper proposed a three-dimensional pose detection method for tomato bunch. The method, named Tomato Pose Method (TPM), is composed of a priori geometric model, a cascaded multi-task network, and a three-dimensional reconstruction process. Based on prior knowledge and agronomic technology, this prior geometric model comprehensively and flexibly describes the spatial location information of tomato bunch. The cascaded multi-task network is designed based on hourglass structure and transfer learning, which is suitable for bounding box and key point prediction of tomato bunches in complex environments. Finally, combining the prior geometric model and the spatial position information of each key point, the tomato bunch is reconstructed. Only a medium training dataset, containing 1800 RGBD images covering changing lighting, occlusion, and various poses, is needed for training. Its success rate of TPM on two-dimensional keypoint detection is 94.02%, the accuracy of 85.77% predicted points are at medium level. And 70.05% tomato bunch with multi-pose can be constructed. More importantly, this method only needs one RGBD image taken by a commercial camera to realize the three-dimensional reconstruction of a single-bunch scenario in 1.0 s, and a multi-bunch scenario in 2.0 s. It provides comprehensive information, and provides data basis for target positioning and path planning of picking robot, which makes the non-destructive harvesting possible.
This paper presents the development and evaluation of a pneumatic finger-like end-effector for cherry tomato harvesting robot. The end-effector is pneumatically controlled and has the capability of picking cherry tomatoes continuously and steadily. The end-effector is compact in overall structure, which consists of finger-like clamping finger, rotating and telescopic cylinders, and RGB-D camera, etc. Another important feature is that the clamping finger uses a combination of clamping and rotating and has good adaptability in dense operating environments. Further, a hand-picking dynamic measurement system is developed to measure and analyze the applied force and disturbance for simplified picking methods during picking. The hand-picking test shows that both the applied force and the disturbance are smaller for rotating compared to pulling, so the method of rotating is chosen for the design of the end-effector in this study. The field test shows the average cycle time of picking single cherry tomato is 6.4 s. The harvest success rates for pickable cherry tomatoes in different directions are 84% (right), 83.3% (back), 79.8% (left), and 69.4% (front), respectively. The failure cases are analyzed and collision and localization failure are the main causes of unsuccessful picking during picking.
The fast and precise detection of dense litchi fruits and the determination of their maturity is of great practical significance for yield estimation in litchi orchards and robot harvesting. Factors such as complex growth environment, dense distribution, and random occlusion by leaves, branches, and other litchi fruits easily cause the predicted output based on computer vision deviate from the actual value. This study proposed a fast and precise litchi fruit detection method and application software based on an improved You Only Look Once version 5 (YOLOv5) model, which can be used for the detection and yield estimation of litchi in orchards. First, a dataset of litchi with different maturity levels was established. Second, the YOLOv5s model was chosen as a base version of the improved model. ShuffleNet v2 was used as the improved backbone network, and then the backbone network was fine-tuned to simplify the model structure. In the feature fusion stage, the CBAM module was introduced to further refine litchi's effective feature information. Considering the characteristics of the small size of dense litchi fruits, the 1,280 × 1,280 was used as the improved model input size while we optimized the network structure. To evaluate the performance of the proposed method, we performed ablation experiments and compared it with other models on the test set. The results showed that the improved model's mean average precision (mAP) presented a 3.5% improvement and 62.77% compression in model size compared with the original model. The improved model size is 5.1 MB, and the frame per second (FPS) is 78.13 frames/s at a confidence of 0.5. The model performs well in precision and robustness in different scenarios. In addition, we developed an Android application for litchi counting and yield estimation based on the improved model. It is known from the experiment that the correlation coefficient R2 between the application test and the actual results was 0.9879. In summary, our improved method achieves high precision, lightweight, and fast detection performance at large scales. The method can provide technical means for portable yield estimation and visual recognition of litchi harvesting robots.