Fruit detection is a crucial task in plant phenotyping but remains challenging due to limited training data, high variability in fruit appearances across different growth stages, and occlusions that hinder accurate detection. To address these issues, we propose an Adjustable Anchor Box Detection Network with Transfer Learning (ADNet_TL) for robust fruit detection under unstructured conditions. Our approach leverages the backbone of an existing detector to extract discriminative fruit regions, integrates two adjustable anchor box mechanisms that align with dataset-specific characteristics, and employs transfer learning to boost performance on small target datasets effectively. A comprehensive analysis is conducted to assess the impact of training sample sizes in both source and target domains. Experimental evaluations on Strawberry, Tomato, and Multi-fruit datasets reveal that ADNet_TL outperforms both the standard ADNet and the classical Single Shot MultiBox Detector (SSD), with up to a 14
Accurate weed identification in cotton fields is crucial for weeding robots in precision agriculture. However, the weed identification task faces challenges from environmental interference, complex backgrounds, inter-class similarity, and intra-class variations. This study proposes the KETK-DenseNet model based on DenseNet and Kolmogorov-Arnold Networks (KANs) for weed identification in cotton fields. Firstly, the model adopts DenseNet-121 as the backbone and incorporates the KAN-SE-enhanced Efficient Multi-scale Attention (KS-EMA) module after each Dense Block to optimize feature weights and enhance precise feature representation. Then, a Transition with Inverted Residual (TIR) module is introduced between adjacent Dense Blocks to enhance feature transformation during network transitions. It reduces the dimension and size of feature maps while preserving more essential information. Subsequently, Instance Normalization (IN) is applied within the Dense Layer to improve robustness to data distribution variations. Furthermore, KAN is utilized as the classification layer to enable more refined species discrimination. The model was trained and evaluated on a dataset derived from the publicly available CottonWeedID15 and CottonWeeds datasets. Experimental results demonstrated that KETK-DenseNet achieved an accuracy of 96.37% and an F1-score of 94.45% on the test set, which were 5.09% higher in accuracy and 7.01% higher in F1-score than DenseNet-121. It also outperformed existing mainstream classification models in overall performance. In addition, KETK-DenseNet exhibited fast and stable convergence, achieving a validation accuracy exceeding 90% at approximately the 10th epoch. These results indicate that KETK-DenseNet can effectively achieve the weed classification task in cotton fields, providing potential support for the development of automated weeding robot systems.
Crop yield prediction is one of the key measures for dynamically adjusting harvest schedules and minimizing production losses. To address the production requirements for multi-object real-time detection and accurate counting of tomatoes in complex greenhouse environments, this study proposes a lightweight “detection–tracking–counting–compensation” cascaded network framework, optimized for deployment on inspection robots. First, a sparse convolution module (G-FConv) and a lightweight multi-scale feature aggregation module (L-MFA) were designed, and a lightweight feature extraction network was constructed to reduce computational load and complexity. Second, a lightweight asymmetric decoupled head incorporating the SimAM attention mechanism was developed to enhance critical feature extraction while further reducing parameter count. Third, a multi-view cooperative detection and compensation method was proposed to optimize multi-object tracking performance in dense and occluded scenes. Finally, TensorRT-based inference acceleration was applied to improve the robot’s real-time processing capability. The proposed model achieves high detection (90.8%) and counting (95.2%) accuracy with a compact size of 3.98 MB and 5.3 GFLOPs, while the proposed MVC method reduces counting errors by 95.6%. This study provides an effective technical solution for precise tomato yield prediction in greenhouses, offering significant practical value for optimizing harvest schedules and reducing production losses.
Automated patrol robots hold great promise for advancing livestock farming by enabling precise monitoring of animal health and environmental conditions. However, their deployment in dense and dynamic environments such as pigsties remains challenging due to the reliance on manually annotated waypoints, which are required to define inspection targets a process that is labor-intensive, inefficient, and lacks flexibility. To address this limitation, we propose a modular multi-sensor semantic SLAM framework that enables fully autonomous navigation through intuitive semantic commands. The system integrates LiDAR, RGB-D camera, and IMU data to construct a unified geometric and semantic representation of the environment in real-time. Key facilities (e.g., troughs, fences, and vents) are automatically detected using a YOLO11n model, and their semantic labels are projected onto the global map frame. An Extended Kalman Filter (EKF) fuses sensor data to ensure accurate coordinate alignment for reliable semantic registration. Comparative experiments against RTAB-Map and DS-SLAM demonstrate that our method achieves superior localization and navigation accuracy. The mean localization error was reduced to 0.6 m, and the navigation deviation averaged 0.7 m representing an improvement of twenty percent over the baseline systems. Moreover, defining navigation goals semantically rather than manually reduces mission configuration time from about five minutes to under 30 s per task. These results highlight the robustness, precision, and operational efficiency of the proposed semantic SLAM system, offering a practical solution for autonomous robotic inspection and livestock management in complex farm environments.
As the burden of herbicide resistance grows and the environmental repercussions of excessive herbicide use become clear, new ways of managing weed populations are needed. This is particularly true for cereal crops, like wheat and barley, that are staple food crops and occupy a globally significant portion of agricultural land. Even small improvements in weed management practices across these major food crops worldwide would yield considerable benefits for both the environment and global food security. Blackgrass is a major grass weed which causes particular problems in cereal crops in north-west Europe, a major cereal production area, because it has high levels of herbicide resistance. Moreover, detecting blackgrass in grass crops is challenging due to its visual similarity to these crops. Despite this, a systematic review of the literature on weed recognition in wheat and barley, included in this study, highlights that blackgrass-and grass weeds more broadly-have received less research attention compared to certain broadleaf weeds. With the use of machine vision and multispectral imaging, we investigate the effectiveness of state-of-the-art methods to identify blackgrass in wheat and barley crops. As part of this work, we present the Eastern England Blackgrass Dataset, a large dataset with which we evaluate several key aspects of blackgrass weed recognition. Firstly, we determine the performance of different CNN and transformer-based architectures on images from unseen fields. Secondly, we demonstrate the role that different spectral bands have on the performance of weed classification. Lastly, we evaluate the role of dataset size in classification performance for each of the models trialled. All models tested achieved an accuracy greater than 80%. Our best model achieved 89.6% and that only half the training data was required to achieve this performance. Our dataset is available at: https://lcas.lincoln.ac.uk/wp/research/data-sets-software/eastern-england-blackgrass-dataset/.
Autonomous navigation in agricultural environments is challenged by varying field conditions that arise in arable fields. State-of-the-art solutions for autonomous navigation in such environments require expensive hardware such as RTK-GNSS. This paper presents a robust crop row detection algorithm that withstands such field variations using inexpensive cameras. Existing datasets for crop row detection does not represent all the possible field variations. A dataset of sugar beet images was created representing 11 field variations comprised of multiple grow stages, light levels, varying weed densities, curved crop rows and discontinuous crop rows. The proposed pipeline segments the crop rows using a deep learning-based method and employs the predicted segmentation mask for extraction of the central crop using a novel central crop row selection algorithm. The novel crop row detection algorithm was tested for crop row detection performance and the capability of visual servoing along a crop row. The visual servoing-based navigation was tested on a realistic simulation scenario with the real ground and plant textures. Our algorithm demonstrated robust vision-based crop row detection in challenging field conditions outperforming the baseline.
Cassava is the third largest source of carbohydrates for human consumption worldwide; however, it is highly susceptible to viral and bacterial diseases, which pose a significant threat to food security. The advancement of deep learning algorithms in precision agriculture holds the key to enabling the early classification of plant diseases, thereby leading to enhanced crop yields and ultimately stabilizing food security. In the coarse-grained label discrimination task of weak supervision learning, high-quality semantic features contain abundant semantic description information, which plays a crucial role in constructing a precise description of plant disease discrimination in tanglesome field circumstances and directly influences the performance of neural networks. Thus, a multiattention IBN anti-aliasing neural network (MAIANet) was proposed to improve the classification accuracy of cassava leaf disease classification by improving the feature quality in the coarseness label classification task. The proposed MAIANet neural network includes two innovative approaches. First, the multiattention method was designed to scale the feature signals twice to adjust the angular frequency of the feature signals in the residual branch for optimal feature fitting within the residual unit. Second, the anti-aliasing block extracts the high-frequency component feature and optimizes the quantization result of the pooling operation to depress the aliasing signal in the down-sampled feature maps. When the proposed method was tested and validated on the cassava dataset, the results showed that the prediction accuracy of the proposed method significantly improved, with an accuracy of 95.83 %, a loss of 1.720, and an F1-score of 0.9585, outperforming V2-ResNet101, EfficientNet-B5, RepVGG-B3g4, and AlexNet with significant margins. Based on the above experimental results, the proposed algorithm is suitable for classifying cassava leaf diseases.
Spectral technology is a scientific method used to study and analyze substances. In recent years, the role of spectral technology in the non-destructive testing (NDT) of fruits has become increasingly important, and it is expected that its application in the NDT of fruits will be promoted in the coming years. However, there are still challenges in terms of dataset collection methods. This article aims to enhance the effectiveness of spectral technology in NDT of citrus and other fruits and to apply this technology in orchard environments. Firstly, the principles of spectral imaging systems and chemometric methods in spectral analysis are summarized. In addition, while collecting fruit samples, selecting an experimental environment is crucial for the study of maturity classification and pest detection. Subsequently, this article elaborates on the methods for selecting regions of interest (ROIs) for fruits in this field, considering both quantitative and qualitative perspectives. Finally, the impact of sample size and feature size selection on the experimental process is discussed, and the advantages and limitations of the current research are analyzed. Therefore, future research should focus on addressing the challenges of spectroscopy techniques in the non-destructive inspection of citrus and other fruits to improve the accuracy and stability of the inspection process. At the same time, achieving the collection of spectral data of citrus samples in orchard environments, efficiently selecting regions of interest, scientifically selecting sample and feature quantities, and optimizing the entire dataset collection process are critical future research directions. Such efforts will help to improve the application efficiency of spectral technology in the fruit industry and provide broad opportunities for further research.
The current mainstream approaches for plant organ counting are based on convolutional neural networks (CNNs), which have a solid local feature extraction capability. However, CNNs inherently have difficulties for robust global feature extraction due to limited receptive fields. Visual transformer (ViT) provides a new opportunity to complement CNNs' capability, and it can easily model global context. In this context, we propose a deep learning network based on a convolution-free ViT backbone (tea chrysanthemum-visual transformer [TC-ViT]) to achieve the accurate and real-time counting of TCs at their early flowering stage under unstructured environments. First, all cropped fixed-size original image patches are linearly projected into a one-dimensional vector sequence and fed into a progressive multiscale ViT backbone to capture multiple scaled feature sequences. Subsequently, the obtained feature sequences are reshaped into two-dimensional image features and using a multiscale perceptual field module as a regression head to detect the overall scale and density variance. The resulting model was tested on 400 field images in the collected TC test data set, showing that the proposed TC-ViT achieved the mean absolute error and mean square error of 12.32 and 15.06, with the inference speed of 27.36 FPS (512 x 512 image size) under the NVIDIA Tesla V100 GPU environment. It is also shown that light variation had the greatest effect on TC counting, whereas blurring had the least effect. This proposed method enables accurate counting for high-density and occlusion objects in field environments and this perception system could be deployed in a robotic platform for selective harvesting and flower phenotyping.
Weed and crop segmentation is becoming an increasingly integral part of precision farming that leverages the current computer vision and deep learning technologies. Research has been extensively carried out based on images captured with a camera from various platforms. Unmanned aerial vehicles (UAVs) and ground-based vehicles including agricultural robots are the two popular platforms for data collection in fields. They all contribute to site-specific weed management (SSWM) to maintain crop yield. Currently, the data from these two platforms is processed separately, though sharing the same semantic objects (weed and crop). In our paper, we have proposed a novel method with a new deep learning-based model and the enhanced data augmentation pipeline to train field images alone and subsequently predict both field images and UAV images for weed segmentation and mapping. The network learning process is visualized by feature maps at shallow and deep layers. The results show that the mean intersection of union (IOU) values of the segmentation for the crop (maize), weeds, and soil background in the developed model for the field dataset are 0.744, 0.577, 0.979, respectively, and the performance of aerial images from an UAV with the same model, the IOU values of the segmentation for the crop (maize), weeds and soil background are 0.596, 0.407, and 0.875, respectively. To estimate the effect on the use of plant protection agents, we quantify the relationship between herbicide spraying saving rate and grid size (spraying resolution) based on the predicted weed map. The spraying saving rate is up to 90% when the spraying resolution is at 1.78×1.78 cm2. The study shows that the developed deep convolutional neural network could be used to classify weeds from both field and aerial images and delivers satisfactory results. To achieve this performance, it is crucial to perform preprocessing techniques that reduce dataset differences between two distinct domains.
Large Language Models (LLMs) have exhibited remarkable capabilities in understanding and interacting with natural language across various sectors. However, their effectiveness is limited in specialized areas requiring high accuracy, such as plant science, due to a lack of specific expertise in these fields. This paper introduces PLLaMa, an open-source language model that evolved from LLaMa-2. It's enhanced with a comprehensive database, comprising more than 1.5 million scholarly articles in plant science. This development significantly enriches PLLaMa with extensive knowledge and proficiency in plant and agricultural sciences. Our initial tests, involving specific datasets related to plants and agriculture, show that PLLaMa substantially improves its understanding of plant science-related topics. Moreover, we have formed an international panel of professionals, including plant scientists, agricultural engineers, and plant breeders. This team plays a crucial role in verifying the accuracy of PLLaMa's responses to various academic inquiries, ensuring its effective and reliable application in the field. To support further research and development, we have made the model's checkpoints and source codes accessible to the scientific community. These resources are available for download at .
Picking cherry tomatoes is a time-consuming and labour-intensive task, and robots are an alternative solution to address this issue. The end effector is a key component of the harvesting robot, and it is crucial for achieving automated harvesting of cherry tomatoes. To develop efficient end effectors for picking robots, this study proposes a method to aid end effector design by evaluating and analysing picking patterns of cherry tomatoes. Based on manual picking methods, four potential robot picking patterns are proposed: pressing–breaking combination, pulling, pulling–rotating combination and twisting. A dynamic measurement system based on multi-sensor fusion was developed to measure applied forces and angles during the picking process. Based on the selected picking patterns, two pneumatically controlled picking end effectors, namely, a vacuum end effector and a rotating end effector, were designed. The results of the dynamic measurement experiment and the picking pattern evaluation indicated that the recommended order of picking patterns was twisting, pulling, pulling–rotating combination and pressing–breaking combination in descending order. The picking performance test results of the end effector revealed that for the vacuum end effector, the picking success rate was 66.3 %, whereas the detachment failure was the main reason for picking failure. For the rotating end effector, the picking success rate was 70.1 %, whereas localisation failure and collision were the main reasons for picking failure. This study provides a valuable reference and theoretical analysis basis for the development of cherry tomato picking robots and the design of the end effector in the future.
Weeds pose a persistent threat to farmers’ yields, but conventional methods for controlling weed populations, like herbicide spraying, pose a risk to the surrounding ecosystems. Precision spraying aims to reduce harms to the surrounding environment by targeting only the weeds rather than spraying the entire field with herbicide. Such an approach requires weeds to first be detected. With the advent of convolutional neural networks, there has been significant research trialing such technologies on datasets of weeds and crops. However, the evaluation of the performance of these approaches has often been limited to the standard machine learning metrics. This paper aims to assess the feasibility of precision spraying via a comprehensive evaluation of weed detection and spraying accuracy using two separate datasets, different image resolutions, and several state-of-the-art object detection algorithms. A simplified model of precision spraying is proposed to compare the performance of different detection algorithms while varying the precision of the spray nozzles. The key performance indicators in precision spraying that this study focuses on are a high weed hit rate and a reduction in herbicide usage. This paper introduces two metrics, namely, weed coverage rate and area sprayed, to capture these aspects of the real-world performance of precision spraying and demonstrates their utility through experimental results. Using these metrics to calculate the spraying performance, it was found that 93% of weeds could be sprayed by spraying just 30% of the area using state-of-the-art vision methods to identify weeds.
Multiple source images acquired from diverse sensors mounted on unmanned aerial vehicles (UAVs) offer valuable complementary information for ground vegetation analysis. However, accurately aligning heterogeneous UAV images poses challenges due to differences in geometry, intensity, and noise resulting from varying imaging principles. This paper presents a two-stage registration method aimed at fusing visible RGB and multispectral images for cotton leaf lesion grading. The coarse alignment stage utilizes Scale Invariant Feature Transform (SIFT), while the refined alignment stage employs a novel correlation coefficient-based template matching. The proposed method first employs the EfficientDet network to detect infected cotton leaves with lesions in RGB images. Subsequently, lesion leaves in multiple spectral imagery (red, green, red edge, and near-infrared bands) are located using the perspective transformation matrix derived from SIFT and the coordinates of lesion leaves in RGB images. Refined registration between RGB and multispectral imagery is achieved through template matching with the new correlation coefficient. The registered reflectance data from the different spectral bands and RGB components are utilized to classify pixels in each infected leaf into lesion, healthy, and soil parts. The lesion grade is determined based on the ratio of lesion pixels to the total corresponding leaf area. Experimental results, compared with manual assessment, demonstrate a lesion leaves detection model with a mAP@0.5 of 91.01% and a leaf lesion grading accuracy of 92.01%. These results validate the suitability of the proposed method for UAV RGB and multispectral image registration, enabling automated cotton leaf lesion grading.
Usage of purely vision based solutions for row switching is not well explored in existing vision based crop row navigation frameworks. This method only uses RGB images for local feature matching based visual feedback to exit crop row. Depth images were used at crop row end to estimate the navigation distance within headland. The algorithm was tested on diverse headland areas with soil and vegetation. The proposed method could reach the end of the crop row and then navigate into the headland completely leaving behind the crop row with an error margin of 50 cm.
Effective detection of potato late blight (PLB) is an essential aspect of potato cultivation. However, it is a challenge to detect late blight in asymptomatic biotrophic phase in fields with conventional imaging approaches because of the lack of visual symptoms in the canopy. Hyperspectral imaging can capture spectral signals from a wide range of wavelengths also outside the visual wavelengths. Here, we propose a deep learning classification architecture for hyperspectral images by combining 2D convolutional neural network (2D-CNN) and 3D-CNN with deep cooperative attention networks (PLB-2D-3D-A). First, 2D-CNN and 3D-CNN are used to extract rich spectral space features, and then the attention mechanism AttentionBlock and SE-ResNet are used to emphasize the salient features in the feature maps and increase the generalization ability of the model. The dataset is built with 15,360 images (64x64x204), cropped from 240 raw images captured in an experimental field with over 20 potato genotypes. The accuracy in the test dataset of 2000 images reached 0.739 in the full band and 0.790 in the specific bands (492 nm, 519 nm, 560 nm, 592 nm, 717 nm and 765 nm). This study shows an encouraging result for classification of the asymptomatic biotrophic phase of PLB disease with deep learning and proximal hyperspectral imaging.
Vision-based mobile robot navigation systems in arable fields are mostly limited to in-row navigation. The process of switching from one crop row to the next in such systems is often aided by GNSS sensors or multiple camera setups. This paper presents a novel vision-based crop row-switching algorithm that enables a mobile robot to navigate an entire field of arable crops using a single front-mounted camera. The proposed row-switching manoeuvre uses deep learning-based RGB image segmentation and depth data to detect the end of the crop row, and re-entry point to the next crop row which would be used in a multi-state row switching pipeline. Each state of this pipeline use visual feedback or wheel odometry of the robot to successfully navigate towards the next crop row. The proposed crop row navigation pipeline was tested in a real sugar beet field containing crop rows with discontinuities, varying light levels, shadows and irregular headland surfaces. The robot could successfully exit from one crop row and re-enter the next crop row using the proposed pipeline with absolute median errors averaging at 19.25 cm and 6.77{\deg} for linear and rotational steps of the proposed manoeuvre.
Tea chrysanthemum detection at its flowering stage is one of the key components for selective chrysanthemum harvesting robot development. However, it is a challenge to detect flowering chrysanthemums under unstructured field environments given the variations on illumination, occlusion and object scale. In this context, we propose a highly fused and lightweight deep learning architecture based on YOLO for tea chrysanthemum detection (TC-YOLO). First, in the backbone component and neck component, the method uses the Cross-Stage Partially Dense Network (CSPDenseNet) as the main network, and embeds custom feature fusion modules to guide the gradient flow. In the final head component, the method combines the recursive feature pyramid (RFP) multiscale fusion reflow structure and the Atrous Spatial Pyramid Pool (ASPP) module with cavity convolution to achieve the detection task. The resulting model was tested on 300 field images, showing that under the NVIDIA Tesla P100 GPU environment, if the inference speed is 47.23 FPS for each image (416 * 416), TC-YOLO can achieve the average precision (AP) of 92.49% on our own tea chrysanthemum dataset. In addition, this method (13.6M) can be deployed on a single mobile GPU, and it could be further developed as a perception system for a selective chrysanthemum harvesting robot in the future.
Fruit detachment is one of the essential tasks of cherry tomato harvesting. The harvesting effects of cherry tomato picking robots are greatly influenced by different picking patterns. In this study, to find feasible robotic picking patterns for cherry tomatoes, four potential robotic picking patterns are proposed. A hand-picking dynamic measurement system is developed to measure the applied force, angle variation, displacement, etc. during picking. A series of trials are conducted in a greenhouse to compare and analyze applied force, angle variation, displacement, picking time, calyx retention rate, damage rate, etc. In addition, the compression test is conducted on the ability of the cherry tomato to resist deformation, and the detachment effect of the fruit at high operating speed is tested with a designed air-suction end-effector and a rotating end-effector. The greenhouse picking trials show that pattern 1 (bending) is the standard manual picking method with better indicator results but needs to identify the pedicel and precisely localize the abscission layer for robotic picking. Pattern 2 (pulling) is a simple and effective picking method with large disturbances. Pattern 4 (twisting) is simpler and more suitable for dense environments compared to pattern 3 (pulling with twisting) but requires a large rotation angle. Besides, the compression test results indicate that cherry tomatoes are more resistant to deformation in the axial direction than in the radial direction. The high-speed picking test shows that increasing the detachment speed of the fruit can reduce the disturbance and rotation angle.
Agricultural datasets for crop row detection are often bound by their limited number of images. This restricts the researchers from developing deep learning based models for precision agricultural tasks involving crop row detection. We suggest the utilization of small real-world datasets along with additional data generated by simulations to yield similar crop row detection performance as that of a model trained with a large real world dataset. Our method could reach the performance of a deep learning based crop row detection model trained with real-world data by using 60% less labelled real-world data. Our model performed well against field variations such as shadows, sunlight and growth stages. We introduce an automated pipeline to generate labelled images for crop row detection in simulation domain. An extensive comparison is done to analyze the contribution of simulated data towards reaching robust crop row detection in various real-world field scenarios.