Agriculture is undergoing a significant transformation process with the help of robots. Weeding robots have made their way into the market, and they play a crucial role in automating the weeding process in the field. This study introduces a generalized concept of autonomy levels for currently available weeding robots in the field as well as a comprehensive rating system that allows for a comparison of different weeding robots, irrespective of their developmental stages. We examine the different abilities of market-available robots, tractor implements, and smart weeding systems when it comes to navigating and recognizing crops and weeds in the field. A technological rating system is employed to rate the robots based on their advantages and critical aspects. To achieve this, we introduce a comparison system based on a measurable ability scale for the three important robot skills navigation, recognition, and target specificity. To demonstrate its applicability, we apply this system of robotic capability clustering to different available weeding system: the market-available self-propelled Farmdroid FD 20, the Farming GT, the experimental self-propelled research platform Bonn Bot, the market-available smart tractor implement Ecorobotix Ara, and the Bosch BASF Smart Sprayer. We discuss the outlook of interaction models from remote sensing and robots starting from swarm robot aspects to the spot farming, advantages, and limitations of GNSS- and vision-based robots, as well as current challenges for the use of robots in the field and try to answer the question how robots can support farmers in existing workflows.
Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing both visual growth dynamics and per-plant measurement labels are scarce. In this paper, we introduce a novel, annotated image time-series dataset of 691 sweet pepper plants monitored over two growing seasons, comprising 4837 images with per-plant fruit counts categorized by maturity. We propose a multimodal deep learning framework that fuses high-dimensional image features, extracted using the DinoV3 encoder, with numerical count measurements. Our architecture utilizes a Long Short-Term Memory (LSTM) network to model temporal dependencies and handles irregular sampling intervals common in greenhouse monitoring. Through quantitative experiments, we demonstrate that this multimodal approach reduces RMSE over a persistence baseline by 33
In this manuscript we release two datasets for visual sensing of tomato plants grown in commercial-like settings and acquired using a robot. The first is BUTom21 which consists of still images and manual annotations. The second is BUTom-ST21 which consists of video-based data and semi-automated annotations through AI-based methods, referred to as pseudo-labels. In both cases, we provide pixel-level labels for the ripeness of the fruit. The aim is to provide the research community a challenging set of real-world imagery to explore methods to sense and estimate the state of tomato plants and their fruit, which is an important horticultural crop. Importantly, the spatial-temporal dataset provides individual fruit count and ripeness information enabling researchers to push the boundaries of field-based phenotyping.
Using areal and ground-based robotics systems is one possible component on the way to more sustainable crop production. In this work, we show, how a combination of a UAV and ground robot can be used to automate the process of single plant and even leaf-level inspection of crops in the field, as the farmer usually does manually by physically going to the plants. The work integrates UAV-based plant and weed mapping, UGV path planning and execution, viewpoint planning and control of a robot-mounted arm with a multi-camera head, as well as high-resolution 3D reconstruction.
Accurate monitoring of crop phenotypic traits is essential for efficient farm management and automation in agriculture. Multi-object tracking (MOT) and video instance segmentation (VIS) offer promising approaches to enhance agricultural robotic-vision systems, yet a major limitation is the scarcity of high-quality spatial-temporal datasets. In this paper, we introduce BUP-ST20, a novel weakly labelled spatial-temporal dataset for sweet pepper tracking and segmentation captured on a robotic platform. Our dataset is generated by leveraging still image annotations and utilizing a neural radiance field approach (PAg-NeRF) to automatically obtain consistent object semantics and identities across video sequences. BUP-ST20 contains 16,240 images from 275 sequences, with weak labels for training and validation, and human-annotated ground truth for evaluation. We describe how this pseudo-labelling approach can be adapted to any robotic platform with the required inputs, greatly reducing the annotation requirements for dataset creation, with a focus on agriculture and horticulture. Utilizing BUP-ST20, we evaluate state-of-the-art MOT approaches and propose two novel tracklet matching criteria, enhancing robustness in frame-skipped scenarios and low frame rate cameras. When we decrease the frame rate to approximately 1 frame per second our offline MOT based matching criteria is able to improve performance by an absolute value of 19.63, outlining its validity as a tracklet aggregation technique in this scenario. Our experiments demonstrate the effectiveness of the dataset in benchmarking MOT and VIS techniques within the agricultural domain. This also allows us to highlight challenges such as occlusion, shape variations, and weak-labelling limitations. BUP-ST20 serves as a valuable resource for further advancements in robotic crop monitoring and agricultural automation, while demonstrating the ability to create future weakly labelled datasets using robotic platforms.
As the world population is expected to reach 10 billion by 2050, our agricultural production system needs to double its productivity despite a decline of human workforce in the agricultural sector. Autonomous robotic systems are one promising pathway to increase productivity by taking over labor-intensive manual tasks like fruit picking. To be effective, such systems need to monitor and interact with plants and fruits precisely, which is challenging due to the cluttered nature of agricultural environments causing, for example, strong occlusions. Thus, being able to estimate the complete 3D shapes of objects in presence of occlusions is crucial for automating operations such as fruit harvesting. In this paper, we propose the first publicly available 3D shape completion dataset for agricultural vision systems. We provide an RGB-D dataset for estimating the 3D shape of fruits. Specifically, our dataset contains RGB-D frames of single sweet peppers in lab conditions but also in a commercial greenhouse. For each fruit, we additionally collected high-precision point clouds that we use as ground truth. For acquiring the ground truth shape, we developed a measuring process that allows us to record data of real sweet pepper plants, both in the lab and in the greenhouse with high precision, and determine the shape of the sensed fruits. We release our dataset, consisting of almost 7,000 RGB-D frames belonging to more than 100 different fruits. We provide segmented RGB-D frames, with camera intrinsics to easily obtain colored point clouds, together with the corresponding high-precision, occlusion-free point clouds obtained with a high-precision laser scanner. We additionally enable evaluation of shape completion approaches on a hidden test set through a public challenge on a benchmark server.
Agriculture faces several challenges including climate change and biodiversity loss while, at the same time, the demand for food, feed, biofuels, and fiber is increasing. Sustainable intensification aims to increase productivity and input-use efficiency while enhancing the resilience of agricultural systems to adverse environmental conditions through improved management and technology. Recent advances in sensing, machine learning, modeling, and robotics offer opportunities for novel smart digital technologies to enable sustainable intensification. However, developing smart digital technologies and putting them into agricultural practice, requires closing major research gaps, related in particular to (1) the utilization of multi-scale multi-sensor monitoring in space and time, (2) using artificial intelligence for linking process and data-driven methods, (3) improving decision making and intervention in plant production, and finally (4) modeling conditions and consequences of farmers acceptance. Closing these gaps requires an interdisciplinary approach. Here, we present a research agenda and steps forward to steer research efforts, highlighting research priorities, and identifying required interdisciplinary research collaboration. Following this agenda will leverage the full potential of smart digital technologies for sustainable crop production.
Instance-based semantic segmentation provides detailed per-pixel scene understanding information crucial for both computer vision and robotics applications. However, state-of-the-art approaches such as Mask2Former are computationally expensive and reducing this computational burden while maintaining high accuracy remains challenging. Knowledge distillation has been regarded as a potential way to compress neural networks, but to date limited work has explored how to apply this to distill information from the output queries of a model such as Mask2Former. In this paper, we match the output queries of the student and teacher models to enable a query-based knowledge distillation scheme. We independently match the teacher and the student to the groundtruth and use this to define the teacher to student relationship for knowledge distillation. Using this approach we show that it is possible to perform knowledge distillation where the student models can have a lower number of queries and the backbone can be changed from a Transformer architecture to a convolutional neural network architecture. Experiments on two challenging agricultural datasets, sweet pepper (BUP20) and sugar beet (SB20), and Cityscapes demonstrate the efficacy of our approach. Across the three datasets the student models obtain an average absolute performance improvement in AP of 1.8 and 1.9 points for ResNet-50 and Swin-Tiny backbone respectively. To the best of our knowledge, this is the first work to propose knowledge distillation schemes for instance semantic segmentation with transformer-based models.
In this article, we focus on the critical tasks of plant protection in arable farms, addressing a modern challenge in agriculture: integrating ecological considerations into the operational strategy of precision weeding robots like BonnBot-I. This article presents the recent advancements in weed management algorithms and the real-world performance of BonnBot-I at the University of Bonn's Klein-Altendorf campus. We present a novel Rolling-view observation model for the BonnBot-Is weed monitoring section which leads to an average absolute weeding performance enhancement of 3.4%. Furthermore, for the first time, we show how precision weeding robots could consider bio-diversity-aware concerns in challenging weeding scenarios. We carried out comprehensive weeding experiments in sugar-beet fields, covering both weed-only and mixed crop-weed situations, and introduced a new dataset compatible with precision weeding. Our real-field experiments revealed that our weeding approach is capable of handling diverse weed distributions, with a minimal loss of only 11.66% attributable to intervention planning and 14.7% to vision system limitations highlighting required improvements of the vision system.
Automation in agriculture is a growing area of research with fundamental societal importance as farmers are expected to produce more and better crop with fewer resources. A key enabling factor is robotic vision techniques allowing us to sense and then interact with the environment. A limiting factor for these robotic vision systems is their cross-domain performance, that is, their ability to operate in a large range of environments. In this paper, we propose the use of auxiliary tasks to enhance cross-domain performance without the need for extra data. We perform experiments using four datasets (two in a glasshouse and two in arable farmland) for four cross-domain evaluations. These experiments demonstrate the effectiveness of our auxiliary tasks to improve network generalisability. In glasshouse experiments, our approach improves the panoptic quality of things from 10.4 to 18.5 and in arable farmland from 16.0 to 27.5; where a score of 100 is the best. To further evaluate the generalisability of our approach, we perform an ablation study using the large Crop and Weed dataset (CAW) where we improve cross-domain performance (panoptic quality of things) from 12.8 to 30.6 for the CAW dataset to our novel WeedAI dataset, and 21.2 to 36.0 from CAW to the other arable farmland dataset. Although our proposed approaches considerably improve cross-domain performance we still do not generally outperform in-domain trained systems. This highlights the potential room for improvement in this area and the importance of cross-domain research for robotic vision systems.
Precise scene understanding is key for most robot monitoring and intervention tasks in agriculture. In this work we present PAg-NeRF which is a novel NeRF-based system that enables 3D panoptic scene understanding. Our representation is trained using an image sequence with noisy robot odometry poses and automatic panoptic predictions with inconsistent IDs between frames. Despite this noisy input, our system is able to output scene geometry, photo-realistic renders and 3D consistent panoptic representations with consistent instance IDs. We evaluate this novel system in a very challenging horticultural scenario and in doing so demonstrate an end-to-end trainable system that can make use of noisy robot poses rather than precise poses that have to be pre-calculated. Compared to a baseline approach the peak signal to noise ratio is improved from 21.34dB to 23.37dB while the panoptic quality improves from 56.65% to 70.08%. Furthermore, our approach is faster and can be tuned to improve inference time by more than a factor of 2 while being memory efficient with approximately 12 times fewer parameters.
This paper presents a novel treadmill-based multifunctional testing system (TRUSTS) to execute ground testing for wheeled planetary exploration rover (WPER) in both regular and extreme physical conditions before launching. The TRUSTS with specially-made components featuring compact structure and multifunctional testing items comprises a traction loading subsystem (TLS) and a treadmill-based resistance torque loading subsystem (TRTLS). The former offers the working modes including traction loading, position holding, and range limiting. The latter offers working modes including resistance torque loading and leader–follower tracking. By appropriately allocating the working modes of two subsystems, the TRUSTS is capable of conducting diverse testing items. The customized controllers for the TLS and the TRTLS are proposed to guarantee the operation performance of the TRUSTS in all testing items. Extensive experimental results on a Mars rover prototype in both regular and extreme physical conditions demonstrate that the devised TRUSTS could effectively executing trafficability and maneuverability as well as adaptability tests with reliable scenarios and satisfactory precisions of motion tracking and traction/torque loading, which makes the TRUSTS a favorable choice for comprehensive performance assessment of rover in simulant extraterrestrial environment.
Panoptic segmentation provides both holistic and detailed image parsing information at both the pixel and the instance level. However, the computational burdens restrict its applications in real-time scenarios. A potential approach to learn more efficient models is to employ knowledge distillation. However, previous knowledge distillation schemes have focused mainly on classification with limited attention given to rearession-related tasks which is key for panoptic segmentation. In this paper, we establish a logits-based, a hints-based, and a combination-based scheme for panoptic knowledge distillation by using logits from the final layers and features in the middle layers. Then we explore different combinations of balancing weights for optimal solutions according to different network structures and datasets. To validate our proposed approach, various experiments on different datasets have been conducted and efficient networks with higher performance have been obtained. We show that knowledge distillation can be applied to develop accurate ResNet-34 networks improving their panoptic quality on things by an absolute amount of 4.1 points for sweet pepper (glasshouse environment) and 2.2 points for sugar beet (arable farming environment). These student ResNet-34 networks are able to run inference at faster than a framerate of 53Hz on computing infrastructure similar to PATHoBot (a glasshouse robot). To the best of our knowledge, this is the first work to propose knowledge distillation schemes for panoptic semantic segmentation.
In weed control, precision agriculture can help to greatly reduce the use of herbicides, resulting in both economical and ecological benefits. A key element is the ability to locate and segment all the plants from image data. Modern instance segmentation techniques can achieve this, however, training such systems requires large amounts of hand-labelled data which is expensive and laborious to obtain. Weakly supervised training can help to greatly reduce labelling efforts and costs. We propose panoptic one-click segmentation, an efficient and accurate offline tool to produce pseudo-labels from click inputs which reduces labelling effort. Our approach jointly estimates the pixel-wise location of all N objects in the scene, compared to traditional approaches which iterate independently through all N objects; this greatly reduces training time. Using just 10% of the data to train our panoptic one-click segmentation approach yields 68.1% and 68.8% mean object intersection over union (IoU) on challenging sugar beet and corn image data respectively, providing comparable performance to traditional one-click approaches while being approximately 12 times faster to train. We demonstrate the applicability of our system by generating pseudo-labels from clicks on the remaining 90% of the data. These pseudo-labels are then used to train Mask R-CNN, in a semi-supervised manner, improving the absolute performance (of mean foreground IoU) by 9.4 and 7.9 points for sugar beet and corn data respectively. Finally, we show that our technique can recover missed clicks during annotation outlining a further benefit over traditional approaches.
Monitoring plants and fruits at high resolution play a key role in the future of agriculture. Accurate 3D information can pave the way to a diverse number of robotic applications in agriculture ranging from autonomous harvesting to precise yield estimation. Obtaining such 3D information is non-trivial as agricultural environments are often repetitive and cluttered, and one has to account for the partial observability of fruit and plants. In this paper, we address the problem of jointly estimating complete 3D shapes of fruit and their pose in a 3D multi-resolution map built by a mobile robot. To this end, we propose an online multi-resolution panoptic mapping system where regions of interest are represented with a higher resolution. We exploit data to learn a general fruit shape representation that we use at inference time together with an occlusion-aware differentiable rendering pipeline to complete partial fruit observations and estimate the 7 DoF pose of each fruit in the map. The experiments presented in this paper, evaluated both in the controlled environment and in a commercial greenhouse, show that our novel algorithm yields higher completion and pose estimation accuracy than existing methods, with an improvement of 41 % in completion accuracy and 52 % in pose estimation accuracy while keeping a low inference time of 0.6 s in average.
Monitoring plants and fruits is important in modern agriculture, with applications ranging from high-throughput phenotyping to autonomous harvesting. Obtaining highly accurate 3D measurements under real agricultural conditions is a challenging task. In this letter, we address the problem of estimating the 3D shape of fruits when only a partial view is available. We propose a pipeline that exploits high-resolution 3D data in the learning phase but only requires a single RGB-D frame to predict the 3D shape of a complete fruit during operation. To achieve this, we first learn a latent space of potential fruit appearances that we can decode into an SDF volume. With the pretrained, frozen decoder, we subsequently learn an encoder that can produce meaningful latent vectors from a single RGB-D frame. The experiments presented in this letter suggest that our approach can predict the 3D shape of whole fruits online, needing only 4 ms for inference. We evaluate our approach in controlled environments and illustrate its deployment in greenhouses without modifications.
In agriculture, the majority of vision systems perform still image classification. Yet, recent work has highlighted the potential of spatial and temporal cues as a rich source of information to improve the classification performance. In this letter, we propose novel approaches to explicitly capture both spatial and temporal information to improve the classification of deep convolutional neural networks. We leverage available RGB-D images and robot odometry to perform inter-frame feature map spatial registration. This information is then fused within recurrent deep learnt models, to improve their accuracy and robustness. We demonstrate that this can considerably improve the classification performance with our best performing spatial-temporal model (ST-Atte) achieving absolute performance improvements for intersection-over-union (IoU[%]) of 4.7 for crop-weed segmentation and 2.6 for fruit (sweet pepper) segmentation. Furthermore, we show that these approaches are robust to variable framerates and odometry errors, which are frequently observed in real-world applications.
Cultivation and weeding are two of the primary tasks performed by farmers today. A recent challenge for weeding is the desire to reduce herbicide and pesticide treatments while maintaining crop quality and quantity. In this paper we introduce BonnBot-I a precise weed management platform which can also performs field monitoring. Driven by crop monitoring approaches which can accurately locate and classify plants (weed and crop) we further improve their performance by fusing the platform available GNSS and wheel odometry. This improves tracking accuracy of our crop monitoring approach from a normalized average error of 8.3% to 3.5%, evaluated on a new publicly available corn dataset. We also present a novel arrangement of weeding tools mounted on linear actuators evaluated in simulated environments. We replicate weed distributions from a real field, using the results from our monitoring approach, and show the validity of our work-space division techniques which require significantly less movement (a 50% reduction) to achieve similar results. Overall, BonnBot-I is a significant step forward in precise weed management with a novel method of selectively spraying and controlling weeds in an arable field.
Autonomous navigation of a robot in agricultural fields is essential for every task from crop monitoring to weed management and fertilizer application. Many current approaches rely on accurate GPS, however, such technology is expensive and can be impacted by lack of coverage. As such, autonomous navigation through sensors that can interpret their environment (such as cameras) is important to achieve the goal of autonomy in agriculture. In this paper, we introduce a purely vision-based navigation scheme that is able to reliably guide the robot through row-crop fields using computer vision and signal processing techniques without manual intervention. Independent of any global localization or mapping, this approach is able to accurately follow the crop-rows and switch between the rows, only using onboard cameras. The proposed navigation scheme can be deployed in a wide range of fields with different canopy shapes in various growth stages, creating a crop agnostic navigation approach. This was completed under various illumination conditions using simulated and real fields where we achieve an average navigation accuracy of 3.82cm with minimal human intervention (hyper-parameter tuning) on BonnBot-I.
This paper explores the potential for performing temporal semantic segmentation in the context of agricultural robotics without temporally labelled data. We achieve this by proposing to generate virtual temporal samples from labelled still images. By exploiting the relatively static scene and assuming that the robot (camera) moves we are able to generate virtually labelled temporal sequences with no extra annotation effort. Normally, to train a recurrent neural network (RNN), labelled samples from a video (temporal) sequence are required which is laborious and has stymied work in this direction. By generating virtual temporal samples, we demonstrate that it is possible to train a lightweight RNN to perform semantic segmentation on two challenging agricultural datasets. Our results show that by training a temporal semantic segmenter using virtual samples we can increase the performance by an absolute amount of 4.6 and 4.9 on sweet pepper and sugar beet datasets, respectively. This indicates that our virtual data augmentation technique is able to accurately classify agricultural images temporally without the use of complicated synthetic data generation techniques nor with the overhead of labelling large amounts of temporal sequences.
Nikola Pavešić合作论文数Fakulteta za elektrotehniko;Univerza v Ljubljani3