
Remote intervention is an important function for agricultural robots operating under remote supervision, but video feedback delay can affect both vehicle tracking performance and operator steering behavior. In open-field agricultural vehicle teleoperation, however, field-based evidence linking controlled video delay to operator control behavior remains limited. This study field-validates video-based teleoperated driving of an agricultural robot tractor under controlled glass-to-glass delay conditions and examines how delay affects lateral tracking performance and the frequency-domain characteristics of operator steering behavior. A teleoperation function was integrated into an autonomous robot tractor, and an in-line network emulator imposed additional fixed one-way delay on the upstream video stream. Teleoperated driving trials were conducted on an agricultural road to emulate between-field travel, with eight participants completing four delay conditions. Time-domain tracking performance was evaluated using lateral RMSE, while steering behavior was analyzed using frequency-domain metrics derived from the steering-rate spectrum. Lateral RMSE tended to increase with delay but exhibited substantial inter-individual variability and did not show a significant condition effect. In contrast, the spectral centroid and the 95% roll-off frequency of the steering-rate spectrum decreased significantly with increasing delay, with substantial within-subject effect sizes. Under the present low-speed controlled task, significant decreases relative to the 0.118-s reference were observed in both frequency-domain metrics at the tested 0.318- and 0.418-s conditions. Correlation analyses suggested that the spectral shifts were not consistently explained by tracking-error changes alone. These findings provide field-based evidence that frequency-domain steering metrics can complement conventional tracking-error measures when evaluating the effects of communication delay on open-field agricultural vehicle teleoperation under the tested low-speed conditions and may inform delay-adaptive operator assistance.
Coordinated scheduling between operating tractors and supply vehicles is essential for improving the efficiency of large-scale crop management. This study proposes a framework for scheduling heterogeneous machinery swarms in large-scale farm operations while accounting for operational interferences arising from machine capability variations and operator proficiency. The framework comprises two main components: (i) an environment setup module that performs task generation and tractor-vehicle navigation over field and road networks, and (ii) a decision-making module based on LSTM-PPO for joint task-tractor-supply selection. Experimental results show that the environment setup module can generate executable tasks and routes that guide machinery to complete field operations, rendezvous for pesticide refilling at field access points, and return to the warehouse; field tests further verify the feasibility of the generated within-field paths and road-network routes. Tests show that the generated paths can guide tractors to conduct field operations and travel along the road network to and from the warehouse, or rendezvous with a supply vehicle for pesticide refilling. In addition, time deviations associated with task execution and refilling operations are recorded from 15 complete field operation-transfer-service cycles to provide empirical inputs for disturbance modeling and case-specific evaluation of large-scale machinery scheduling. Considering labor and machinery induced disturbances, LSTM-PPO achieves better scheduling performance than baseline methods, including PPO, GA, and ACO, by reducing and stabilizing total operation time. Compared with the static, open-loop ACO baseline, LSTM-PPO reduces the mean total operation time by 478.19 min (7.97 h; 37.3 %) to 802.20 min (13.37 h) across five independently trained random seeds. The results also demonstrate that the trained policy can adapt dynamic decision, maintaining high operational efficiency of the tractor-vehicle-warehouse system under varying disturbance conditions.
The demand for precise monitoring of wheat plant nitrogen concentration (PNC) has prompted the adoption of remote sensing technology for precision crop management. However, most existing remote sensing studies rely on a single observation scale, making it difficult to fully leverage the synergistic potential of multi-modal data due to inherent scale differences among spectral and texture features. To address this issue, this study developed a multi-scale observation framework to fully explore the representational potential of multi-modal remote sensing data. A data-driven cross-modal feature interaction mechanism was constructed to augment information exchange among heterogeneous features, and a Multi-Branch Intermediate Fusion Deep Neural Network (MBIF-DNN) was proposed to effectively integrate multi-scale and multi-modal information. Transfer learning was also incorporated to support cross-year model generalization. The results demonstrated that spectral features, particularly MS-NIR and RGB-Green bands, were sensitive to scale variations (p < 0.05), while texture features showed even higher sensitivity (p < 0.001). At a single observation scale, multi-modal data fusion significantly improved PNC estimation accuracy. Among all models, MBIF-DNN achieved the best performance, with R2 increasing by 0.087–0.588 and root mean square error (RMSE) decreasing by 0.518–0.837%. Furthermore, integrating optimal features across multiple observation scales further improved model performance, with MBIF-DNN achieving the highest accuracy (R2 = 0.943, RMSE = 0.326%). In cross-year validation, MBIF-DNN combined with transfer learning maintained stable predictive performance (R2 = 0.733, RMSE = 0.560%), demonstrating strong robustness and generalization ability. This study provides an effective framework for wheat PNC monitoring and offers methodological support for precision nitrogen management and agro-ecological sustainability.
Automated fruit counting and yield estimation systems are necessary for efficient orchard management. This study presents a computer vision system based on video multi-object-tracking for fruit load estimation in apple orchards and provides a comprehensive analysis of the system performance under diverse scanning conditions. The system integrates fruit detection, tracking, localization within orchard, and fruit load map generation. Experiments were carried out in an experimental apple orchard containing 420 apple trees. Data was collected with two different RGB-D sensors (Azure Kinect DK and ZED 2) at three different scanning distances (125 cm, 175 cm, and 225 cm) on two different dates prior to the harvest. Comparing the two evaluated sensors, Azure Kinect provided more consistent performance across different dates. Results also show that the longer scanning distance improves accuracy due to seeing the full tree view gives better fruit counts than close partial views. Between the two dates, best results were achieved near harvest due to fruit color at this stage, achieving a Mean Absolute Percentage Error (MAPE) of 6.91 % and a determination coefficient (R2) of 0.733 (using ZED2 sensor at 225 cm distance). Finally, a test comparing scanning from one or both sides of the tree row showed that bilateral scanning improved fruit load estimation at the stretch level by incorporating information from both sides of the canopy. The results of this work demonstrate the effectiveness of the video fruit tracking systems as a useful tool for automating fruit load estimation.
Yield monitor data are widely used for management zone delineation, prescription map development, and on-farm decision making; however, delay in the data caused by sensor response latency can misassign yield measurements to incorrect locations. Existing correction methods often rely on manual tuning, or handcrafted similarity measures that are sensitive to noise, missing data, and irregular harvesting geometry. This study proposes FlowDelayNet, a self-supervised deep learning framework for estimating field-level flow delay by learning delay-aware spatial representations. FlowDelayNet employs a Siamese convolutional encoder trained to discriminate whether two yield patches have similar or dissimilar delays. A dataset of 632 harvested fields was split at the field level into training (70%), validation (15%), and testing (15%). Yield surfaces were rasterized over candidate delays from −15 to +15 logging intervals, and overlapping patches (24 × 24, stride 12) were extracted. A Gaussian-based spatial smoothness score computed using correlation analysis provided the self-supervised supervisory signal. Model optimization combined contrastive loss on Siamese embeddings with regression loss against the normalized smoothness signal. Compared to hard-argmax selection, a temperature-scaled soft-argmax formulation improves delay estimation by computing the expected delay. Across 95 fields, the mean absolute error was 1.24 delay units, with 64 fields predicted within ±1 delay relative to the Gaussian spatial smoothness reference. Embedding analysis showed a mean anchor–positive distance of 0.25 compared to 0.72 for anchor–negative pairs, with 94.8% of sampled triplets satisfying AP < AN. These results demonstrate that FlowDelayNet provides an automated and robust solution for flow-delay estimation across heterogeneous field conditions.