We study a decentralized online estimation problem with additive communication noises over the fixed digraph. Each node has a linear measurement of an unknown parameter with random measurement matrices and runs a continuous-time online estimation algorithm. We transform the convergence analysis of the algorithm into the stability analysis of the non-autonomous linear stochastic differential equation (SDE) with random time-varying coefficients, and develop the asymptotic stability by numerical approximation theory. Based on the stability results, we show that the algorithm gains can be properly designed to ensure mean square convergence if the measurement matrices and the communication graph satisfy the stochastic spatial-temporal persistence of excitation condition. Furthermore, a special case where the measurement matrices contain a Markov chain is investigated, and the theoretical results are demonstrated by a numerical example.
This paper investigates the discrete-time nonlinear zero-sum game problem with alternating dual-impulse control based on adaptive dynamic programming (ADP). Such problems appear in scenarios with scheduled discrete interventions, such as nonlinear oscillators with impulsive excitations or periodic medical dosing. The main objective is to determine equilibrium strategies and corresponding saddle-point solutions in impulsive nonlinear systems. The key challenge lies in handling nonsmooth state transitions and the computational complexity of nonlinear dynamics. To overcome these difficulties, we design a periodic cumulative utility function that focuses computation on impulse instants and develop ADP algorithms for efficient learning. Convergence to saddle-point equilibria is theoretically proven when such equilibria exist, and the approach remains effective even when exact equilibria do not exist. Simulations on Duffing and torsional systems confirm both the accuracy and robustness of the proposed framework.
This paper studies the challenging issue of privacy-preserving formation control for second-order multi-agent systems under stochastic switching topologies, mainly protecting privacy of true initial formation errors. Considering the circumstances of communication noises among agents, distributed intermittent event-triggered privacy-preserving formation algorithms including the fuzzy logic system are given. Meanwhile, a novel privacy-preserving mask function is proposed, where indistinguishable space is increased by 26.2% than ones in existing literatures. Based on stochastic stability analysis and graph theory, sufficient conditions of mean square formation with the bound error are derived. Moreover, optimal controllers for the leader are proposed within the intermittent event-triggered privacy-preserving framework. In order to deal with nonlinear and unknown function, we exploit the idea of adaptive control to parameterize partial derivative of corresponding cost function, which is another different technique rather than the adaptive dynamic programming one. Finally, two simulation examples are given to indicate feasibility of our proposed theoretical results.
Executable approach pose estimation for robotic apple harvesting remains challenging in natural orchards because near-spherical fruit geometry, weak growth-axis observability, and branch-and-leaf occlusion make approach-direction estimation unstable. This study proposes a growth-axis-aware apple pose estimation method that integrates growth-axis representation with spatial geometric constraints. An Oriented Deformable Box (ODB) annotation-level representation is designed from stem-calyx cues to encode the apple growth axis in 2D supervision. AppleAxisNet jointly outputs rotated detections and instance masks, with ARConv and KFIoU incorporated to improve local orientation modeling and angular regression stability. Instance masks are used to construct frustum-constrained RGB-D point clouds, and adaptive DBSCAN with prior seed points is applied to estimate the 3D fruit centroid. The 2D growth-axis orientation is then transformed into a 3D directional constraint and fused with the 3D observation vector to generate executable approach poses. AppleAxisNet achieved mAP50 values of 86.32% and 88.85% for rotated detection and instance segmentation, respectively, with an inference time of 14.2 ms. The mean 2D growth-axis angle error was 12.32 degrees, and the mean 3D centroid errors were 3.9, 4.2, and 5.6 mm at near, medium, and far ranges, respectively. In robotic tests, 94.41% of generated poses fell within the feasible approach range, and harvesting success rates of 91.00% and 80.75% were obtained on a 7-DOF manipulator and a four-arm orchard harvesting robot, respectively. These results demonstrate that the proposed method can generate accurate and executable apple approach poses in complex orchard environments.
This paper investigates privacy-preserving distributed cooperative control for multi-agent systems within the framework of differential privacy. In cooperative control, communication noise is inevitable and is usually regarded as a disturbance that impairs coordination. This work revisits such noise as a potential privacy-enhancing factor. A linear quadratic regulator (LQR)-based framework is proposed for agents communicating over noisy channels, where the noise variance depends on the relative state differences between neighbouring agents. The resulting controller achieves formation while protecting the reference signals from inference attacks. It is analytically proven that the inherent communication noise can guarantee bounded (ε,δ)-differential privacy without adding dedicated privacy noise, while the system cooperative tracking error remains bounded and convergent in both the mean-square and almost-sure sense.
In this study, a moment adaptation evolution strategy based on an aging leader and learned challenger mechanism (MoA-ES-ALLC) is proposed to enhance the optimization performance of the covariance matrix adaptation evolution strategy (CMA-ES). In contrast to the CMA-ES, the proposed algorithm defines a well-performing individual as the leader to retain advantageous information. It achieves the moment adaptation by guiding the mean using an influence coefficient and modifying the covariance matrix with a new meaningful evolution path, thereby accelerating convergence. Learned challengers, generated by a dual-center learning strategy, increase population diversity. A restart strategy, coupled with learned challengers, helps avoid local optima. To further demonstrate the generality of these strategies, they are also integrated into the matrix adaptation evolution strategy (MA-ES), resulting in a second algorithm termed moment MA-ES based on an aging leader and learned challenger mechanism (MoMA-ES-ALLC). Moreover, numerical examples on various landscapes demonstrate that both proposed algorithms perform similarly and outperform the variants of CMA-ES in terms of convergence speed and search accuracy. Subsequently, a comprehensive evaluation is conducted focusing on the MoA-ES-ALLC. The comparison results on the CEC-2022 benchmarks show that it achieves better results than several state-of-the-art algorithms. Finally, the practicability of the proposed MoA-ES-ALLC is verified through applications to eight engineering problems and trajectory planning for a 6-DOF robot manipulator. The proposed MoA-ES-ALLC presents an efficient solution to the critical challenges of convergence speed and local optima in evolutionary computation, making it particularly effective for complex optimization problems.
This article focuses on the finite-time adaptive distributed Nash equilibrium seeking (NES) strategies for networked games of multiagent systems with partial-decision information. Multiplayer optimization and decision-making are widely applied in practical engineering scenarios such as smart grid dispatching, sensor networks, and distributed cooperative control. Existing distributed NES algorithms fail to simultaneously achieve finite-time convergence and adaptive gain tuning under partial-decision information, and often become invalid under switching topologies in dynamic networks. To overcome this challenge, we propose two types of seeking strategies leveraging gradient-based play, leader-following consensus algorithms, and adaptive control laws. First, a finite-time node-based adaptive seeking strategy is developed, where each agent adaptively adjusts its weight based on the consensus error in finite time. Then, a finite-time edge-based seeking strategy is designed by tuning the adaptive parameters combined with the weights of the edges of communication graphs in finite time. Moreover, we show that the proposed finite-time NES algorithms can be extended to switching communication networks, providing a feasible solution for distributed decision-making in large-scale networked systems requiring fast and reliable convergence. Finally, some illustrative examples are presented to verify the proposed algorithms’ effectiveness.
Considering the complexity of lag time and the convergence time, as well as the influence of the observer and the event-triggered mechanism, the prescribed-time lag bipartite consensus (PTLBC) control problem for nonlinear multi-agent systems (MASs) is researched in the paper. First of all, it designs the prescribed-time dynamic observer (PTDO) for followers to get the states of the leader within an arbitrarily prescribed time. Furthermore, to significantly decrease the communication consumption, the innovative event-triggered mechanism is studied for followers. To realize the lag bipartite consensus (LBC) of nonlinear MASs within an arbitrarily prescribed time, a PTLBC control scheme is presented via the aforementioned PTDO and event-triggered mechanism. With Lyapunov stability analysis method, sufficient conditions are obtained and detailed stability analysis is studied, which indicates that the nonlinear MASs can realize PTLBC. Furthermore, the analysis proves that the proposed event-triggered mechanism excludes Zeno behaviour. The theoretical analysis is validated by means of a simulation example.
This paper investigates the existence of Nash equilibria in multi-player nonzero-sum differential games governed by hybrid impulsive systems, where one player applies impulse controls while the remaining N players adopt piecewise-continuous controls. The proposed framework captures complex dynamics in which gradual adjustments coexist with abrupt interventions, with applications in traffic management, power dispatch, and emergency response. First, the necessary conditions for the existence of hybrid-strategy Nash equilibrium are rigorously derived using a variational method. Second, under traditional convexity assumptions, a sufficiency theorem for equilibrium existence is provided. By further introducing pseudoconvexity and invexity, a weak convexity framework with broader applicability is proposed, significantly relaxing the regularity requirements on the cost functions. Finally, numerical simulations, including a power system load dispatch case, demonstrate the validity of the theoretical results.
RGB-D cameras are widely used in indoor robots. However, their ranging capability for agricultural robots under natural lighting still needs to be evaluated. Especially in the field of robotics-based high-throughput crop phenotyping, the measurement accuracy of phenotypic parameters is deeply related to the ranging performances of RGB-D cameras. In this paper, we propose a depth-ranging evaluation framework and an online ranging compensation strategy for RGB-D cameras on phenotyping robots. The goal is to acquire high-quality depth-ranging performances for plant phenotyping tasks. First, we evaluate ranging performances of RealSense D435i and Kinect V2 under typical phenotyping scenes with different lighting conditions, verify their feasibility on different maize organ observations in different growth periods, and give the optimal observation ranging areas. Second, we employ image brightness to reflect the lighting situations, and propose a novel ranging compensation strategy to decrease the lighting influences in real-time. The results of sufficient field experiments show that RealSense D435i has better ranging performances than Kinect V2 for crop phenotyping, especially for open-field, in-row, and close-range observations. The optimal ranging area of RealSense D435i is within a region of [0.16-1.2] m. However, Kinect V2 is not suitable for field phenotyping robots due to significant interference from natural sunlight, limited measurement range, and instability in depth measurements under outdoor conditions. In addition, we also verify that our online depth error compensation strategy can effectively reduce the influences of lighting intensity and target distance on the depth ranging of RGB-D cameras. Although we test and verify our ranging evaluation framework and ranging error compensation strategy with two old-fashion cameras, the framework and strategy are generic and applicable to other new RGB-D cameras.
This article investigates the fuzzy adaptive practical prescribed-time formation control of nonlinear multi-agent systems with unmodeled dynamics based on a prescribed-time observer. At first, a prescribed-time observer is developed for followers to estimate the leader’s states. Then, fuzzy logic systems combined with adaptive methods are employed to approximate unknown continuous nonlinear functions. To compensate for the effects of unmodeled dynamics, a sliding-mode control strategy is incorporated. On these bases, a fuzzy adaptive practical prescribed-time formation control protocol is designed. This protocol guarantees that the formation error converges to a small neighborhood around zero within an arbitrary prespecified time. Finally, simulation results are provided to demonstrate the effectiveness of the proposed approach.
One of the key challenges in orchard robots is accurately localizing occluded fruits in complex environments, especially when the fruit targets are split into multiple isolated regions within images. Traditional single-task network models exhibit limited capability in discerning fragmented targets that belong to the same fruit but are segmented into multiple spatially isolated regions within images. In addition, fruit localization largely relies on high-cost sensors or additional 3-D localization algorithms. To address this issue, we propose a fruit detection and centroid localization method based on a Multi-Task Wavelet-Enhanced YOLO (MT-WavYOLO) to enhance the success rate of robotic operations on occluded fruit targets. Initially, a lightweight semantic segmentation branch was integrated into the YOLOv8 backbone network to precisely segment exposed fruits, while retaining the original object detection branch to fully identify occluded fruits. To address the diminished sensitivity of conventional models to geometric profiles of heavily occluded fruits, a novel feature fusion module, C2f_WTConv, was designed by incorporating wavelet transform convolution, leveraging the multi-frequency robustness of wavelet representations to enhance the model's feature extraction capabilities under complex orchard occlusions. Subsequently, a 3D frustum-based point cloud processing method was proposed, combining the detection results from MT-WavYOLO with the semantic segmentation masks to accurately localize occluded fruits. MT-WavYOLO demonstrated a 2%, 1.5%, and 2.2% improvement in Precision, Recall, and mAP50, respectively, on our custom-built dataset compared to the latest YOLOv10s model. Semantic segmentation performance, measured by Intersection over Union (IoU) and Accuracy, was improved by 5.2% and 3.8%, respectively, over the state-of-the-art Deeplabv3+ network. Compared to the adapted multi-task network YOLOP, MT-WavYOLO achieved a 3.4% increase in mAP50 and a 2.7% improvement in IoU. In addition, MTWavYOLO has a compact footprint of 10.2 M parameters and achieves approximately 27 FPS in real-time inference, thereby meeting the requirements of robotic harvesting operations. The proposed localization method was evaluated through 600 fruit localization tests using six different RGB-D cameras in an orchard environment. The average experimental results demonstrated that the centroid localization and radius estimation errors were reduced by 42.5%, 73.7%, 16.17%, and 11.25%, respectively, compared to traditional 3D bounding box methods and our previous approaches. These results indicate that the MT-WavYOLO combined with the frustum-based method significantly enhances the accuracy of apple localization under complex orchard conditions using consumer-grade sensors, providing a strong practical foundation for non-destructive robotic harvesting.
In this paper, a distributed real-time pricing algorithm based on social welfare maximisation is proposed to institute electricity buying-back schemes for the smart grid that contains multiple providers and integrated with renewable energy (RE) and storage devices. In the proposed model, a profit function is introduced to encourage people to use more RE, and the depreciation of the storage capacity is considered. By dual decomposition, the primal multiseller-multibuyer problem is decoupled into a set of single-buyer and single-sellersingle-time-slot subproblems, through which the relationship between prices of electricity and Lagrangian multipliers is derived. Then, a distributed algorithm is further designed to obtain the optimal solution. The strong duality of the original problem is also demonstrated. With this approach, subproblems are solved by each user and utility company, respectively, which ensures privacy and system scalability. Numerical results show that the proposed method has good performance in reducing peak-time loading and balancing system energy distribution.
This paper investigates the zero-sum (ZS) game problem for discrete-time nonlinear impulsive systems based on adaptive dynamic programming (ADP). First, a nonlinear impulsive system model is constructed, in which two impulse control players alternately trigger impulses at specific times to achieve optimal system performance. Then, transformations are applied to the nonlinear systems and the utility function to enable the ADP framework. Subsequently, the upper and lower iterative impulsive ADP algorithms are designed to approximate the upper and lower optimal performance index functions for this ZS game. Through the iterative process of the ADP algorithm, the system gradually approaches a saddle-point equilibrium under impulse controls. Simulation results confirm the proposed improved ADP method's efficacy.
This article aims to establish a Bayesian Stackelberg game framework for analyzing the incomplete information demand response management with overlapping electricity sales areas, and further provide the corresponding equilibrium strategies. Considering that the satisfaction parameters of power users are private, a Bayesian game model is constructed among these power users, and a non-cooperative game model is established due to the price competition of microgrids. To ensure the sequential interactions of demand response, a Stackelberg game is developed by assuming that the microgrids are leaders and the power users are followers, and the Bayesian Nash equilibrium and Stackelberg equilibrium are proved to exist and are unique under some conditions. In addition, the Bayesian Nash equilibrium for power users is obtained using the fictitious play method in the symmetrical case, and an iterative algorithm is presented for determining the Stackelberg equilibrium. Finally, the numerical simulations are provided showing the effectiveness and convergence of the iterative algorithm, which indicates that the proposed approach can enhance profits for microgrids while ensuring power supply and demand balance.
Global fruit production costs are increasing amid intensified labor shortages, driving heightened interest in robotic harvesting technologies. Although multi-arm coordination in harvesting robots is considered a highly promising solution to this issue, it introduces technical challenges in achieving effective coordination. These challenges include mutual interference among multi-arm mechanical structures, task allocation across multiple arms, and dynamic operating conditions. This imposes higher demands on task coordination for multi-arm harvesting robots, requiring collision-free collaboration, optimization of task sequences, and dynamic re-planning. In this work, we propose a framework that models the task planning problem of multi-arm operation as a Markov game. First, considering multi-arm cooperative movement and picking sequence optimization, we employ a two-agent Markov game framework to model the multi-arm harvesting robot task planning problem. Second, we introduce a self-attention mechanism and a centralized training and execution strategy in the design and training of our deep reinforcement learning (DRL) model, thereby enhancing the model’s adaptability in dynamic and uncertain environments and improving decision accuracy. Finally, we conduct extensive numerical simulations in static environments; when the harvesting targets are set to 25 and 50, the execution time is reduced by 10.7% and 3.1%, respectively, compared to traditional methods. Additionally, in dynamic environments, both operational efficiency and robustness are superior to traditional approaches. The results underscore the potential of our approach to revolutionize multi-arm harvesting robotics by providing a more adaptive and efficient task planning solution. We will research improving the positioning accuracy of fruits in the future, which will make it possible to apply this framework to real robots.
We study the distributed nonconvex resource allocation with Limited Communication Data Rate (LCDR) over a communication network. Each node in the network has its own private cost function and determines the optimal resource allocation through interactions solely with its neighboring nodes. The nodes need to cooperatively minimize the total cost function to achieve the optimal resource allocation under the constraint of constant total resources. First, we consider exact communication. By Lagrange dual method, we propose a successive convex approximation-based distributed dual gradient tracking algorithm to solve the distributed nonconvex resource allocation problems. Then, we consider the case of digital communication among nodes based on LCDR. The information transmission among nodes is based on the Dynamic Encoding and Decoding (DED) with finite-level uniform quantization. We propose a successive convex approximation-based distributed dual gradient tracking algorithm with LCDR and conduct numerical simulations. The numerical results show that the algorithm converges based on merely one-bit quantizers when appropriate step sizes and scaling functions are chosen.
For heterogeneous multi-robot systems with unknown parameters and input constraints, the problem of simultaneous target tracking and collision avoidance is studied in this paper. A novel auxiliary system is proposed to generate the corresponding compensation signal to tackle the difficulty of controller design resulting from input constraints. For each heterogeneous robot system, a unique collision avoidance controller is constructed utilizing the repulsive potential field method. Then, the adaptive collisions-free tracking control strategy is presented under the control framework of the proposed auxiliary system. It can be proved via Lyapunov stability theory that the designed controllers can allow the heterogeneous multi-robot systems to track the reference target with collision avoidance and guarantee that all the closed-loop signals in the multi-robot systems are bounded. Subsequently, the validity of our control scheme is further verified through a common simulation example.
This paper investigates a pursuit-evasion game involving multiple pursuers and a single evader with unknown nonlinear dynamics under switching topologies. A distributed neural network-based estimator is designed to reconstruct the evader’s dynamics. An augmented system is constructed by integrating estimation errors with pursuer–evader relative errors, and a discounted performance index is established to derive optimal control strategies for both sides. To address unknown nonlinearities, an integral reinforcement learning-based actor–critic architecture is proposed for online approximation of the optimal value function and control strategies. The uniform ultimate boundedness of the closed-loop error system and network weight estimation error is rigorously proven. Simulation results verify the effectiveness of the proposed method.
In this paper, a novel primal-dual adaptive dynamic programming (PDADP) method is developed to solve finite-horizon optimal control problems (OCPs) with isoperimetric constraints for continuoustime nonlinear systems. The OCP with isoperimetric constraints is approximated by a series of linear quadratic time-varying OCPs with isoperimetric quadratic constraints, which are solved by the primal- dual (PD) method. The convergence of the PDADP method is proven. Furthermore, the optimality of the solution is analyzed by deriving the necessary optimality conditions via Pontryagin's principle and proving that the limiting values of the iterations satisfy the necessary optimality conditions. Finally, simulation experiments are provided to show the effectiveness of the present PDADP method. (c) 2024 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.