In the study of complex networks, the identification of key nodes is crucial but challenging. Traditional centrality indicators (such as degree, betweenness, and closeness centrality) are difficult to achieve a balance between accuracy, effectiveness and time complexity, time-consuming and resource-consuming, or unable to accurately locate key nodes. In order to solve these problems, based on the commonly used Laplacian centrality (LC), a two-hop LC (TLC) index is proposed, which aims to improve the index recognition performance by introducing second-order neighbor node information. The two-hop degree centrality (TDC) is proposed and combined with LC to derive the TLC, whose superiority is verified through multi-dimensional comparative analysis. Meanwhile, the application of higher-order hop degree centrality in the derivation of LC is deduced, and it is proved that it is not feasible to apply three-hop and above degree centrality to the derivation of LC.
In this paper, an off-policy hierarchical reinforcement learning (HRL) algorithm is proposed to solve the collision avoidance problem for a class of multi-agent systems. The collision avoidance refers to maintaining a predefined formation pattern and avoiding collisions with obstacles while driving each agent to the target state, which is formulated as a differential game. We leverage the idea of divide and conquer to artificially decompose the problem into three corresponding subtasks: target state attraction, neighbor agent repulsion, and static obstacle repulsion, to cope with the complex external environment. The off-policy HRL algorithm is designed based on the original policy iteration algorithm and implemented in real-time using only measured data to cope with the problem of completely unknown system information. Compared with the traditional least-square and gradient descent approach, critic and action neural networks of each subtask are simultaneously added to a broad learning system (BLS). It is worth noting that the pseudo-inverse operation of BLS allows us to achieve a faster and better approximate solution of the weight using global data online. The uniform ultimate bounded stability of the closed-loop system is proved based on the Lyapunov approach. Finally, a simulation example is given to demonstrate the effectiveness of the developed algorithm.
This paper investigates cooperative formation tracking control strategy of a multiple quadrotor unmanned aerial vehicles (multi-QUAVs) systems subject to mismatched disturbances and actuator saturation, aiming to ensure that each follower tracks the reference trajectory of the leader with a desired geometric configuration within a user-prescribed time. To address singularity problem that arises from the differentiation of scaling functions, a non-scaling virtual control law construction strategy is proposed, in which the time-varying scaling function is directly injected into the control channel as an external gain. Furthermore, RBFNNs are utilized as online approximators, and an anti-windup compensator is designed to reduce adverse effects of truncation errors induced by actuator saturation. Additionally, an event-triggered control is incorporated to effectively save communicational and computational resources, and exclusion of the Zeno phenomenon is rigorously proved. Stability analysis shows that all signals of the closed-loop system are bounded, and tracking error of every quadrotor unmanned aerial vehicle (QUAV) converges to an arbitrarily small neighborhood of origin within prescribed time, with the convergence being independent of initial conditions. Finally, two simulation examples of multi-QUAVs formation flight are presented to validate effectiveness and robustness of this work.
In multiagent mixed-motive games, the dynamic adaptation among agent policies induces a nonstationary environment, posing challenges to stable learning and convergence toward Pareto-optimal equilibrium. To address this, this article proposes a reciprocity-based actor-critic (AC) framework by embedding a mutual help mechanism, which encourages agents to balance self-interest with the preferences of others. Our approach utilizes a centralized critic to estimate Q-values and infer expected policies of other agents, while decentralized actors concurrently update individual policies through selective altruistic adjustments. By integrating expected policy computation and redefining the advantage function, it promotes agent cooperation without sacrificing individual interests, thus mitigating local optimality. A policy improvement lower bound is established to ensure monotonic convergence with theoretical guarantees for Pareto improvement. Furthermore, we design a hybrid neural architecture that combines a shared backbone network for global knowledge transfer with personalized branches for individual policy optimization. This structure balances coordination and adaptation by synthesizing extrinsic environmental feedback with intrinsic agent motivation. Experimental results validate the effectiveness of our method in enhancing behavior coordination and learning efficiency in decentralized multiagent settings. Internal feedback, External feedback
Perception plays a critical role in developing effective policies for multi-task embodied manipulation, in which visual comprehension and task interpretation are essential. Existing methods typically rely on multi-view 2D representations for visual perception, aiming to build computation-friendly perception modules through imitation learning from extensive collections of high-quality robot trajectories. However, these approaches face significant challenges when expert demonstrations are limited or tasks are highly complex, resulting in inefficiencies. To address these limitations, we propose Temporal Consistent Multi-View Perception (TMVP), a sample-efficient two-stage framework for robot manipulation that integrates temporal information into multi-view representations. Specifically, TMVP employs contrastive learning to extract meaningful, task-relevant features from visual inputs, enhancing temporal consistency and alignment with task instructions. This results in visual representations that are temporally coherent and grounded in task trajectories, enabling the model to better comprehend and execute complex manipulation tasks from diverse perspectives. Experiments conducted on RLBench demonstrate that TMVP outperforms baseline models across a wide range of tasks, achieving superior multi-task performance and few-shot training efficiency. These results highlight the potential of TMVP as an efficient and effective solution for embodied manipulation.
This paper proposes a saturated impulsive control method for bearing-constrained UAV swarm. Unlike conventional approaches that rely on continuous communication and sensing, the proposed method ensures effective formation control even when information is acquired intermittently, with random interruption intervals. This significantly reduces the performance demands on communication and sensing electronic devices. Moreover, a practical distance estimation technique, coupled with a collision avoidance assistance strategy, eliminates the need for initial position constraints typically required in traditional bearing-based methods, thereby improving the adaptability of the algorithm across diverse scenarios. In contrast to standard impulsive control methods, which are generally restricted to first-order single-agent systems and seldom address the issue of input saturation, the proposed approach is applicable to second-order multi-agent systems and effectively mitigates input saturation through a specified matrix constraint formulation. This extension broadens the approach to freight UAV swarm deployments, ensuring safe, reliable operation within hardware constraints. The effectiveness of the proposed strategy is demonstrated through a series of simulation results.
This paperaddresses the resilient consensus problem of networked systems under hybrid attacks by co-designing the network and control layers. To ensure system stability and enhance robustness, the approach involves constructing hidden nodes and edges, with three critical steps: detection, isolation, and control. Firstly, auxiliary variables are employed as information carriers to estimate neighboring states and detect stealthy attacks, following the mechanism of safe diffusion. Secondly, compromised nodes are isolated (removed) after identification to prevent the propagation of attacks, with an emphasis on the connectivity conditions for achieving resilient consensus under various attack scenarios. Thirdly, a secure control algorithm is developed for unaffected nodes to maintain system consensus post-isolation. Finally, simulations validate the algorithm's effectiveness while demonstrating its advantages due to highly tolerant isolation conditions and robust anti-attack performance. Overall, this study provides a solution for resilient control of networked systems under single or hybrid complex cyber-attacks.
This article focuses on the practically adaptive safety tracking issue for the motor turntable of anti-drone systems subject to stochastic disturbances. Unlike existing research methods, the constraints issue of current for the motor turntable is fully considered to ensure the safety of the experimental environment. Specifically, the tangent function is employed for coordinate transformation as a means to obtain unconstrained variables, addressing current constraints. Meanwhile, chattering problems and stochastic disturbances are inevitable in the physical environment, potentially leading to wear and tear on actual systems and even instability. Furthermore, a novel adaptive fast finite-time safety algorithm has been proposed to address those problems for anti-drone systems, enabling precise strikes against the target. On this basis, an innovative fast finite-time controller has been successfully constructed that integrates the neural networks with special piecewise functions via the backstepping techniques. Therefore, the proposed controller can not only ensure the boundedness of all closed-loop system signals in probability but also avoid the singularity of the control signals. Finally, the superior performance of the proposed tracking algorithm is thoroughly demonstrated through a combination of hardware experiments and numerical simulations.
Mobile Internet of Things (MIoT) has emerged as a key enabler of the low-altitude economy. However, task-driven dynamic interactions among mobile devices introduce significant security risks due to malware attacks. Despite extensive studies on malware defense, identifying effective patch allocation strategies to enhance MIoT resilience remains challenging. To fill this gap, we first develop a temporal multilayer model that integrates device mobility with decentralized patch dissemination. Second, we establish a microscopic Markov chain approximation-based malware-patch propagation model and derive a defense-dependent resilience boundary. This boundary characterizes whether an initial malware introduction will die out under a given defense regime. Third, we formulate a finite-horizon resource-constrained optimization problem that minimizes cumulative infections subject to network bandwidth and maintenance budgets, and propose a Dynamic Deployable Optimization (DDO) algorithm for patch allocation. Extensive experiments demonstrate that the DDO algorithm outperforms topology-based heuristics. The results further reveal distinct time-scale dependencies in defense. Specifically, bandwidth optimization is critical for rapid short-term suppression, whereas long-term resilience benefits from maintenance policies that maximize sustained immunization. We also find that the impact of device mobility on malware propagation is not strictly amplifying or suppressing, as it reshapes the spatial heterogeneity of device distribution among stations. Overall, this work provides a theoretically grounded framework for adaptive decentralized patching, enabling resilience enhancement in evolving cyber-physical systems.
This paper investigates the predefined-time formation control of quadrotor unmanned aerial vehicles (QUAVs) under input saturation constraints. To mitigate the effects of input saturation, an auxiliary system is meticulously designed. In contrast to other studies which merely attain the asymptotic stability of the auxiliary signals, the auxiliary system proposed ensures the predefined-time stability, thus contributing the stability of entire system. Subsequently, based on the predefined-time stability criterion, a distributed formation controller is developed, thus guaranteeing the predefined-time stability and robustness of formation under input saturation. Finally, simulations validate the effectiveness of the proposed method.
The work intends to study the optimized leader-following consensus control with fault-tolerant ability for a class of second-order nonlinear multi-agent systems (MASs) subjected to actuator failures. Since MAS is an important topic in the field of low-altitude of technology and engineering, this work can contribute to the development of this field. To realize optimized control, reinforcement learning (RL) strategy is employed for avoiding the derivation of analytical solution of Hamilton–Jacobi–Bellman (HJB) equation. Nevertheless, second-order MASs need to simultaneously regulate both position and velocity states to reach consensus, which inevitably increases the complexity of optimization algorithms. In this context, to endow the optimal control scheme with fault-tolerant capability against actuator faults, an adaptive estimation algorithm for unknown actuator fault parameters is integrated with the RL-based optimal control strategy. For making the combination smoothly, a simplified RL algorithm is obtained by taking the negative gradient of a simple positive function, which is equivalent to HJB equation. Finally, both theoretical analysis and numerical simulations verify the feasibility of the proposed control.
In this study, a hybrid air-sea swarm control methodology is introduced to achieve resilience against denial-of-service (DoS) attacks, where such attacks intermittently block communication links to disrupt coordination. We start by establishing a swarm system comprising uncrewed aerial vehicles (UAVs) and uncrewed surface vehicles (USVs) with distinct dynamic properties and degrees of connectivity. Unlike existing studies with homogeneous communication links, we consider heterogeneous (hybrid) UAV-to-UAV, USV-to-USV, and UAV-to-USV communication links. A cooperative-competitive interaction strategy is studied, where the swarm is divided into two clusters based on vehicle type. To face cyber vulnerabilities, an event-triggering security-oriented control law is designed, which is proven to counteract attacks while balancing communication and control resource usage. Furthermore, the relation between communication weights and system resilience to DoS attacks is analyzed. Numerical simulations demonstrate the efficacy and resilience of the proposed approach.
This paper proposes an adaptive fast finite-time tracking consensus protocol for high-order nonlinear multi-agent systems. To overcome the limitation of finite-time stability, where the convergence speed slows down when the initial state is far from the origin, the fast finite-time stability theory is incorporated into the multi-agent systems to ensure rapid convergence of the tracking error. Furthermore, the power integrator technique is integrated into the backstepping framework to address the inherent singularity issues in high-order systems. Meanwhile, neural networks are used as online approximators to model unknown nonlinear functions, with the tanh(.) function adopted to mitigate the impact of approximation errors effectively. The developed dynamic event-triggered controller can reduce the frequency of control updates, effectively saving communication resources. Finally, two simulation examples demonstrate the effectiveness of the proposed strategy.
Deep Non-negative Matrix Factorization (DMF) holds immense potential for learning hierarchical data representations. However, the commonly used pre-training and fine-tuning paradigm may suffer from an optimization inconsistency: the reconstruction-driven fine-tuning stage can degrade the hierarchical representations learned during pre-training. To resolve this, we propose Hierarchical Dynamic Self-Supervised Learning (HDSSL), a framework for training DMF models with intermediate self-supervised guidance. Unlike traditional methods that rely solely on a single global objective, HDSSL introduces a dynamic self-supervision mechanism that acts as a hierarchical structural regularizer. It iteratively refines target representations by projecting global data into the current latent space and updates network parameters via a hierarchical gradient aggregation algorithm. This approach explicitly injects consistent intermediate guidance, effectively mitigating optimization conflicts and reducing the risk of representation degradation during optimization. Extensive experiments on benchmark datasets demonstrate that HDSSL achieves consistently competitive clustering performance and favorable optimization behavior across different settings. Moreover, it shows faster convergence and improved stability in our experiments, while eliminating the need for layer-wise pre-training.
In the field of artificial intelligence, the wide adoption of third-party data has heightened the risk of backdoor attacks based on data poisoning. Although post-training defenses can mitigate such attacks, aggressive strategies often degrade the model’s performance on its main task. To address this, a novel method that combines backdoor detection and elimination through machine unlearning is proposed. Specifically, unlearning perturbation is first defined to capture the parameter variation induced by forgetting a subset of samples. Subsequently, experiments confirm that backdoor samples exhibit lower sensitivity to perturbations generated from normal samples. In addition, a learning-dynamics analysis attributes this discrepancy to unlearning sensitivity, which is defined as the inner product between the gradients of normal and backdoor samples. This analysis further demonstrates that this metric quantifies the extent to which backdoor removal perturbs the model’s main task. Leveraging this insight, an orthogonality-constrained gradient projection method projects the unlearning gradient onto the null space of the normal-sample gradient, thereby eliminating the aforementioned unlearning sensitivity and preserving the accuracy of normal samples. The proposed method is evaluated across six backdoor attack scenarios and two network architectures, reducing the average attack success rate by 96.34 percentage points and improving robust accuracy by 83.68 percentage points, while maintaining the model’s performance on the main task.
Aiming at the problems of insufficient coordination accuracy and poor adaptability to dynamic environments caused by complex coupling relationships in large-scale heterogeneous multi-agent systems, this paper proposes a cooperative decision-making method that balances rationality and adaptability. Firstly, To address the limitation that uniform weights in traditional mean field games cannot characterize the differentiated contributions of heterogeneous agents, this paper proposes a learnable heterogeneous weight mechanism. The weights are dynamically generated and adaptively updated by the environmental perception network to explicitly reflect the influence of different types of agents on group behavior. On this basis, we further present a heterogeneous weighted mean field game model, which quantifies the differentiated impacts of different types of agents on group behaviors through type-level dynamic weights, breaking through the limitations of uniform weights in traditional mean field theory. Secondly, a dynamic adaptive mean field decision-making framework is designed; an environment perception module is introduced to update the reward function and state transition parameters in real time, and combined with reinforcement learning, the Dynamic Adaptive Mean Field Reinforcement Learning (DAFRL) algorithm is constructed to achieve real-time tracking of equilibrium solutions. Finally, experiments conducted on the scenario of red-blue UAV swarm confrontation demonstrate that the proposed heterogeneous weighted modeling method effectively addresses the coupling problem of heterogeneous agents, and the dynamic adaptive framework significantly enhances environmental robustness. DAFRL exhibits comprehensive advantages in efficiency, accuracy and stability, thereby providing theoretical and technical support for multi-agent cooperative decision-making in complex scenarios.
Effective policy optimization in multiagent reinforcement learning (MARL) necessitates extensive exploration of high-dimensional state-action spaces. However, such exploration may not only trigger unsafe states but also compromise system stability, posing significant challenges for deployment in safety-critical systems. To address this challenge, this article proposes a safety-stability layer that integrates robust control barrier functions (RCBFs) and input-to-state stable control Lyapunov functions (ISS-CLFs) for multiagent systems operating in unknown environments with uncertain dynamics. Furthermore, by integrating safety-stability constraints with a MARL framework, during the training phase, we exclusively focus on goal-reaching objectives to expand the policy network’s exploration space, while in the deployment phase, policy outputs are filtered through a real-time safety-stability layer. In addition, an event-triggered mechanism for action compensation calculation is designed based on safety condition assessments to conserve computational resources. Finally, the effectiveness of the proposed method is validated through simulation experiments in dynamic multiunicycle environments. The results demonstrate that our approach not only ensures strict adherence to safety constraints but also significantly enhances the task execution efficiency of multiagent systems.
In this article, we investigate the adaptive safety-constrained control problem for quadrotor unmanned aerial vehicle (QUAV) clusters in dense forest environments to achieve adaptive navigation and obstacle avoidance. Compared to traditional methods, obstacle avoidance constraints are introduced for the first time, and the limitations of fixed formations and the need for prior data are eliminated. First, a cluster constraint mechanism is developed to constrain the distance between QUAVs within the cluster and the distance between the QUAV and the desired trajectory. Then, considering the lack of targeted obstacle avoidance constraint mechanisms in previous methods and the extensive prior data required by learning-based approaches, an obstacle constraint model is established to ensure that the QUAV maintains a safe distance from obstacles to avoid collisions. Finally, an adaptive safety control strategy for QUAV clusters is proposed by combining constraint conditions and stability criteria. Under the proposed control strategy, the QUAV clusters can achieve stable, safe, and efficient navigation and obstacle avoidance, and all constraints will always be satisfied. Furthermore, a numerical simulation experiment on a QUAV cluster navigation demonstrates the effectiveness and flexibility of this strategy.
Safety is paramount for deploying reinforcement learning (RL) in autonomous unmanned systems operating in complex environments. Conventional safe RL approaches often rely on discounted safety costs, which may misalign with real-world finite-horizon safety requirements. This paper presents Constrained Risk-Aware Policy Optimization (CRAPO), an off-policy safe RL algorithm designed to enhance adaptability and stability under undiscounted safety constraints. CRAPO jointly optimizes policies using gradients on discounted safety costs while adaptively adjusting safety weights based on undiscounted measurements, ensuring compliance with true operational limits. To improve stability in constraint enforcement, we integrate the Modified Differential Method of Multipliers (MDMM), effectively damping oscillations in Lagrange multiplier updates. Risk sensitivity is incorporated through a sample-based Conditional Value at Risk (CVaR) estimation, avoiding distributional approximations and reducing computational overhead. Evaluations on challenging navigation tasks in the Safety Gym suite demonstrate that CRAPO achieves superior risk control, stable constraint satisfaction, and competitive performance compared to state-of-the-art safe RL methods. These results highlight CRAPO’s practicality and robustness for safety-critical spatial intelligence in unmanned urban systems..
In this paper, an analysis-definition-processing (ADP) framework is proposed to search positive-incentive noise in continuous action iterated dilemma (CAID). We analyze the influence of communication noise on the cooperative behavior of players in the system and introduce the concept of positive-incentive noise in CAID. We design a global cost function to ensure convergence of the system can be achieved and strive to improve the final level of cooperation. An optimal CAID control method is proposed to derive the deterministic optimal learning rate in analytical form, avoiding the variability and uncertainty brought about by neural network fitting or parameter adjustment. On this basis, the convergence of the dynamic model is further analyzed by using the Lyapunov function instead of the Jacobian matrix. Additionally, an adaptive filtering mechanism is designed to dynamically ensure that only positive-incentive noise affects the system, effectively reducing the impact of negative noise and enhancing system stability. The framework is validated through simulations involving triple classical game models, including the hawk-dove game, the stag hunt game, the chicken game on networks, and a straightforward illustrative example.