In this paper, we study attitude synchronization on $SO(3)$ under switching communication topologies within a hybrid framework. A hybrid model is constructed to capture continuous attitude dynamics and discrete topology switching, and a distributed controller based on relative attitude information is adopted. Due to switching, the Lyapunov function may be discontinuous at switching instants. To address this issue, a trajectory-based bounding technique is developed. Under a minimum dwell-time condition, semiglobal exponential synchronization is established. Numerical simulations validate the theoretical results.
Renewable energy has attracted widespread attention as a vital component in the global transition to low-carbon energy systems. Renewable-integrated energy systems (R-IESs) effectively combine diverse renewable energy (e.g., solar, wind, hydro) with carbon-free nuclear power to maximize energy efficiency and operational flexibility, an integration facilitated by the modularity and passive safety of small modular reactors (SMRs). Within such systems, optimization frameworks are typically designed to maximize system performance while minimizing operational costs by adjusting the decision variables. However, optimizing R-IESs remains highly challenging due to the complex interaction between stochastic generation and strict safety constraints. To address these limitations, artificial intelligence (AI) has emerged as an influential approach for enhancing system safety, reliability, and sustainability. This review discusses advanced methodologies including reinforcement learning (RL) and large language models (LLMs), detailing their applications and state-of-the-art algorithms. Furthermore, this review summarizes the efficacy of RL-based and LLM-based optimization frameworks in managing multi-energy coupling, optimizing generation, and ensuring system reliability. Finally, a forward-looking perspective is presented on emerging AI paradigms, such as physics-informed AI, LLM-based agent, embodied intelligence and multimodal LLMs, highlighting the capacity to realize a profound cyber-physical convergence in R-IESs.
This article analyzes the vulnerability of the circuit system in the charge pump phase-locked loop (CPPLL) in inverter-based distributed energy systems by designing an attack strategy that can exploit the tracking characteristics and closed-loop bandwidth of the CPPLL at non-nominal grid frequencies. Meanwhile, considering attack resource constraints, the CPPLL-based attack is analyzed across different regions to study the selection of attack targets. A multi-objective optimization problem is formulated to maximize the attack effect, such as renewable energy revenue loss, power dispatch error and voltage deviation rate. To solve this dynamic and high-dimensional optimization problem, an improved Deep Deterministic Policy Gradient (DDPG) combined with robust optimization is proposed to integrate the multi-objective into a scalarized reward function and learn the optimal policy with an Actor-Critic network. Simulation studies, conducted on a modified IEEE-33 bus system, reveal that the CPPLL-based attack with the proposed attack strategy can result in: 1) larger loss to renewable energy revenue, power dispatch, and voltage deviation rate at the low-nominal grid frequency and 2) maximized comprehensive loss in the distributed energy system by attacking cluster inverters compared to distributed inverters.
Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback. Existing approaches typically mitigate performance degradation caused by observation delays by constructing augmented states or predicting the true states. However, these methods often overlook the inherent discrepancy between delayed state and true states induced by stochastic MDP. We theoretically prove the existence of such a discrepancy and show that it leads to the degradation of the optimal policy. To address this challenge, we propose Diffusion Guided Uncertainty Aware Delayed Policy Optimization (DUPO). Our method explicitly models the relationship between delayed state message and the current state using a diffusion model, and leverages the resulting discrepancy estimates to weight delayed policies. Extensive experiments on continuous robotic control tasks with multiple stochastic delays demonstrate that DUPO consistently outperforms existing methods and remains effective even under long and random delay scenarios.
This article investigates the security issue caused by false data injection attacks in distributed estimation, wherein each sensor can construct two types of residues based on local estimates and neighbor information, respectively. The resource-constrained attacker can select partial channels from the sensor network and arbitrarily manipulate the transmitted data. We derive necessary and sufficient conditions to reveal system vulnerabilities, under which the attacker is able to diverge the estimation error while preserving the stealthiness of all residues. We propose two defense strategies with mechanisms of exploiting the Euclidean distance between local estimates to detect attacks, and adopting the coding scheme to protect the transmitted data, respectively. It is proven that the former has the capability to address the majority of security loopholes, while the latter can serve as an additional enhancement to the former. By employing the time-varying coding matrix to mitigate the risk of being cracked, we demonstrate that the latter can safeguard against adversaries injecting stealthy sequences into the encoded channels. Hence, drawing upon the security analysis, we further provide a procedure to select security-critical channels that need to be encoded, thereby achieving a trade-off between security and coding costs. Finally, some numerical simulations are conducted to demonstrate the theoretical results.
Open-Set Domain Adaptation for Semantic Segmentation (OSDA-SS) presents a significant challenge, as it requires both domain adaptation for known classes and the distinction of unknowns. Existing methods attempt to address both tasks within a single unified stage. We question this design, as the annotation imbalance between known and unknown classes often leads to negative transfer of known classes and underfitting for unknowns. To overcome these issues, we propose SATS, a Separating-then-Adapting Training Strategy, which addresses OSDA-SS through two sequential steps: known/unknown separation and unknown-aware domain adaptation. By providing the model with more accurate and well-aligned unknown classes, our method ensures a balanced learning of discriminative features for both known and unknown classes, steering the model toward discovering truly unknown objects. Additionally, we present hard unknown exploration, an innovative data augmentation method that exposes the model to more challenging unknowns, strengthening its ability to capture more comprehensive understanding of target unknowns. We evaluate our method on public OSDA-SS benchmarks. Experimental results demonstrate that our method achieves a substantial advancement, with a +3.85
Online learning-based control is a promising approach to control uncertain systems, where unknown components are identified during operation to improve control performance. However, resource-intensive online learning algorithms introduce non-negligible computational delays, especially when executed on systems with limited local computational resources. To mitigate this, an in-network online learning-based control structure is employed by deploying the learning-based controller on a remote computation node and connecting it via a communication channel. In this paper, control performance guarantee is first established by deriving tracking error bound for the in-network control architecture, while accounting for computational delays. The derived tracking error bound allows for diverse communication and computation strategies under a specific condition, including time-/event-triggered mechanisms. Additionally, the trade-off between communication and computation performances is shown for a given desired control performance. Furthermore, to enhance the efficiency in both communication and computation, an efficient control framework with an asynchronous event-triggered mechanism in both control and online learning is devised under the existence of computational delay. The proposed event-triggered strategy is proven to achieve the same control performance as time-triggered scenario while excluding Zeno behavior. Finally, we derive an explicit expression of the proposed event-trigger condition for exponentially stabilizable systems, and demonstrate its effectiveness through simulations.
This article addresses the data-driven stabilization problem for a class of time-scale-type (TST) sampled-data networked control systems (NCSs) with unknown model parameters. A novel data-driven sampled-data (DDSD) control protocol is developed that relies solely on offline-collected datasets, thus eliminating the need for explicit model knowledge. Furthermore, to enlarge the maximum allowable sampling interval (ASI) while ensuring admissibility of all sampling instants, a TST matrix exponential gain is embedded into the DDSD control protocol, and a backward-jump sampling operator is introduced to address the discontinuities inherent in time scales. By integrating the direct data-driven methods, the theory of TST dynamic equations, and a generalized Halanay-like inequality, we establish a set of sufficient conditions that ensure the exponential stability of TST sampled-data NCSs under the proposed DDSD control protocol. We further employ a convex optimization framework to maximize the ASIs while satisfying stability constraints. Finally, the effectiveness of the proposed approach is demonstrated through numerical simulations and a case study involving an operational amplifier circuit with unknown model parameters.
This article investigates target-attacker-defender (TAD) differential games comprising an active target, multiple attackers, and defenders, including some anomalous agents within the defenders' team. First, we introduce two types of anomalous defenders, characterized by coefficients in the performance metrics, termed "greedy" and "fearful." Then, we explore two distinct scenarios: first, when the attackers possess unlimited observation capabilities while visibility limitations constrain the defenders and target, and second, when all agents are subject to imperfect information, meaning they only possess limited visibility. The interactions among agents are modeled as nonzero-sum games to analyze their optimal decision-making processes. Due to the agents' visibility constraints, a corresponding visibility network emerges during their interactions. To address this problem, we employ inverse game theory to derive Nash equilibrium strategies with adaptive state feedback for agents. Furthermore, we analyze the impact of anomalous defenders on the victory conditions and capture time of the defenders' team. Finally, we validate the effectiveness of our results through numerical simulations.
Opinion dynamics models that elucidate the evolution and formation of opinions conventionally focus on pairwise interactions within graphs, often overlooking the complex higher-order interactions that arise in real-world social networks, such as online meetings and group chats. In this article, a continuous-time dynamical system is developed to study opinion-forming processes over higher-order networks associated with undirected hypergraphs. The proposed model introduces a novel diffusion-like interaction function to characterize interactions of different orders over hypergraphs. The convergence and stability of the dynamical systems are further examined in both the presence and absence of stubborn individuals. Building on traditional opinion dynamics models and integrating weak-tie theory, we emphasize the critical role of higher-order interactions in shaping individual opinions, enhancing network communication efficiency, and mitigating opinion polarization. Finally, all theoretical results are extensively investigated and empirically validated through numerical experiments on both synthetic and real-world network datasets.
This paper studies multi-agent reinforcement learning with submodular team utilities for online distributed task allocation. In this setting, each agent selects one action from a local categorical policy, so feasible joint actions form a partition matroid over agent-action pairs. Classical multilinear extensions use independent Bernoulli sampling and therefore do not match the categorical policies executed by decentralized agents. To address this mismatch, we introduce the Partition Multilinear Extension (PME), a continuous relaxation whose value equals the expected team utility under factorized categorical policies. We prove that submodular difference rewards provide unbiased PME marginal-gradient information and yield a stagewise score-function policy-gradient estimator. Based on this connection, we propose SubMAPG, a centralized-training decentralized-execution policy-gradient framework with masked categorical policies and submodular difference-reward training signals. For the associated PME marginal-space projected stochastic-gradient dynamics, we prove a stagewise 1/2-approximation guarantee and sublinear dynamic regret in slowly varying environments, measured by the path length of the optimal PME marginals. To handle open systems with time-varying agents and targets, we instantiate SubMAPG with graph neural network policies. Experiments on multi-robot coverage and multi-target tracking show that SubMAPG outperforms local greedy and shared-reward baselines and is competitive with centralized myopic greedy strategies.
This article investigates reach-avoid games involving defenders equipped with capture radii, where both defenders and attackers have different speeds. The main challenge lies in using geometric methods to analyze different speed ratios, construct barriers, or defensive advantage angles, and divide the state space into defensive and offensive advantage regions. This article proposes optimal analytical strategies for players based on the corresponding payoff functions, depending on the attacker's position within different winning regions under various speed ratios. In addition, we demonstrate the existence of a unique optimal target point within the offensive advantage region. Unlike numerical methods, which are limited by computational complexity and real-time application capabilities, the proposed method allows for the precise calculation of barriers and real-time updates in nonpoint capture scenarios. Finally, simulation results validate the effectiveness of the constructed barriers in multiplayer reach-avoid games.
This paper studies policy learning for distributed task allocation in open multi-agent systems, where agents may join and leave in a time-varying fashion, with submodular stage team utilities. At each time, the active agents select actions from local categorical policies such that the feasible joint agent-action pairs form a partition matroid. Standard continuous relaxations of submodular set functions are based on independent Bernoulli sampling, making them inconsistent with agents' policies.To solve this mismatch, we propose the partition multilinear extension (PME), a policy-based relaxation whose continuous support matches feasible actions under categorical policies.We prove that the marginal gains of the stage utility provide an unbiased estimator of the gradient of the PME and that maximizing the PME over action distributions is equivalent to maximizing the stage utilities over agent actions, which are critical to devise principled policy gradient.Building on this, we design SubMAPL, a centralized-training decentralized-execution KL-mirror policy-learning method that uses local marginal gains as stochastic PME gradients during training. KL-mirror updates preserve categorical feasibility without Euclidean projection.In the case where agents run tabular-softmax policies, we introduce open policy migration and an open-system KL tracking variation to handle agent arrivals and departures. Using dynamic regret analysis, we establish a lower bound on the cumulative utility which accounts for the openness of the environment and for the gap between optimal stage-wise and global utilities. Simulations on multi-agent coverage demonstrate that SubMAPL outperforms policy-gradient and online-learning baselines.
This article addresses critical challenges in multiagent containment control (MACC) under adversarial disturbances and hard constraints by integrating zero-sum differential games, collision avoidance, energy efficiency, and formation preservation into a unified framework. First, a terminal-cloud collaborative control framework is built. For the cloud agent, a barrier-function-augmented reinforcement learning (RL) mechanism ensures collision-free path planning, and an integral RL algorithm via policy iteration solves the two-player zero-sum differential game for optimal control against adversarial disturbances. For the terminal agents, the event-driven distributed RL controllers optimize energy consumption while maintaining formation tracking under topology switching. Furthermore, the hierarchical architecture enables cross-layer algorithmic decoupling and event-driven adaptive RL. Rigorous theoretical analysis establishes the asymptotic stability of containment errors and uniform ultimate boundedness of actor-critic weight estimation errors, even under resource limitations and dynamic obstacle environments. Simulations of satellite swarms validate the framework's ability to navigate environments with dense obstacles, tolerate node failures, and adapt to intermittent connectivity.
This article addresses the problem of cooperative path following for wheeled mobile robots (WMRs) under system constraints and external bounded disturbances within a switching communication network, by proposing a robust distributed model predictive control (DMPC) strategy. First, the cooperative path-following task is decoupled into two subtasks using a modified virtual structure: a cooperative task involving virtual reference robots and an individual path-following task between each actual robot and its corresponding virtual reference. A time-like path parameter is introduced to generate predefined path information for the virtual reference robot in advance, enabling dynamic formation tracking. Subsequently, discrete-time error dynamics subject to external bounded disturbances are derived for each robot, and a centralized predictive control problem is formulated as a baseline. A nominal DMPC strategy is then developed for the disturbance-free case, followed by an extension to a robust DMPC formulation that accounts for nonzero disturbances. In this context, a stability constraint is incorporated to ensure closed-loop stability without relying on neighboring agents’ real-time information. Theoretical analysis confirms the feasibility of the proposed scheme and guarantees the convergence of system trajectories to a disturbance invariant set. Finally, simulation and experimental results validate the effectiveness of the proposed strategy in cooperative path-following scenarios involving WMRs.
With the rapid advancement of artificial intelligence, multi-agent systems (MASs) are evolving from classical paradigms toward architectures built upon large foundation models (LFMs). This survey provides a systematic review and comparative analysis of classical MASs (CMASs) and LFM-based MASs (LMASs). First, within a closed-loop coordination framework, CMASs are reviewed across four fundamental dimensions: perception, communication, decision-making, and control. Beyond this framework, LMASs integrate LFMs to lift collaboration from low-level state exchanges to semantic-level reasoning, enabling more flexible coordination and improved adaptability across diverse scenarios. Then, a comparative analysis is conducted to contrast CMASs and LMASs across architecture, operating mechanism, adaptability, and application. Finally, future perspectives on MASs are presented, summarizing open challenges and potential research opportunities.
3D editing—the task of locally modifying the geometry or appearance of a 3D asset—has wide applications in immersive content creation, digital entertainment, and AR/VR. However, unlike 2D editing, it remains challenging due to the need for cross-view consistency, structural fidelity, and fine-grained controllability. Existing approaches are often slow, prone to geometric distortions, or dependent on manual and accurate 3D masks that are error-prone and impractical. To address these challenges, we advance both the data and model fronts. On the data side, we introduce 3DEditVerse, the largest paired 3D editing benchmark to date, comprising 116,309 high-quality training pairs and 1,500 curated test pairs. Built through complementary pipelines of pose-driven geometric edits and foundation model-guided appearance edits, 3DEditVerse ensures edit locality, multi-view consistency, and semantic alignment. On the model side, we propose 3DEditFormer, a 3D-structure-preserving transformer. By enhancing image-to-3D generation with dual-guidance attention and time-adaptive gating, 3DEditFormer disentangles editable regions from preserved structure, enabling precise and consistent edits without requiring auxiliary 3D masks. Extensive experiments demonstrate that our framework outperforms state-of-the-art baselines both quantitatively and qualitatively, establishing a new standard for practical and scalable 3D editing. Dataset and code will be released. Project: https://anonymousresearch37.github.io/3DEditFormer/
In this article, we propose a planning-operation coordinated mitigation scheme for load redistribution (LR) attacks to overcome the deficiencies of separately designed phase shifting transformer-based mitigation strategies. Specifically, the interactions amongst the defender, attacker, and system are formulated as a trilevel optimization, where the deployment of defense devices and phase shift angles can be optimized according to possible operation state. Based on the proposed load similarity metric, a clustering-based approximate solution is designed to reduce the computational complexity caused by the integration of planning and operation stages. Simulation results on the IEEE 14-bus and 30-bus test systems verify the performance of the proposed mitigation scheme and the clustering-based approximate solution method.
Probabilistic time series forecasting plays a crucial role in supporting decision-making across various domains. While most existing methods focus on modeling the raw time series, modeling step-wise differences (Deltas) offers a complementary perspective, as Deltas tend to be more stationary and easier to learn. However, modeling the Deltas and reconstructing the original series through recursive accumulation may lead to error propagation, impairing prediction accuracy. In this work, we propose DeltaDiffusion, a diffusion-based framework that is built on modeling Deltas instead of raw series to better exploit their stationarity. To mitigate error accumulation, we leverage a pretrained point prediction model as an anchor for calibration. To enhance the modeling of Deltas, a variance-aware noise scheduling strategy is introduced to capture uncertainty. Furthermore, a cyclic denoising network that injects phase-aware priors into the attention layers is designed to extract periodic patterns in Deltas. Extensive experiments demonstrate the effectiveness of our method across various forecasting horizons and evaluation metrics.
This paper presents an interpretable reward design framework for reinforcement learning based constrained optimal control problems with state and terminal constraints. The problem is formalized within a standard partially observable Markov decision process framework. The reward function is constructed from four weighted components: a terminal constraint reward, a guidance reward, a penalty for state constraint violations, and a cost reduction incentive reward. A theoretically justified reward design is then presented, which establishes bounds on the weights of the components. This approach ensures that constraints are satisfied and objectives are optimized while mitigating numerical instability. Acknowledging the importance of prior knowledge in reward design, we sequentially solve two subproblems, using each solution to inform the reward design for the subsequent problem. Subsequently, we integrate reinforcement learning with curriculum learning, utilizing policies derived from simpler subproblems to assist in tackling more complex challenges, thereby facilitating convergence. The framework is evaluated against original and randomly weighted reward designs in a multi-agent particle environment. Experimental results demonstrate that the proposed approach significantly enhances satisfaction of terminal and state constraints and optimization of control cost.