To achieve trajectory tracking control of a robotic manipulator system, it is crucial to establish precise kinematic and dynamic models. Considering the complex characteristics of robotic manipulator, a dynamic modeling method based on virtual decomposition control is proposed. Considering the interaction force between the end effector of the robotic manipulator and the environment, an adaptive controller-nonlinear extended state observer combination control is designed through virtual decomposition control theory. This method virtually decomposes the complex robotic manipulator system into several subsystems, achieving adaptive controller-nonlinear extended state observer combination control for each subsystem. By introducing the concept of virtual power flow into the controller and designing Lyapunov functions based on nonlinear extended state observer error dynamics, the stability of this redundant system is proven.
This article presents an adaptive fuzzy tracking control protocol for a category of uncertain nonlinear multiagent systems (MASs) subject to deception attacks and multiple constraints (time-varying asymmetry constraints on tracking error and system states). Fuzzy logic systems are employed to approximate the unknown nonlinear dynamics in MASs. Since deception attacks over the sensor network make the actual states of the MASs unavailable, this article employs the compromised states for feedback control. Based on the relationship between the compromised states and actual states, the problem of satisfying the original state constraints boils down to the new constraints on compromised states. The use of a nonlinear mapping together with the dynamic surface control strategy can ensure that the multiple constraints are solved with the constraint boundaries of system output being freely selected by the user, the feasibility requirement on the virtual controller is removed, and the computational explosion is mitigated. Under this control protocol, the tracking performance of MASs subject to deception attacks is achieved, while all constraints are not violated. Ultimately, simulation results verify the effectiveness and advantage of the developed control method.
Parameter sharing is a central design choice in cooperative multi-agent reinforcement learning, yet it fundamentally conflicts with the need for role specialization in heterogeneous cooperative environments. Existing role-based methods typically learn monolithic role representations, which often suffer from gradient interference and fail to capture the compositional structure of complex behaviors. Inspired by Trait Theory, we propose DEcompose and COnstruct Roles (DECOR), a framework that models agent roles as dynamic compositions of orthogonal behavioral traits. DECOR introduces an orthogonal Mixture-of-Experts architecture to decompose behaviors into independent traits, mitigating destructive gradient interference under parameter sharing, and a group-consensus guided mechanism to extract team-level tactical intents that guide role composition.Experiments on multiple benchmarks demonstrate that DECOR consistently improves sample efficiency and overall performance over existing related methods.
The optimal control method is a valuable technique for improving the safety and reliability of the aeration process in the wastewater treatment process (WWTP). However, due to the complex and time-varying characteristics in WWTP, it is difficult to design the adaptive optimal controller to guarantee the system stability and reduce operating energy consumption. Therefore, a neural network approximation-based adaptive fuzzy dynamic optimal tracking control method is designed to real-time control the dissolved oxygen concentration (DOCN). First, due to the time-varying nonlinearity of WWTP, an inverse fuzzy dynamic MOEA/D algorithm is designed to acquire the optimal setpoints of DOCN. Second, an actor network is utilized to identify uncertain dynamic situations, and a critic network is utilized to minimize the utility function, which can improve the identification accuracy. Meanwhile, an adaptive auxiliary signal, specifically calibrated to dead-zone parameters, is constructed to counteract the influence of an asymmetric dead-zone on control performance. Finally, the adaptive optimal controller is used to achieve precise tracking control of DOCN in a WWTP. The stability of the control system is proved, and the effectiveness is evaluated in benchmark simulation model I.
Offline reinforcement learning enables policy learning solely from fixed datasets, without costly or risky environment interactions, making it highly valuable for real-world applications. While Transformer-based approaches have recently demonstrated strong sequence modeling capabilities, they typically learn from complete trajectories conditioned on final returns. To mitigate this limitation, we propose the Peak-Return Greedy Slicing (PRGS) framework, which explicitly partitions trajectories at the timestep level and emphasizes high-quality subtrajectories. PRGS first leverages an MMD-based return estimator to characterize the distribution of future returns for state-action pairs, yielding optimistic return estimates. It then performs greedy slicing to extract high-quality subtrajectories for training. During evaluation, an adaptive history truncation mechanism is introduced to align the inference process with the training procedure. Extensive experiments across multiple benchmark datasets indicate that PRGS significantly improves the performance of Transformer-based offline reinforcement learning methods by effectively enhancing their ability to exploit and recombine valuable subtrajectories.
Value decomposition (VD) methods have achieved remarkable success in cooperative multi-agent reinforcement learning (MARL). However, their reliance on the max operator for temporal-difference (TD) target calculation leads to systematic Q-value overestimation. This issue is particularly severe in MARL due to the combinatorial explosion of the joint action space, which often results in unstable learning and suboptimal policies. To address this problem, we propose QSIM, a similarity weighted Q-learning framework that reconstructs the TD target using action similarity. Instead of using the greedy joint action directly, QSIM forms a similarity weighted expectation over a structured near-greedy joint action space. This formulation allows the target to integrate Q-values from diverse yet behaviorally related actions while assigning greater influence to those that are more similar to the greedy choice. By smoothing the target with structurally relevant alternatives, QSIM effectively mitigates overestimation and improves learning stability. Extensive experiments demonstrate that QSIM can be seamlessly integrated with various VD methods, consistently yielding superior performance and stability compared to the original algorithms. Furthermore, empirical analysis confirms that QSIM significantly mitigates the systematic value overestimation in MARL. Code is available at https://github.com/MaoMaoLYJ/pymarl-qsim.
For nonlinear state-dependent constrained multi-input-multi-output (MIMO) systems subject to input saturation, a novel reinforcement learning-based adaptive neural network (NN) optimal control strategy is presented in this work. By reformulating the system dynamics via nonlinear state-dependent mappings (NSDMs), which not only avoids the dependence on predefined assumptions of virtual controller boundedness, but also eliminates the hypothesis of unknown control gain functions. Under the unified backstepping technology framework, an adaptive feed-forward controller is developed to eliminate the nonsmooth factors arising from input saturation constraints. Subsequently, the solution of the Hamilton-Jacobi-Bellman (HJB) is approximated by NN. By minimizing the discounted performance function (DPF), the optimal regulation control of the nonlinear MIMO systems is realized. Finally, a simulation verification on the cascade continuous stirred tank reactor (C-CSTR) systems is implemented to confirm the stability and robustness of the developed control strategy.
Coal gasification technology plays a pivotal role in chemical production as a key process for efficiently converting coal into liquid fuels and chemical feedstocks. During gasification, high-temperature reactions generate syngas, and optimizing its operational parameters is essential for improving syngas quality, carbon efficiency and liquid fuel yield. However, the intricate chemical reactions and heat transfer mechanisms in gasification necessitate costly simulations or experimental testing, making it an expensive multi-objective optimization problem. To address this challenge, this paper proposes a Knee Point-guided Heterogeneous Surrogate-assisted Evolutionary Algorithm (KG-HSEA) that integrates Kriging and Feedforward Neural Networks (FNN) to construct a heterogeneous surrogate model, leveraging their complementary strengths to reduce computational costs while maintaining predictive accuracy. By incorporating a knee point-guided search mechanism, the method prioritizes solutions that embody critical trade-offs among conflicting objectives. Moreover, an adaptive sampling strategy combined with dual-archive management is employed to dynamically update the surrogate model, ensuring it adapts to unstable operating conditions while maintaining robust convergence-diversity balance in coal gasification processes. Experimental results show that KG-HSEA achieved a 71.9% superiority rate with 23 optimal solutions out of 32 benchmark problems, highlighting its potential for efficient and feasible coal gasification optimization.
The wastewater treatment process (WWTP) operates under diverse conditions, and precise control of dissolved oxygen (DO) and nitrate nitrogen (NO3-N) concentrations is crucial for satisfying water quality standards. This paper proposes an adaptive fuzzy neural network (FNN) constrained control method based on an event-triggered mechanism. Firstly, the unknown nonlinear dynamic functions encountered in WWTP are effectively approximated using the FNN's strong adaptive capability. Secondly, an event-triggered control (ETC) strategy is introduced to reduce the communication burden, with trigger conditions designed based on the control signal error. Subsequently, a time-varying asymmetric barrier Lyapunov function (BLF) is used to construct controllers for DO and NO3-N concentrations, ensuring variables remain within time-varying constraint ranges. Meanwhile, a fault-tolerant control (FTC) approach is introduced to cope with potential actuator failures during the WWTP. Finally, simulations using Benchmark Simulation Model 1 (BSM1) are performed to verify the effectiveness of the proposed method. Note to Practitioners-The wastewater treatment process (WWTP) is a typical complex industrial process characterized by strong non-linearities, large time delays, frequent disturbances, and multiple operating conditions. The inevitability of faults in aeration and reflux equipment, alongside stringent demands for resource efficiency, places heightened requirements on system stability and control precision. Considering communication resource constraints and equipment fault risks, the core control objective of this paper is to design a multivariable event-triggered constrained control method for WWTP under varying operating conditions, which aims to enhance operational performance while ensuring stable effluent quality that meets standards. Specifically, a mathematical model applicable to the WWTP is established, and a fuzzy neural network (FNN) is used in the controller to approximate unknown system dynamics and disturbances. Furthermore, a time-varying asymmetric barrier Lyapunov function (BLF) is introduced to mitigate process risks arising from excessive dissolved oxygen (DO) and nitrate nitrogen (NO3-N) concentrations, like microbial community instability and non-compliant effluent quality. However, with prolonged operation, critical equipment such as aeration blowers and internal recirculation pumps experience accelerated mechanical wear, including the performance degradation of fan blades and pump bearings. This paper employs an event-triggered mechanism to reduce high-frequency updates of control commands, enabling smoother operation of aeration and reflux equipment. This minimizes unnecessary operations, thereby lowering operation and maintenance costs and equipment failure risks. Finally, the effectiveness of the proposed constrained strategy is evaluated on a simulation platform for WWTP, and the results demonstrate a significant improvement in its operational performance.
To enhance the control effectiveness and operational efficiency of the wastewater treatment process (WWTP), this article proposes a multivariable optimal control scheme based on an identifier-critic-actor reinforcement learning (RL) framework with an event-triggered mechanism (ETM) for dissolved oxygen (DO) and nitrate nitrogen (NO) concentrations. First, the first fuzzy neural network (FNN) is used to estimate the unknown dynamics in WWTP, and the second FNN is implemented within the critic-actor optimization framework. Second, the RL algorithm is applied to design optimal controllers for DO and NO concentrations by constructing a tangent barrier Lyapunov function. Moreover, a dynamic ETM based on the adaptive threshold strategy is proposed to balance the control performance and energy consumption in the wastewater system. Finally, stability analysis and the benchmark simulation model no. 1 are conducted to verify that the control scheme proposed in this article demonstrates effectiveness and enhanced performance.
In this article, consider the fault-tolerant control method of nonlinear heterogeneous multi-agent systems (NHMASs) with state constraints and actuator faults, and propose a state-triggered control method based on adaptive fuzzy strategy. For the unknown nonlinear dynamics in the NHMASs, the fuzzy logic systems (FLSs) are used for effective approximation, and the multiple faults is solved by introducing a damping term in the middle control laws. Considering the computational burden that may be caused by frequent information interactions in NHMASs, we design a state-triggered mechanism (STM) to update the control inputs in a nonperiodic manner, which significantly reduces the consumption of communication and computational resources. In addition, by constructing the nonlinear transformation function (NTF), the complex state-constrained problem is transformed into an easy-to-handle bounded problem, which guarantees the feasibility and stability of the NHMASs. Finally, simulation results prove the feasibility of the proposed control method.
In this paper, we present a controller approach to solve the problem of bearing-constrained autonomous aerial vehicle (AAV) reaching a target with predefined accuracy, also avoiding multiple obstacles and coping with possible actuator failure in the process. First, to ensure that AAV with weak perception capabilities still maintains high performance, we propose a bearing-based AAV localization model to accurately locate targets and obstacles while only obtaining angular information, which effectively solves the problem of maintaining high performance of AAV with limited loading capacity and cost. Second, an obstacle avoidance assistance system is designed in combination with the localization model, so that the AAV can avoid obstacles perfectly even with local perception. Third, we also employ an asymmetric time-varying Lyapunov barrier function (BLF) to enable the AAV to reach the target with a predefined accuracy, and use neural networks (NNs) to approximate unknown functions in the system while being able to cope with actuator failures. Finally, simulation experiments verify that the AAV can perfectly achieve the control goal in various environments.
This paper studies the problem of safe trajectory tracking control for a virtual decomposition manipulator system with joint disturbances and joint position constraints. For the joint and link subsystems based on virtual decomposition control (VDC), a robust tracking control strategy integrating parameter adaptive and saturation function sliding mode terms is proposed to enhance the system’s ability to suppress external disturbances. Secondly, stability analyses were conducted for the subsystems and the entire machine system, respectively. Based on this, in combination with the core idea of virtual decomposition, input-to-state safety high-order control barrier functions (ISSf-HOCBF) were designed for each subsystem with different joint position constraints. This approach ensures that the safety sets of each subsystem after disturbance expansion remain forward invariant under the conditions of external disturbances and differentiated joint state space constraints. In addition, a quadratic programming (QP) framework for optimizing control input was constructed, with system safety as the primary constraint. Finally, the effectiveness of the proposed safety-critical control strategy was verified through a simulation experiment of a three-degree-of-freedom robotic manipulator based on a virtual decomposition model.
In high-dimensional time-series analysis, it is essential to have a set of key factors (namely, the style factors) that explain the change of the observed variable. For example, volatility modeling in finance relies on a set of risk factors, and climate change studies in climatology rely on a set of causal factors. The ideal low-dimensional style factors should balance significance (with high explanatory power) and stability (consistent, no significant fluctuations). However, previous supervised and unsupervised feature extraction methods can hardly address the tradeoff. In this paper, we propose Style Miner, a reinforcement learning method to generate style factors. We first formulate the problem as a Constrained Markov Decision Process with explanatory power as the return and stability as the constraint. Then, we design fine-grained immediate rewards and costs and use a Lagrangian heuristic to balance them adaptively. Experiments on real-world financial data sets show that Style Miner outperforms existing learning-based methods by a large margin and achieves a relatively 10% gain in R-squared explanatory power compared to the industry-renowned factors proposed by human experts.
In a wastewater treatment process (WWTP), which covers a variety of physical, chemical, and biological treatment steps, it is a challenging control task to maintain proper dissolved oxygen (DO) and nitrate nitrogen (NO) concentrations to meet effluent standards. This paper proposes an event-triggered adaptive control method using a self-organizing fuzzy neural network (ETSOFNN) to regulate DO and NO concentrations. Moreover, a self-organizing fuzzy neural network (SOFNN) based on the maximum correlation entropy-induced criterion identifies and approximates the nonlinear function, which further enables the dynamic adjustment of the controller structure, including the addition or deletion of parameters. In addition, a dynamic event-triggered mechanism with a relative threshold strategy is introduced into the controller, and the trigger conditions are designed according to the tracking error. Utilizing Lyapunov stability theory, the stability of the control system is demonstrated. Finally, simulations based on the benchmark simulation model 1 (BSM1) platform are conducted to verify the effectiveness of the ETSOFNN method.
Recently, researchers have elevated the performance of algorithms in multi-agent reinforcement learning (MARL) to new heights by leveraging sequence models and Transformer architecture. However, introducing Transformer architecture into MARL can naturally lead to centralized strategies that receive information from all agents and issue instructions uniformly. Since a centralized controller is impractical in many real-world tasks, we introduce an online distillation architecture, OLEN, to extend the excellent performance to fully decentralized scenarios. Our approach is applicable to any centralized multi-agent reinforcement learning method. We use online distillation simultaneously with the policy improvement process to reduce interaction costs, enhance data efficiency, and facilitate student-aware distillation. The utilization of hypernetwork enhances both conciseness and scalability in our approach. Additionally, we incorporate a certainty metric and a dynamic coefficient to facilitate meaningful learning. To the best of our knowledge, this work is the first to introduce online distillation into MARL problems. Experiments demonstrate the desirable performance of our method.
This paper introduces an event-triggered saturation-tolerant prescribed control (STPC) framework for rigid spacecraft subject to actuator faults and actuator saturation via a fixed-time disturbance observer (FTDO). An FTDO with time-varying observer gains is initially developed to reconstruct the lumped perturbations caused by external disturbances, parameter uncertainties, and actuator faults. A fixed-time auxiliary system is employed to counter the adverse effects of actuator saturation. Additionally, asymmetric prescribed performance and shift functions are skillfully incorporated to handle arbitrary bounded initial conditions. Subsequently, a novel FTDO-based event-triggered STPC strategy is formulated, ensuring that attitude-tracking errors converge to predefined performance bounds within a predetermined time while minimizing unnecessary control signal updates. The practical fixed-time stability of all closed-loop signals is validated, with the strict avoidance of Zeno behavior. Finally, simulation studies are conducted to verify the accuracy and effectiveness of the proposed approach.
To obtain the effective purification performance in wastewater treatment process (WWTP), the optimal control is an important method to guarantee the effluent quality reaching the standard and improve the treatment efficiency. The concentrations of dissolved oxygen (DO) and nitrate nitrogen (NO3-N) are primary metrics that impact effluent quality, which is needed to be stably tracking controlled for achieving optimal performance in WWTP. Therefore, the double-layer fuzzy neural network (FNN)-based optimal control method with multivariable is proposed. First, considering the dynamic characteristic of WWTP, the FNN-based actor network is exploited to approximate the unknown dynamic information. Subsequently, the FNN-based critic network is integrated to minimize the cost function of DO and NO3-N concentrations, which is composed of the control error and the control variable. Then, to guarantee the stability of the optimal controller, the Lyapunov function is constructed through backstepping method to analyze the control system performance. Finally, the optimality and effectiveness of the control system with multivariable are verified via the simulation experiments in benchmark simulation model 1.
Mixed Integer Linear Programming (MILP) is a type of NP-hard problem that is widely used in industries but challenging to solve. Recently, there has been increasing interest in using Deep Learning (DL) to accelerate solving MILPs, particularly through learning to branch. However, the validity of DL models is based on the assumption that the test set and the training set are drawn from the same distribution, while in MILPs, the distribution shift often appears. To improve the generalization of the branching policy, we introduce the domain generalization method and propose an Adaptive Policy Switching (APS) framework. The experiments demonstrate that APS outperforms the previous state-of-the-art (SOTA) method, reducing the number of nodes by 18
For cooperative multi-agent reinforcement learning, various methods have been proposed to enhance the collaborative strategy capabilities of agents. However, when agents make decisions, humans have no knowledge of their subsequent decision-making intentions or sub-goals. This lack of understanding hinders human comprehension of agent strategies and further research on agents. Currently, there are limited relevant studies. To address this problem, we propose a novel framework which can generate the decision intention of agents. We first formalize this problem and use states crucial to the task to express the decision intentions of agents. Then, we introduce the polarization index to measure the importance of states and select them for training. Finally, we learn the decision intentions through a diffusion model with rapid generation capability and generate them during the decision-making process. This study sheds light on the problem of agent decision intention and enhances the transparency of agent strategies, facilitating deeper research on agents. The experimental results demonstrate the effectiveness of our approach.