In this paper, the self-learning robust safe control of continuous-time safety-critical nonlinear systems with exter nal disturbances is investigated via adaptive dynamic programming. In order to obtain the accurate disturbance information, a disturbance observer is developed. Subsequently, by integrating the Hamilton-Jacobi-Bellman equation with a robust control barrier function, a constrained optimization problem is formulated to reflect the relationship between the safety and optimality of safety-critical nonlinear systems. Moreover, based on the convex optimization theory and the ADP technique, a Lagrangian function is constructed to formulate the Karush-Kuhn-Tucker optimality condition leading to the derivation of the optimal safe control law. To approximately solve the constrained optimization problem, an actor-critic framework is established to learn the approximate optimal safe control law online. Furthermore, theoretical analysis shows that the proposed self-learning robust safe control approach guarantees the stability and safety of the safety-critical nonlinear systems with external disturbances. Finally, a single-link manipulator and a flexible manipulator are utilized to verify the validity of the present method.
In this paper, an event-triggered decentralized adaptive critic learning (ACL) control method is proposed for interconnected systems with nonlinear inequality state constraints. First, by introducing a slack function, the nonlinear inequality state constraints of original isolated subsystem are transformed into equality forms, and then the original isolated subsystem is augmented to an unconstrained one. Then, by establishing a cost function with discount factors for each isolated subsystem, a local policy iteration-based decentralized control law is developed by solving the Hamilton-Jacobi-Bellman equation with the help of a local critic neural network (NN) for each isolated subsystem. Through developing a novel event-triggering mechanism for each isolated subsystem, the decentralized control policy is updated at the triggering instants only, which assists to save the computational and communication resources. Hereafter, the event-triggered decentralized control law of isolated subsystem is derived. Then, the overall optimal control for the entire interconnected system is derived by constituting an array of developed event-triggered decentralized control laws. Furthermore, the closed-loop nonlinear interconnected system and the weight estimation errors of local critic NNs are guaranteed to be uniformly ultimately bounded. Finally, the effectiveness of the proposed method is validated through two comparative simulation examples.
In this paper, an optimal asynchronous control method via value iteration (VI) technique is proposed for Markov jump systems. Since the exact mode of the system is assumed to be unknown, a hidden Markov model (HMM) is constructed to describe the probability distribution of the system mode. By integrating the probability distribution into the design of cost function, the coupled algebraic Riccati equation (CARE) is established. Then, a VI algorithm is proposed to solve the CARE and obtain the optimal asynchronous control. Moreover, theoretical analysis is performed to confirm the convergence of the proposed VI scheme and the stochastic stability of the closed-loop system. By applying the proposed optimal asynchronous control to a power system with single-machine connected to a infinite bus model, the simulation results demonstrate its effectiveness regardless of mode transitions.
This paper proposes a dynamic event-triggered prescribed-time optimal attitude control method for two-degree-of-freedom helicopters by using adaptive dynamic programming (ADP). A transformed tracking error is introduced by combining the attitude tracking error with a time-varying scaling function that explicitly embeds the desired convergence time and accuracy. This transformation enables the prescribed-time control objective to be regraded as an optimal control problem for the transformed tracking error dynamics. Then, the optimal control law is obtained through the ADP framework. To mitigate communication and computational burdens, a dynamic event-triggered mechanism is designed to determine the updating instants of optimal control law. Theoretical analysis confirms that the closed-loop system achieves the prescribed-time stability while excluding Zeno behavior. Simulation studies illustrate the effectiveness of the proposed method.
Recent advances have shown that sequential fine-tuning (SeqFT) of pre-trained vision transformers (ViTs), followed by classifier refinement using approximate distributions of class features, can be an effective strategy for class-incremental learning (CIL). However, this approach is susceptible to distribution drift, caused by the sequential optimization of shared backbone parameters. This results in a mismatch between the distributions of the previously learned classes and that of the updated model, ultimately degrading the effectiveness of classifier performance over time. To address this issue, we introduce a latent space transition operator and propose Sequential Learning with Drift Compensation (SLDC). SLDC aims to align feature distributions across tasks to mitigate the impact of drift. First, we present a linear variant of SLDC, which learns a linear operator by solving a regularized least-squares problem that maps features before and after fine-tuning. Next, we extend this with a weakly nonlinear SLDC variant, which assumes that the ideal transition operator lies between purely linear and fully nonlinear transformations. This is implemented using learnable, weakly nonlinear mappings that balance flexibility and generalization. To further reduce representation drift, we apply knowledge distillation (KD) in both algorithmic variants. Extensive experiments on standard CIL benchmarks demonstrate that SLDC significantly improves the performance of SeqFT. Notably, by combining KD to address representation drift with SLDC to compensate distribution drift, SeqFT achieves performance comparable to joint training across all evaluated datasets.
In this article, a fixed-time deep reinforcement learning (DRL) formation hunting control problem is investigated for a multi-marine surface vehicle (MSV) system. First, considering the lack of dynamic adaptability caused by the conventional deep neural network (DNN) framework, an online adaptive DNN method is proposed for the high-dimensional multi-MSV system. Second, a novel DRL framework is developed for designing fixed-time formation hunting controllers, which integrates the online adaptive DNNs method with the actor-critic-based reinforcement learning (RL) algorithm. Finally, a nonsmooth fixed-time stability analysis is established for the nonsmooth closed-loop system induced by the DRL-based structure, which rigorously demonstrates that all signals converge within a fixed-time interval independent of initial states. The simulation example demonstrates the practical viability of the presented scheme.
In neural architecture search (NAS) methods based on latent space optimization (LSO), a deep generative model is trained to embed discrete neural architectures into a continuous latent space. In this case, different optimization algorithms that operate in the continuous space can be implemented to search neural architectures. However, the optimization of latent variables is challenging for gradient-based LSO since the mapping from the latent space to the architecture performance is generally non-convex. To tackle this problem, this paper develops a convexity regularized latent space optimization (CR-LSO) method, which aims to regularize the learning process of latent space in order to obtain a convex architecture performance mapping. Specifically, CR-LSO trains a graph variational autoencoder (G-VAE) to learn the continuous representations of discrete architectures. Simultaneously, the learning process of latent space is regularized by the guaranteed convexity of input convex neural networks (ICNNs). In this way, the G-VAE is forced to learn a convex mapping from the architecture representation to the architecture performance. Hereafter, the CR-LSO approximates the performance mapping using the ICNN and leverages the estimated gradient to optimize neural architecture representations. Experimental results on three popular NAS benchmarks show that CR-LSO achieves competitive evaluation results in terms of both computational complexity and architecture performance.
In this article, a dynamic feedback (DF)-based dynamic event-triggered (DET) control method for unknown nonaffine systems (UNSs) is developed by using reinforcement learning (RL). Through introducing a DF signal as a virtual control input, the UNS is augmented into a partially unknown affine system (PUAS). Subsequently, by designing a novel cost function that reflects the system states, and the actual and virtual control inputs, the DET optimal control (OC) problem of UNS is transformed into a DET OC problem of PUAS. To relax the requirement of PUAS dynamics, a neural network (NN)-based observer is established by using the measured system data. Moreover, a novel DET condition is established based on the static event-triggered (SET) rule, and the relationship of the triggering interval between SET and DET is revealed. In order to solve the DET Hamilton-Jacobi-Bellman equation (HJBE), a critic NN is constructed with the concurrent learning method to release the persistence of excitation (PE) condition. Furthermore, according to Lyapunov's direct method, the stability of the closed-loop system is guaranteed under the developed DF-based DET control strategy. Finally, simulation results of two examples demonstrate the effectiveness of the present DF-based DET method.
In this article, a fractional-order online policy iteration (FOOPI) algorithm-based approximate optimal control scheme is developed for fractional-order nonlinear systems (FONSs). Using the fractional Taylor series and the property of Hadamard product, the fractional Hamilton-Jacobi-Bellman (FHJB) equation corresponding to FONSs is formulated. Motivated by the procedure of solving the optimal control problem of integer-order nonlinear systems (IONSs), the design of the fractional-order (FO) controller for FONSs is transformed into solving the FHJB equation utilizing a FOOPI algorithm. By constructing a critic neural network (NN) and using the fractional calculus theory, the FO derivatives of the value function are established, followed by the approximate fractional Hamiltonian, which is obtained via the FOOPI algorithm. Hereafter, the FO control policy is obtained approximately to guarantee the stability of closed-loop FONSs via the FO Lyapunov's direct method. Simulation results guarantee the effectiveness of the proposed FOOPI-based approximate optimal control scheme.
This article investigates the prescribed-performance control (P2C) of variable-order fractional-order nonlinear systems (VOFONSs) through a control barrier function (CBF)-embedded adaptive dynamic programming (ADP) method. Owing to the time dependence and the global memory characteristics of the variable-order fractional calculus, the existing analysis methods for traditional integer-order systems cannot be directly applied. By using the variable-order fractional calculus and the integration-by-parts formula, the VOFONS is approximately transformed into a time-varying integer-order nonlinear system. To address the P2C problem, a prescribed-performance function is designed to reformulate the P2C problem as a constrained control problem of the time-varying integer-order nonlinear system. Unlike conventional function transformation-based methods, the proposed CBF-embedded ADP control method incorporates the safety constraints into the ADP-based optimal control framework without requiring explicit state transformation. Since the constructed discounted cost function is time-varying due to the involvement of time-varying CBF, it is approximated by a neural network with a time-varying activation function. Then, a time-varying policy iteration algorithm is developed to solve the Hamilton-Jacobi-Bellman (HJB) equation. Hereafter, the CBF-based control policy is derived to guarantee that the tracking error satisfies the predefined constraint. Simulation results verify the effectiveness of the present CBF-embedded ADP control scheme.
Class-incremental learning (CIL) with Vision Transformers (ViTs) faces a major computational bottleneck during the classifier reconstruction phase, where most existing methods rely on costly iterative stochastic gradient descent (SGD). We observe that analytic Regularized Gaussian Discriminant Analysis (RGDA) provides a Bayes-optimal alternative with accuracy comparable to SGD-based classifiers; however, its quadratic inference complexity limits its use in large-scale CIL scenarios. To overcome this, we propose Low-Rank Factorized RGDA (LR-RGDA), a scalable classifier that combines RGDA's expressivity with the efficiency of linear classifiers. By exploiting the low-rank structure of the covariance via the Woodbury matrix identity, LR-RGDA decomposes the discriminant function into a global affine term refined by a low-rank quadratic perturbation, reducing the inference complexity from 𝒪(Cd^2) to 𝒪(d^2 + Crd^2), where C is the class number, d the feature dimension, and r ≪ d the subspace rank. To mitigate representation drift caused by backbone updates, we further introduce Hopfield-based Distribution Compensator (HopDC), a training-free mechanism that uses modern continuous Hopfield Networks to recalibrate historical class statistics through associative memory dynamics on unlabeled anchors, accompanied by a theoretical bound on the estimation error. Extensive experiments on diverse CIL benchmarks demonstrate that our framework achieves state-of-the-art performance, providing a scalable solution for large-scale class-incremental learning with ViTs. Code: https://github.com/raoxuan98-hash/lr_rgda_hopdc.
The architecture encoding has achieved competitive performance in neural architecture search (NAS) due to its effective enhancement of downstream architecture search efficiency. Computation-aware Transformer-based encodings use dependency masks to capture architectural contexts, while their stochastic and irreversible nature may lead to the loss of or inaccurate information. To this end, we propose a latent spatial computation-aware Transformer-based encoding, which is more reasonable to efficiently search for the optimal neural architecture in a continuous latent space. Furthermore, a surrogate-assisted evolutionary algorithm is employed to accelerate the search process in the latent space. The experiments show that the proposed surrogate-assisted NAS with latent spatial computation-aware Transformer-based encoding achieves competitive performance on the NAS-benchmarks. Moreover, the architectures with an average test accuracy of 97.45% found by our NAS method from DARTS space on CIFAR-10 with approximately 0.02 GPU days demonstrate both the effectiveness and efficiency of the proposed method.
In this paper, the self-triggered approximate optimal control of nonlinear systems subject to unknown actuator saturation is developed via adaptive dynamic programming. To begin with, an optimal control law for the nominal system is obtained by adopting the critic-only architecture, where a singular value maximization-based current learning mechanism is employed to relax the persistence of excitation condition. Then, a feedforward neural network (NN) is constructed to reduce the influence of the unknown actuator saturation. To save the communication resource, the computational cost and hardware, a self-triggered mechanism is designed to predict the updating time instants of the critic NN and the overall control law. Furthermore, both the approximation error of the critic NN weights and the closed-loop system states are proven to be uniformly ultimately bounded through the Lyapunov stability analysis. Finally, simulation results are provided to verify the effectiveness of the proposed method.
This paper reviews the application of adaptive dynamic programming (ADP) in controlling constrained nonlinear systems, with an emphasis on integrating ADP to address optimization and constraint issues in complex systems. First, the conventional ADP algorithm for tackling the optimal control problems of unconstrained nonlinear systems is introduced. Next, the general solutions and recent advances of ADP for controlling nonlinear systems with various constraints, mainly including input constraints, state constraints, output constraints and cost constraints, are elaborated. Moreover, several typical real control applications for constrained nonlinear systems with respect to ADP are summarized, particularly in the fields of aerospace systems, robots, autonomous systems and energy systems. Finally, some possible future prospects are explored, and the conclusions of this paper are presented. Overall, ADP plays a significant role in guaranteeing system safety and optimizing performance with broad application prospects, and the comprehensive investigation demonstrates the tremendous potential of ADP for controlling constrained nonlinear systems in the current eras of artificial intelligence, control science and engineering, and systems science.
This paper develops a dynamic self-triggered prescribed-time optimal tracking control scheme for helicopters by employing adaptive dynamic programming (ADP). A transformed tracking error is constructed by integrating the tracking error with a time-varying auxiliary function that incorporates both the prescribed convergence time and accuracy. Subsequently, the prescribed-time control problem is reformulated as an optimal control problem for the dynamics of transformed tracking error, and the control performance is enhanced through an improved cost function. Then, the optimal control problem is solved under the ADP framework, and the prescribed-time optimal tracking control law is obtained. To reduce the communication and computational burdens, a novel dynamic self-triggered mechanism is proposed for the optimal control law on the basis of dynamic event-triggered control method. This mechanism determines the updating time instants based solely on the current system information, thereby the continuous state monitoring which is required by conventional dynamic event-triggered approaches, is avoided. Finally, theoretical analysis guarantees the prescribed-time stability of the tracking error, and avoids the Zeno behavior, while simulation results demonstrate the effectiveness of the proposed control method. Note to Practitioners-Helicopters play critical roles in military transportation and surveillance operations. For such practical applications, convergence time and control accuracy are two critical performance metrics, while the control cost remains a key consideration in the controller design. Moreover, although existing static and dynamic event-triggered control strategies can reduce communication and computational burdens, they typically require dedicated hardware to monitor the system information continuously, which increases the implementation cost and complexity. To overcome this limitation, we develop a dynamic self-triggered prescribed-time optimal tracking control scheme for helicopters by using adaptive dynamic programming. The developed method inherently eliminates the need for continuous state monitoring by designing a dynamic self-triggering mechanism, thereby reducing hardware dependence and simplifying control algorithm deployment. Overall, this method ensures prescribed-time convergence with a pre-assigned accuracy and optimizes a predefined cost function, which make it particularly suitable for time-sensitive missions while reducing the control cost.
This article investigates the secure containment control problem for switched nonlinear stochastic multi-agent systems (MASs) by using the mode-dependent average dwell time (MDADT) method. As a class of cyber-physical systems, MASs are vulnerable to denial of service (DoS) attacks, which can greatly degrade the control performance or even influence the stability of controlled systems. By combining the backstepping method, a secure containment control method is developed with the event-triggering mechanism. Based on the connectivity information between communication links, this method eliminates the effects of DoS attacks. In the backstepping recursive design process, the problem of “explosion of complexity” caused by the derivation of virtual control is addressed by introducing a command filter. Additionally, the event-triggering mechanism is designed in the backstepping process to effectively save the communication resources. Then, the event-triggered secure containment control policy is derived to ensure the stability of single follower and MASs under synchronous and asynchronous switchings. The followers led by multiple leaders eventually enter and keep moving in the convex hull composed of leaders. A simulation example demonstrates the effectiveness of the present secure containment control scheme.
Nonconvex Activated Fuzzy Zeroing Neural Network-based (NAFZNN) and Nonconvex Activated Fuzzy Noise-Tolerant Zeroing Neural Network-based (NAFNTZNN) models are devised and analyzed, drawing inspiration from the classical ZNN/NTZNN-based model for online addressing Time-Varying Quadratic Programming Problems (TVQPPs) with Equality and Inequality Constraints (EICs) in noisy circumstances, respectively. Furthermore, the proposed NAFZNN model and NAFNTZNN model are considered as general proportion-differentiation controller, along with general proportion-integration-differentiation controller. Besides, theoretical results demonstrate the global convergence of both the NAFZNN and NAFNTZNN models for TVQPPs with EIC under noisy conditions. Moreover, numerical results illustrate the efficiency, robustness, and ascendancy of the NAFZNN and NAFZNN models in addressing TVQPPs online, exhibiting inherent noise tolerance. Ultimately, an application example to plant leaf disease identification is conducted to support the feasibility and efficacy of the designed NAFNTZNN model, which shows its potential practical value in the field of image recognition.
This paper investigates the distributed fault-tolerant consensus problem for nonlinear multi-agent systems (MASs). A novel distributed fault-tolerant control protocol is proposed under a zero-sum differential game framework, where the consensus problem is reformulated as a minimax optimization between the control inputs of agents and the actuator faults through a local cost function. A critic neural network is trained online to solve the coupled Hamilton-Jacobi-Isaacs (HJI) equation, where the optimal control and upper bound of fault compensation are simultaneously derived from the Nash equilibrium condition. Leveraging the Lyapunov stability theorem, it is rigorously proved that the designed distributed fault-tolerant consensus control law guarantees the uniform ultimate boundedness (UUB) of the closed-loop systems. Simulation results validate the effectiveness of the present method.
Filter pruning effectively compresses the neural network by reducing both its parameters and computational cost. Existing pruning methods typically rely on pre-designed pruning criteria to measure filter importance and remove those deemed unimportant. However, different layers of the neural network exhibit varying filter distributions, making it inappropriate to implement the same pruning criterion for all layers. Additionally, some approaches apply different criteria from the set of pre-defined pruning rules for different layers, but the limited space leads to the difficulty of covering all layers. If criteria for all layers are manually designed, it is costly and difficult to generalize to other networks. To solve this problem, we present a novel neural network pruning method based on the Criterion Learner and Attention Distillation (CLAD). Specifically, CLAD develops a differentiable criterion learner, which is integrated into each layer of the network. The learner can automatically learn the appropriate pruning criterion according to the filter parameters of each layer, thus the requirement of manual design is eliminated. Furthermore, the criterion learner is trained end-to-end by the gradient optimization algorithm to achieve efficient pruning. In addition, attention distillation, which fully utilizes the knowledge of unpruned networks to guide the optimization of the learner and improve the pruned network performance, is introduced in the process of learner optimization. Experiments conducted on various datasets and networks demonstrate the effectiveness of the proposed method. Notably, CLAD reduces the FLOPs of ResNet-110 by about 53% on the CIFAR-10 dataset, while simultaneously improves the network’s accuracy by 0.05%. Moreover, it reduces the FLOPs of ResNet-50 by about 46% on the ImageNet-1K dataset, and maintains a top-1 accuracy of 75.45%.
A novel control design problem for a class of non-strict feedback multi-agent systems (MAS) in discrete-time form is studied based on reinforcement learning (RL) and applied to multi-marine vehicles (MMV). Firstly, for this kind of discrete-time MAS, a novel system transformation, which can not only solve the noncausal problem that exists in the backstepping method but also reduce the computational complexity, is proposed. Secondly, the algebraic-loop problem inherent in the conventional controller design is solved by compensating the dynamics and using the property of NN. Thirdly, the multi-gradient recursive (MGR) RL scheme is developed for the sake of designing the optimal controller. Finally, the stability analysis is presented, and all signals are ensured to be semi-global uniformly ultimately bounded (SGUUB) in the Lyapunov's sense. Besides, this scheme is applied to the MMV which can be described the design in the non-strict feedback form to extend the application of the designed controller. The MMV simulation demonstrates the validation of this scheme.