To address the challenge of coordinated optimization between energy consumption and thermal comfort in electric vehicle thermal management systems under low-temperature conditions, this article proposes an intelligent control framework integrating a channel-temporal attention mechanism with knowledge distillation. A channel-temporal attention-enhanced dual-delay deep deterministic policy gradient (DDPG) algorithm [twin delayed deep deterministic policy gradient with channel and temporal attention (CT-TD3)] is developed to enable dynamic feature selection and long-term temporal dependency modeling of thermal states across coupled subsystems-including the battery, motor, and cabin-thereby significantly improving algorithmic convergence speed and enhancing the foresight of control decisions. Furthermore, response-layer knowledge distillation is employed to construct a lightweight thermal comfort prediction network, reducing computational latency to 0.514 ms while preserving high prediction accuracy, thus enabling real-time optimization under comfort constraints. Simulation results demonstrate that, at a low ambient temperature of -10 degrees C, the proposed method reduces total system energy consumption by 5.27% compared to the baseline strategy while increasing the proportion of occupant thermal comfort ratio to 75.03%.
Battery Electric Vehicles (BEVs) experience significant range degradation and reduced thermal comfort in low-temperature environments, underscoring the importance of energy-efficient Heat Pump Air Conditioning (HPAC) systems as a core thermal management technology. However, existing Thermal Management Strategies (TMSs) struggle with HPAC's strong coupling and nonlinearity, and rarely consider joint health optimization of batteries and motors. To address these gaps, this study proposes a Deep Reinforcement Learning (DRL)-based TMS integrating component health awareness. This study proposes an optimization framework considering both battery State of Health (SOH) and motor aging, and improves the original Proximal Policy Optimization (PPO) by embedding the Beta Distribution (BD) to enhance its adaptability to bounded continuous action spaces. Simulations and Hardware-in-the-Loop (HIL) tests validate the strategy: compared with the original PPO-based TMS, the compressor energy consumption is reduced by approximately 11.26%. Relative to existing TMSs, battery energy consumption and health degradation show respective reductions of 4.20% and 6.90%, while motor health degradation is reduced by 5.64%. The proposed TMS achieves stable temperature control of battery, motor, and cabin under different initial temperatures, demonstrating superior energy efficiency, robustness, and engineering deployability.
In recent years, unmanned ground vehicles (UGVs) have advanced rapidly, attracting significant attention for their applications in modern military operations, particularly as target vehicles. Their ability to perform realistic combat maneuvers relies heavily on formation control, a key technology in this domain. This paper presents a novel formation control framework aimed at improving the accuracy and stability of dynamic formations for such vehicles. Building upon an enhanced leader-follower structure, the proposed method generates virtual dynamic targets for tracking by each follower vehicle. Based on the kinematic model of skid-steering vehicles, a Linear Quadratic Regulator (LQR) controller is then designed for precise tracking of these virtual targets during formation maneuvers. To further optimize the LQR controller's performance, the Proximal Policy Optimization (PPO) algorithm is incorporated into an asynchronous training framework inspired by the Asynchronous Advantage Actor-Critic (A3C) method. The PPO networks are trained in a simulation environment to enable real-time prediction of optimal LQR parameters during formation control. Compared with end-to-end deep reinforcement learning (DRL) methods and other related approaches, the combination of the PPO network and LQR controller results in a compact network size suitable for deployment on embedded systems. The effectiveness of the proposed methodology is validated through comparative experiments with alternative controllers in both simulation and real-vehicle tests. This approach is particularly well suited for low-cost UGVs with limited computational resources and for applications demanding dynamic motion control with high tracking accuracy and stability.
In high-temperature environments, thermal management systems are essential to improving comfort, ensuring the safe operation of electric vehicles, and reducing energy consumption. This study introduces a multi-agent control framework that leverages the Deep Deterministic Policy Gradient (MADDPG) algorithm to optimize temperature regulation and energy efficiency in an integrated thermal management system (ITMS). In recognition of the variability of air conditioning (AC) system energy efficiency under different vehicle operating conditions, the coordination agent dynamically plans target temperature trajectories for the cabin and battery, whereas the execution agent follows these trajectories and regulates motor system temperatures, ensuring coordinated energy optimization. To enhance the robustness and mitigate the risks of local optima caused by hyperparameter sensitivity, the Cross-Entropy Method (CEM) is integrated into MADDPG, forming the CEM-MADDPG control strategy. This hybrid method combines the global search capability of evolutionary algorithms with the sample efficiency of deep reinforcement learning, achieving a balance between stability and efficiency. Experimental results show that the proposed method effectively controls the temperatures of the battery, cabin, and motor, reducing the overall energy consumption of the ITMS by 19.8% and 7.2% respectively, compared to rule-based and model predictive control (MPC) methods.
Coordinating a platoon of connected and automated vehicles significantly improves traffic efficiency and safety. Current platoon control methods prioritize consistency and convergence performance but overlook the inherent interdependence between the platoon and the the non‐connected leading vehicle. This oversight constrains the platoon's adaptability in car‐following scenarios, resulting in suboptimal optimization performance. To address this issue, this paper proposed a platoon control framework based on multi‐agent reinforcement learning, aiming to integrate cooperative optimization with platoon tracking behavior and internal coordination strategies. This strategy employs a bidirectional cooperative optimization mechanism to effectively decouple the platoon's tracking behavior from its internal coordination control, and then recouple it in a multi‐objective optimized manner. Additionally, it leverages long short‐term memory networks to accurately capture and manage the platoon's dynamic nature over time, aiming to achieve enhanced optimization outcomes. The simulation results demonstrate that the proposed method effectively improves the platoon's cooperative effect and car‐following adaptability. Compared to the consensus control strategy, it reduces the average spacing error by 8.3%. Furthermore, the average length of the platoon decreases by 19.1%.
Ecological driving (eco-driving) is a crucial technique for electric vehicles to reduce energy consumption while maintaining driving quality. This paper proposes an improved PPO-based eco-driving strategy (EDS) for a dualmotor electric vehicle (DMEV) with multiple modes in three-lane highway scenarios. The technique of weighted experience replay (WER) is employed to prioritize data of collision and boundary-violation (CB) to enhance safety awareness, and a mechanism of value function decoupling (VFD) is proposed to improve the accuracy of estimation of state value function. Simulation results demonstrate that the improved PPO method enables the agent to learn a strategy achieving superior performance in driving safety, stability, ride comfort, and energy efficiency, and the performance markedly surpasses those of the baseline methods. The EMS in the proposed model can achieve 98.85 % of energy economy of that of dynamic programming (DP) in the tested traffic flows, and the model has high adaptability to flows with different densities and different levels of battery SoC.
In the pursuit of decarbonization within the transportation sector, battery electric vehicles (BEVs) have emerged as a pivotal solution due to their zero-emission characteristics. However, the thermal management systems of BEVs have become a critical bottleneck, particularly under extreme temperatures, where dual energy loads from battery cooling and cabin air-conditioning reduce driving range by up to 37%. To address this issue, this study proposes a soft actorcritic (SAC) based integrated thermal management strategy (TMS). By integrating the stochastic strategy optimization of SAC with thermodynamic principles, the proposed TMS efficiently manages the synergistic operation of battery-motorcabin loops. To enhance the global search efficiency of SAC, the cross entropy method (CEM) is introduced, thereby avoiding local optima. Simulation results demonstrate that proposed evolutionary SAC-based TMS outperforms conventional rule-based and other DRL-based TMSs, achieving a 22.78% reduction in energy consumption and superior temperature control of the battery, motor, and cabin.
Optical interference phase measurement is a crucial technology for measuring the edge height of segments during the co-phased adjustment stage of giant astronomical telescopes equipped with segmented primary mirrors. For the Chinese Giant Solar Telescope (CGST), achieving optical interferometric measurements with a range of 10 µm or more is a critical challenge that must be addressed to integrate the the co-focus and phasing adjustment processes. Given the unique requirements of solar observation, CGST intends to implement multi-wavelength technology to tackle the measurement range issue. However, this multi-wavelength measurement approach encounters the problem of edge jumps, and merely extending the exposure time does not effectively resolve this issue, which could compromise the telescope's diffraction-limited observational capabilities. The study indicates that the relative measurement error between two wavelengths, caused by atmospheric turbulence, is the primary factor leading to edge jumps. To address this issue, the paper proposes a dual-wavelength synchronous measurement technique. An experiment conducted on a segmented-mirror system demonstrates that, under turbulent conditions and with an exposure time of one second, the probability of edge jumps is negligible. By employing dual-wavelength synchronous technology, each measurement and adjustment takes only a few seconds, allowing the co-phased adjustment of CGST to be completed in just two to three rounds of measurement and adjustment.
The integrated thermal management system (TMS) expends a considerable quantity of energy in hightemperature environments to maintain the battery and motor systems within a reasonable temperature range, while also affecting the overall vehicle's safety and stable operation. In light of these considerations, a hierarchical reinforcement learning approach for integrated TMS is proposed. Given the low-frequency operation of the thermal system, the study employs the Deep Deterministic Policy Gradient (DDPG) with experience replay technical (DDPG-E) to investigate the battery-motor TMS control strategy under car-following behavior. This approach transforms the trade-off between high-frequency planning and low-frequency execution into a Markov Decision Process (MDP), thereby enabling coordinated optimization of multiple objectives. Specifically, in the pre-optimization layer, based on DDPG-E the heat generation optimization (HGO-DDPG-E) is introduced as a multi-objective optimization criterion to achieve "active load reduction" for the battery-motor system. Subsequently, the battery-motor temperature difference and energy consumption of TMS ancillary components are employed as constraints at the integrated control layer for all TMS components, based on the pre-optimization results. The results of the simulation demonstrate that the proposed method achieves an optimization of 15.3% in heat generation and a 14.1% reduction in TMS energy consumption.
Trajectory tracking is crucial in vehicle control, as it ensures stable driving along a predefined path. This paper proposes a deep reinforcement learning (DRL)-tuning hierarchical trajectory tracking framework, aiming to improve the tracking accuracy of traditional kinematic model predictive control (MPC) methods in uncertain environments. The proposed hierarchical vehicle trajectory tracking framework consists of two layers: the upper layer serves as a compensation layer for the vehicle side-slip angle (VSA), designed using bidirectional long shortterm memory (BiLSTM); while the lower layer is the trajectory tracking layer, in which the improved kinematic MPC is enhanced by integrating the twin delayed deep deterministic policy gradient (TD3) algorithm with an external attention (EA) mechanism. The contribution in artificial intelligence is improving the TD3 algorithm with the EA mechanism, enhancing its ability to capture contextual information and improve adaptability. The contribution in engineering applications is implementing the EA-TD3-tuned hierarchical kinematic MPC framework in the field of vehicle trajectory tracking. With 95 % confidence, compared to traditional kinematic MPC controller, the proposed hierarchical vehicle trajectory tracking framework reduces the average lateral error by 33 % (confidence interval, CI: [0.0616, 0.0850]), the average heading angle error by 34 % (CI: [0.01173, 0.0157]), the average yaw rate variation by 31 % (CI: [0.0244, 0.0346]), and the average front wheel steering angle variation by 28 % (CI: [0.0244, 0.0346]).
Surface electromyographic (sEMG) signals contain rich motion information and have a wide range of applications in prosthetics, medical rehabilitation, and muscle status assessment. Feature extraction of sEMG signals is a commonly used and effective approach in the analysis and application of sEMG. This article proposes a novel stratified feature extraction method aimed at capturing the feature information of sEMG signals under different time-domain (TD) intensity levels. The method divides the sEMG signal into different amplitude ranges to obtain the muscle activity frequency information under different intensity levels. First, the threshold value of each layer of each channel is updated in real time to obtain the sEMG signals under the corresponding TD intensity. Second, the number of zero crossing (ZC) features that satisfy the intensity requirements is extracted and combined sequentially with the number of ZC features extracted in the unstratified case to form a feature dataset. Finally, the sEMG signals are divided into several layers to explore the impact of each layer on accuracy. Experiments on a self-built dataset and the publicly available DB5 dataset reveal that accuracy rises initially and then declines gradually with increasing layers. On the self-built dataset, the highest accuracy of 98.22% is achieved at the fourth layer compared to the unstratified extraction with an improvement of 8.74%. In the E2 and E3 of the DB5 dataset, the highest accuracies of 92.16% and 91.37% are achieved, respectively, at the third layer. This indicates that the proposed innovative stratified feature extraction method can significantly improve the accuracy of gesture recognition and effectively expand the applicability of sEMG signal-based action recognition methods.
To address the need for accurate and real-time force measurement in dexterous hand grasping, this article proposes a novel force sensor based on the optical reflection principle, integrated into the fingertips of a biomimetic dexterous hand. The sensor utilizes an elastic structure that induces changes in the optical reflection path under mechanical deformation, modulating the received light intensity for force perception. The sensor's composition is detailed, and theoretical models are developed to analyze the relationships between force, deformation, reflected light intensity, and output voltage. The key parameters are selected, and theoretical response curves are derived. Independent tests validate the sensor's basic performance, with a resolution better than 0.2 N and a sensitivity of 22.44 mV/N. Upon integration into the dexterous hand, the sensor enables real-time acquisition and visualization of fingertip contact forces. Comprehensive evaluation confirms the feasibility of the sensor for force sensing applications in dexterous hands.
Under the impetus of transportation electrification and intelligentization, distributed four-wheel drive electric vehicles (4WD-EVs) have emerged as a pivotal new energy vehicle architecture. However, high-frequency dynamic coordination of multiple motors leads to challenges in energy and thermal management. Traditional energy management strategies and thermal management strategies have been independently optimized, neglecting the intricate coupling relationship between them. This study elucidates the nonlinear coupling between energy management strategies and thermal management strategies in 4WD-EVs and proposes an improved soft actor-critic (SAC) integrated with an evolutionary strategy (ES) to develop an integrated thermal and energy management strategy (ITEMS). The collaborative learning framework enables joint optimization of power distribution and thermodynamic regulation. Compared to model predictive control-based ITEMS, simulation results show the proposed improved SAC based ITEMS reduces energy consumption by 5.52 % for energy management tasks and 15.94 % for thermal management tasks. Meanwhile, it achieves more precise temperature control for motors, battery, and the cabin among these comparison ITEMSs. Moreover, the embedded ES effectively lowers the power consumption of thermal management accessories while improving motor efficiency. This study provides a novel learning framework for breaking through the performance ceiling of 4WD-EVs.
Thermal management systems (TMSs) are crucial for driving safety, mileage, and comfort in battery electric vehicles. To maximize the potential of the integrated TMSs in terms of temperature control and energy saving, the deep deterministic policy gradient (DDPG) is utilized to design a learning thermal management methodology (TMM) for it. Considering the key challenges faced, linear mapping trick and gated recurrent unit are employed to improve the original DDPG. The former empowers the DDPG agents with the ability to make decisions in the discrete-continuous hybrid action space while helping it to avoid the ’curse of dimensionality’. The latter provides the DDPG agents with beneficial historical information, which enhances its decision-making quality. Simulation results show that the proposed TMM decreases the convergence episode to 77 while increasing the convergence reward to 553.4. Further adaptive tests demonstrate that the suggested TMM promptly stabilizes the motor and battery temperatures at the desired values. Simultaneously, energy consumption decreases by 14.71% and 11.45% for two cases compared to the conventional rule-based TMM. In conclusion, the proposed method provides a theoretical foundation for addressing the hybrid action space optimization problem in integrated TMSs.
This paper proposes a deep reinforcement learning (DRL)-based ecodriving strategy for distributed drive electric vehicles (DDEVs) to enhance energy efficiency and dynamic adaptability in complex traffic scenarios. Conventional DRL approaches often employ fixed exploration strategies that are insensitive to dynamic environmental states, leading to suboptimal and unstable performance when handling the unique sensitivity of DDEVs to stochastic disturbances. To address this limitation, the proposed method adaptively adjusting exploration noise based on environmental states by using state-dependent exploration (SDE). A high-fidelity 7-degree-of-freedom vehicle dynamics model and a traffic simulation environment are employed to validate the method. Results demonstrate that SDE-SAC achieves superior performance compared to baseline methods, with a 0.56 kWh power consumption, and a 27.2 m/s mean velocity, alongside improved motion stability. Furthermore, the method exhibits strong generalization in high-density traffic scenarios, maintaining energy efficiency and safety. This work advances the development of adaptive energy management systems for DDEVs by bridging the gap between environmental uncertainty and vehicular control.
To overcome the limitations of conventional force sensors in micro-force detection and structural adaptability, a miniature high-resolution force sensor based on the principle of laser displacement detection is proposed and developed. The sensor employs a stainless-steel elastic plate as the primary sensing element, with external loads applied via a central stud to induce minute axial deformation. A laser diode, rigidly coupled to the stud, transfers this deformation into a displacement of the laser spot on a Position-Sensitive Detector (PSD), thereby establishing a mapping from applied force to spot displacement. Finite element analysis (FEA) is further conducted to investigate the stress and displacement characteristics of the diaphragm under various loads. Experimental validation is carried out through the application of standard micro-weights, including full-range calibration and low-load resolution testing. The results demonstrate excellent repeatability and linear response of the proposed sensor, achieving a minimum detectable force of approximately 2 mN and a sensitivity of 0.1153 mm/N. Overall, the sensor features a compact configuration, high sensitivity, and straightforward assembly, making it particularly suitable for applications in precision manufacturing, contact force monitoring, and other scenarios requiring accurate micro-force measurement. Copyright (c) 2025 The Authors. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/)
Deep reinforcement learning (DRL)-based methods have become predominant in the field of energy management strategy (EMS) today. However, DRL methods are often burdened by extensive training periods and a predisposition towards local optima. This work uses Imitation Learning (IL) to fully exploit the enormous training-guide potential of optimization-based methods, as a way to address the aforementioned disadvantages of DRL, building an IL-embedded DRL framework for EMS. Firstly, the off-line globally optimal trajectory is extracted, and then IL is utilized to imitate the trajectory. Subsequently, the network learned from IL is used as the initial policy network for the DRL algorithm to start training. The EMS based on the framework considers hydrogen consumption, fuel cell degradation, and power battery aging, with the goal of reducing the total driving cost. Behavioral Cloning and Proximal Policy Optimization (PPO) are used as the algorithms of IL and DRL parts, respectively in this study, to experiment the effect of the framework. Simulation results show that, compared to standalone PPO, the proposed framework can reduce the training steps by 51.69% while improving the total reward by 4.57%. Within the test cycle, the proposed framework attains 95.59% of the global optimal driving cost, exceeding standalone PPO by 5.79%.
The method of hand gesture recognition based on surface electromyography (sEMG) signals has garnered increasing attention due to its real-time, convenient, and non-invasive characteristics. This paper presents the design of sEMG signal acquisition system aimed at capturing subtle sEMG signals from the skin surface for applications in human-computer interaction and remote control. Firstly, an analysis of the characteristics of sEMG signals and their sources of interference noise is conducted. Subsequently, a circuit for processing sEMG signals is designed. Additionally, a hardware structure for forearm wearable devices is developed, incorporating a rotating toothed tube to enhance the tightness between the electrodes and the skin surface. Finally, experiments based on the anatomical structure of the forearm muscles are designed to verify the system.
While traditional dictionary learning models under one-shot learning always have the problem of low accuracy of texture recognition, in this paper we propose a novel adaptive multilayer dictionary learning (AMDL) algorithm. An adaptive dictionary generation model is creatively proposed and designed, based on which each sample is able to learn the idiosyncratic dictionary, thus reducing the interference between different material samples. A method of learning label groups for each test sample is proposed to ensure that the test sample learns the correct material category. Support vector guided dictionary learning with locality constraints algorithm is proposed to achieve accurate texture classification. The results of validation experiments on the LMT-108 dataset show that the AMDL algorithm proposed in this paper obtains a better performance in one-shot texture recognition with a recognition rate of 96.71% compared with other similar dictionary learning algorithms.
Surface electromyography (sEMG) signals are biological signals generated during skeletal muscle contraction, which can be captured by electrodes placed on the skin surface. In order to facilitate the acquisition of forearm sEMG signals for hand movement intention recognition, a six-channel wearable EMG sensor is designed and developed in this paper. The architecture of the armband is described in detail. A total of three hand motions from 10 healthy subjects are captured. Firstly we preprocess the original sEMG signals and extract four time-domain features and two frequency-domain features. Secondly, to obtain the optimal recognition results, we sequentially arrange the six features to obtain 43 feature combinations. Finally, the experiment evaluates the dataset with all feature combinations by four models to get the accuracy of hand motion recognition. The experimental results indicate that the average recognition accuracy of the five-feature combination RMS+WL+ZC+MPF+MF is the highest, reaching 93.45%. Therefore, utilizing the electromyography signals collected by the armband developed in this study can effectively achieve hand motion recognition.