
Abstract In this study, we propose a knowledge-selective transfer reinforcement learning method that simultaneously achieves heterogeneous domain transfer and knowledge selection in autonomous robots. In recent years, autonomous robots capable of recognition, decision, and action in their environments have been developed, and their utilization is advancing in a wide range of fields such as disaster response and logistics support. Reinforcement learning (RL) enables adaptive learning in unknown environments and is instrumental in realizing technologies such as autonomous robots. However, RL requires extensive exploration, leading to the issue of long training times. To address this, transfer reinforcement learning (TRL) has been introduced to reduce the training time by reusing previously learned knowledge. Nevertheless, the transfer effectiveness depends on the knowledge reused, creating a need for appropriate knowledge selection. Therefore, we propose heterogeneous domain SAP-net (HDSAP-net), which enables knowledge transfer by utilizing heterogeneous-domain knowledge sets acquired from various agents and tasks, thus extending the existing Spreading Activation Policy Network (SAP-net). In its algorithm design, HDSAP-net incorporates inter-task mapping using linear interpolation to bridge discrepancies between the heterogeneous domains. Furthermore, by analyzing the behavior of HDSAP-net, we formulated a network design method using optimal transport cost and implemented a new activation-value control method, thereby improving its performance. This enables TRL, wherein knowledge sets derived from various robots are shared among various agents, which was challenging with the conventional SAP-net. Verification experiments using the physics simulator Webots involving a mobile robot, robotic arm, and drone demonstrated that HDSAP-net significantly improves the learning efficiency when compared with that of conventional RL. The proposed method reduced the learning time by 32.1% to 82.0%, confirming its capability of autonomously discovering and utilizing effective knowledge across various domains.
Autonomous vehicles (AVs) have demonstrated significant potential in revolutionizing transportation, yet ensuring their safety and reliability remains a critical challenge, especially when exposed to dynamic and unpredictable environments. Real-world testing of an Autonomous Driving System (ADS) is both expensive and risky, making simulation-based testing a preferred approach. In this paper, we propose Scenetic , a Reinforcement Learning (RL)-based approach to generate plausible critical scenarios for testing ADSs in simulation environments. To capture the complexity of driving scenarios, Scenetic comprehensively represents the environment by both the internal states of an ADS under-test (e.g., the status of the ADS’s core components, speed, or acceleration) and the external states of the surrounding factors in the simulation environment (e.g., weather, traffic flow, or road condition). Scenetic trains the RL agent to configure the simulation environment that places the AV in dangerous situations and potentially leads it to collisions. We introduce a diverse set of actions that allows the RL agent to systematically configure both environmental conditions and traffic participants. Additionally, based on established safety requirements, we enforce heuristic constraints to promote the challenge and practical relevance of the generated test scenarios. Scenetic is evaluated on two popular simulation maps with four different road configurations. Our results show Scenetic ’s ability to outperform the state-of-the-art approach by generating 30 Scenetic achieves up to 275 Scenetic in enhancing the safety testing of AVs through the generation of comprehensive and plausible critical scenarios.
Neural network-based quantum Monte Carlo (NNQMC), an emerging method for solving many-body quantum systems with high accuracy, has been mainly applied to small systems owing to demanding computation requirements. Here we introduce a framework based on local pseudopotentials to break through such limitation, improving the computational efficiency and scalability of NNQMC. The incorporation of local pseudopotentials reduces the number of electrons treated in neural network and also achieves better relative energy accuracy than all electron NNQMC calculations for complex systems. This counterintuitive outcome is made possible by the distinctive characteristics inherent to NNQMC. Notably, by avoiding costly integration terms, this approach is also substantially more efficient than its widely used semilocal counterparts. Our approach enables the reliable treatment of large and challenging systems, such as the Fe 4 S 4 ( SCH 3 ) 4 iron-sulfur cluster. Overall, our findings demonstrate that the synergy between NNQMC and local pseudopotentials substantially expands the scope of accurate ab initio calculations.
Modern logical reasoning with LLMs primarily relies on employing complex interactive frameworks that decompose the reasoning process into subtasks solved through carefully designed prompts or requiring external resources (e.g., symbolic solvers) to exploit their strong logical structures. While interactive approaches introduce additional overhead, hybrid approaches depend on external components, which limit their scalability. A non-interactive, end-to-end framework enables reasoning to emerge within the model itself – improving generalization while preserving analyzability without any external resources. In this work, we introduce a non-interactive, end-to-end framework for reasoning tasks. We show that introducing structural information into the few-shot prompt activates a subset of attention heads that patterns aligned with logical reasoning operators. Building on this insight, we propose Attention-Aware Intervention (AAI), an inference-time intervention method that reweights attention scores across selected heads identified by their logical patterns. AAI offers an efficient way to steer the model's reasoning toward leveraging prior knowledge through attention modulation. Extensive experiments show that AAI enhances logical reasoning performance across diverse benchmarks and model architectures, while incurring negligible additional computational overhead. Code is available at https://github.com/phuongnm94/aai_for_logical_reasoning.
We theoretically and computationally investigate long-memory processes based on the Markovian lifts of affine jump-diffusion processes. A nominal superposition process consisting of an infinite number of interacting affine processes is considered, along with its finite-dimensional version and associated generalized Riccati equations. We propose a splitting scheme suited to the Markovian lifts where jump and diffusion parts are dealt with separately based on recently developed exact discretization methods. We examine the computational performance of the scheme through comparisons with the analytical results. We also numerically investigate a more complex model arising in the environmental sciences and some extended cases in which superposed processes belong to a class of nonlinear processes that generalize affine processes.