The multiple-input multiple-output dual functional radar communication (MIMO-DFRC) system is a promising platform for future integrated sensing and communication applications. Ensuring reliable performance of both radar and communication functions, the beam selection is a critical technology in MIMO-DFRC systems. However, the beam selection problem is known to be NP-hard, and efficiently addressing it remains an open issue, especially in distributed systems. In this paper, we address the beam selection problem for a MIMO-DFRC system by formulating it as a semi-Markov decision process and propose a novel hierarchical reinforcement learning (HRL) algorithm. In our approach, codebook-based beam selection for transmitting and receiving BS is controlled by an agent deployed in the cloud. Inspired by the mechanism of hierarchical codebook beam training, we employ an option-based policy that enables the agent to explore different layers of the codebook and extract context information across multiple discrete time steps. We utilize an invalid action masking technique to overcome the dynamic action space problem caused by the option-based policy. Simulation results demonstrate that the HRL-based algorithm outperforms existing beam selection methods and achieves remarkable performance even under conditions of a high probability of false alarm and low signal-to-noise ratio. Furthermore, we find that the proposed algorithm exhibits promising capabilities to learn a more efficient policy beyond the full hierarchical codebook training trajectory.
更多
查看译文
关键词
MIMO,Array signal processing,Radar,Heuristic algorithms,Training,Vectors,Prediction algorithms,Covariance matrices,Signal to noise ratio,Data mining,Integrated sensing and communication,beam selection,deep reinforcement learning,hierarchical codebook