In recent years, unmanned ground vehicles (UGVs) have advanced rapidly, attracting significant attention for their applications in modern military operations, particularly as target vehicles. Their ability to perform realistic combat maneuvers relies heavily on formation control, a key technology in this domain. This paper presents a novel formation control framework aimed at improving the accuracy and stability of dynamic formations for such vehicles. Building upon an enhanced leader-follower structure, the proposed method generates virtual dynamic targets for tracking by each follower vehicle. Based on the kinematic model of skid-steering vehicles, a Linear Quadratic Regulator (LQR) controller is then designed for precise tracking of these virtual targets during formation maneuvers. To further optimize the LQR controller's performance, the Proximal Policy Optimization (PPO) algorithm is incorporated into an asynchronous training framework inspired by the Asynchronous Advantage Actor-Critic (A3C) method. The PPO networks are trained in a simulation environment to enable real-time prediction of optimal LQR parameters during formation control. Compared with end-to-end deep reinforcement learning (DRL) methods and other related approaches, the combination of the PPO network and LQR controller results in a compact network size suitable for deployment on embedded systems. The effectiveness of the proposed methodology is validated through comparative experiments with alternative controllers in both simulation and real-vehicle tests. This approach is particularly well suited for low-cost UGVs with limited computational resources and for applications demanding dynamic motion control with high tracking accuracy and stability.