Online learning has drawn attention to the problem of adaptive learning path recommendations. The reinforcement learning (RL) algorithm has become an important tool in this research field due to its effectiveness in solving the multi-sequence decision-making problem during dynamic environment interaction. This study enhances the cognitive structure enhanced framework for adaptive learning (CSEAL) and constructs a new adaptive learning path recommendation (ALPR) framework to address problems like incomplete characterization of dynamic learning environments and sparse and delayed reward design incentives. First, the framework incorporates the core dynamic features of the domain model into dynamic learning environment characterization with increased completeness and accuracy. Second, the reward function is redesigned based on the idea of reward shaping to make the performance of the agent more stable while exploring. The adaptive learning algorithms involved in the ALPR framework are described in detail. Finally, experiments on appropriate datasets show that the learning path recommended by the ALPR method is better than those recommended by other advanced baseline methods. The proposed framework raises the standard of learning path recommendations in a dynamic learning environment, which not only enhances the theory of adaptive learning but also encourages the creation and use of individualized online learning programs.