Surrogate model–assisted multi-objective optimization has emerged as a leading approach for solving computationally expensive multi-objective optimization problems. Conventional methods typically rely on either a fixed single surrogate or multiple models for each objective in model management, often underutilizing the complementary strengths of different models across problems and, when an ensemble is used to enhance approximation, incurring a higher training burden. We propose a reinforcement learning-driven model selection framework designed to maximize the cumulative reward based on the approximation errors of both the current surrogate and the newly sampled point. The agent operates in two states defined by recent reward trends and autonomously selects a single model at each step to approximate the objective function. A set of reference vectors guides the optimization of the selected models toward the Pareto front. In infill sampling, an informative solution is chosen from the final population for expensive evaluations based on the convergence distance and the historical sample distribution associated with the reference vectors. Extensive experiments on benchmark suites, involving DTLZ and WFG, as well as three real-world problems, demonstrate that the proposed algorithm is competitive with six state-of-the-art SAEAs.