In this paper, we propose a lightweight multiscale reference frame generation network for VVC inter coding, named LMRFG. Unlike the previous work [1], LMRFG does not employ the high-performance operation point (HOP) network as preprocessing for frame enhancement. LMRFG replaces the bidirectional motion estimation network with a dualbranch coordinated attention motion estimator (DBCA-ME), which integrates X-axis and Y-axis optical flows for accurate optical flow estimation. Moreover, LMRFG uses depthwise over-parameterized convolutional layer (DO-Conv) to reduce model complexity and minimize bitstreams while maintaining video quality. LMRFG adopts a quantization parameter (QP) distance-based training strategy that takes compressed data at higher QP as input and compressed data at lower QP as label for training, thus addressing the imbalanced QP gap between the compressed input and its uncompressed label. LMRFG is embedded between DPB and RPL to generate a new reference frame and replace the original reference frame in RPL. As shown in Table 1, LMRFG achieves average BD-rate gains of {RA: 4.31% (Y), 6.54% (U), 7.05% (V)} and $\{\text{LDB}: 3.81 \%(\mathrm{Y}), 8.90 \%(\mathrm{U}), 8.77 \%(\mathrm{V})\}$ over the VTM_11.0-NNVC_10.0 (NN-tools ON) anchor, achieving state-of-the-art performance.