Steady-state visual evoked potentials (SSVEPs) are widely used in brain-computer interfaces (BCIs) due to their stability and high signal-to-noise ratio. However, decoding short-time SSVEPs remains a key bottleneck that limits system performance. We propose a framework that integrates time series forecasting with SSVEP decoding to extend the effective data length and improve recognition accuracy for short recordings. Specifically, we introduce the Time-Frequency Fusion Network (TFFNet), a deep-learning-based forecasting model that forecasts short-time SSVEP signals for up to 40 stimulus classes. We then combine TFFNet with the state-of-the-art ensemble Task-Related Component Analysis by classifying extended segments formed by concatenating the original signal with its forecasting signal. Experiments on two public SSVEP datasets show that the proposed method significantly improves short-time SSVEP recognition and outperforms state-of-the-art baselines across all tested lengths; at 0.4 s, it achieves information transfer rates of 262.35 and 175.61 bits/min on the two datasets, respectively. Correlation analyses further corroborate the effectiveness of the forecasting strategy. To our knowledge, this is the first study to incorporate time series forecasting into SSVEP decoding, offering a generalizable approach to enhance SSVEP-BCI performance under short recording intervals.