Compared to conventional image content analysis tasks, visual emotion analysis is perceived as a complex, abstract, and potentially culturally dependent endeavor. The accuracy of automatic image-based emotion recognition remains a challenge, and the most significant obstacles are the affective gap and scarcity of data, particularly labeled data. To address these challenge, this paper proposes a saliency-guided masked image modeling approach. Specifically, the proposed framework employs multi-modal large model to generate more emotional images, thereby reducing the impact of data scarcity on model performance. Subsequently, neuroimaging and behavioral studies have demonstrated that human visual attention is attracted by the emotional relevance of a stimulus. In light of this, our model employs a saliency-guided masking strategy to identify emotion-related regions for masking sampling to fit the affective gap. In contrast to the conventional approach of using the original pixel values for the reconstruction target, our model eliminates high-frequency components from the pixels, thus enhancing the generalizability of the model. The use of this unsupervised representation learning approach enables the model to exhibit outstanding recognition performance in downstream emotion recognition tasks on three standard emotion datasets. Furthermore, ablation experiments, robustness test, and visualization experiments corroborate the effectiveness of the proposed method.