Understanding how the human brain encodes complex natural scenes remains a central problem in computational neuroscience and artificial intelligence. Existing visual encoding models often rely on a single dominant feature representation and may insufficiently characterize how saliency-guided spatial information and high-level semantic context jointly contribute to cortical response prediction. To address this issue, this study proposes a saliency-guided multimodal visual encoding model, termed SMG-MVEM, to predict voxel-wise cortical responses to natural scene stimuli. The model integrates image features, saliency cues, and text-derived semantic representations through a hierarchical fusion architecture, followed by a Transformer-based brain mapper. Experiments on the Natural Scenes Dataset (NSD) show that SMG-MVEM improves prediction performance over representative neural encoding baselines and internal control variants, with the average PCC increasing from [Formula: see text] for the best-performing baseline to [Formula: see text]. Regional analyses further show that saliency contributed more strongly to early visual areas, whereas semantic features provided greater benefits in higher-order regions. Representational analyses also suggest that the model-predicted responses preserved aspects of hierarchical and category-related organization across the visual cortex. These findings indicate that structured integration of saliency and semantic context can improve cortical response prediction and provide interpretable representational patterns for natural vision.