Multimodal stance detection leverages multimodal information to identify an author’s stance towards a specific target. Existing approaches typically focus on coarse alignment between image and text representations, overlooking fine-grained modality alignment oriented by the target. Therefore, we propose a novel Target-oriented Consistent Cross-modal Alignment (TCCA) method. Our approach first designs a Target-oriented Prompt as Visual-Language Query module, which leverages the target to activate visual and textual representations, guiding the model to focus on stance-relevant visual regions and text spans. This enables more efficient and precise cross-modal semantic fusion. Furthermore, we devise a Target-oriented Cross-modal Alignment module to align the target-activated image and text representations. This facilitates tighter target-oriented cross-modal information and enhances the model’s ability to capture stance-specific information. By orienting both modalities towards the target, TCCA effectively simplifies the alignment process and improves the model’s capacity to capture fine-grained stance cues. Finally, we construct the Multiview Decoding module that enables more effective extraction of rich stance cues from diverse perspectives. The experimental results demonstrate that TCCA outperforms existing approaches in most metrics, validating its effectiveness for multimodal stance detection.