Weakly supervised remote sensing semantic segmentation aims to achieve pixel-level prediction with limited annotation costs. Recently, CLIP-based methods have shown promising potential by leveraging vision-language alignment for semantic localization; however, they often suffer from ambiguous semantic representations and incomplete pseudo-labels in complex remote sensing scenes. To address these issues, we propose a spatial cue-guided framework for weakly supervised remote sensing semantic segmentation. Specifically, a confidence-aware prototype alignment module is designed to extract reliable semantic cues from high-confidence regions, enhance feature discrimination through contrastive learning, and improve object completeness by exploiting structural information. An adaptive pseudo-label completion strategy is developed to progressively improve pseudo-label coverage while reducing noise propagation. In addition, a structure-aware heterogeneous multi-head attention decoder is introduced to effectively fuse global semantic context, local spatial details, and cue information for refined prediction. Experimental results on two benchmark remote sensing datasets demonstrate the superiority of the proposed method.