Existing Audio-visual saliency prediction (AVSP) methods often overlook the importance of spatiotemporal alignment of audio-visual features, leading models to over-reliance on visual signals. As a result, audio-visual features cannot be fully utilized, and blindly fusing misaligned audio-visual features may lead to performance decline. To address this challenge, we propose an AVSP method that combines adversarial learning and co-attention. Specifically, to achieve spatiotemporal alignment of audio-visual features, the frame-wise collaborative attention (FWCA) module is introduced. This module exploits latent correlations between audio-visual features to align spatiotemporal information using a frame-wise collaborative cross-modal attention mechanism. Additionally, it accumulates the weights of the previous frame during frame-wise propagation, thereby enhancing the model’s ability to utilize audio-visual information. To combat the interference of background noise and the limitations on the size of audio-visual datasets, we design a spatiotemporal adversarial learning (STAL) module. This further ensures the spatiotemporal consistency of audio-visual features, guides the model to balance attention to audio-visual information, and eliminates the reliance on pre-training with video datasets, significantly improving the model’s training efficiency. Through extensive experiments on several benchmark datasets, our proposed method not only achieves competitive performance but also potentially provides deep insights into the audio-visual fusion mechanism.