Decoding the brain’s visual neural activity is crucial for understanding visual mechanisms and advancing brain-computer interface (BCI) technology. Existing methods often rely on static alignment or non-aligned strategies, making it difficult to fully utilize the semantic information in brain activity, resulting in suboptimal decoding performance, especially in complex tasks. To address this, we propose a Dynamic Aligned Visual Decoding Model (DA-VDM), based on a generative language model, employing a dynamic alignment strategy implemented via progressively weighted training. By dynamically adjusting the training weights, this strategy gradually shifts the model’s focus from low-level feature mapping to high-level semantic generation, thereby enhancing the semantic association between brain activity features and image-text representations and improving semantic expression capabilities. Additionally, we introduce Prompt-based techniques, leveraging the generative power of language models to tackle cross-subject and cross-task decoding challenges, enabling the model to adapt to different subjects within a unified framework while predicting both the category and textual description of visual stimuli. Experiments on the Natural Scenes Dataset (NSD) demonstrate that DA-VDM achieves an accuracy of 0.685 in category decoding tasks, outperforming existing models; in text decoding tasks, it also leads in metrics such as METEOR and ROUGE. Ablation studies confirm that the dynamic alignment strategy significantly enhances category and text decoding accuracy, while Prompt-based techniques exhibit strong generalization capabilities in cross-subject and cross-task decoding.