This study proposes an optimized method for immersive and interactive dramatic narratives, grounded in integrated design principles that center on multimodal perception, dynamic decision-making, and collaborative generation. The proposed approach establishes an adaptive decision-making mechanism driven by both emotion and behavior, thereby moving beyond the conventional optimization paradigm based on isolated modules. Specifically, the method leverages Adaptive Reinforcement Learning (ARL) and Multimodal Generative Adversarial Networks (MM-GAN). A multimodal adaptive neural network is first employed to model sequences of user actions, speech, and visual behaviors, while an integrated emotion analysis module predicts real-time emotional states, enabling accurate capture of multidimensional interaction demands. Subsequently, an ARL model incorporating Double Deep Q-Network (Double DQN) and Prioritized Experience Replay (PER) is adopted to maximize the long-term reward of user experience, thus facilitating dynamic branching decisions in the plot. Third, MM-GAN is utilized to generate personalized narrative content across text, image, and speech modalities, with a multi-task joint optimization framework ensuring collaborative training across all components. Experimental results on a self-built VR-IID dataset demonstrate that the proposed method outperforms mainstream baseline models in terms of plot coherence (BERTScore-F1 = 0.897), behavioral response accuracy (92.8%), and immersion score (4.6 points). The system achieves an average response time of 132 ms, satisfying the real-time interaction requirements of VR environments, and reaches training convergence within 58 rounds. Furthermore, this study introduces a narrative adaptability calculation model to quantify the matching degree between user states and plot content, offering a computable theoretical foundation for the intelligent optimization of interactive narratives.