In video captioning, current methods often struggle with generating accurate descriptions due to challenges in modeling object interactions, especially predicates that rely on both object dynamics and motion patterns. To address this, we introduce the Behavior Modeling-Aware Video Captioning Network (BMVCap), which enhances video captioning by capturing fine-grained details of interactions, co-occurring objects, and contextual background using a transformer-based encoder for each modality stream. The model integrates these streams through an adaptive fusion mechanism in the transformer decoder, allowing it to generate more precise captions. Additionally, BMVCap employs a caption length control mechanism and optimized reinforcement learning, maximizing rewards from multiple evaluation metrics. Extensive experiments on the MSVD, MSR-VTT, and VATEX datasets show that BMVCap significantly improves captioning accuracy by better modeling complex interactions and activity attributes. The results demonstrate that emphasizing complex interactions and activity attributes leads to substantial improvements in the accuracy and reliability of video captioning.