The exponential growth of multimedia data necessitates advanced cross-modal retrieval methods capable of bridging the semantic gap between heterogeneous modalities, such as images and text. Although deep cross-modal hashing techniques have demonstrated strong performance by jointly integrating feature extraction and hash code generation, existing approaches often struggle to generate genuinely unified hash codes. These limitations mainly arise from modality-specific feature discrepancies and insufficient mechanisms for robust semantic alignment. To address these challenges, we propose a Cross-Attention Aware Fusion-based Cross-Modal Hashing (CAFH) method. The CAFH framework introduces a cross-attention aware fusion module that effectively captures and integrates shared semantic information across modalities, thereby producing coherent representations of semantically related data points. Unified hash codes are subsequently generated from these integrated features, improving retrieval accuracy in cross-modal hashing tasks. Furthermore, the model incorporates semantic similarity learning to enhance the separability of dissimilar samples, thereby improving both robustness and retrieval precision. Extensive experiments on three benchmark datasets, namely MIRFLICKR25K, NUSWIDE-10K, and MSCOCO, demonstrate the superior performance of the proposed method in both image-to-text and text-to-image retrieval tasks. By achieving state-of-the-art results and addressing key limitations of existing methods, CAFH provides a robust and effective solution for large-scale multimedia retrieval.