Both a deep understanding of visual cues and their contextual importance are demanded by effective image captioning. However, seamlessly integrating balanced contextual information continues to be a substantial challenge. In this paper, we present FreConCap, a novel Frequency-guided Con-textual Image Captioning framework, to overcome the challenge using high-frequency and background features, along with object-level region features. We transform grid features into frequency domain and filter out low-frequency components by a cutoff ratio that enhances fine details critical for detailed visual understanding. Multi-Stream Cross Attention is developed to reduce the modality gap between vision and language, and to capture the interaction of text features with high-frequency local features, objects, context, and their relationships. Our experiments on the MS COCO image captioning benchmark show the superiority of our approach as compared with existing methods for enhanced image captions with more contextual information.