Contextualized image captioning is a task that 001 extends beyond generating a purely visual de-002 scription of the image content and aims to pro-003 duce a caption that is influenced by the con-004 text and informed by the real world knowl-005 edge. In this paper, we present an approach 006 to knowledge-aware image captioning, with a 007 specific focus on the temporal domain. We 008 propose a way to identify relevant information 009 in external data sources, such as geographic 010 databases and common knowledge bases, and 011 then encode it in a way that is most useful for 012 the captioning network. We develop an end-013 to-end caption generation system that incorpo-014 rates external knowledge into the captioning 015 process at several stages. The system is trained 016 and tested on our novel temporal knowledge-017 aware captioning dataset, achieving significant 018 improvements over multiple baselines across 019 standardly used metrics. We demonstrate that 020 our approach is effective for generating highly 021 contextualized captions with both relevant and 022 accurate temporal facts. 023