Digital Image Correlation (DIC) serves as a non-contact optical metrology technique for quantifying object deformation, with widespread deployment in industrial inspection and intelligent manufacturing. Conventional DIC approaches commonly exploit the spatial–temporal continuity of object deformation to boost measurement precision and robustness. While existing deep learning-based DIC methods have effectively leveraged spatial continuity, delivering marked improvements in accuracy, efficiency, and anti-interference capability, they remain limited by underutilization of temporal continuity. Specifically, current DIC networks only take a reference image and a deformed image as input, disregarding historical deformation sequences. To overcome this constraint, this work presents Temporal-DICnet, a sequential-image-input network that incorporates temporal information for multi-frame deformation field estimation. The proposed framework comprises three core modules: a feature encoder that extracts image features and constructs 4D correlation cost volumes between the reference and each deformed frame; a motion encoder that enables inter-frame motion feature propagation and fusion; and a deformation prediction module with convolutional residual blocks used to refine the output deformation field. Experimental results on synthetic datasets reveal that Temporal-DICnet reduces the mean absolute error (MAE) by more than 50% relative to the ALDIC method. Physical uniaxial tensile tests further validate the accuracy and generalization of the proposed network in real-world deformation measurement scenarios.