Recently, speaker verification algorithms have been of great interest because of their wide applicability in several fields such as speech communication, domestic services, security and access control, and intelligent terminals. Today we have interactive devices such as smartphone assistants and smart speakers that are designed to understand basic voice commands. However, the performance of current speaker verification programs degrades significantly when analyzing short utterances, and larger speech segments are needed to achieve high performance. Although CNNs, LSTM networks, MFCCs, CMVN, and PLDA have been thoroughly studied for speaker recognition, their performance is still limited in the case of extremely brief speech, due to the lack of speaker-discriminative information in short utterances. The uniqueness of this work is not in proposing a new neural architecture but in presenting an integrated framework that integrates CMVNC-normalized acoustic data with complementary CNN and LSTM-based temporal modeling for robust short-utterance speaker verification. Unlike prior studies that predominantly utilise traditional MFCC features or assess a sole deep learning architecture, this study systematically explores the contribution of CMVNC features across two deep learning models, and juxtaposes them against the traditional i-vector/PLDA baseline under varying training and testing utterance lengths.