Speaker recognition models are typically enrolled using neutral speech, yet real users rarely speak under emotionally neutral conditions. Emotion alters prosody, spectral structure, articulation, and speaking rate, shifting utterances away from the acoustic distribution observed during enrolment. Conventional augmentation with noise or speed variation introduces generic acoustic variability but does not explicitly model these coordinated, emotion-specific changes. This study therefore investigates whether synthetic emotional speech representations can improve speaker recognition when authentic emotional enrolment recordings are unavailable. Three conditional generative adversarial networks translate neutral mel spectrograms into angry, happy, and sad variants. The translation models are trained using lexically parallel neutral and emotional recordings from ESD and SAVEE. The generated spectrograms are subsequently added to neutral training data, while evaluation is conducted exclusively on genuine recordings from the independent RAVDESS and CREMA-D corpora. This protocol assesses whether learned emotional transformations transfer across speakers, lexical content, and recording conditions. Combined emotional augmentation improves speaker identification and verification on both corpora, with a larger effect on RAVDESS. Identification accuracy increases from 65.00% to 71.83% on RAVDESS and, by a smaller margin, from 56.04% to 60.20% on CREMA-D. The experiments further indicate that augmentation effectiveness depends on the target emotion and corpus, and that single-emotion augmentation can degrade recognition on some corpora.