Experiments on non-native phonemic contrast learning typically use either lexical or perceptual training. However, few studies have compared their efficiency using comparable protocols. In our previous study (Pattamadilok, Welby, & Tyler, 2022), we applied a lexical training protocol to examine the contribution of articulatory and orthographic cues to new word learning. Native speakers of French were trained to associate unknown objects with minimal-pairs of new English pseudowords differing in the initial /f/ versus /θ/ phoneme, a contrast that does not exist in French (e.g., fedge, thedge). The presence of an orthographic cue during training led to an overnight improvement of both retention of new word-object associations and the ability to perceive the /f/-/θ/ contrast. Here, we employed the same material and assessment tasks while employing perceptual, rather than lexical training; that is, the visual cues were presented during a phoneme discrimination training. The benefits of the visual cues differed from those of the initial study. While both articulatory and orthographic cues enhanced phoneme discrimination ability during the training phase, only the articulatory cue yielded an overnight improvement in perceptual ability without generalizing this benefit to the retention of word-object associations. Together, these two datasets demonstrate that, even in well-matched experimental protocols, different training methods can yield distinct learning outcomes. Optimal training protocols must consider the complex interplay between the level of language representations targeted by the training, the differential benefits of input modality, and the short- versus long-term training effects on different processing levels.