Two experiments investigate the effectiveness of audiovisual (AV) speech cues (cues derived from both seeing and hearing a talker speak) in facilitating perceptual learning of spectrally distorted speech. Speech was distorted through an eight channel noise-vocoder which shifted the spectral envelope of the speech signal to simulate the properties of a cochlear implant with a 6 mm place mismatch. Experiment 1 found that participants showed significantly greater improvement in perceiving noise-vocoded speech when training gave AV cues than when it gave auditory cues alone. Experiment 2 compared training with AV cues with training which gave written feedback. These two methods did not significantly differ in the pattern of training they produced. Suggestions are made about the types of circumstances in which the two training methods might be found to differ in facilitating auditory perceptual learning of speech.
The aim of this paper was to examine the effect of regional accent on speechreading accuracy and the utility of contextual cues in reducing accent effects. Study 1: Participants were recruited from Nottingham (n=24) and Glasgow (n=17). Their task was to speechread 240 visually presented sentences spoken by 12 talkers, half with a Glaswegian accent, half a Nottingham accent. Both participant groups found the Glaswegian talkers less intelligible (p<0.05). A significant interaction between participant location and accent type (p<0.05) indicated that both participant groups showed an advantage for speechreading talkers with their own accent over the opposite group. Study 2: Participants were recruited from Nottingham (n=15). The same visual sentences were used, but each one was presented with a contextual cue. The results showed that speechreading performance was significantly improved when a contextual cue was used (p<0.05). However the Nottingham observers still found the Glaswegian talkers less intelligible than the Nottingham talkers (p<0.05). The findings of this paper suggest that accent type may have an influence upon visual speech intelligibility and as such may impact upon the design, and results, of tests of speechreading ability.
The visible movement of a talker's face is an influential component of speech perception. However, the ability of this influence to function when large areas of the face (~50%) are covered by simple substantial occlusions, and so are not visible to the observer, has yet to be fully determined. In Experiment 1, both visual speech identification and the influence of visual speech on identifying congruent and incongruent auditory speech were investigated using displays of a whole (unoccluded) talking face and of the same face occluded vertically so that the entire left or right hemiface was covered. Both the identification of visual speech and its influence on auditory speech perception were identical across all three face displays. Experiment 2 replicated and extended these results, showing that visual and audiovisual speech perception also functioned well with other simple substantial occlusions (horizontal and diagonal). Indeed, displays in which entire upper facial areas were occluded produced performance levels equal to those obtained with unoccluded displays. Occluding entire lower facial areas elicited some impairments in performance, but visual speech perception and visual speech influences on auditory speech perception were still apparent. Finally, implications of these findings for understanding the processes supporting visual and audiovisual speech perception are discussed.
The ability to recognize mental states from facial expressions is essential for effective social interaction. However, previous investigations of mental state recognition have used only static faces so the benefit of dynamic information for recognizing mental states remains to be determined. Experiment 1 found that dynamic faces produced higher levels of recognition accuracy than static faces, suggesting that the additional information contained within dynamic faces can facilitate mental state recognition. Experiment 2 explored the facial regions that are important for providing dynamic information in mental state displays. This involved using a new technique to freeze motion in a particular facial region (eyes, nose, mouth) so that this region was static while the remainder of the face was naturally moving. Findings showed that dynamic information in the eyes and the mouth was important and the region of influence depended on the mental state. Processes involved in mental state recognition are discussed.
Aim - To determine whether single-word speechreading is sensitive to prior phrase context. Background - We used a well-established Event Related Potential (ERP) paradigm to elicit the N400, a negative wave around 400 ms. The N400 reflects activation by lexical semantic features and is modulated by contextual integration. Methods - Seventeen adult native English speakers participated. They were shown written three-word contexts followed by a final word (cloze) presented as a silent video clip. Context and cloze formed typical phrases (60 trials, e.g. “No smoke without fire”) and anomalous phrases (120 trials, e.g. “Two of a trade”). Subjects were instructed to identify the cloze at the end of each trial. Typicality (typical/anomalous) and accuracy were used to classify trials. ERPs were computed off-line by averaging epochs aligned to the onset of articulatory movements. Results and discussion - For all four conditions, reliable evoked activity emerged only after 500 ms postonset. Topographical analysis revealed a symmetrical sustained negative field at the frontal sites and positive field at the posterior sites. Statistical comparisons showed that only the waveform for the typical-correct condition differed from the other three. The difference, however, was not compatible with an N400. This finding suggests that the speechreading process integrates lexical semantic features differently from spoken and written language perception. We speculate on those cognitive mechanisms which may be responsible for the observed effects. Index Terms: speechreading, ERP, linguistic context 1. Aim This experiment sought to determine whether the process of single-word speechreading, as indexed by ERP responses, is sensitive to prior phrase context. If visual speech is processed in a similar way to written and spoken (auditory) language, then the characteristic ERP signature for out-of-context words should be present also in the visual speech modality. Topographical analysis and statistical comparison of the ERP waveforms were the main measures used to test the hypothesis.
Recent research had shown that concurrent visual speech modulates the cortical event-related potential N1/P2 to auditory speech. Audiovisually presented speech results in an N1-P2 that is reduced in peak amplitude and with shorter peak latencies than unimodal auditory speech [11]. This effect on the N1/P2 is consistent with a model in which visual speech integrates with auditory speech at an early processing stage in the auditory cortex by suppressing auditory cortical activity. We examined the effects of audiovisual temporal synchrony in producing modulations in the N1/P2. With the visual stream presented in synchrony with the auditory stream our results replicated the basic findings of reduced peak amplitudes in the N1/P2 compared to a unimodal auditory condition. With the visual stream temporally mismatched with the auditory stream (so that the auditory speech signal was presented 200 ms before its recorded position) the recorded N1/P2 was similar to unimodal auditory speech. The results are discussed in terms of Wassenhove’s ‘analysis-bysynthesis model’ of audiovisual integration.
Exposure to audiovisually presented vocoded speech is more effective than exposure to auditory-only vocoded speech in improving the subsequent ability to understand vocoded speech [1]. In addition, improvements in the audiovisual training condition were more rapid and greater in magnitude than in the auditory-only condition. The current study was conducted to establish whether exposure to concurrent textual information also results in improvements in the ability to recognize vocoded speech. Baseline measures of identification performance with auditory-only vocoded speech sentences were assessed for 45 participants. Participants then performed a speech identification task, where they were exposed to vocoded speech in either audiovisual (Group 1), auditory-only (Group 2), or auditory-only with concurrent text conditions (Group 3). Following exposure, participants were tested again on identification performance with auditory-only vocoded speech. Exposure to concurrent text improved subsequent understanding of vocoded speech, to a level similar to that seen with audiovisual speech exposure. In a second experiment, groups of normal hearing adults were exposed to vocoded non-lexical nonsense words in auditory-only with text and audiovisual speech presentation conditions. Exposure to both nonsense audiovisual and concurrent text conditions improved subsequent understanding of lexical auditoryonly vocoded speech, and there was no difference between the levels of improvement. In summary, to computerbased audiovisual and concurrent text exposure improves the ability to recognize vocoded speech over exposure to auditory stimuli alone. This effect does not appear to be dependent on exposure to lexical items.
effect of accent (pronunciation of speech sounds determined by a speaker's regional or national location) on auditory speech comprehension has been well documented, but research is lacking as to its effects on visual speech understanding. In order to address this, the present study examined the effect of regional accent variation on speechreading performance. The aim was to determine if familiarity with an accent would prove to be an advantage when speechreading, and if certain regional accents would be associated with a greater clarity of visual signal than others. The study examined both the identification of accent based on visual or auditory sentences and the effect of accent variation on auditory and visual speech comprehension. Of particular interest was the effect of accent on speechreading accuracy. The two British accents chosen for comparison were Nottingham and Glaswegian. Results showed that accent discrimination is possible using only visual speech. Greater accuracy was achieved when participants differentiated between the accents based on auditory sentences, but with visual speech, performance was also greater than chance (p. > .05). This suggests that the two accents have sufficiently distinct patterns of visual articulation and auditory pronunciation to allow participants to discriminate between them. Further, the participants visual and auditory speech recognition scores were adversely affected by the use of an unfamiliar accent, with keyword recognition accuracy being reduced. These findings both replicate previous findings on the effects of auditory accent, and indicate that accent can also impact the understanding of visual speech.
Lateralized displays are used widely to investigate hemispheric asymmetry in language perception. However, few studies have used lateralized displays to investigate hemispheric asymmetry in visual speech perception, and those that have yielded mixed results. This issue was investigated in the current study by presenting visual speech to either the left hemisphere (LH) or the right hemisphere (RH) using the face as recorded (normal), a mirror image of the normal face (reversed), and chimeric displays constructed by duplicating and reversing just one hemiface (left or right) to form symmetrical images (left-duplicated, right-duplicated). The projection of displays to each hemisphere was controlled precisely by an automated eye-tracking technique. Visual speech perception showed the same, clear LH advantage for normal and reversed displays, a greater LH advantage for right-duplicated displays, and no hemispheric difference for left-duplicated displays. Of particular note is that perception of LH displays was affected greatly by the presence of right-hemiface information, whereas perception of RH displays was unaffected by changes in hemiface content. Thus, when investigated under precise viewing conditions, the indications are not only that the dominant processes of visual speech perception are located in the LH but that these processes are uniquely sensitive to right-hemiface information.
The production of speech involves an individual's control of their various articulators (lips, tongue, larynx etc.) to produce auditory speech signals [1].These movements can be utilised in the processing of visual speech and form the basis of speechreading.However, the production of speech by different talkers can be variable; physiology, accent and speech rate can all change the appearance of the visual signal.The focus of this report is an investigation into the effects of language and accent variation on speechreading, an area previously lacking in systematic research.Results from two experiments indicate, firstly, that the visual differences between French and English, (both accent and language) can be discriminated through visual speech.Secondly, in a comparison of speechreading performance, English sentences produced using a French accent were found to be significantly more difficult to speechread by English observers than those produced in an English accent.This research indicates the importance of further study into the effects of accent on speechreading.
Seeing a talker's face influences auditory speech recognition, but the visible input essential for this influence has yet to be established. Using a new seamless editing technique, the authors examined effects of restricting visible movement to oral or extraoral areas of a talking face. In Experiment 1, visual speech identification and visual influences on identifying auditory speech were compared across displays in which the whole face moved, the oral area moved, or the extraoral area moved. Visual speech influences on auditory speech recognition were substantial and unchanging across whole-face and oral-movement displays. However, extraoral movement also influenced identification of visual and audiovisual speech. Experiments 2 and 3 demonstrated that these results are dependent on intact and upright facial contexts, but only with extraoral movement displays.
The advantage for words in the right visual hemifield (RVF) has been assigned parallel orthographic processing by the left hemisphere and sequential by the right. However, an examination of previous studies of serial position performance suggests that orthographic processing in each hemifield is modulated by retinal eccentricity. To investigate this issue, we presented words at eccentricities of 1, 2, 3, and 4 degrees. Serial position performance was measured using the Reicher-Wheeler task to suppress influences of guesswork and an eye-tracker controlled fixation location. Greater eccentricities produced lower overall levels of performance in each hemifield although RVF advantages for words obtained at each eccentricity (Experiments 1 and 2). However, performance in both hemifields revealed similar U-shaped serial position performance at all eccentricities. Moreover, this performance was not influenced by lexical constraint (high, low; Experiment 2) or status (word, nonword; Experiment 3), although only words (not nonwords) produced an RVF advantage. These findings suggest that although each RVF advantage was produced by left-hemisphere function, the same pattern of orthographic analysis was used by each hemisphere at each eccentricity.
The anatomical arrangement of the human visual system offers considerable scope for investigating functional asymmetries in hemispheric processing. In particular, because each hemisphere receives information initially from the contralateral visual hemifield, visual stimuli presented to the left of a central fixation point can be projected directly to the right hemisphere and visual stimuli presented to the right of a central fixation point can be projected directly to the left hemisphere. Numerous studies using displays of this type suggest that, for the vast majority of individuals, written words produce different patterns of performance when presented to different hemifields and these findings have inspired considerable debate about the processes available for word recognition in each hemisphere.
D. Briihl and A. W. Inhoff (1995; see record 1995-20036-001) found that exterior letter pairs showed no privileged status in reading when letter pairs were presented as parafoveal primes. However, T. R. Jordan, S. M. Thomas, G. R. Patching, and K. C. Scott-Brown (2003; see record 2003-07955-013) used a paradigm that (a) allowed letter pairs to exert influence at any point in the reading process, (b) overcame problems with the stimulus manipulations used by Briihl and Inhoff (1995), and (c) revealed a privileged status for exterior letter pairs in reading. A. W. Inhoff, R. Radach, B. M. Eiter, and M. Skelly (2003; see record 2003-07955-014) made a number of claims about the Jordan, Thomas, et al. study, most of which focus on parafoveal processing. This article addresses these claims and points out that although studies that use parafoveal previews provide an important contribution, other techniques and paradigms are required to reveal the full role of letter pairs in reading.
Exterior letter pairs (e.g., d_k in dark) play a major role in single-word recognition, but other research (D. Briihl & A. W. Inhoff, 1995) indicates no such role in reading text. This issue was examined by visually degrading letter pairs in three positions in words (initial, exterior, and interior) in text. Each degradation slowed reading rate compared with an undegraded control. However, whereas degrading initial and interior pairs slowed reading rate to a similar extent, degrading exterior pairs slowed reading rate most of all. Moreover, these effects were obtained when letter identities across pair positions varied naturally and when they were matched. The findings suggest that exterior letter pairs play a preferential role in reading, and candidates for this role are discussed.
Perception of visual speech and the influence of visual speech on auditory speech perception is affected by the orientation of a talker’s face, but the nature of the visual information underlying this effect has yet to be established. Here, we examine the contributions of visually coarse (configural) and fine (featural) facial movement information to inversion effects in the perception of visual and audiovisual speech. We describe two experiments in which we disrupted perception of fine facial detail by decreasing spatial frequency (blurring) and disrupted perception of coarse configural information by facial inversion. For normal, unblurred talking faces, facial inversion had no influence on visual speech identification or on the effects of congruent or incongruent visual speech movements on perception of auditory speech. However, for blurred faces, facial inversion reduced identification of unimodal visual speech and effects of visual speech on perception of congruent and incongruent auditory speech. These effects were more pronounced for words whose appearance may be defined by fine featural detail. Implications for the nature of inversion effects in visual and audiovisual speech are discussed.
A major issue in the study of word perception concerns the nature (perceptual or nonperceptual) of sentence context effects. The authors compared effects of legal, word replacement, nonword replacement, and transposed contexts on target word performance using the Reicher-Wheeler task to suppress nonperceptual influences of contextual and lexical constraint. Experiment 1 showed superior target word performance for legal (e.g., "it began to flap/flop") over all other contexts and for transposed over word replacement and nonword replacement contexts. Experiment 2 replicated these findings with higher constraint contexts (e.g., "the cellar is dark/dank") and Experiment 3 showed that strong constraint contexts improved performance for congruent (e.g., "born to be wild") but not incongruent (e.g., mild) target words. These findings support the view that the very perception of words can be enhanced when words are presented in legal sentence contexts.
Illumination of only a few key points on a moving human body or face is enough to convey a compelling perception of human motion. A full understanding of the perception of biological motion from point-light displays requires accurate comparison with the perception of motion in normal, fully illuminated versions of the same images. Traditionally, these two types of stimuli (point-light and fully illuminated) have been filmed separately, allowing the introduction of uncontrolled variation across recordings. This is undesirable for accurate comparison of perceptual performance across the two types of display. This article describes simple techniques, using proprietary software, that allow production of point-light and fully illuminated video displays from identical recordings. These techniques are potentially useful for many studies of motion perception, by permitting precise comparison of perceptual performances across point-light displays and their fully illuminated counterparts with accuracy and comparative ease.
Research has shown that auditory speech recognition is influenced by the appearance of a talker's face, but the actual nature of this visual information has yet to be established. Here, we report three experiments that investigated visual and audiovisual speech recognition using color, gray-scale, and point-light talking faces (which allowed comparison with the influence of isolated kinematic information). Auditory and visual forms of the syllables /ba/, /bi/, /ga/, /gi/, /va/, and /vi/ were used to produce auditory, visual, congruent, and incongruent audiovisual speech stimuli. Visual speech identification and visual influences on identifying the auditory components of congruent and incongruent audiovisual speech were identical for color and gray-scale faces and were much greater than for point-light faces. These results indicate that luminance, rather than color, underlies visual and audiovisual speech perception and that this information is more than the kinematic information provided by point-light faces. Implications for processing visual and audiovisual speech are discussed.
Previous investigations of how ERP repetition effects are modulated by stimulus content have typically used words or non-words as stimuli. The extent to which these findings can be generalized to pictorial stimuli was examined in the present study. ERPs were recorded while participants viewed a series of pictures repeated over 5-7 intervening items. One set of pictures comprised colored depictions of a variety of objects and scenes, while other similar pictures were distorted to appear as abstract patterns of colored blocks. Repeated items were characterized by two main modulations of the ERP. An early positivity for second presentations as compared to first presentations appeared at around 200-300 ms, and was present for normal and distorted picture stimuli. A later, more sustained positivity emerged for the second presentations of normal pictures only. This later component had an onset latency of about 300 ms, lasting for some 400 ms, The results suggest that previously observed modulations of ERP repetition effects reflect the operation of processes which are not specific to verbal memory but may be observed for a wider range of visual stimuli.