Rapid advancements in technology have made the acoustic assessment of voice more convenient and less costly; thus, there are few reasons for speech language pathologists (SLPs) not to use acoustic measures to supplement perceptual ratings. Smartphones have been found to be comparable to external microphones in recording quality, and free software programs are available to download into computers to obtain acoustic analyses results. Suggestions on a protocol for capturing and analyzing voice signals using smartphones and computer freeware or smartphone applications will be provided.
This chapter considers at what stage of planning Foreign Accent Syndrome (FAS) takes its toll in speech production. Consideration of the scientific literature leads to the conclusion that FAS is a post-phonological process that has an effect at the stage of phonetic outputting in the speech production planning.
Objectives. Finding measures that track disease progression and determine treatment efficacy is vital for appropriate management in Friedreich ataxia (FA). The purpose of this study was to determine which cepstral- and spectral-based measures extracted from prolonged vowels using Analysis of Dysphonia in Speech and Voice (ADSV) program discriminate between those who have FA and normal voice (NV) peers.Study Design. This is a descriptive, prospective study.Methods. Initial 2 seconds of prolonged /a/, /i/, and /o/ were analyzed through ADSV from 20 individuals diagnosed with FA and 20 NV individuals. ADSV measures used were cepstral peak prominence (CPP), cepstral peak prominence standard deviation (CPP SD), low/high spectral ratio (L/H ratio), low/high spectral ratio standard deviation (L/H ratio SD), and the Cepstral/Spectral Index of Dysphonia (CSID).Results. L/H ratio SD was the only measure where significant differences were found across all vowels between groups. Comparing measures per vowel, the vowel /o/ was significantly different between groups on four of five measures. Discrimination analysis revealed 100% of those in the FA group were classified correctly (sensitivity), whereas 95% of NV members were correctly identified (specificity) when all ADSV measures, with the exception of L/H ratio, were entered.Conclusions. Unstable periods of phonation, such as initiations of voice production in vowels, may yield robust acoustic cues in the FA population. ADSV provides measures that, when considered together, have excellent sensitivity and very good specificity. Vowels yielded differing results on ADSV measures; analysis of different vowel types is recommended.
Recently, a growing number of studies have been published involving phonetic and acoustic analyses on the rare motor-speech disorder known as Foreign Accent Syndrome (FAS). These studies have relied on pre- and post-trauma speech samples to investigate the acoustic and phonetic properties of individual cases of FAS speech. This study presents detailed acoustic analyses of the speech characteristics of two new cases of FAS using identical pre- and post-recovery speech samples, thus affording a new level of control in the study of Foreign Accent Syndrome. Participants include a 48-year-old female who began speaking with an "Eastern European" accent following a traumatic brain injury, and a 45-year-old male who presented with a "British" accent following a subcortical cerebral vascular accident (CVA). The acoustic analysis was based on 18 real words comprised of the stop consonants /p/, /t/, /k/; /b/, /d/, /g/ combined with the peripheral vowels /i/, /a/ and /u/ and ending in a voiceless stop. Computer-based acoustic measures included: (1) voice onset time (VOT), (2) vowel durations, (3) whole word durations, (4) first, second and third formant frequencies, and (5) fundamental frequency. Formant frequencies were measured at three points in the vowel duration: (a) 20%, (b) 50%, and (c) 80% to assess differences in vowel 'onglides' and 'offglides'. The acoustic analysis allowed precise quantification of the major phonetic features associated with the foreign quality of participants' FAS speech. Results indicated post-recovery changes in both duration and frequency measures, including a tendency toward more normal VOT production of voiced stops, changes in average vowel durations, as well as evidence from formant frequency values of vowel backing for both participants. The implications of this study for future research and clinical applications are also considered.
Purpose It is unclear whether the production and perception of speech movements are subserved by the same brain networks. The purpose of this study was to investigate neural recruitment in cortical areas commonly associated with speech production during the production and visual perception of speech. Method This study utilized functional magnetic resonance imaging (fMRI) to assess brain function while participants either imitated or observed speech movements. Results A common neural network was recruited by both tasks. The greatest frontal lobe activity in Broca’s area was triggered not only when producing speech but also when watching speech movements. Relatively less activity was observed in the left anterior insula during both tasks. Conclusion These results support the emerging view that cortical areas involved in the execution of speech movements are also recruited in the perception of the same movements in other speakers.
In the present study, voice onset time (VOT) measurements were compared between a group of individuals with moderate Alzheimer's disease (AD) and a group of healthy age- and gender-matched peers. Participants read a list of consonant-vowel-consonant (CVC) words, which included the six stop consonants. The VOT measurements were made from oscillographic displays obtained from the Brown Laboratory Interactive Speech System (BLISS) implemented on an IBM-compatible computer. VOT measures for the participants' six stop consonant productions were subjected to statistical analysis. The results indicated that VOT values in speakers with Alzheimer's disease were not statistically different from those for the normal control speakers.
A new case of Foreign Accent Syndrome is described. This American woman presented with a British- or Australian- sounding accent after stroke, which resulted in a lacunar infarct in the left internal capsule. The atypical etiology and apparent changes in lexical use are described. It is hypothesized that an abnormally tense vocal tract posture may account for phonetic changes in vowel quality and a higher average fundamental frequency.
In speech perception, there are three predominant models (i.e. top-down, bottom-up or a combination processing model). Recent discussion has focused on use of a combination model by bilingual speakers. Normative information is essential when applied to bilingual participants, in order to clarify what type of language input will best facilitate learning or recovery of specific language abilities. The purpose of this investigation was to study the identification of code-mixed words among Taiwanese–English speaking bilingual individuals. This study posed the following two questions: Is there a difference in the identification of an English or Taiwanese code-mixed word in fluent bilingual speakers (i.e. Taiwanese–English)? Is there a difference in the listener's perception of the code-mixed stimuli according to length of residence (LOR), or to the amount of time a speaker has lived in the United States? Thirty-two Taiwanese–English fluent bilinguals with no reported speech or hearing deficits participated in this experiment. The participants were divided into three subgroups according to LOR (i.e. short LOR, 2–6 years; middle LOR, 7–10 years; and late LOR, 11–20 years). The participants heard sentences with the last word broken up into segments increasing in length (gates). Each targeted word was contained in a sentence presented under four different language conditions. The results of this study indicated that bilingual listeners were able to distinguish differences according to language but there were no significant differences by subgroup or differences in language by group interactions. The listeners were not able to differentiate items with regard to the specific language conditions being heard. However, a profile plot analysis indicated that differences did occur with regard to the language conditions. In addition, a profile of different English language acquisition was noted by the short, middle or long length of residence groups.
In an earlier study (Ryalls, Zipprer and Baldauff, 1997) significant differences in Voice Onset Time (VOT) production were found between younger males and younger females, as well as between African-American and Caucasian-American speakers. In this study we attempted to replicate these significant effects for gender and ethnic background in a group of older speakers. Participants in this study were healthy normal speakers between 50 and 70 years of age and included 10 African-Americans and 10 Caucasian-Americans with 10 males and 10 females. Speakers in this study produced real words with each of the 6 stop consonants of English (/p/, /t/, /k/; /b/, /d/, /g/) combined with the three extreme vowels (/i/, /a/, /u/) in monosyllabic words ending with a voiceless stop. Each of these 18 stimuli were then read at least five times in random order within a repetition and recorded onto Digital Audio Tape (DAT). The Brown University Laboratory Interactive Speech System (BLISS, Mertus, 1999) was used to display waveforms and measure VOT for three of the repetitions for a total of 54 measures per speaker. An Analysis of Variance indicated no significant effects for either ethnic background or gender. However, a comparison with the previous VOT data from younger speakers revealed a significant difference with the present data from older speakers. The results of this study are discussed in the context of past findings in the published literature.
Acoustic measures of (1) syllable duration, (2) voice onset time, (3) fundamental frequency (F0) and (4) the first three formant frequencies (F1, F2, F3) of vowels were provided for a speech corpus produced by 10 French-speaking children (five boys, five girls, with an average age of 9;4 years) with profound hearing impairment (i.e. pure tone averages of more than 100 dB HL.) These data were then compared with previously published data for children with moderate-to-severe hearing impairment (i.e. pure tone averages of more than 60 dB HL) and normally hearing children. The children with profound hearing loss were significantly different in all measures from peers with moderate-to-severe hearing loss and normally hearing peers, while children with moderate-to-severe hearing impairment were not significantly different from their normally hearing peers.
The purpose of this study was to determine whether voice onset time (VOT) values of persons with dysphagia differed from those of a person with normal swallow function. Five male subjects with dysphagia (average age = 80.6 years) and a control subject (age = 79 years) read 18 consonant–vowel–consonant words in quasi-random order. These syllables began with the voiced and voiceless cognates from the three stop places of articulation (i.e., bilabial, alveolar, and velar). These consonants were followed by the vowels /i/, /a/, and /u/. Digital audio tape recordings were performed and speech was digitized onto disk. Measurements were completed using BLISS software (Mertus J: BLISS User's Manual. Providence: Department of Cognitive and Linguistic Sciences, Brown University, 1989) implemented on a 486 microcomputer. Averages and standard deviations of the VOT measures for the six stop consonants were compared between the two experimental groups. For the dysphagic speakers, average VOT values for voiceless stops were shorter, and there were larger negative VOT values for voiced stops. Standard deviations for the VOT productions pf the dysphagic subjects were smaller. Statistical comparisons showed significant differences between individual dysphagic speakers and the normal control for three of the five subjects. These preliminary data suggest that dysphagia affects the fine motor control required for accurate VOT production in speech.