This paper presents novel data from a Yorùbá language game called Ẹnà, an iterative affixation game that typically involves copying of vowels and tones onto a dummy syllable. Yorùbá VV sequences are all analyzed as disyllabic in existing literature, yet we find that Ẹnà treats them differently depending on their provenance: underlying VV sequences, those created from pronouns, and those derived through floating tone are treated as a single locus of insertion, as variably are those that share tone, while VV sequences derived through consonant deletion or compounding are generally treated as two separate loci. We argue that the difference indicates that Yorùbá in fact has long vowels, contrary to previous assumptions. We analyze the pattern in Optimality Theory, following Krämer & Vogt’s (2018) analysis of reduplicative language games but adding a reduplicative template and back-copying, and we consider the implications of the pattern and analysis to the study of Yorùbá and the study of language games.
This paper presents novel data from a Yorùbá language game called Ẹnà, an iterative affixation game that typically involves copying of vowels and tones onto a dummy syllable. Yorùbá VV sequences are all analyzed as disyllabic in existing literature, yet we find that Ẹnà treats them differently depending on their provenance: underlying VV sequences, those created from pronouns, and those derived through floating tone are treated as a single locus of insertion, as variably are those that share tone, while VV sequences derived through consonant deletion or compounding are generally treated as two separate loci. We argue that the difference indicates that Yorùbá in fact has long vowels, contrary to previous assumptions. We analyze the pattern in Optimality Theory, following Krämer & Vogt’s (2018) analysis of reduplicative language games but adding a reduplicative template and back-copying, and we consider the implications of the pattern and analysis to the study of Yorùbá and the study of language games.
Smiling is a social signal that can be both seen and heard. Smiling can increase speech amplitude and raise F0 and formants. However, experimental research on the role of larynx height in smiled speech is limited. 21 English speakers (6 M) repeated words in a carrier phrase with a neutral face or while smiling. The participants were recorded with audio, video and laryngeal ultrasound. F0, F1 and F2 were extracted for the duration of target vowels /i/, /u/ and /a/. Ultrasound images of laryngeal position were measured using Optical Flow. The laryngeal and acoustic data were analyzed in R with linear mixed models with smiling condition, timepoint-in-vowel, and gender as fixed effects. There was a significant effect of timepoint-in-vowel for larynx height (raising towards the end) and a smile-timepoint interaction effect (the larynx raised more at the end for smiling condition). Acoustically, smiling led to significantly higher F0 across vowels, and significantly higher F1 and F2 for /a/ but not /i/ or /u/. F2 timepoints were significant for all three vowels (F2 trajectories differed) across smile conditions. Results indicate smiling has a consistent effect on larynx height and variable effect on specific speech sounds.
Abstract Background/Aims: Both music and language impose constraints on fundamental frequency (F0) in sung music. Composers are known to set words of tone languages to music in a way that reflects tone height but fails to include tone contour. This study tests whether choral singers add linguistic tone contour information to an unfamiliar song by examining whether Cantonese singers make use of microtonal variation. Methods: 12 native Cantonese-speaking non-professional choral singers learned and sang a novel song in Cantonese which included a minimal set of the Cantonese tones to probe whether everyday singers add in missing contour information. Results: Cantonese singers add in a rising F0 contour of less than a semitone when singing syllables with lexical rising tones. This microtonal variation is not observed when singing in a lower register. Conclusion: Cantonese singers use microtonal contours to reflect rising contours of Cantonese linguistic tones.
Representations of tones in Cantonese songs have been studied extensively from a compositional point of view. In contrast, the perception of lexical tones in sung Cantonese music has received significantly less attention. Existing research has looked at lexical comprehension, but perceptual acceptability has not been examined. This study compares Cantonese native speakers’ acceptability judgments for short melodies that contain either matches or mismatches between the musical spoken and sung melodies. Results show that native Cantonese listeners prefer musical contours that match the tonal contour of the spoken tones of the language.
The field of articulatory phonetics is concerned primarily with the question of how speech is realized through movements of body structures. While articulatory phoneticians have often described speech sounds using terms that refer to an inventory of body parts (e.g., “tongue,” “lips,” “velum,” etc.), a core challenge of articulatory phonetics is to understand how such structures function and interact to produce speech sounds. A complete understanding of articulatory phonetics thus requires that we define ways of mapping our descriptive terms onto groupings of nerves and muscles that our brains and bodies can use to produce speech movements. This chapter explores some of the ways articulatory phoneticians describe speech sounds, suggesting that a more fully embodied approach can provide novel insights into speech articulation and can help to understand links between speech and other functions of the human
Advances in virtual reality (VR) and avatar technologies have created new platforms for face-to-face communication in which visual speech information is presented through avatars using simulated articulatory movements. These movements are typically generated in real time by algorithmic response to acoustic parameters. While the communicative experience in VR has become increasingly realistic, the visual speech articulations remain intentionally imperfect and focused on synchrony to avoid uncanny valley effects [1]. While considerable previous research has demonstrated that listeners can incorporate visual speech information produced by computer-simulated faces with precise and pre-programmed articulations [2], it is unknown whether perceivers can make use of such underspecified and at times misleading simulated visual cues to speech. The current study investigates whether reliable segmental information can be extracted from visual speech algorithmically-generated through a popular VR platform. We focused on the platform’s most consistent and easily perceived articulator movements: bilabial closure in consonants; and lip rounding, lip spreading, and jaw lowering in vowels (see Figure 1). We report on an experiment using a speech-in-noise task with audiovisual stimuli in two conditions (with articulator movement and without) to ask the following questions: 1) whether the visual information from an avatar improves identification of target words, and 2) whether that visual information improves categorization of the target segment.
Course description: The goal of this course is to provide students with broad training in the nomenclature, theory, and practice of articulatory phonetics. Students in this course will learn about the organs of the vocal tract used in speech production, the muscles controlling them, and their coordination. We will also discuss factors thought to affect speech articulation, including speaker-internal factors (e.g. aerodynamics, coarticulation, speech rate), speaker-external factors (e.g. social information), as well as other factors, such as syllable structure and prosody. Along the way, students will learn about some of the techniques available for quantifying articulation, as well as what methods exist for analyzing those data. Students will design and be assessed on a proposal for novel research related to this topic.
Previous research has shown that the sensation of airflow causes bilabial stop closures to be perceived as aspirated even when paired with silent articulations rather than an acoustic signal [Bicevskis et al. 2016, JASA 140(5): 3531–3539]. However, some evidence suggests that perceivers integrate this cue differently if the silent articulations come from an animated face [Keough et al. 2017, Canadian Acoustics 45(3):176–177] rather than a human one. Participants shifted from a strong initial /ba/ bias to a strong /pa/ bias by the second half of the experiment, suggesting the participants learned to associate the video with the aspirated articulation through experience with the airflow. One explanation for the above findings is methodological: participants saw a single video clip while previous work exposed participants to multiple videos. The current study reports two experiments using a single clip with a human face (originally from Bicevskis et al. 2016). We found no evidence of a bias shift, indicating that the findings reported by Keough et al. are not attributable to the use of a single video. Instead, our findings suggest that aero-tactile cues shift consonant perception regardless of the number of recordings presented as long as the speaking face is human.
Bodomo (1997) describes Dagaare (Gur; Ghana) as having a single low vowel, [a], which is neutral to ATR harmony. This paper presents acoustic data from a study of Dagaare which is inconsistent with this description. A list of sentences was elicited from five native speakers of Dagaare. Each sentence contained in one of four verbal particles situated in one of four contexts: ATR _ ATR, ATR _ RTR, RTR _ ATR, and RTR _ RTR. Formants of the low vowel were measured and compared across contexts. Results showed a substantial, significant difference in F1 values and a smaller but still significant difference in F2 values in contexts where is followed by an ATR word compared to when it is followed by an RTR word. All speakers and all particles showed the same pattern. We conclude that, contrary to previous claims, the Dagaare low vowel is not neutral to harmony, but rather has acoustically distinct variants in RTR versus ATR contexts. Bodomo, A. (1997). The structure of Dagaare. California: CSLI publications. [Funded by SSHRC.]
Previous research on multimodal speech perception with hearing-impaired individuals focused on audiovisual integration with mixed results. Cochlear-implant users integrate audiovisual cues better than perceivers with normal hearing when perceiving congruent [Rouger et al. 2007, PNAS, 104(17), 7295–7300] but not incongruent cross-modal cues [Rouger et al. 2008, Brain Research 1188, 87–99), leading to the suggestion that early auditory exposure is required for typical speech integration processes to develop (Schorr 2005, PNAS, 102(51), 18748–18750). If a deficit of one modality does indeed lead to a deficit in multimodal processing, then hard of hearing perceivers should show different patterns of integration in other modality pairings. The current study builds on research showing that gentle puffs of air on the skin can push individuals with normal hearing to perceive silent bilabial articulations as aspirated. We report on a visual-aerotactile perception task comparing individuals with congenital hearing loss to those with normal hearing. Results indicate that aerotactile information facilitated identification of /pa/ for all participants (p < 0.001) and we found no significant difference between the two groups (normal hearing and congenital hearing loss). This suggests that typical multi-modal speech perception does not require access to all modalities from birth. [Funded by NIH.]
Speech sounds have been shown to adapt quickly but imperfectly under bite block perturbation, supporting opposing acoustic vs. articulatory compensation mechanisms [Gay et al. JASA 69: 802. 1981; Flege et al. JASA 83: 212. 1998]. The present study considers whether lingual bracing may provide insight into these apparently conflicting findings. Tongue bracing against the teeth or palate is a pervasive posture maintained during normal speech [Gick et al. JSLHR. 60:494. 2017]; we aim to test whether the tongue adapts its bracing position rather than adapting each speech movement individually, providing a single, postural parametric mechanism for responding to jaw perturbation. Results of an experiment will be presented in which native English-speaking participants read aloud passages normally and under bite block conditions translating the jaw in forward, backward or lateral directions, and to varying degrees of opening. Coronal ultrasound imaging results will be reported, measuring positions of the lateral tongue for indications of stable bracing postures. Implications of these findings will be discussed for models of speech production. [Funding from NSERC.]
Bodomo (1997) describes intervocalic velar [g] in Dàgáárè as fricative [ɣ]. With 42 tokens of intervocalic [g] from a native speaker of Dàgáárè, we investigated the acoustic and articulatory features of Dàgáárè intervocalic velar [g] using ultrasound images, waveforms, spectrograms, and palatogram. The results of the study suggest that Dàgáárè intervocalic [g] is not a fricative but a velar with strong taplike features, a previously unattested sound in natural language (Ladefoged 1990). Following from this, we conclude that Dàgáárè intervocalic velar [g] is not a fricative but a tap.
This paper provides a novel Optimality Theoretic analysis of the 19th century French secret language Largonji. While Largonji is a reversal game, we show that it is a type not previously described, in which the first onset that is not an /l/ reverses, even if it is not at an edge. Thus, traditional approaches to reversal games, such as cross-anchoring, do not work for Largonji. However, our account does not require direct reference to onsets. Instead, it is based on preservation of moraic structure, combined with alignment of a Largonji-specific prefix. Though suprasegmental faithfulness has been noted previously in language games, the present account implements it in Optimality Theory for the first time. Further, in analyzing the Largonji affix as a prefix that is sometimes realized as an infix, we suggest that Largonji provides additional evidence that language games can reflect cross-linguistic patterns not present in the base language.
Ultrasound overlay videos involve the superposition of ultrasound imaging of the tongue onto facial profile videos in order to serve as instructional materials (Abel et al., 2015). Bliss et al. (2016) used this technique to develop instructional and cultural materials for Indigenous communities by creating custom overlay videos of community members which highlight difficult sound contrasts in the languages for learners. Building on this work, this paper reports on the creation of this type of video for Han, a Dene/Athabaskan language of Eagle, Alaska and Dawson City, Yukon with 6-7 native speakers remaining. In the paper, we explore the challenges behind a new possibility of collecting the ultrasound data recorded for these instructional videos to serve a dual-purpose: instructional/cultural and linguistic/scientific. Some acoustic work has been done on Han (Manker 2012), but never articulatory. As Han is known for its large phonemic inventory, being tied for first as the language with the most affricates and containing a 5-6 way contrast in the coronal region, articulatory work on the language is of interest for phonetic and phonological theory. However, ultrasound work has quite strict methodological standards, which can be impractical to include in many field situations, such as precise head and probe stabilization, and fully controlled phonological environments and speaker groups. These standards by necessity and design could not be fully adhered to for the Han recordings. But, despite the methodological limitations involved in dual-purpose fieldwork, we argue that it is important to consider the possibility of drawing linguistic insights from data collected for instructional purposes. Otherwise, these insights simply wouldn’t exist as work with these communities is limited. References J. Abel, B. Allen, S. Burton, M. Kazama, M. Noguchi, A. Tsuda, N. Yamane, and B. Gick. Ultrasound-Enhanced Multimodal Approaches to Pronunciation Teaching and Learning. Canadian Acoustics , 43 (3). 124-125, 2015. Bliss, H., Burton, S., and Gick, B. (2016). Ultrasound Overlay Videos and Their Application in Indigenous Langauge Learning and Revitalization. Canadian Acoustics, 44 (3). Manker, Jonathan. (2012). An Acoustic Study of Stem Prominence in Han Athabaskan. Master’s thesis, University of Alaska Fairbanks.
Lateral bracing refers to contact of the sides of the tongue along the upper molars or palate; evidence from articulatory analysis of native English speakers as well as 3D biomechanical simulations suggests that bracing involves mechanical support which occurs consistently throughout speech [Gick et al. 2017. J Speech Lang Hear Res. 60(3):494-506]. Release of lateral bracing occurs only during some lateral consonants and low vowels. The current study tests for the presence of active lateral bracing in seven languages: Cantonese, English, Korean, Mandarin, Portuguese, Spanish, and Turkish. Ten native speakers of these languages (2 each for English, Mandarin and Korean and one each for the other languages) read aloud passages of the North Wind and the Sun [Handbook of the IPA, 1999] while a coronal ultrasound video of their tongue was recorded. Tracings were made from still images of the M-mode ultrasound videos, and measurements of the vertical motion of the tongue midline and both edges were taken. The percentage of time the tongue is not laterally braced was calculated. Active lateral bracing is implicated if the left and right edges of the tongue are less variable in vertical motion than midline and/or positioned at a stable baseline height for a larger percentage of time than they are lowered. Preliminary analysis supports the hypothesis that tongue bracing in speech exists regardless of language.
: Music has played a significant role in the study and analysis of lexical tone. This paper presents a brief summary of some of the ways scholars have used and adapted conventions of Western musical notation to further their study of tone. It also examines the possible influence of cultural perceptions of melody on Western perspectives of tone.
Recent research has shown that aero-tactile cues influence speech perception without the presence of an acoustic signal (Bicevskis, Derrick & Gick, 2016); when participants viewed a bilabial articulation that co-occurred with a puff of air felt on the skin, they were significantly more likely to perceive it as aspirated. These results and others (Gick & Derrick, 2009, etc.) suggest that this integration is relatively automatic, enough so that it does not require the physical presence of the source to arise. However, it may be that perceivers are willing to extend physical capabilities to these non-present sources because they are human and therefore possible sources of the aero-tactile cue. The current study examines whether aero-tactile information from an impossible source—a computer-animated face on a computer monitor—can affect perception of aspirated consonants. Sixteen native English speakers are shown an animated video of a computer-animated head performing a bilabial plosive but hear only babble noise through headphones. Some of the presentations are accompanied by a light, synchronous puff of air on the neck. They are asked to identify the syllable as either /ba/ or /pa/. Analysis of this two-alternative forced choice response task will be presented. Evidence of integration from an impossible source would support the idea that visual-tactile integration is an automatic process that occurs even in the absence of an interlocutor capable of producing the stimuli.