In situations of telecommunication and proximal artistic performance, human groups have engineered diverse, ingenious formats of non-voiced auxiliary speech: (i) modulating the vocal tract to enhance selected acoustic features of a sound source alternative to the vocal cords; or (ii) adapting musical instruments to simulate aspects of the spoken phonetic signal. Both types of ‘surrogate language’ (Nketia 1971) employ techniques of ‘speech abridgment’ (Stern 1957), and coin novel ‘acoustic icons’ (Sebeok and Umiker-Sebeok 1976). Auxiliary speech phenomena attest non-literate traditions of text transmission and implicit mental capacities of natural language processing. This chapter provides a comparative view of these traditional special speech types worldwide. It shows that their various expressions depend on (a) the acoustic nature of the modified speech form, (b) the grammar of the base language, (c) the expertise of the performers, and (d) but also whether they imitate spoken dialogues, declamative verbal arts, or musical sung forms of speech. Case studies explore more details in the second part of the chapter.
The study of the whistled speech of several languages of the world has already provided an alternative point of view on many aspects of language. Whistles, which convey the linguistic information in the equivalent of spoken speech, are remarkably adapted to the perceptive capacities of human beings and to the natural environment in which they have helped social communication in everyday life. Whistled languages seem to have naturally focused on key elements of language to transmit the essential part of linguistic information. They can be seen as phonetical descriptions of local languages. As they are strongly linked to the traditional way of life, the great majority of them is endangered. Despite this situation, new approaches are now developed in order to combine the international scientific study and the local cultural transmission of these particular languages.
Deemed one of the world’s most representative whistled languages, the Canary Islands’ whistled Spanish, locally known as Silbo, has long attracted linguistic research. However, most studies have adopted linguistic, ethnological, or bioacoustic perspectives, overlooking the potential of computational methods within the digital humanities. This work advances the computational study of Silbo by presenting the first automated approach to Speaker Identification (SI)—i.e., the process of determining the speaker of a given utterance by computational means—in a closed-set configuration for this language. The proposal leverages standard feature extraction methods as well as pre-trained Speech Recognition models to extract representative embeddings and incorporates class-balancing mechanisms to mitigate biases arising from uneven representation of whistlers in the data—i.e., label imbalance. The results obtained on the only existing dataset specifically designed for computational analysis of Silbo, comparing three representative feature extraction methods, three oversampling policies, and five classification strategies, validate the proposal, achieving F _1 scores close to 90
In this study, we investigated the transfer of musical skills to speech perception by analyzing the perception and categorization of consonants produced in whistled speech, a naturally modified speech form. The study had two main objectives: (i) to explore the effects of different levels of musical skill on speech perception, and (ii) to better understand the type of skills transferred by focusing on a group of high-level musicians, playing various instruments. Within this high-level group, we aimed to disentangle general cognitive transfers from soundspecific transfers by considering instrument specialization, contrasting general musical knowledge (shared by all instruments) with instrument-specific ones. We focused on four instruments: voice, violin, piano and flute. Our results confirm a general musical advantage and suggest that only a small amount of musical experience is sufficient for musical skills to benefit whistled speech perception. However, higher-level musicians reached better performances, with differences for specific consonants. Moreover, musical expertise appears to enhance rapid adaptation to the whistled signal throughout the experiment and our results highlight the specificity of instrument expertise. Consistent with previous research showing the impact of the instrument played, the differences observed in whistled speech processing among high-level musicians seem to be primarily due to instrument-specific expertise.
In this paper, we looked at the impact of musical experience on whistled vowel categorization by native French speakers. Whistled speech, a natural, yet modified speech type, augments speech amplitude while transposing the signal to a range of fairly high frequencies, i.e. 1 to 4 kHz. The whistled vowels are simple pitches of different heights depending on the vowel position, and generally represent the most stable part of the signal, just as in modal speech. They are modulated by consonant coarticulation(s), resulting in characteristic pitch movements. This change in speech mode can liken the speech signal to musical notes and their modulations; however, the mechanisms used to categorize whistled phonemes rely on abstract phonological knowledge and representation. Here we explore the impact of musical expertise on such a process by focusing on four whistled vowels (/i, e, a, o/) which have been used in previous experiments with non-musicians. We also included inter-speaker production variations, adding variability to the vowel pitches. Our results showed that all participants categorize whistled vowels well over chance, with musicians showing advantages for the middle whistled vowels (/a/ and /e/) as well as for the lower whistled vowel /o/. The whistler variability also affects musicians more than nonmusicians and impacts their advantage, notably for the vowels /e/ and /o/. However, we find no specific training advantage for musicians over the whole experiment, but rather training effects for /a/ and /e/ when taking into account all participants. This suggests that though musical experience may help structure the vowel hierarchy when the whistler has a larger range, this advantage cannot be generalized when listening to another whistler. Thus, the transfer of musical knowledge present in this task only influences certain aspects of speech perception.
In this paper, we explore the effect of musical expertise on whistled word perception by naive listeners. In whistled words of nontonal languages, vowels are transposed to relatively stable pitches, while consonants are translated into pitch movements or interruptions. Previous behavioral studies have demonstrated that naive listeners can categorize isolated consonants, vowels, and words well over chance. Here, we take an interest in the effect of musical experience on words while focusing on specific phonemes within the context of the word. We consider the role of phoneme position and type and compare the way in which these whistled consonants and vowels contribute to word recognition. Musical experience shows a significant and increasing advantage according to the musical level achieved, which, when further specified according to vowels and consonants, shows stronger advantages for vowels over consonants for all participants with musical experience, and advantages for high-level musicians over nonmusicians for both consonants and vowels. By specifying high-level musician skill according to one's musical instrument expertise (piano, violin, flute, or singing), and comparing these instrument groups to expert users of whistled speech, we observe instrument-specific profiles in the answer patterns. The differentiation of such profiles underlines a resounding advantage for expert whistlers, as well as the role of instrument specificity when considering skills transferred from music to speech. These profiles also highlight differences in phoneme correspondence rates due to the context of the word, especially impacting "acute" consonants (/s/ and /t/), and highlighting the robustness of /i/ and /o/.
In this paper, we explore whistled word perception by naive French speakers.In whistled words of non-tonal languages, vowels are transposed to relatively stable pitches, which contrast with consonant movements or interruptions.Previous studies on whistled speech with naive listeners have tested vowels and consonants separately.Other studies on spoken word recognition have found that vowels and consonants contribute differently to intelligibility, where the role of vowels was highly mediated by the context.Here, naive participants recognize disyllabic whistled words above chance, and vowels are shown to contribute differently than consonants.When focusing on the role of vowels, we found different scales of performance between the vowels tested, mediated by their position in the word.We also highlighted the importance of the vowels' relative frequency difference (called 'interval') in the word.
We explore whistled vowel categorization by untrained listeners, focusing specifically on the impact of the different vocalic frequency ranges of two whistlers (for the vowels /i/, /e/, /a/, /o/) and the effect of training on performance. In the experiment, we included stimuli that show inter-individual and intra-individual variations of production. In the analyses, we looked at the whistler identity effect and at the learning effect throughout the experiment for the studied vowels. The results showed an effect of the whistler, where the larger vocalic range led to improved categorization, and highlighted the robustness of the vowel recognition hierarchy. There was no general learning effect, albeit for one vowel and for the whistler with a narrower vocalic range. This study provides insight into representations of the vowel space in non-tonal languages.
Dans cette étude nous avons cherché à comprendre l'effet de la pratique instrumentale sur la perception et la catégorisation de la parole sifflée.Nous nous sommes intéressés à la spécificité instrumentale avec une focalisation sur 4 instruments : la voix, le violon, le piano et la flûte.Bien que le bénéfice de la pratique musicale sur la perception de la parole modifiée soit vérifié dans nos résultats, il apparait clairement que l'instrument pratiqué ainsi que le niveau de pratique ont un effet sur la perception de la parole sifflée.Ces résultats suggèrent que, lors de ce processus de catégorisation, les effets observés s'expliquent plus par un traitement modifié du signal sonore, grâce à une familiarisation acoustique spécifique chez les musiciens expérimentés, plutôt que par des fonctions générales (fonctions exécutives, mémoire ou attention) plus performantes.
Whistled speech is a form of modified speech where, in non-tonal languages, vowels and consonants are augmented and transposed to whistled frequencies, simplifying their timbre. According to previous studies, these transformations maintain some level of vowel recognition for naive listeners. Here, in a behavioral experiment, naive listeners' capacities for the categorization of four whistled consonants (/p/, /k/, /t/, and /s/) were analyzed. Results show patterns of correct responses and confusions that provide new insights into whistled speech perception, highlighting the importance of frequency modulation cues, transposed from phoneme formants, as well as the perceptual flexibility in processing these cues.
Humans use whistled communications, the most elaborate of which are commonly called “whistled languages” or “whistled speech” because they consist of a natural type of speech. The principle of whistled speech is straightforward: people articulate words while whistling and thereby transform spoken utterances by simplifying them, syllable by syllable, into whistled melodies. One of the most striking aspects of this whistled transformation of words is that it remains intelligible to trained speakers, despite a reduced acoustic channel to convey meaning. It constitutes a natural traditional means of telecommunication that permits spoken communication at long distances in a large diversity of languages of the world. Historically, birdsong has been used as a model for vocal learning and language. But conversely, human whistled languages can serve as a model for elucidating how information may be encoded in dolphin whistle communication. In this paper, we elucidate the reasons why human whistled speech and dolphin whistles are interesting to compare. Both are characterized by similar acoustic parameters and serve a common purpose of long distance communication in natural surroundings in two large brained social species. Moreover, their differences – e.g., how they are produced, the dynamics of the whistles, and the types of information they convey – are not barriers to such a comparison. On the contrary, by exploring the structure and attributes found across human whistle languages, we highlight that they can provide an important model as to how complex information is and can be encoded in what appears at first sight to be simple whistled modulated signals. Observing details, such as processes of segmentation and coarticulation, in whistled speech can serve to advance and inform the development of new approaches for the analysis of whistle repertoires of dolphins, and eventually other species. Human whistled languages and dolphin whistles could serve as complementary test benches for the development of new methodologies and algorithms for decoding whistled communication signals by providing new perspectives on how information may be encoded structurally and organizationally.
Whistled forms of languages are distributed worldwide and survive only in some of the most remote villages on the planet. They are not limited to a given continent, language family, or language structure, but they have been detected only sporadically by researchers and travelers, partly because they can be taken for nonlinguistic phenomena, such as simple signaling. Whistled speech consists of speaking while whistling to communicate at a long distance. The result is a melody that imitates modal speech and that remains intelligible for the interlocutors. This review proposes a typology of this special, little-known, natural speech type and takes socio-environmental and linguistic aspects into consideration. The amazing potential of this phenomenon to provide an alternative point of view into language diversity and speech offers a unique occasion to revisit human language with original insights embracing the adaptive flexibility that characterizes speech production and perception.
The Gavião, a native Amazonian group in Rondônia, Brazil, use three different traditional musical instruments that they identify as “speaking” ones and that are characterized by a very tight music-lyric relation through similar pitch patterns: a flute (called kotiráp), a pair of mouth bows (iridináp), and three large bamboo clarinets (totoráp), played by three different players, each one playing a single-note clarinet. They show in different ways the relation of acoustic iconicity which exists between the words of the songs’ lyrics and the music played on such instruments to “sing” the songs. Linguistic analysis makes it possible to understand the phonetic and phonological nature of the iconicity. The sung speech form, being intermediate between the spoken and the instrumental forms, is useful for both learning and explaining the musical notes. In a language with distinctive tone and length, such as Gavião of Rondônia, the first question about speech that is played by musical instruments is the relation between the melodies and the supersegmental phonology of the corresponding words in sung speech and in modal spoken speech. It is influenced by the phonological possibilities of the spoken form and by the musical possibilities of the instrumental form. The description and analysis of Gavião instrumental speech and song practices are found to be a noteworthy contribution to the typology of instrumental language surrogates associated with a tone language, one that calls for a reexamination of hypotheses about which aspects of the phonological/phonetic structure can be transposed in instrumental speech and how this can be done. The role of this kind of instrumental sung speech is artistic and also practical as it contributes to maintain the oral heritage. Such practice represents a little-studied and threatened cultural heritage of the traditional substratum of the cultures of Amazonia.
Whistled speech is a form of modified speech where some frequencies of vowels and consonants are augmented and transposed to whistling, modifying the timbre and the construction of each phoneme. These transformations cause only some elements of the signal to be intelligible for naive listeners, which, according to previous studies, includes vowel recognition. Here, we analyze naive listeners’ capacities for whistled consonant categorization for four consonants: /p/, /k/, /t/ and /s/ by presenting the findings of two behavioral experiments. Though both experiments measure whistled consonant categorization, we used modified frequencies — lowered with a phase vocoder — of the whistled stimuli in the second experiment to better identify the relative nature of pitch cues employed in this process. Results show that participants obtained approximately 50% of correct responses (when chance is at 25%). These findings show specific consonant preferences for “s” and “t” over “k” and “p”, specifically when stimuli is unmodified. Previous research on whistled consonants systems has often opposed “s” and “t” to “k” and “p”, due to their strong pitch modulations. The preference for these two consonants underlines the importance of these cues in phoneme processing.
Human languages have the flexibility to be acoustically adapted to the context of communication, such as in shouting or whispering. Drummed forms of languages represent one of the most extreme natural expressions of such speech adaptability. A large amount of research has been conducted on drummed languages in anthropology or linguistics, particularly in West African societies. However, in spite of the clearly rhythmic nature of drumming, previous studies have largely neglected exploring systematically the role of speech rhythm. Here, we explore a unique corpus of the Bendre drummed speech form of the Mossi people, transcribed published in the 80's by the anthropologist Kawada Junzo. The analysis of this large database in Moore language reveals that the rhythmic units encoded in the length of pauses between drumbeats match more closely with vowel-to-vowel intervals than with syllable parsing. Meanwhile, we confirm for the first time a result found recently on the drummed speech tradition of the Bora Amazonian language. However, the complex acoustic structure of the Bendre skin drum required much more attention than the simple two pitch hollow log drum of the Bora. Thus, we also present here results on how drummed Bendre timbre encodes tones of Moore language.