Human languages have the flexibility to be acoustically adapted to the context of communication, such as in shouting or whispering. Drummed forms of languages represent one of the most extreme natural expressions of such speech adaptability. A large amount of research has been conducted on drummed languages in anthropology or linguistics, particularly in West African societies. However, in spite of the clearly rhythmic nature of drumming, previous studies have largely neglected exploring systematically the role of speech rhythm. Here, we explore a unique corpus of the Bendre drummed speech form of the Mossi people, transcribed published in the 80's by the anthropologist Kawada Junzo. The analysis of this large database in Moore language reveals that the rhythmic units encoded in the length of pauses between drumbeats match more closely with vowel-to-vowel intervals than with syllable parsing. Meanwhile, we confirm for the first time a result found recently on the drummed speech tradition of the Bora Amazonian language. However, the complex acoustic structure of the Bendre skin drum required much more attention than the simple two pitch hollow log drum of the Bora. Thus, we also present here results on how drummed Bendre timbre encodes tones of Moore language.
The present study compares the perceptual categorization of four CV syllables /ta, da, ka, ga/ in two different speech registers -modal speech and whistled speech -of Tashlhiyt Berber used in the Moroccan High Atlas.Whistled speech in a non-tonal language such as Tashlhiyt is a special speech register used for long distance dialogues that consists of the natural production of vocalic and consonantal qualities in a simple modulated whistled signal.The technique of whistling imposes various restrictions on speech articulation, which result in a simplification of the phonetics of spoken speech into a 'whistled formant'.Here, we describe this simplification for Tashlhiyt syllables /ta, da, ka, ga/ and use them as stimuli in a behavioral experiment.We analyze and compare the perceptual categorization obtained from native Tashlhiyt listeners (trained since childhood in whistled speech) for both speech registers on these 4 syllable types.Results show that whistled stimuli were fairly well identified (~42%) above chance (25%), though less well than spoken ones (~84%).The detailed analysis of confusions between CVs enabled us to understand better how whistled consonants are perceived, highlighting the phonological contrasts that are best perceived and retained from spoken to whistled speech in this language.
To increase the range of modal speech in natural ambient noise, individuals increase their vocal effort and may pass into the 'shouted speech' register. To date, most studies concerning the influence of distance on spoken communication in outdoor natural environments have focused on the 'productive side' of the human ability to tacitly adjust vocal output to compensate for acoustic losses due to sound propagation. Our study takes a slightly different path as it is based on an adaptive speech production/perception experiment. The setting was an outdoor natural soundscape (a plane forest in altitude). The stimuli were produced live during the interaction: each speaker adapted speech to transmit French disyllabic words in isolation to an interlocutor/listener who was situated at variable distances in the course of the experiment (30m, 60m, 90m). Speech recognition was explored by evaluating the ability of 16 normal-hearing French listeners to recognize these words and their constituent vowels and consonants. Results showed that in such conditions, speech adaptation was rather efficient as word recognition remained around 95% at 30m, 85% at 60m and 75% at 90m. We also observed striking differences in patterns of answers along several lines: different distances, speech registers, vowels and consonants.
Many drum communication systems around the world transmit information by emulating tonal and rhythmic patterns of spoken languages in sequences of drumbeats. Their rhythmic characteristics, in particular, have not been systematically studied so far, although understanding them represents a rare occasion for providing an original insight into the basic units of speech rhythm as selected by natural speech practices directly based on beats. Here, we analyse a corpus of Bora drum communication from the northwest Amazon, which is nowadays endangered with extinction. We show that four rhythmic units are encoded in the length of pauses between beats. We argue that these units correspond to vowel-to-vowel intervals with different numbers of consonants and vowel lengths. By contrast, aligning beats with syllables, mora or only vowel length yields inconsistent results. Moreover, we also show that Bora drummed messages conventionally select rhythmically distinct markers to further distinguish words. The two phonological tones represented in drummed speech encode only few lexical contrasts. Rhythm thus appears to crucially contribute to the intelligibility of drummed Bora. Our study provides novel evidence for the role of rhythmic structures composed of vowel-to-vowel intervals in the complex puzzle concerning the redundancy and distinctiveness of acoustic features embedded in speech.
Whistled speech in a non-tonal language consists of the natural emulation of vocalic and consonantal qualities in a simple modulated whistled signal. This special speech register represents a natural telecommunication system that enables high levels of sentence intelligibility by trained speakers and is not directly intelligible to naïve listeners. Yet, it is easily learned by speakers of the language that is being whistled, as attested by the current efforts of the revitalization of whistled Spanish in the Canary Islands. To better understand the relation between whistled and spoken speech perception, we look herein at how Spanish, French, and Standard Chinese native speakers, knowing nothing about whistled speech, categorized four Spanish whistled vowels. The results show that the listeners categorized differently depending on their native language. The Standard Chinese speakers demonstrated the worst performance on this task but were still able to associate a tonal whistle to vowel categories. Spanish speakers were the most accurate, and both Spanish and French participants were able to categorize the four vowels, although not as accurately as an expert whistler. These results attest that whistled speech can be used as a natural laboratory to test the perceptual processes of language.
Whistled speech in a non tonal language consists of the natural emulation of vocalic and consonantal qualities in a simple modulated whistled signal. This special speech register represents a natural telecommunication system that enables high levels of sentence intelligibility by trained speakers. It is not directly intelligible to naive listeners. Yet, it is easily learned by speakers of the language that is being whistled, as attested by current efforts of revitalization of whistled Spanish in the Canary Islands. To understand better the relation between whistled and spoken speech perception, we looked here at how Spanish native speakers knowing nothing about whistled speech categorized four Spanish whistled vowels. The results show that naive participants were able to categorize these vowels, although not as accurately as a native whistler.
Listening abilities in humans have developed in rural environments which are the dominant setting for the vast majority of human evolution. Hence, the natural acoustic constraints present in such ecological soundscapes are important to take into account in order to study human speech. Here, we measured the impact of basic properties of a typical 'natural quiet' and non reverberant soundscape on speech recognition. A behavioural experiment was implemented to analyze the intelligibility loss in spoken word lists with variations of Signal-to-Noise Ratio corresponding to different speaker-to-listener distances in a typical low-level natural background noise recorded in a plain dirt open field. To highlight clearly the impact of such noise on recognition in spite of its low level, we contrasted the 'noise + distance' condition with a 'distance only' condition. The recognition performance for vowels and consonants and for different classes of consonants is also analyzed.
The lightning detection performance of the very low frequency VLF long-range Sferics Tracking and Ranging Network STARNET was evaluated by comparison with a simultaneous data collection made by the Sistema de Proteção da Amazônia Lightning Detection Network SIPAM-LDN. The study period was 110 days between August 2008 and March 2009, corresponding to a stable period of the STARNET network when it was operating with five sensors to cover the eastern Amazon region. The selected area of study corresponded to a circle of 130 km diameter within a homogeneous zone of detection efficiency DE of the SIPAM network. A method of coinciding flash identification time window of 1 ms and spatial range of 50 km was used for these two networks. The coinciding cloud-to-ground CG flashes were discriminated by polarity and peak current values from the SIPAM-LDN measurements. The total number of coincident CG flashes represented about 9.7% of the CG SIPAM data set. Moreover, 94 % of the coinciding CG flashes had a time error <300 µs. The spatial error of the coincident CG flashes yielded a mean of around 16 km. The final result shows that the relative detection efficiency RDE of the coincident CG flashes decreased with the value of their peak currents. RDE values were below 10% for peak currents lower than 20 kA, between 10% and 30% for peak currents between 20 and 40 kA, and above 30% for peak currents greater than 40 kA.
In the real world, human speech recognition nearly always involves listening in background noise. The impact of such noise on speech signals and on intelligibility performance increases with the separation of the listener from the speaker. The present behavioral experiment provides an overview of the effects of such acoustic disturbances on speech perception in conditions approaching ecologically valid contexts. We analysed the intelligibility loss in spoken word lists with increasing listener-to-speaker distance in a typical low-level natural background noise. The noise was combined with the simple spherical amplitude attenuation due to distance, basically changing the signal-to-noise ratio (SNR). Therefore, our study draws attention to some of the most basic environmental constraints that have pervaded spoken communication throughout human history. We evaluated the ability of native French participants to recognize French monosyllabic words (spoken at 65.3 dB(A), reference at 1 meter) at distances between 11 to 33 meters, which corresponded to the SNRs most revealing of the progressive effect of the selected natural noise (-8.8 dB to -18.4 dB). Our results showed that in such conditions, identity of vowels is mostly preserved, with the striking peculiarity of the absence of confusion in vowels. The results also confirmed the functional role of consonants during lexical identification. The extensive analysis of recognition scores, confusion patterns and associated acoustic cues revealed that sonorant, sibilant and burst properties were the most important parameters influencing phoneme recognition. . Altogether these analyses allowed us to extract a resistance scale from consonant recognition scores. We also identified specific perceptual consonant confusion groups depending of the place in the words (onset vs. coda). Finally our data suggested that listeners may access some acoustic cues of the CV transition, opening interesting perspectives for future studies.
This study presents a new methodology adapted to the analysis of word rhythmic cues in drummed forms of languages. The semi-automatic beat detection procedure applied to the Manguare drummed form of the Bora language enabled us to measure inter-beat durations. These were found to correspond to Vowel-to-Vowel intervals (V-to-V) of the associated speech utterances and to differ as a function of the vowel duration and of the presence/absence of consonant(s) in the V-to-V cluster.
Whistled speech consists of a phonetic emulation of the sounds produced in spoken voice. This style of speech is the result of the adaptation of the human productive and perceptive intelligence to a language behavior. In the typology of whistled forms of languages, Spanish is among the languages for which the whistled strategy emulates primarily segmental acoustic cues of vowels and consonants. The present study tests the perception of four Spanish whistled vowels by French non-whistlers. The results show that French non-whistlers were able to categorize these vowels without any learning, although not as accurately as native whistlers.