Earlier multi-speaker, multi-language measurements found that in purposeful fluent speech, total voice duration was half that of the whole – a maximum Shannon entropy criterion for two state signals. The present aim is to discover whether other species’ vocalisations are similarly associated with maximum timing entropy control. Data selection has been based on the use of courtship songs since they are purposeful and fluent. These songs are intended for con-specific females and, at present, often too complex for reliable acoustic signal analysis. In consequence it has been necessary, for this preliminary study, to find courtship songs in which the echemes are well demarcated from silence – this has led to a curtailment of potential signal resources. The vocalisation timings in examples of the courtship songs of Nightingale, Parotid Wasp and Hawaiian Planthopper have been analysed. The resulting measurements are compared with those for voice timing in purposeful fluent, human speech. There are three primary findings: ### Competing Interest Statement The authors have declared no competing interest.
Current approaches to voice diagnosis involve a clinician examining the patient, listening to their voice and in some cases, using additional measurements of the larynx such as EGG. Here we train a feedforward convolutional neural network on a database of normal healthy drama students recorded speaking passages in English, to reconstruct the associated EGG (Lx) waveform. We then use the network to predict the EGG from the acoustic speech signal on a different set of speakers, including ones that exhibit laryngeal pathologies. We show the predicted EGG is very similar to the actual recorded EGG and, as such, can provide a useful indication of voice pathology. Importantly, the network is able predict the pathological EGG waveforms even though it was never trained on pathological speech.
Die vorliegende Studie untersuchte die Rolle von stimmhaften Intervallen (d.h. Intervalle laryngaler Aktivität) rhythmische Charakteristika im Sprachsignal zu kodieren. Die Dauercharakteristika stimmhafter und stimmloser intervalle (%VO, deltaUV, VarcoUV, VarcoVO, n-PVI_VO, r-PVI_UV) wurden analysiert. Aufgrund der untersuchten Sprachen konnten wir zeigen, dass stimmhafte Dauercharakteristika effektiv zu einer Klassifizierung von Sprachen führen, die einer auditorischen Klassifizierung der Sprachen in Rhythmusklassen (akzentzählend, silbenzählend) entspricht. Weiterhin fanden wir Variation zwischen den Sprechern einer Sprache (Deutsch). Wir argumentieren, dass unsere Methode direkt verwandt mit der möglicherweise auditiv hervortretensten Komponente der menschlichen Stimme (das Stimmsignal) ist. Methodische Vorteile sind, dass die stimmlichen Dauercharakteristika verlässlich automatisch aufgrund des Stimmsignals berechnet werden können. Implikationen unserer Befunde zum Erwerb prosodischer Phänomene und zur Wahrnehmung von Sprache durch Neugeborene werden diskutiert.
Introduction Speech perception requires a receiver to make decisions both about trends and also about the language patterns of a speaker. Hearing people use the speaker’s speech signal as the direct sensory evidence. In the special case of visual speech perception, otherwise known as speechreading, this sensory evidence is derived from the visible articulation movements of speech. Unfortunately, many important articulation movements are invisible. On the segmental level, some articulation positions are difficult or impossible to distinguish. Even for those sounds that are visible, few have unique visual cues. At the suprasegmental level, the basic speech elements intonation and lexical tone which are mainly conveyed by the vibrating frequencies of the larynx, are totally missing in visible facial gestures. Therefore, a successful speechreader has to compensate for the resulting shortage of sensory data by taking maximal advantage of prior knowledge of language structure and of the linguistic and situational context. This process is demanding, however, and even the most competent speechreaders have to accept a substantial probability of error.
Although voice disorder is ordinarily first detected by listening, hearing is little used in voice measurement. Auditory critical band approaches to the quantitative analysis of dysphonia are compared with the results of applying cycle-by-cycle time based methods and the results from a listening test. The comparisons show that quite large rough/smooth differences, that are readily perceptible, are not as robustly measurable using either peripheral human hearing based GammaTone spectrograms, or a cepstral prominence algorithm, as they may be when using cycle-by-cycle based computations that are linked to temporal criteria. The implications of these tentative observations are discussed for the development of clinically relevant analyses of pathological voice signals with special reference to the analytic advantages of employing appropriate auditory criteria.
This brief study has two essential aims. First, it is directed toward the measurement of changes in voice control that may be consequent on the overnight deactivation of cochlear implants (CIs) by individual young children in a residential school for the deaf. Second, the work is based on the exploratory use of a set of voice analytic procedures that, although developed in the first instance for work on connected speech with hearing-impaired children, have subsequently been applied extensively in voice clinic environments. Acoustic and electrolaryngograph speech recordings have been made and analyzed for a group of children with CIs, early in the morning with acoustic and CI aids switched off and at the end of a normal day's use. Special attention has been paid to the analysis of perceptually relevant physical aspects of pitch, intonation, and voice quality. Differences in voice control between these conditions of implant use have been found for all of the children.
Voice is a dominant component of everyday speech in all languages. The possibility is examined that its use may have evolved so that its timing in connected speech is ideal from the point of view of information theory-with voicing taking up 50% of the total speaking time. Initial measurements have been made of voice timing proportions using Laryngograph (R) (EGG) signals as the basis of timing analyses. The results of these analyses for data from two groups of speakers are reported: single native speakers of each of 8 different languages; and 56 speakers of British English. The average 51% and 52% voice timing proportions that were found closely approximate the ideal of 50%. Implications of this finding for voice evolution are briefly discussed.
This study investigated whether congenital amusia, a neuro-developmental disorder of musical perception, also has implications for speech intonation processing. In total, 16 British amusics and 16 matched controls completed five intonation perception tasks and two pitch threshold tasks. Compared with controls, amusics showed impaired performance on discrimination, identification and imitation of statements and questions that were characterized primarily by pitch direction differences in the final word. This intonation-processing deficit in amusia was largely associated with a psychophysical pitch direction discrimination deficit. These findings suggest that amusia impacts upon one's language abilities in subtle ways, and support previous evidence that pitch processing in language and music involves shared mechanisms.
Speech rhythm classes can be distinguished acoustically and perceptually from the variability of consonantal durations and the relative durations of vocalic intervals. The present research investigated whether this distinction can be made robustly, simply on the basis of voice timing, by measuring the durational characteristics of voiced and voiceless intervals in fluent speech. We show that voice patterns — in terms of vocal fold vibration — provide an effective basis for classification and that they can be automatically processed for large datasets. The possible implications that this finding can have on the ability of infants to distinguish between languages of different rhythmic classes are discussed.
Applications of the use of connected speech material for the objective assessment of two primary physical aspects of voice quality are described and discussed. Simple auditory perceptual criteria are employed to guide the choice of analysis parameters for the physical correlate of pitch, and their utility is investigated by the measurement of the characteristics of particular examples of the normal-speaking voice. This approach is extended to the measurement of vocal fold contact phase control in connected speech and both techniques are applied to pathological voice data.
Quantitative clinical voice analysis is discussed with special reference to four factors: 1) measurement criteria that are based on well established auditory parameters; 2) voice material that is modelled on the connected speech of ordinary spoken communication rather than sustained vowels; 3) direct monitoring so as to provide both acoustic and vocal fold contact signals; and 4) phonetic structural similarities across what are ordinarily regarded as highly dissimilar languages. These factors have motivated the development and clinical application of physical analyses that provide measurements related both to vocal fold function and to the perceptual attributes of pitch, loudness, and an important aspect of voice quality.
It has been demonstrated that speech rhythm classes (e.g. stress-timed, syllable-timed) can be distinguished acoustically and perceptually on the basis of the variability of consonantal and vocalic interval durations. It has moreover been shown that even infants are able to use these cues to distinguish between languages from different rhythm classes. Here we demonstrate that the same classification is possible in the acoustic domain based simply on the durational variability of voiced and voiceless intervals in speech. The advantages of such a procedure will be discussed and we will argue that 'voice' possibly offers a more plausible cue for infants to distinguish between languages of different rhythmic class.
Four examples of the use of vocal fold contact phase measurement are discussed for unilateral paresis. In each case this aspect of voice quality is of greater importance than the physical measurement of loudness and pitch related parameters. For three of the cases electro-stimulation has been used as a main part of the treatment. Phonation in both connected speech and, for comparison, in sustained sound production has been used with electro-laryngograph / egg signals providing the basis for measurement. The main new descriptors that have been found to be useful relate to: vocal fold closure and closure duration regularities and distributions; but reference is also made to related measures of peak acoustic amplitude. The new measures described give, in some cases, quite striking results that are of auditory significance and potentially of clinical value.
Isabel Trancoso合作论文数Instituto Superior Tecnico, University of Lisbon2