PURPOSE:We studied the role of gender in metacognition of voice emotion recognition ability (ERA), reflected by self-rated confidence (SRC). To this end, we guided our study in two approaches: first, by examining the role of gender in voice ERA and SRC independently and second, by looking for gender effects on the ERA association with SRC. METHOD:We asked 100 participants (50 men, 50 women) to interpret a set of vocal expressions portrayed by 30 actors (16 men, 14 women) as defined by their emotional meaning. Targets were 180 repetitive lexical sentences articulated in congruent emotional voices (anger, sadness, surprise, happiness, fear) and neutral expressions. Trial by trial, the participants were assigned retrospective SRC based on their emotional recognition performance. RESULTS:A binomial generalized linear mixed model (GLMM) estimating ERA accuracy revealed a significant gender effect, with women encoders (speakers) yielding higher accuracy levels than men. There was no significant effect of the decoder's (listener's) gender. A second GLMM estimating SRC found a significant effect of encoder and decoder genders, with women outperforming men. Gamma correlations were significantly greater than zero for women and men decoders. CONCLUSIONS:In spite of varying interpretations of gender in each independent rating (ERA and SRC), our results suggest that both men and women decoders were accurate in their metacognition regarding voice emotion recognition. Further research is needed to study how individuals of both genders use metacognitive knowledge in their emotional recognition and whether and how such knowledge contributes to effective social communication.
The effect of bilingualism on verbal learning and memory was explored in different studies. Different researchers assume that the Arabic diglossia, represents a case of bilingualism in the lingual context. Hence, the current study aimed to investigate the impact of diglossia in Arabic on the phonological working memory among beginner readers. Forty-one Arabic first graders (M = 7.13, SD = .73) were administered three tasks of phonological working memory in two versions (i.e., spoken and standard language); Two tasks were designed to test verbal retrieval and one task was designed to test remembering of instructions. The participants showed significant diglossic differences between spoken and standard stimuli in verbal retrieval tasks while no such significant differences appeared in remembering of instructions' task, especially, when the processing demands increased. In addition, the findings may shed light on the importance of developing research tools and tasks with a higher level of sensitivity in order to examine the diglossic effect on memory functions in general and verbal working memory in particular. The results were discussed considering the impact of the Arabic diglossia on cognitive and memory processing skills.
Expression and perception of emotions by voice are fundamental for basic mental health stability. Since different languages interpret results differently, studies should be guided by the relationship between speech complexity and the emotional perception. The aim of our study was therefore to analyze the efficiency of speech stimuli, word vs. sentence, as it relates to the accuracy of four different categories of emotions: anger, sadness, happiness, and neutrality. To this end, a total of 2,235 audio clips were presented to 49 females, native Hebrew speakers, aged 20–30 years (M = 23.7; SD = 2.13). Participants were asked to judge audio utterances according to one of four emotional categories: anger, sadness, happiness, and neutrality. Simulated voice samples were consisting of words and meaningful sentences, provided by 15 healthy young females Hebrew native speakers. Generally, word vs. sentence was not originally accepted as a means of emotional recognition of voice; However, introducing a variety of speech utterances revealed a different perception. Thus, the emotional conveyance provided new, even higher precision to our findings: Anger emotions produced a higher impact to the single word (χ2 = 10.21, p < 0.01) as opposed to the sentence, while sadness was identified more accurately with a sentence (χ2 = 3.83, p = 0.05). Our findings resulted in a better understanding of how speech types can interpret perception, as a part of mental health.
The human voice signal carries much information in addition to direct linguistic semantic information. This information can be perceived by computational systems. In this work, we show that early diagnosis of Parkinson's disease is possible solely from the voice signal. This is in contrast to earlier work in which we showed that this can be done using hand-calculated features of the speech (such as formants) as annotated by professional speech therapists. In this paper, we review that work and show that a differential diagnosis can be produced directly from the analog speech signal itself. In addition, differentiation can be made between seven different degrees of progression of the disease (including healthy). Such a system can act as an additional stage (or another building block) in a bigger system of natural speech processing. For example it could be used in automatic speech recognition systems that are used as personal assistants (such as Iphones' Siri, Google Voice), or as natural man-machine interfaces. We also conjecture that such systems can be extended to monitoring and classifying additional neurological diseases and speech pathologies. The methods presented here use a combination of signal processing features and machine learning techniques.
Purpose Motor speech abnormalities are highly common and debilitating in individuals with idiopathic Parkinson's disease (IPD). These abnormalities, collectively termed hypokinetic dysarthria (HKD), have been traditionally attributed to hypokinesia and bradykinesia secondary to muscle rigidity and dopamine deficits. However, the role of rigidity and dopamine in the development of HKD is far from clear. The purpose of the present study was to offer an alternative view of the factors underlying HKD. Method The authors conducted an extensive, but not exhaustive, review of the literature to examine the evidence for the traditional view versus the alternative view. Results The review suggests that HKD is a highly complex and variable phenomenon including multiple factors, such as scaling and maintaining movement amplitude and effort; preplanning and initiation of movements; internal cueing; sensory and temporal processing; automaticity; emotive vocalization; and attention to action (vocal vigilance). Although not part of the dysarthria, nonmotor factors, such as depression, aging, and cognitive-linguistic abnormalities, are likely to contribute to the overall speech symptomatology associated with IPD. Conclusion These findings have important implications for clinical practice and research.
Nearly 90% of individuals with PD will also develop swallowing disorders (dysphagia) at some point (9). Dysphagia symptoms in PD include diffi culty with lingual motility, reduced initiation of swallow, diffi culty with bolus formation, delayed pharyngeal response, and decreased pharyngeal contraction (9-11). These symptoms are often accompanied by weight loss and lack of enjoyment of eating. Aspiration pneumonia is not uncommon, especially in the later stages and can be a cause of death in PD (12). Dysphagia and abnormalities in tongue and esophageal motility have been shown to be present in early-stage PD (13-16).
To examine the role of morphology in verbal working memory. Forty nine children, all native speakers of Arabic from the same region and of the same dialect, performed a Listening Word Span Task, whereby they had to recall Arabic uninflected words (i.e., base words), inflected words with regular (possessive) morphology, or inflected words with irregular (broken plural) morphology. Each of these words was at the end of a sentence (henceforth, target word). The participant's task was to listen to a series of sentences and then recall the target words. Recall of inflected words was significantly poorer than uninflected words, and recall of words with regular morphology was significantly poorer than recall of words with irregular morphology. These findings, albeit preliminary, suggest a role of morphology in verbal working memory. They also suggest that, at least in Arabic, regular morphological forms are decomposed into their component elements and hence impose an extra load on the central executive and episodic buffer components of working memory. Furthermore, in concert with findings from other studies, they suggest that the effect of morphology on working memory is probably language-specific. The clinical implications of the present findings are addressed.
Background: Parkinson' disease (PD) is a slowly progressive and highly debilitating CNS disease. By the time it is firmly diagnosed via routine neurological examination, there is already substantial damage to the CNS. Dysarthria is present in 70–/INS;90% of individuals with PD. It has been suggested that subtle signs of the dysarthria, detectable only by acoustic methods, might serve as biomarkers to help detect the presence of the disease in its early stages.
Purpose To assess the feasibility and effectiveness of a newly developed assistive technology system, Lee Silverman Voice Treatment Companion (LSVT ® Companion™, hereafter referred to as “Companion”), to support the delivery of LSVT ® LOUD, an efficacious speech intervention for individuals with Parkinson disease (PD). Method Sixteen individuals with PD were randomized to an immediate ( n = 8) or a delayed ( n = 8) treatment group. They participated in 9 LSVT LOUD sessions and 7 Companion sessions, independently administered at home. Acoustic, listener perception, and voice and speech rating data were obtained immediately before (pre), immediately after (post), and at 6 months post treatment (follow-up). System usability ratings were collected immediately post treatment. Changes in vocal sound pressure level were compared to data from a historical treatment group of individuals with PD treated with standard, in-person LSVT LOUD. Results All 16 participants were able to independently use the Companion. These individuals had therapeutic gains in sound pressure level, pre to post and pre to follow-up, similar to those of the historical treatment group. Conclusions This study supports the use of the Companion as an aid in treatment of hypokinetic dysarthria in individuals with PD. Advantages and disadvantages of the Companion, as well as limitations of the present study and directions for future studies, are discussed. Supplemental Material https://doi.org/10.23641/asha.14963514
Recent advances in neuroscience have suggested that exercise-based behavioral treatments may improve function and possibly slow progression of motor symptoms in individuals with Parkinson disease (PD). The LSVT (Lee Silverman Voice Treatment) Programs for individuals with PD have been developed and researched over the past 20 years beginning with a focus on the speech motor system (LSVT LOUD) and more recently have been extended to address limb motor systems (LSVT BIG). The unique aspects of the LSVT Programs include the combination of (a) an exclusive target on increasing amplitude (loudness in the speech motor system; bigger movements in the limb motor system), (b) a focus on sensory recalibration to help patients recognize that movements with increased amplitude are within normal limits, even if they feel "too loud" or "too big," and (c) training self-cueing and attention to action to facilitate long-term maintenance of treatment outcomes. In addition, the intensive mode of delivery is consistent with principles that drive activity-dependent neuroplasticity and motor learning. The purpose of this paper is to provide an integrative discussion of the LSVT Programs including the rationale for their fundamentals, a summary of efficacy data, and a discussion of limitations and future directions for research.
Abstract Background: The aim of this study was to examine developmental trends in rate change detection of auditory rhythmic signals (repetitive sinusoidally frequency modulated tones). Methods: Two groups of children (9–10 years old and 11–12 years old) and one group of young adults performed a rate change detection (RCD) task using three types of stimuli. The rate of stimulus modulation was either constant (CR), raised by 1 Hz in the middle of the stimulus (RR1) or raised by 2 Hz in the middle of the stimulus (RR2). Results: Performance on the RCD task significantly improved with age. Also, the different stimuli showed different developmental trajectories. When the RR2 stimulus was used, results showed adult-like performance by the age of 10 years but when the RR1 stimulus was used performance continued to improve beyond 12 years of age. Conclusions: Rate change detection of repetitive sinusoidally frequency modulated tones show protracted development beyond the age of 12 years. Given evidence for abnormal processing of auditory rhythmic signals in neurodevelopmental conditions, such as dyslexia, the present methodology might help delineate the nature of these conditions.
acoustic analysis of speech is a powerful, noninvasive, and cost effective tool to study different aspects of motor speech disorders such as the dysarthria associated with pd. in this presentation we will discuss the rationale for using acoustic analysis, its advantages and disadvantages, and methods to overcome these disadvantages. as an example, we will address the use of vowel space area (vsa) in the study of dysarthric vowel articulation in pd. although the vsa is theoretically driven, it is highly sensitive to inter-speaker variability, which, statistically speaking, introduces noise. this noise can mask important differences that do exist between speakers with and without pd. some of this noise can be reduced by logarithmic transformation of the formant frequencies. however, even with this transformation, some statistical noise might be still present. recently sapir and colleagues introduced two acoustic metrics-the vowel articulation index (vai) and its inverse, the formant centralization ratio (fcr)-that are theoretically driven and empirically tested. these metrics show promise as they effectively reduce inter-speaker variability noise while maintaining high sensitivity to vowel centralization (the latter reflecting abnormally reduced (hypokinetic) articulatory movements in pd). data will be presented of 38 individuals with parkinson's disease and 14 healthy controls whose speech was effectively differentiated by the vai, but not the vsa, yet the logarithmically scaled vsa (lnvsa) did significantly differentiate between dysarthric and normal speech, although not as strongly as the vai.
Speech and voice disorders are key elements in the diagnosis and management of individuals with Parkinson disease (PD). This chapter reviews the classic symptoms of speech and voice disorders in PD and their assessment. The complex origin of these disorders is described in relation to motor problems (hypokinesia/bradykinesia reflecting reduced muscle activation and abnormal scaling or maintenance of the gain of movement amplitude), sensory processing problems (abnormal gating of the somatosensory cortex, abnormal gating of the auditory cortex via feed-forward mechanisms, and abnormal perception of one's own voice), cueing problems (reflecting deficits in internal/implicit cueing), and neuropsychological problems (impaired attention to action, vocal vigilance, and self-regulation of vocal output). The impact of medical treatment (neuropharmacological, neurosurgical) on speech and voice is reviewed. Speech treatment is described with special emphasis on LSVT® LOUD, a scientifically tested, efficacious treatment consistent with principles that drive activity-dependent neural plasticity. The recommendation is made for early referral to a speech clinician for optimum management of speech and voice disorders in PD.
Parkinson's disease (PD) is a slowly progressive and highly debilitating disease of the central nervous system, affecting 8,000,000 or more people the world over. By the time the disease is diagnosed, 60% of nerve cells in the substantia nigra are degenerated and 80% of dopamine is depleted in the striatum. There is an urgent need for cost-effective methods to detect the disease in its early phases, to differentiate it from other diseases, and to monitor its progression and its response to treatment. Parkinsonian speech is characterized by abnormally low voice intensity, with vocal decay, poor voice quality, reduced prosodic pitch and loudness inflection, imprecise vowels and consonants, dysrhythmia and short rushes of speech, mumbling, and reduced speech intelligibility
Advances in neuroscience have led to an expanded and improved understanding of neurobiological changes associated with rehabilitation and exercise in Parkinson’s disease (PD). This knowledge has led to a direct clinical impact of increased referral for early and continuous exercise programs for individuals with PD (physical, occupational, speech therapy and general exercise programs) and an increased research focus on the impact of such approaches in humans with PD. The purpose of this article is to examine the role of speech therapy in the landscape of exercise-based interventions for individuals with PD. We will specifically focus on the intensive voice treatment protocol, Lee Silverman Voice Treatment, as an example therapy. This article will briefly review the literature on the characteristics and features of speech and voice disorders in individuals with PD, and will discuss the impact of pharmacological and surgical treatment techniques on these disorders. This will be followed by a focus on behavioral speech treatment, specifically Lee Silverman Voice Treatment, including development of the treatment approach, documenting efficacy, discovery of unexpected outcomes and insights into the mechanism of speech disorders in PD gained from treatment-related changes. This research will be placed in the context of other previous and current speech treatment approaches in development for individuals with PD, and will highlight future directions for research.
Purpose: The vowel space area (VSA) has been used as an acoustic metric of dysarthric speech, but with varying degrees of success. In this study, the authors aimed to test an alternative metric to the VSA-the formant centralization ratio (FCR), which is hypothesized to more effectively differentiate dysarthric from healthy speech and register treatment effects.Method: Speech recordings of 38 individuals with idiopathic Parkinson's disease and dysarthria (19 of whom received 1 month of intensive speech therapy [Lee Silverman Voice Treatment; LSVT LOUD]) and 14 healthy control participants were acoustically analyzed. Vowels were extracted from short phrases. The same vowel-formant elements were used to construct the FCR, expressed as (F2u + F2A + F1i + F1u)/ (F2i + F1A), the VSA, expressed as ABS([F1i x (F2A-F2u) + F1A x (F2u-F2i) + F1u x (F2i-F2A)] /2), a logarithmically scaled version of the VSA (LnVSA), and the F2i/F2u ratio.Results: Unlike the VSA and the LnVSA, the FCR and F2i/F2u ratio robustly differentiated dysarthric from healthy speech and were not gender sensitive. All metrics effectively registered treatment effects and were strongly correlated with each other.Conclusion: Albeit preliminary, the present findings indicate that the FCR is a sensitive, valid, and reliable acoustic metric for distinguishing dysarthric from unimpaired speech and for monitoring treatment effects, probably because of reduced sensitivity to interspeaker variability and enhanced sensitivity to vowel centralization.