Background/Objectives: Most patients with alcohol use disorder (AUD) suffer from mild cognitive decline, which does not meet the diagnostic criteria of the severe form of alcohol-related cognitive impairment (ARCI). ARCI is associated with executive abnormalities in addictive behaviors and therefore influences relapse and daily functioning. Abnormalities in speech production reflect cognitive disturbances. The aim of this study was to examine the temporal speech parameters (TSPs) in ARCI. Methods: The TSPs were measured with the S-GAP Test® on 34 AUD patients with intact cognitive functions and 31 age- and gender-matched control participants. Results: Ten out of fifteen parameters of TSPs were significantly different between the AUD and healthy groups. Speech tempo and the total pause duration rate have significant classification potential. Conclusions: Our exploratory study revealed that filled pause-related temporal parameters appear to be particularly altered in ARCI and indicated that marked TSP alterations could serve as early indicators of cognitive deficits.
Narrative speech production (NSP), i.e., the conceptualization, linguistic formulation, and articulation of a story, is a multifaceted process underpinned by cognitive functions and mentalization ability, often impaired in individuals with borderline personality disorder (BPD). This study examines differences in coherence and temporal parameters between individuals with BPD and healthy controls (HCs) and explores associations between these factors in both groups. Spontaneous speech of 33 BPD and 31 HC individuals was recorded in three task types (telling their previous day, retelling a story, picture sequences), tapping different cognitive functions. Local and global coherence were extracted with contextual sentence vectors, while temporal parameters were extracted with automatic speech recognition. A series of linear mixed-effects models revealed that NSP of individuals with BPD is mainly characterized by significantly lower global coherence and speech rate and higher number of silent and filled pauses than HCs’. Global coherence displayed significant between-group differences only in picture tasks and correlated with picture arrangement. Spearman correlation matrix showed a significant negative association between global coherence and speech rate within the BPD group and an opposite tendency among HCs. Findings indicate that individuals with BPD might benefit from speaking at a slower pace to improve global coherence in their narratives.
BACKGROUND:Narrative speech production (NSP), i.e., the conceptualization, linguistic formulation, and articulation of a story, is a multifaceted process underpinned by cognitive functions and mentalization ability, often impaired in individuals with borderline personality disorder (BPD). This study examines differences in linguistic formulation between individuals with BPD and healthy controls (HCs), and explores how task type influences linguistic formulation, as well as how linguistic formulation relates to temporal parameters of speech uniquely in BPD. METHODS:Speech of 33 BPD and 31 HC individuals was recorded in three task types: telling their previous day, retelling a story, and picture sequences. Features of linguistic formulation were extracted with natural language processing methods, while temporal parameters were extracted using automatic speech recognition. Hypothesis-driven generalized linear mixed-effects models (GLMMs) were applied to test predefined group differences in four linguistic features (content words, first- and third-person singular verbs, and syntactic complexity). Additional exploratory GLMMs examined other linguistic features and task effects. Within-group Spearman correlations assessed associations between linguistic and temporal measures, controlling for task. RESULTS:Hypothesis testing showed that the NSP in BPD is characterized by fewer content words, more first-person singular verbs, and lower syntactic complexity than that of HCs. Exploratory analyses revealed that individuals with BPD used pronouns more frequently than HCs, particularly demonstrative pronouns (e.g., this) and first-person singular pronouns (e.g., I). In BPD, higher first-person singular reference (pronouns and verbs) correlated with fewer silent pauses, while greater syntactic complexity correlated with more filled pauses. Task modulated verbosity and the use of other pronoun types. CONCLUSIONS:Findings suggest that NSP in BPD is characterized by dominant self-referential thought content, reflected in elevated first-person singular reference, and by qualitatively impoverished language use, marked by reduced content word production, increased pronoun use, and lower syntactic complexity. Heightened self-focus may hinder the efficient allocation of cognitive resources required for cohesive, listener-oriented NSP.
Multiple Sclerosis is a chronic inflammatory disease of the central nervous system. Over time, people with MS may experience significant changes in cognition, language and speech processes. In this study we investigate speech utterances recorded over the course of three years for 16 MS subjects and 12 healthy controls. Our examination is based on speaker category classification (healthy or MS) using wav2vec2 embeddings as features. We found that subject classification performance improved over time: the 0.745-0.844 AUC values from year one increased to 0.891-0.979 in the third year. By analyzing the posterior estimates, we measured a statistically significant improvement in the scores corresponding to the third year for the MS category, while for the control subjects there was no such tendency. This, in our view, indicates that the change is due to a subtle deterioration in the condition of MS patients, which was detected by our machine learning workflow.
Alcohol is a progressive central nervous system depressant. Increased alcohol consumption leads to alterations in cognitive processes and also affects speech production. In this study we present a corpus of n=35 patients diagnosed with Alcohol Dependency Syndrome (ADS) and n=35 matched healthy controls, and attempt to automatically distinguish the two speaker groups based on their spontaneous speech. By using wav2vec 2.0 embeddings as features, we were able to identify the two speaker categories with quite high accuracy (EER scores between 9
Multiple sclerosis (MS) is a chronic autoimmune neurodegenerative disease, affecting the central nervous system. The disease can induce various symptoms, such as adversarily affecting the speech of the subject in various ways, therefore allowing the use of automatic speech analysis for the detection of MS and for monitoring the condition of the patient. Owing to data scarcity, however, deep neural networks are usually not employed for this task as classifiers, but are used as feature extractors. This is the case for self-supervised networks such as wav2vec 2.0 as well, where a straightforward source of embeddings (used as features) are the last layers of the convolutional (lower) and fine-tuned (higher) blocks. In this study we investigate whether extracting the embeddings from some other, inner layer of the fine-tuned (transformer) block can help improve MS detection performance. Tested on two speech tasks, we found that the lowest one-third of the 24 fine-tuned layers proved to be the most suitable for feature extraction, which led to statistically significant improvements in the AUC scores for both speech tasks.
In the past few years, self-supervised learning has revolutionalized automatic speech recognition. Self-supervised models such as wav2vec2, due to their generalization ability on huge unannotated audio corpora, were claimed to be state-of-the-art feature extractors in paralinguistic and pathological applications as well. In this study we test embeddings extracted from a wav2vec 2.0 model fine-tuned on the target language as features on a multiple sclerosis audio corpus, using three speech tasks. After comparing the resulting classification performances with traditional features such as ComParE functionals, ECAPA-TDNN and activations of a HMM/DNN hybrid acoustic model, we found that wav2vec2-based models, surprisingly, only produced a mediocre classification performance. In contrast, the decade-old ComParE functionals feature set consistently led to high scores. Our results also indicate that the number of features correlates surprisingly well with classification performance.
Our aim was to find out whether speech-related temporal parameters (SRTPs) are sensitive indicators of the clinical outcome in acetylcholinesterase (AChE) inhibitor therapy with donepezil, compared to the standard cognitive Alzheimer's Disease Assessment Scale-Cognitive Subscale (ADAS-Cog) used in clinical trials. In this 24-week-long, naturalistic, self-control, open-labeled, prospective pilot study with 10 mg donepezil on 20 mild AD patients, cognitive functions were evaluated using 15 different SRTPs analyzed by automatic speech recognition in the Speech-Gap Test® compared to ADAS-Cog test results. Among the SRTPs, the filled pause duration ratio significantly improved after 12 weeks of donepezil treatment. During the 24-week follow-up, additional SRTPs such as the filled pause count ratio and the filled pause frequency showed significant benefits. ADAS-Cog total scores showed a slight but not significant improvement compared to baseline after 12 and 24 weeks of donepezil treatment. Among the ADAS-Cog subtests, only orientation improved significantly after 24 weeks of donepezil treatment. Our results indicate that subtle changes in SRTPs measured by the Speech-Gap Test® could be considered as sensitive indicators of the efficacy of the pharmacotherapy in mild AD. According to our data, other cognitive domains did not show improvement in response to donepezil therapy rating by ADAS-Cog. Based on all of this, it is likely that examining and evaluating speech parameters may play an important role in determining the effects of pharmacological treatment of mild AD. The novelty of our study is that it applies the measurement of linguistic parameters as primary outcomes during a drug trial of mild AD in scientific research for the first time.
Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system which, in addition to affecting motor and cognitive functions, may also lead to specific changes in the speech of patients. Speech production, comprehension, repetition and naming tasks, as well as structural and content changes in narratives, might indicate a limitation of executive functions. In this study we present a speech-based machine learning technique to distinguish speakers with relapsing-remitting subtype MS and healthy controls (HC). We exploit the fact that MS might cause a motor speech disorder similar to dysarthria, which, with our hypothesis, might affect the phonetic posterior estimates supplied by a Deep Neural Network acoustic model. From our experimental results, the proposed posterior posteriorgram-based feature extraction approach is useful for detecting MS: depending on the actual speech task, we obtained Equal Error Rate values as low as 13.3%, and AUC scores up to 0.891, indicating a competitive and more consistent classification performance compared to both the x-vector and the openSMILE 'ComParE functionals' attributes. Besides this discrimination performance, the interpretable nature of the phonetic posterior features might also make our method suitable for automatic MS screening or monitoring the progression of the disease. Furthermore, by examining which specific phonetic groups are the most useful for this feature extraction process, the potential utility of the proposed phonetic features could also be utilized in the speech therapy of MS patients.
Multiple Sclerosis (MS) is a chronic disease affecting over 2.5 million people worldwide. Its early detection is crucial for the management and treatment of the disease. Here we present an approach for automatic MS screening based on encoded speech representations. Our methods rely on Wav2Vec2 models to extract relevant traits from speech recordings of patients, which are then fed into a Support Vector Machine. Besides employing Wav2Vec2 models pre-trained on large public corpora, we also fine-tune them on 85 hours of the target language (Hungarian) in two distinct ways: for ASR and for speaker identification. Both variations outperformed the original models and conventional methods (ComParE functionals, x-vectors, and ECAPA-TDNN). Our findings suggest that fine-tuning for the actual speaker provides more advantages than the typical approach of fine-tuning for ASR purposes. Still, we improved our best MS discrimination performance when we fused features from our two fine-tuned models.
Dementia is a chronic or progressive clinical syndrome, mainly characterized by the deterioration of memory, thinking, reasoning and language. In Mild cognitive impairment (MCI), often considered as the prodromal stage of dementia, there is also a subtle deterioration of these functions, but they do not affect the daily life of the patient. However, due to the slight nature of the changes, it is quite hard to diagnose MCI. In this study, we employ sequence-to-sequence deep autoencoders in order to extract compact, robust and efficient attributes from the spontaneous speech of 25 MCI subjects and 25 healthy controls. From our results, this approach gives a competitive performance, as we significantly outperformed x-vectors even though they were trained on more data. Our additional efforts to identify mild Alzheimer's (mAD) subjects as well were less successful; but since the focus is on the early detection of dementia, this is not a limitation of the methodology from a practical point of view.
Dementia is a chronic or progressive clinical syndrome, characterized by the deterioration of problem-solving skills, memory and language. In Mild Cognitive Impairment (MCI), which is often considered to be the prodromal stage of dementia, there is also a subtle deterioration of these cognitive functions; however, it does not affect the patients’ ability to carry out simple everyday activities. The timely identification of MCI could provide more effective therapeutic interventions to delay progression, and to postpone the possible conversion to dementia. Since language changes in MCI are present even before the manifestation of other distinctive cognitive symptoms, a non-invasive way of early automatic screening could be the use of speech analysis. Earlier, our research team developed a set of temporal speech parameters that mainly focus on the amount of silence and hesitation, and demonstrated its applicability for MCI detection. However, for the automatic extraction of these attributes, the execution of a full Automatic Speech Recognition (ASR) process is necessary. In this study we propose a simpler feature extraction approach, which still quantifies the amount of silence and hesitation in the speech of the subject, but does not require the application of a full ASR system. We experimentally demonstrate that this approach, operating directly on the frame-level output of a HMM/DNN hybrid acoustic model, is capable of extracting attributes as useful as the ASR-based temporal parameter extraction workflow was able to. That is, on our corpus consisting of 25 healthy controls, 25 MCI and 25 mild AD subjects, we achieve a (three-class) classification accuracy of 70.7%, an F-measure score of 89.6 and a mean AUC score of 0.804. We also show that this approach can be applied on simpler, context-independent acoustic states with only a slight degradation of MCI and mild Alzheimer’s detection performance. Lastly, we investigate the usefulness of the three speaker tasks which are present in our recording protocol.
Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system. It affects cognitive and motor functions, and the limitation of executive functions can also manifest itself in speech production. Due to this, automatic speech analysis might serve as an effective technique for assessing MS, or for monitoring the status of the patient. However, choosing the features to be extracted from the recordings is not straightforward. In the past few years, general feature extractors such as i-vectors, d-vectors and x-vectors have found their way into automatic speech analysis. In this study we show that there is no need to employ a special neural network architecture such as x-vectors to calculate effective features, but (even more) indicative features can be derived on the basis of a standard Deep Neural Network acoustic model. From our results, these features could effectively be used to distinguish MS subjects from healthy controls, as we measured AUC scores up to 0.935. We found that classification performance depended only slightly on the choice of the hid-den layer used to extract our features, but the speech task per-formed by the subject turned out to be an important factor.
Introduction: The earliest signs of cognitive decline include deficits in temporal (time-based) speech characteristics. Type 2 diabetes mellitus (T2DM) patients are more prone to mild cognitive impairment (MCI). The aim of this study was to compare the temporal speech characteristics of elderly (above 50 y) T2DM patients with age-matched nondiabetic subjects. Materials and Methods: A total of 160 individuals were screened, 100 of whom were eligible (T2DM: n=51; nondiabetic: n=49). Participants were classified either as having healthy cognition (HC) or showing signs of MCI. Speech recordings were collected through a phone call. Based on automatic speech recognition, 15 temporal parameters were calculated. Results: The HC with T2DM group showed significantly shorter utterance length, higher duration rate of silent pause and total pause, and higher average duration of silent pause and total pause compared with the HC without T2DM group. Regarding the MCI participants, parameters were similar between the T2DM and the nondiabetic subgroups. Conclusions: Temporal speech characteristics of T2DM patients showed early signs of altered cognitive functioning, whereas neuropsychological tests did not detect deterioration. This method is useful for identifying the T2DM patients most at risk for manifest MCI, and could serve as a remote cognitive screening tool.
Background: The development of automatic speech recognition (ASR) technology allows the analysis of temporal (time-based) speech parameters characteristic of mild cognitive impairment (MCI). However, no information has been available on whether the analysis of spontaneous speech can be used with the same efficiency in different language environments. Objective: The main goal of this international pilot study is to address the question of whether the Speech-Gap Test (R) (S-GAP Test (R)), previously tested in the Hungarian language, is appropriate for and applicable to the recognition of MCI in other languages such as English. Methods: After an initial screening of 88 individuals, English-speaking (n = 33) and Hungarian-speaking (n = 33) participants were classified as having MCI or as healthy controls (HC) based on Petersen's criteria. The speech of each participant was recorded via a spontaneous speech task. Fifteen temporal parameters were determined and calculated through ASR. Results: Seven temporal parameters in the English-speaking sample and 5 in the Hungarian-speaking sample showed significant differences between the MCI and the HC groups. Receiver operating characteristics (ROC) analysis clearly distinguished the English-speaking MCI cases from the HC group based on speech tempo and articulation tempo with 100% sensitivity, and on three more temporal parameters with high sensitivity (85.7%). In the Hungarian-speaking sample, the ROC analysis showed similar sensitivity rates (92.3%). Conclusion: The results of this study in different native-speaking populations suggest that changes in acoustic parameters detected by the S-GAP Test (R) might be present across different languages.
This study presents a novel approach for the early detection of mild cognitive impairment (MCI) and mild Alzheimer's disease (mAD) in the elderly. Participants were 25 elderly controls (C), 25 clinically diagnosed MCI and 25 mAD patients, included after a clinical diagnosis validated by CT or MRI and cognitive tests. Our linguistic protocol involved three connected speech tasks that stimulate different memory systems, which were recorded, then analyzed linguistically by using the PRAAT software. The temporal speech-related parameters successfully differentiate MCI from mAD and C, such as speech rate, number and length of pauses, the rate of pause and signal. Parameters pauses/duration and silent pauses/duration linearly decreased among the groups, in other words, the percentage of pauses in the total duration of speech continuously grows as dementia progresses. Thus, the proposed approach may be an effective tool for screening MCI and mAD.
Mild Cognitive Impairment (MCI) is a heterogeneous clinical syndrome, often considered as the prodromal stage of dementia. It is characterized by the subtle deterioration of cognitive functions, including memory, executive functions and language. Mainly due to the tenuous nature of these impairments, a high percentage of MCI cases remain undetected. There is evidence that language changes in MCI are present even before the manifestation of other distinctive cognitive symptoms, which offers a chance for early recognition. A cheap noninvasive way of early screening could be the use of automatic speech analysis. Earlier, our research team developed a set of speech temporal parameters, and demonstrated its applicability for MCI detection. For the automatic extraction of these attributes, a Hungarian -language ASR system was employed to match the native language of the MCI and healthy control (HC) subjects. In practical applications, however, it would be convenient to use exactly the same tool, regardless of the language spoken by the subjects. In this study we show that our temporal parameter set, consisting of articulation rate, speech tempo and various other attributes describing the hesitation of the subject, can indeed be reliably extracted regardless of the language of the ASR system used. For this purpose, we performed experiments both on English-speaking and on Hungarian-speaking MCI patients and healthy control subjects, using English and Hungarian ASR systems in both cases. Our experimental results indicate that the language on which the ASR system was trained only slightly affects the MCI classification performance, because we got quite similar scores (67-92%) as we did in the monolingual cases (67-92% as well). As our last investigation, we compared the proposed attribute values for the same utterances, utilizing both the English and the Hungarian ASR models. We found that the articulation rate and speech tempo values calculated based on the two ASR models were highly correlated, and so were the attributes corresponding to silent pauses; however, noticeable differences were found regarding the filled pauses (still, these attributes remained indicative for both languages). Our further analysis revealed that this is probably due to a difference regarding the annotation of the English and the Hungarian ASR training utterances. (c) 2021 Published by Elsevier Ltd.
Schizophrenia is a heterogeneous chronic and severe mental disorder. There are several different theories for the development of schizophrenia from an etiological point of view: neurochemical, neuroanatomical, psychological and genetic factors may also be present in the background of the disease. In this study, we examined spontaneous speech productions by patients suffering from schizophrenia (SCH) and bipolar disorder (BD). We extracted 15 temporal parameters from the speech excerpts and used machine learning techniques for distinguishing the SCH and BD groups, their subgroups (SCH-S and SCH-Z) and subtypes (BD-I and BD-II). Our results indicated, that there is a notable difference between spontaneous speech productions of certain subgroups, while some appears to be indistinguishable for the used classification model. Firstly, SCH and BD groups were found to be different. Secondly, the results of SCH-S subgroup were distinct from BD. Thirdly, the spontaneous speech of the SCH-Z subgroup was found to be very similar to the BDI, however, it was sharply distinct from BD-II. Our detailed examination highlighted the indistinguishable subgroups and led to us to make our S and Z theory more clarified.
Háttér és célkitűzésekA szakkádikus szemmozgási paraméterek biomarkerként történő önálló használata egyes degeneratív neuropszhichiátriai kórképek felismerésében egyelőre kérdéses. Jelen vizsgálatunk célkitűzése egy olyan szakkádikus szemmozgásvizsgálati protokoll megvalósítása, amely a nemzetközi klinikai kutatásoknak megfelelően kialakított vizsgálati paraméterekkel jellemezhető. Továbbá egészséges vizsgált személyek szakkádparamétereinek a nemzetközileg publikált adatokkal való összehasonlítása, illetve a Boston Sütilopás feladatban használt kép vizuális tesztkörnyezetbe történő beépítésének disztraktor-hatásvizsgálata a szakkádikus paraméterekre.MódszerVizsgálatainkat egészséges önkéntes alanyokon végeztük Tobii Pro X3-120 berendezéssel, két eltérő vizuális környezetet tartalmazó, egyéb tekintetben teljesen megegyező felépítésű, proszakkád és antiszakkád tesztek gap és overlap feladataiban. Az egyik csoport vizuális tesztkörnyezete standard szürke háttérből és fekete stimulusokból állt (STD tesztcsoport), a másiké a Boston Afázia Teszt Sütilopás képleírási feladatban használt képet és zöld-piros színű stimulusokat tartalmazott (BSL tesztcsoport).EredményekA proszakkád és antiszakkád overlap latencia mindkét csoportban nagyobb volt, mint a gap feladatban. A vizsgálati alanyok életkora közepes erősségű módon korrelált az STD csoport proszakkád gap latenciájával, illetve a BSL csoport proszakkád overlap latencia, antiszakkád gap/overlap latencia és overlap időtartam értékével. A csoportok közötti életkori eltérés statisztikai kontrollja mellett nem találtunk különbséget a proszakkád és antiszakkád latencia, illetve csúcssebesség tekintetében. A szakkád időtartamok szignifi kánsan rövidebbek voltak a BSL csoportban. Az iránytévesztési ráta a proszakkád tesztben megegyezett; az antiszakkád gap időtartam, az overlap időtartam, illetve a gap iránytévesztési ráta szignifi káns csoportkülönbséget mutatott.KövetkeztetésekA nemzetközi sztenderdeknek megfelelő újszerű vizsgálati protokoll alkalmasnak látszik hagyományos neuropszichológiai tesztekkel és kutatócsoportunk közelmúltban kidolgozott automatikus beszédfelismerő algoritmusokat alkalmazó tesztjeivel kombinálva egészséges és demencia szindrómával élő személyek multimodális klinikai vizsgálatában történő felhasználásra.Background and aimsThe use of saccadic parameters as specifi c biomarkers in the diagnosis of degenerative neuropsychiatric disorders is still problematic. The aim of the current study was to 1) establish a protocol for saccadic eye movement measurements that is concordant with international clinical investigations; 2) compare saccadic parameters of healthy subjects with internationally published data, and 3) integrate the picture from Boston Cookie Theft test into the visual test environment and evaluate its distractor effects on saccadic parameters.MethodsHealthy volunteers were assessed with Tobii Pro X3-120 eye tracker in two distinct visual settings, but otherwise identical test environment evaluating prosaccades and antisaccades in gap and overlap conditions. One group was assessed in a visual test environment based on a traditional uniform grey background with black stimuli (STD test group), while the visual test environment of the other group contained the Cookie Theft picture of the Boston Diagnostic Aphasia Examination with green and red stimuli (BSL group).ResultsProsaccade and antisaccade latencies were signifi cantly longer in the overlap condition both in the BSL and STD groups. After controlling for age, the STD and BSL group did not differ in terms of prosaccade latency, antisaccade latency or peak velocity; only saccade durations were shorter in the BSL group. Saccadic direction error rate in the prosaccade task was identical in both groups, while the antisaccade gap duration, overlap duration and gap direction error rates showed signifi cant group differences.The present, newly developed protocol conforms to international standards, and may be useful for multi-modal clinical studies assessing healthy subjects and people suffering from dementia, combined with traditional neuropsychological tasks and the speech recognition task recently developed by our research group.
One of the world’s chronic neuro-degenerative diseases, Alzheimer’s Disease (AD), leads its sufferers, among other symptoms, to suffer from speech difficulties. In particular, the inability to recall vocabulary which makes patients’ speech different. Furthermore, Mild Cognitive Impairment (MCI) is usually considered as a prodromal neuro-degenerative state of AD. The key to abate the progress of both disorders is their early diagnosis. However, actual ways of diagnosis are costly and quite time-consuming. In this study, we propose the extraction of features from speech through the use of the i-vector approach, by which we seek to model the speech pattern of the three mental conditions from the subjects. To the best of our knowledge, no previous studies have utilized i-vector features to assess Alzheimer’s before. These i-vectors are extracted from Mel-Frequency Cepstral Coefficients (MFCCs), then they are given to a SVM classifier in order to identify the speech in one of the following manners: AD - Alzheimer Disease, MCI - Mild Cognitive Impairment, HC - Healthy Control. We tested these i-vector features by performing a 5-fold cross-validation and we achieved an F1-score of 79.2%.