Advances in automated language and speech analysis using machine learning have validated digital biomarkers for non-invasive detection of subtle cognitive changes. While distinguishing Alzheimer's Disease (AD) from Normal Controls (NC) is straightforward, classifying Mild Cognitive Impairment (MCI) remains challenging, due to its potential progression to AD or its association with other factors such as affective disorders, requiring detailed expert evaluation. Expanding upon prior research, this study assesses LANGaware's biomarker capabilities on recent data for (a) binary classification of AD versus NC, (b) multi-class classification into AD, NC, and MCI, and (c) binary classification of Depression (D) versus NC. Diagnoses have been collected for the above cases, provided by medical experts following standardized protocols. Participants were engaged in simple elicitation tasks, such as describing a picture or narrating an event. Digital biomarkers that reflect linguistic, speech, and acoustic features were extracted from recorded audio and corresponding transcripts. These biomarkers were then used as input for classification tasks, employing a tailored neural network and an XGBoost model to perform binary and multi-class (three-class) classifications. The methodology was designed to be easily adaptable to multiple languages. For cognitive classification, 15827 elicitation tasks (5173 AD, 7660 MCI, 2994 NC) from English- and Greek-speaking participants were analyzed using nested cross-validation. The binary classifier (AD vs. Healthy) achieved an average F1-macro score of 89.01%, while the three-class one (AD vs. MCI vs. NC) attained a score of 73.0%. For depression classification, 4421 elicitation tasks (1481 D, 2940 NC) from English-speaking participants were evaluated, achieving an average F1-macro score of 70.38%. The results obtained confirm the strong discriminative capability of the proposed biomarkers for the early detection of cognitive decline. The findings support the applicability of automated assessment, facilitating early diagnosis and timely intervention. Additionally, the depression classification experiment further complements the cognitive analysis, given the established link between depression and MCI. This study is crucial for ensuring the quality and reliability of the LANGaware product as it is increasingly adopted by diagnostic centers and hospitals. We express our gratitude to all organizations that contributed valuable data to this work.
Background Most common forms of dementia, including Alzheimer's disease, are associated with alterations in spoken language. Objective This study explores the potential of a speech-based machine learning (ML) approach in estimating cognitive impairment, using inputs of speech audio recordings. Methods We develop an automatic ML pipeline that ingests multimodal inputs of audio and transcribed text, mapping speech and language to domain-specific biomarkers optimized for high explainability and predictive ability. The resulting features are fed through a multi-stage pipeline to determine efficient classification configurations. Results We evaluated the system on large real-world datasets, achieving above 90% and 70% weighted average F1 scores for two-class (AD versus normal controls) and three-class (AD versus mild cognitive impairment versus normal controls) classification tasks, respectively. Model performance remains stable across different population characteristics. Conclusions The study introduces a robust, non-invasive method for gauging the cognitive status of AD and MCI patients from speech samples, with the potential of generalizing effectively to multiple types of diseases/disorders which may burden language.
Recent advancements in automatic language and speech analysis, coupled with machine learning (ML) methods, showcase the effectiveness of digital biomarkers in non-invasively detecting subtle changes in cognitive status. While successfully distinguishing between Alzheimer's Disease (AD) and Normal Control (NC) individuals, classifying Mild Cognitive Impairment (MCI) proves to be a more challenging task. MCI can progress to AD or result from various factors, including affective disorders, necessitating multiple expert examinations for accurate detection. Building upon previous research, we create an experimental setup to assess LANGaware’s biomarkers pool on three objectives: a) binary separation into Dementia and NC cohorts, b) broad three-class separation into Dementia, NC, and MCI groups, c) binary differentiation into Depression coupled with Anxiety disorder and NC cohorts. Patient audio recordings and ASR-generated transcripts were fed into LANGaware’s multimodal ML pipeline, extracting hundreds of linguistic and audio features, distilled into interpretable categories with a neural network assigning weights. These categorical values served as inputs to a final neural network layer generating probabilities for target labels (Dementia, NC). Similar methodologies were applied to our second (Dementia vs MCI vs NC) and third discrimination task (Depression/Anxiety vs NC), where the neural network allocated varying weights to input features for each of the aforementioned cases. In all scenarios, data were split into a 70% training set and a 30% testing set, validated against medical expert diagnosis. For binary separation, with 2927 Dementia and 815 NC instances, the model demonstrated 89% accuracy and an 85% macro-averaged F1 score. For three-class separation (3752 Dementia, 1117 NC, 5993 MCI instances), the model achieved 70% accuracy and a 71% F1 score. Discriminating affective disorders (1016 Depression/Anxiety, 1630 NC instances) resulted in 71% accuracy and a 71% F1 score. The assessment suggests that our modelling approach aptly discerns language and speech patterns, distinguishing individuals with MCI from those with Dementia or in optimal health (NC). These outcomes contribute significantly to automatic evaluation, offering early diagnosis and timely treatment access. Our third experiment showcases the methodology's applicability in detecting affective disorders, specifically Depression and Anxiety, which may co-occur with or precede MCI.
A wide range of neurodegenerative disorders impair communication faculties, manifested as subtle changes in speech and language. Following trends in Machine Learning (ML), Text and Multimedia Analysis, recent works pursue classification of such disorders by developing biomarkers useful for cognitive status description. However, investigations that gauge biomarker applicability on cases of radically different impairment / symptom severity have been under-investigated. In this work, we build upon previous research and evaluate LANGaware’s biomarker suite in the drastically different tasks of Dementia and Depression classification. For both cases, we use appropriate Healthy Controls and evaluate the proposed biomarker workflow under a multi-language experimental investigation. We utilize audio recordings and transcripts from patient responses to verbal cognitive assessment tasks. These are analyzed with LANGaware’s multimodal biomarker pool, mapping raw data to biomarkers scores that quantify vocal, linguistic and grammatical usage proficiency, structure and patterns, by applying both statistical analysis and explainable template matching. Biomarkers activations are fed to ML workflows composed of gradient boosting learners that employ feature selection, filtering and ensemble-based learning to arrive at configurations best suited for discriminating the disorder of interest. The pipeline is evaluated using a standard train-test and cross-validation setup. Experimental results indicate that our method achieves weighted F1 test scores of 82.49% for Greek Dementia classification, using a train / test dataset of 1271 / 624 instances and 80.28% on 570 / 243 English data. We use the same pipeline and biomarker pool for Depression classification, reaching 74.03% and 71.36% performance scores, obtained from 73 / 31 and 652 / 42 available train-test data, for Greek and English respectively. The above findings show that the proposed biomarker pool generalizes across neurodegenerative and affective disorders of radically different severity. As a result, they constitute valuable decision support tools for accurate, automatic and explainable early diagnosis, facilitating proactive care, improved symptom management and better patient quality of life. Future research efforts include extending our experimental evaluation of LANGaware biomarkers to additional disorders and diseases, as well as expanding the biomarker set to enable analysis of additional modalities.
Multiple neurodegenerative and psychiatric diseases can affect speech, manifesting as subtle changes in spoken language. We perform a rigorous evaluation of a Machine Learning (ML) pipeline on a large‐scale multiaxial experimental setting, to investigate its capacity for detecting early signs of cognitive decline as manifested in speech. Using a diverse dataset of speech samples spanning different languages, patient cohorts and cognitive status diagnoses, we propose that the developed pipeline showcases robust performance, facilitating efficient decision support for early cognitive decline detection.
In the current evaluation we verified our conjectures regarding the strong capacity of speech to predict cognitive decline. Audio analysis and machine learning are proven to be invaluable tools in the prediction of early signs of cognitive decline, which are coupled with a wide spectrum of neurodegenerative and psychiatric diseases. [1] Boschi, Veronica, et al., Frontiers in psychology 8 (2017): 269. [2] Vassiliki Rentoumi et al., Alzheimer's & Dementia, Wiley, volume 16, 2020. [3] Alberdi, Ane et al., Artificial intelligence in medicine 71 (2016): 1-29.
There are multiple neurodegenerative diseases that directly affect speech [1]. However, its utilization as a robust indicator for cognitive impairment is under-investigated. In many cases, mild cognitive decline progresses to a neurodegenerative disease and its detection is of utmost importance, since it is at this stage that treatment is most effective. One of our core goals is developing techniques for differentiating between patients with cognitive decline and healthy cohorts, by utilizing only speech samples [2]. Such samples are obtained from verbal elicitation tasks designed for cognitive assessment, e.g. picture descriptions and narration of everyday activities. Audio recordings from cognitive assessment tasks are fed through our platform to a Natural Language Processing and Machine Learning pipeline, employing an automatic discovery procedure of predictive salient biomarkers, to train an advanced classification system. The biomarker collection includes features that characterize voice, speech, language structure, composition and usage, and are engineered to highlight symptoms of neurodegenerative disorders. Biomarkers undergo multiple stages of filtering, processing and transformation to train and fine-tune the final model. Diagnostic performance of the output classifier is obtained on an unseen test set, to ensure a robust generalization of the platform. Our platform utilizes the set of automatically selected, cross-linguistic digital biomarkers to obtain sensitivity and specificity scores of 81% and 84% respectively, compared against medical expert diagnosis. The most salient biomarkers with respect to identifying our pathological cohort, relate to feature categories of Syntactic Complexity, Content Word usage, Lexical Repetition, Syntactic Errors and Function Words, with weight contributions of ∼ 16%, 14%, 13%, 12% and 11% respectively. Our platform provides additional, detailed population and patient-based descriptive analytics to enhance transparency and explainability of the results. Early detection of cognitive decline facilitates early intervention, treatment and proactive care, delaying disease progression and reducing symptom severity. We believe that our platform provides an effective solution for risk factor estimation and our findings incentivize further research into speech analysis techniques for the prediction of cognitive decline. [1] Boschi, Veronica, et al.,Frontiers in psychology 8 (2017): 269. [2] Vassiliki Rentoumi et al., Alzheimer's & Dementia, Wiley, volume 16, 2020.
Various neurodegenerative and psychiatric diseases are related with changes in spoken language [1], although they have seldom been investigated. We evaluate the effectiveness of our language‐agnostic Machine Learning (ML) system to detect subtle changes in spoken language that manifest early signs of cognitive decline, thus assisting with its diagnosis. We evaluate our methodology using recordings of speech samples from multiple languages obtained from patient cohorts in early stages of cognitive decline and matched healthy controls.
AbstractBackgroundAlthough Alzheimer's disease (AD) is associated with changes in spoken language, these have seldom been subjected to systematic analysis on a large scale. We evaluated the effectiveness of LangAware to detect the language indicators that are coupled with early AD, thus assisting with diagnosis. We evaluated LangAware using recordings of speech samples obtained from AD patients and matched healthy controls (NC) derived from various elicitation tasks in two languages, English and Greek*.MethodEnglish and Greek datasets were analyzed employing feature selection techniques to choose the most prominent multi‐level linguistic analysis features differentiating the AD from the NC group in both languages. The platform’s diagnostic performance was evaluated on its ability to classify "unseen" audio recordings employing these salient features.ResultEvaluation results indicated that LangAware achieved equally high classification scores for both English and Greek. Most significantly, these scores were achieved by employing a custom set of LangAware‐developed cross‐linguistic markers.ConclusionThe current evaluation verified the robustness of the platform’s predictive models using audio datasets in two languages. Based on the findings, we conclude that LangAware could provide a time and cost‐effective platform for cognitive screening across languages and across tasks pertaining to neurodegenerative diseases in a range of clinical settings. Such findings also advocate in favor of the robustness of LangAware platform in pursuing cognitive assessment on spontaneous speech across languages. *Peter Garrard and Antonis A. Mougias contributed to this work by providing data for the evaluation of the LangAware platform.
The vision of IASIS project is to turn the wave of big biomedical data heading our way into actionable knowledge for decision makers. This is achieved by integrating data from disparate sources, including genomics, electronic health records and bibliography, and applying advanced analytics methods to discover useful patterns. The goal is to turn large amounts of available data into actionable information to authorities for planning public health activities and policies. The integration and analysis of these heterogeneous sources of information will enable the best decisions to be made, allowing for diagnosis and treatment to be personalised to each individual. The project offers a common representation schema for the heterogeneous data sources. The iASiS infrastructure is able to convert clinical notes into usable data, combine them with genomic data, related bibliography, image data and more, and create a global knowledge base. This facilitates the use of intelligent methods in order to discover useful patterns across different resources. Using semantic integration of data gives the opportunity to generate information that is rich, auditable and reliable. This information can be used to provide better care, reduce errors and create more confidence in sharing data, thus providing more insights and opportunities. Data resources for two different disease categories are explored within the iASiS use cases, dementia and lung cancer.
In the present study, we analyzed written samples obtained from Greek native speakers diagnosed with Alzheimer’s in mild and moderate stages and from age-matched cognitively normal controls (NC). We adopted a computational approach for the comparison of morphosyntactic complexity and lexical variety in the samples. We used text classification approaches to assign the samples to one of the two groups. The classifiers were tested using various morphosyntactic and lexical features. The proposed method excels in discerning AD patients in mild and moderate stages from NC leading to the in depth understanding of language deficits in this neurodegenerative disease. 1. Εισαγωγή Οι νευροεκφυλιστικές ασθένειες, όπως η Νόσος Αλτσχάιμερ (εφεξής ΝΑ) συνδέονται με αλλαγές στον προφορικό και γραπτό λόγο, οι οποίες δεν έχουν μελετηθεί εκτενώς. Οι αλλαγές αυτές στις δύο τροπικότητες συνδέονται με προβλήματα μνήμης αλλά και γλωσσικής επεξεργασίας. Τα προβλήματα γλωσσικής επεξεργασίας είναι εμφανή από τα πρώτα στάδια της νόσου (Kempler et al. 1987, Martin & Fedio 1983). Τα γλωσσικά προβλήματα έχουν εντοπιστεί, κυρίως, στη λεξική και σημασιολογική γνώση των ατόμων με ΝΑ και ειδικότερα στην κατονομασία ρημάτων και ουσιαστικών, στην εύρεση και ανάκληση λέξεων, καθώς και στη σημασιολογική ευχέρεια, δηλαδή στην κατονομασία λέξεων που ανήκουν στην ίδια σημασιολογική κατηγορία (π.χ. ζώα, φρούτα) (Altmann et al. 2001, Chertkow et al. 1989). Έχει υποστηριχθεί ότι η ανομία, δηλαδή η δυσκολία εύρεσης/ανάκλησης των κατάλληλων λέξεων, αποτελεί το πιο συχνό χαρακτηριστικό της νόσου, ακόμα και σε πρώιμο στάδιο (Altmann et al. 2001). Οι ασθενείς, συχνά, υποκαθιστούν τη λέξη-στόχο με αντωνυμία ή χρησιμοποιούν σημασιολογικά σχετικές λέξεις (σημασιολογικές παραφασίες) (Tang-Wai & Graham 2008) ή ακόμα και περιφράσεις. Η δυσκολία στην εύρεση των κατάλληλων λέξεων και τα προβλήματα κατονομασίας που παρατηρούνται στα άτομα με ΝΑ έχουν ως αποτέλεσμα έναν κενό περιεχομένου λόγου χωρίς συνοχή (Ripich and Terell 1988, Kavé & Levy 2003). Συχνές είναι οι επαναλήψεις και τα μεταγλωσσικά σχόλια που δυσχεραίνουν τη συνοχή του κειμένου. Σχετικά με το έλλειμμα στη μορφοσυντακτική γνώση, τα πορίσματα των ερευνών είναι αντιφατικά με με άλλες έρευνες να υποστηρίζουν ότι η μορφοσύνταξη είναι διατηρημένη σε σχέση με τη λεξική και σημασιολογική ικανότητα και άλλες να επισημαίνουν ορισμένες διαταραχές στη μορφοσύνταξη ακόμα και σε πρώιμο στάδιο της νόσου (π.χ. δυσκολίες στην παραγωγή λέξεων κλειστής τάξης, στη συμφωνία αριθμού και συμφωνία υποκειμένου-ρήματος). Εξίσου αντιφατικά είναι και τα αποτελέσματα για τη συντακτική ικανότητα, με κάποιες μελέτες να υποστηρίζουν ότι τα άτομα με ΝΑ μπορούν να ερμηνεύσουν συντακτικά πολύπλοκες δομές, παρά την ελλειμματική μνήμη εργασίας, και άλλες να εντοπίζουν προβλήματα συντακτικής κατανόησης, κυρίως, στην κατανόηση σύνθετων δομών. Παρόλο που οι γλωσσικές ικανότητες στη NA έχουν μελετηθεί σε κάποιο βαθμό, υπάρχουν αρκετοί περιορισμοί στις έρευνες που έχουν γίνει για τη μελέτη της γλώσσας. Ένας βασικός περιορισμός είναι ότι η γλωσσική ανάλυση των δεδομένων γίνεται χειροκίνητα, γεγονός που αποτελεί μια χρονοβόρα και πολλές φορές υποκειμενική διαδικασία. Σε αυτή τη μελέτη επιχειρείται μια αυτόματη, υπολογιστική, γλωσσική ανάλυση, με τη χρήση της μηχανικής μάθησης, σε δείγματα γραπτού λόγου φυσικών ομιλητών της Ελληνικής που βρίσκονται σε ήπιο ή και μεσαίο στάδιο άνοιας αλλά και υγιών ηλικιωμένων αντιστοιχισμένων ως προς την ηλικία και την εκπαίδευση με την πειραματική ομάδα. Με την εφαρμογή ποσοτικών μεθόδων ανάλυσης διερευνώνται οι διαφορές στα γλωσσικά χαρακτηριστικά των ατόμων με άνοια και των υγιών ηλικιωμένων. τόχος είναι η εύρεση των σημαντικότερων διαχωριστικών κριτηρίων και γλωσσικών δεικτών για τις δύο ομάδες. Με τον εντοπισμό αυτών των διακριτών, γλωσσικών χαρακτηριστικών, αλλά και κάποιων γλωσσικών δομών που ίσως αποκλίνουν από την νόρμα των υγιών, πιθανόν να μπορέσουμε να βοηθήσουμε στην πρώιμη διάγνωση της ΝΑ και των άλλων μορφών άνοιας, διευκολύνοντας, με αυτόν τον τρόπο, την κλινική διαδικασία των ιατρών. 1.1. Υπολογιστικές μέθοδοι στην ανάλυση λόγου ατόμων με Νόσο Αλτσχάιμερ (ΝΑ) Την τελευταία πενταετία αρκετοί ερευνητές παγκοσμίως έχουν στραφεί στη μελέτη του αυθόρμητου ή συνεχούς λόγου για την ανεύρεση γλωσσικών χαρακτηριστικών που μπορούν να διακρίνουν τον λόγο ατόμων με ΝΑ, ήπια γνωστική διαταραχή ή άλλους τύπους άνοιας, από τον λόγο υγιών ομιλητών. Οι περισσότερες από αυτές τις μελέτες εφαρμόζουν υπολογιστικές μεθόδους εξαγωγής γλωσσικών χαρακτηριστικών που μπορούν να διαφοροποιήσουν τους ασθενείς από τους υγιείς ομιλητές, με στόχο την πρώιμη διάγνωση νευρολογικών διαταραχών μέσω της ανάλυσης λόγου. Πιο συγκεκριμένα, οι de Lira et al. (2011) μελετώντας τον αφηγηματικό λόγο στη ΝΑ, έδειξαν ότι οι ασθενείς είχαν περισσότερα λεξικά λάθη σε σχέση με την ομάδα ελέγχου. Στα λεξικά λάθη υπολογίστηκαν η δυσκολία στην εύρεση λέξεων, οι επαναλήψεις, οι φωνημικές παραφασίες και οι σημασιολογικές υποκαταστάσεις. Επιπλέον, εξετάστηκε ο ρόλος της συντακτικής πολυπλοκότητας και παρατηρήθηκε ότι οι ασθενείς χρησιμοποιούν λιγότερες παρατακτικές και ελλειπτικές δομές. Η μειωμένη χρήση των ελλειπτικών δομών ήταν το χαρακτηριστικό που διαφοροποίησε τις δύο ομάδες μεταξύ τους. Ομοίως, οι Orimaye et al. (2017) παρατήρησαν ότι οι ασθενείς με ΝΑ είχαν δυσκολία στην παραγωγή συντακτικά πολύπλοκων δομών και ελλειπτικών προτάσεων. Επιπρόσθετα, σημειώθηκε σημαντική διαφορά ως προς την παραγωγή του αριθμού των κατηγορημάτων μεταξύ των ασθενών και της ομάδας ελέγχου. Ο αριθμός των κατηγορημάτων ήταν σημαντικά μικρότερος στους ασθενείς συγκριτικά με τους υγιείς. Τέλος, σε ό,τι αφορά τα λεξικά χαρακτηριστικά, παρουσιάστηκε σημαντική διαφορά στη χρήση επαναλήψεων, στην αντικατάσταση λέξεων και στην εμφάνιση μη ολοκληρωμένων λέξεων ανάμεσα στις δύο ομάδες. Η ομάδα των ασθενών χρησιμοποίησε πολύ περισσότερες επαναλήψεις, αντικαταστάσεις λέξεων και ανολοκλήρωτες λεξικές επιλογές. Οι Roark et al. (2011) εφάρμοσαν υπολογιστικές μεθόδους ανάλυσης του προφορικού λόγου για να διακρίνουν άτομα με Ήπιο Γνωστικό Έλλειμμα (Mild Cognitive Impairment) από υγιείς ομιλητές. Συγκεκριμένα, μέτρησαν χαρακτηριστικά γλωσσικής πολυπλοκότητας, όπως λέξεις ανά φράση και πυκνότητα περιεχομένου, αλλά και προσωδιακά χαρακτηριστικά, όπως συχνότητα, μήκος παύσης, συνολικό χρόνο παύσεων και φώνησης. Συμπέραναν ότι ένας συνδυασμός υπολογιστικών μεθόδων μπορεί να διακρίνει τις δύο ομάδες με βάση τις μετρήσεις της γλωσσικής πολυπλοκότητας τους. Οι Garrard et al. (2014) χρησιμοποίησαν μεθόδους μηχανικής μάθησης (όπως τους ταξινομητές Νaive Bayes Gaussian (NBG) και Νaive Bayes Μultinomial (NBM)) για να διακρίνουν τα δείγματα λόγου ατόμων με σημασιολογική άνοια από τον υγιή πληθυσμό, βασιζόμενοι σε λεξικά χαρακτηριστικά απομαγνητοφωνημένων αφηγήσεων. Εντόπισαν ότι στο λεξιλόγιο των ατόμων με σημασιολογική άνοια κυριαρχούν γενικοί και δεικτικοί όροι, όπως «κάτι», «αυτό», καθώς και μεταφηγηματικά εκφωνήματα, όπως «γνωρίζω», «θυμάμαι». Αντιθέτως, στις περιγραφές των υγιών παρατηρούνται λέξεις χαμηλής συχνότητας με υψηλό πλούτο περιεχομένου, όπως τα ουσιαστικά «ακτή», «γρασίδι» και όχι υψηλής συχνότητας και μικρού σημασιολογικού φορτίου, όπως οι γενικοί όροι και οι αόριστες/δεικτικές αντωνυμίες. Οι Rentoumi et al. (2014) χρησιμοποίησαν υπολογιστικές μεθόδους σε τι είδος κειμένου; για την αξιολόγηση ορισμένων λεξικών ποιοτικών και ποσοτικών χαρακτηριστικών (είδος και συχνότητα λέξεων) αλλά και της συντακτικής πολυπλοκότητας για να διακρίνουν τον λόγο ασθενών με μεικτή αγγειακή άνοια από τον λόγο ασθενών με καθαρή άνοια. Συμπέραναν ότι η ομάδα των μεικτών ανοϊκών παρουσιάζειι μειωμένη λεξική ποικιλία και συντακτική πολυπλοκότητα στον προφορικό λόγο σε σχέση με την ομάδα των καθαρά ανοϊκών. Oι Fraser et al. (2016) εφάρμοσαν μια προσέγγιση μηχανικής μάθησης για να μελετήσουν γλωσσικά χαρακτηριστικά στη ΝΑ, όπως σημασιολογικές υποκαταστάσεις, συντακτική πολυπλοκότητα, μήκος ονοματικών, ρηματικών και επιθετικών φράσεων, ποσοστό μερών του λόγου, λεξικό πλούτο, πληροφοριακό περιεχόμενο, επαναλήψεις, ακουστικά/φωνολογικά λάθη, χαρακτηριστικά που μπορούν να εκμαιευτούν αυτόματα από ψηφιακά δείγματα συνεχούς λόγου. Τα αποτελέσματα έδειξαν ένα σημασιολογικό έλλειμμα με αυξημένη χρήση επαναλήψεων, αντωνυμιών και περιορισμένη λεξική ποικιλία. Επιπλέον, εντοπίστηκε συντακτική δυσκολία ως προς την παραγωγή βοηθητικών ρημάτων, γερουνδίων και μετοχών, ενώ στις περιγραφές των ασθενών παρουσιάστηκε χαμηλό πληροφοριακό περιεχόμενο. Τέλος, οι Kavé & Dassa (2018) . χρησιμοποίησαν αυτόματα εργαλεία ανάλυσης κειμένου και ανέλυσαν 10 λεξικά και γραμματικά χαρακτηριστικά: τον συνολικό αριθμό λέξεων, το ποσοστό των λέξεων περιεχομένου σε σχέση με το συνολικό αριθμό λέξεων, τον λόγο των αντωνυμιών, τον λόγο άπαξ λεγομένων και λεξικών τύπων (type-token ratio), τον μέσο όρο στη συχνότητα λέξεων, το ποσοστό των ενεστωτικών ρημάτων, καθώς και των πιο συχνών ρηματικών τύπων, των προθέσεων αλλά και τους δείκτες δευτερεύουσας πρότασης (subordination markers) σε σχέση με το συνολικό αριθμό των λέξεων. Επιπλέον, αναλύθηκαν οι ενότητες πληροφοριακού περιεχομένου (information units), για παράδειγμα, οι ενέργειες στο δείγμα κειμένου που προέκυψε από την περιγραφή της εικόνας . Βρέθηκαν διαφορές μεταξύ των δύο ομάδων σε σχέση με τον συνολικό αριθμό των λέξεων που παρήχθησαν. Συγκεκριμένα, οι ασθενείς με ΝΑ παρήγαγαν πολύ περισσότερες λέξεις σε σύγκριση με τουςυγιείς, αλλά χωρίς πληροφοριακό περιεχόμενο Επίσης, σημειώθηκε υπερβολική χρήση αντωνυμιών σε σχέση με τα ουσιαστικά και οι ασθενείς εμφάνισαν μικρότερο λόγο άπαξ λεγομένων και λεξικών τύπων (type-token ratio), χρησιμοποιώντας τις πιο συχνές λέξεις. 2. Μεθοδολογία 2.1 Συμμετέχοντες Στη μελέτη συμμετείχαν 30 ασθενείς με ΝΑ, ηλικίας 60-85 ετών σε ήπιο ή μεσαίο στάδιο της νόσου (MMSE: 10-25/30). Οι ασθενείς αξιολογήθηκαν με βάση τα διαγνωστικά κριτήρια των McKhann et al. (1984, 2011). Ο εντοπισμός και η διάγνωση του γνωστικού ελλείμματος πραγματοποιήθηκε μέσω της λήψης ιστ
iASiS envisions the transformation of clinical, biological and pharmacogenomic big data into actionable knowledge for personalized medicine and decision makers. Within the context of iASiS this is achieved by integrating and analyzing data from disparate sources, including genomics, electronic health records, and bibliography. The clinical heterogeneity of Alzheimer's disease (AD) presents difficulties for the diagnosis, as well as for the assessment of response to, and therefore evaluation of, new treatments. iASiS paves the way towards personalized medicine for AD patients, by harnessing the potential of big data sources to produce evidence-based clinical knowledge in highly novel and potentially powerful ways. iASiS offers a novel methodology for identifying or confirming associations, responses to treatment, prognosis, and outcomes in AD. The iASiS framework foresees the deployment of an AD-specific knowledge graph. The latter will result from the retrieval, integration and the analysis of AD related biomedical data from heterogeneous resources. iASiS users will be informed on available knowledge relevant to the subject of study. Employing novel inference techniques the iASiS knowledge graph will be able to acquire new knowledge by combining pieces of information that may not be apparent when examining each source separately. The final iASiS system will be a uniquely rich and up to date source of information, which would otherwise be fragmented into different sources. The initial IASIS components tested against a rich data set of biomedical literature comprising more than 150000 textual AD related sources yielded very accurate results. The iASiS components when asked to provide appropriate treatment for AD patients based on the patients’ genetic (allelic) status managed to accurately identify alleles of AD risk with related treatments according to the current bibliography as well as information related to current policies. iASiS will allow the generation of knowledge that will support precision medicine and more effective treatments for Alzheimer's disease. The iASiS project invites health research centers, hospitals, and health organizations to contribute both with pharmacogenomics and clinical data, and to profit from the data processing and analytics that the iASiS platform will offer.
Alzheimer's disease (AD) and other types of dementia are associated with changes in spoken language, but the impact of these changes on different languages has not been extensively examined or compared. In this direction iASiSintends to pave the way towards personalized medicine for AD patients integrating and analyzing data from various sources and languages. In the present study which is performed within the context of iASiS we analysed samples obtained from native speakers of Greek and English who were at mild and moderate stages of probable Alzheimer's disease. We evaluated differences in spoken language between AD patients and normal controls using novel quantitative methods. We searched for the most important sources of variation between the groups in the two languages. Most importantly, we tried to identify AD-induced language characteristics that are either cross-linguistic or language-specific. We adopted a computational approach for the comparison of morpho-syntactic complexity and lexical variety in digitised transcripts of speech produced by AD patients at various stages and by age-matched cognitively normal controls (NC). We used text classification approaches to assign the samples to one of the two groups. The classifiers were tested using various features: morpho-syntactic, lexical as well as complex statistical characteristics. Preliminary findings indicate that syntactic and lexical complexity can be markers of linguistic change in both languages. The method succeeded in finding linguistic characteristics which differentiated AD patients from NC in mild stages in both languages. The cross-linguistic comparison can contribute to a deeper understanding of language deficits in Alzheimer's disease, potentially leading to the development of a cross-linguistic diagnostic tool.
We used a computational linguistic approach, exploiting machine learning techniques, to examine the letters written by King George III during mentally healthy and apparently mentally ill periods of his life. The aims of the study were: first, to establish the existence of alterations in the King’s written language at the onset of his first manic episode; and secondly to identify salient sources of variation contributing to the changes. Effects on language were sought in two control conditions (politically stressful vs. politically tranquil periods and seasonal variation). We found clear differences in the letter corpus, across a range of different features, in association with the onset of mental derangement, which were driven by a combination of linguistic and information theory features that appeared to be specific to the contrast between acute mania and mental stability. The paucity of existing data relevant to changes in written language in the presence of acute mania suggests that lexical, syntactic and stylometric descriptions of written discourse produced by a cohort of patients with a diagnosis of acute mania will be necessary to support the diagnosis independently and to look for other periods of mental illness of the course of the King’s life, and in other historically significant figures with similarly large archives of handwritten documents.
In the present study, we analyzed written samples obtained from Greek native speakers diagnosed with Alzheimer's in mild and moderate stages and from age-matched cognitively normal controls (NC). We adopted a computational approach for the comparison of morpho-syntactic complexity and lexical variety in the samples. We used text classification approaches to assign the samples to one of the two groups. The classifiers were tested using various features: morpho-syntactic and lexical characteristics. The proposed method excels in discerning AD patients in mild and moderate stages from NC leading to the in-depth understanding of language deficits.
Mixed vascular and Alzheimer-type dementia and pure Alzheimer's disease are both associated with changes in spoken language. These changes have, however, seldom been subjected to systematic comparison. In the present study, we analyzed language samples obtained during the course of a longitudinal clinical study from patients in whom one or other pathology was verified at post mortem. The aims of the study were twofold: first, to confirm the presence of differences in language produced by members of the two groups using quantitative methods of evaluation; and secondly to ascertain the most informative sources of variation between the groups. We adopted a computational approach to evaluate digitized transcripts of connected speech along a range of language-related dimensions. We then used machine learning text classification to assign the samples to one of the two pathological groups on the basis of these features. The classifiers' accuracies were tested using simple lexical features, syntactic features, and more complex statistical and information theory characteristics. Maximum accuracy was achieved when word occurrences and frequencies alone were used. Features based on syntactic and lexical complexity yielded lower discrimination scores, but all combinations of features showed significantly better performance than a baseline condition in which every transcript was assigned randomly to one of the two classes. The classification results illustrate the word content specific differences in the spoken language of the two groups. In addition, those with mixed pathology were found to exhibit a marked reduction in lexical variation and complexity compared to their pure AD counterparts.
Owen and Davidson coined the term ‘Hubris Syndrome’ (HS) for a characteristic pattern of exuberant self-confidence, recklessness, and contempt for others, shown by some individuals holding substantial power. Meaning, emotion and attitude are communicated intentionally through language, but psychological and cognitive changes can be reflected in more subtle ways, of which a speaker remains unaware. Of the fourteen symptoms of HS, four imply lexical choices: use of the third person/‘royal we’; excessive confidence; exaggerated self-belief; and supposed accountability to God or History. One other feature (recklessness) could influence language complexity if impulsivity leads to unpredictability. These hypotheses were tested by examining transcribed spoken discourse samples produced by two British Prime Ministers (Margaret Thatcher and Tony Blair) who were said to meet criteria for HS, and one (John Major) who did not. We used Shannon entropy to reflect informational complexity, and temporal correlations (words or phrases whose relative frequency correlated negatively with time in office) and keyness values to identify lexical choices corresponding to periods during which HS was evident. Entropy fluctuated in all three subjects, but consistent (upward) trends in HS-positive subjects corresponded to periods of hubristic behaviour. The first person pronouns ‘I’ and ‘me’ and the word ‘sure’ were among the strongest positive temporal correlates in Blair's speeches. Words and phrases that correlated in the speeches of Thatcher and Blair but not in those of Major included the phrase ‘we shall’ and ‘duties’ (both negative). The keyness ratio of ‘we’ to ‘I’ was clearly higher throughout the terms of office of Thatcher and Blair that at any point in the premiership of Major, and this difference was particularly marked in the case of Blair. The findings are discussed in the context of historical evidence and ideas for enhancing the signal to noise ratio put forward.
Advances in automatic text classification have been necessitated by the rapid increase in the availability of digital documents. Machine learning (ML) algorithms can 'learn' from data: for instance a ML system can be trained on a set of features derived from written texts belonging to known categories, and learn to distinguish between them. Such a trained system can then be used to classify unseen texts. In this paper, we explore the potential of the technique to classify transcribed speech samples along clinical dimensions, using vocabulary data alone. We report the accuracy with which two related ML algorithms [naive Bayes Gaussian (NBG) and naive Bayes multinomial (NBM)] categorized picture descriptions produced by: 32 semantic dementia (SD) patients versus 10 healthy, age-matched controls; and SD patients with left- (n = 21) versus right-predominant (n = 11) patterns of temporal lobe atrophy. We used information gain (IG) to identify the vocabulary features that were most informative to each of these two distinctions.In the SD versus control classification task, both algorithms achieved accuracies of greater than 90%. In the right- versus left-temporal lobe predominant classification, NBM achieved a high level of accuracy (88%), but this was achieved by both NBM and NBG when the features used in the training set were restricted to those with high values of IG. The most informative features for the patient versus control task were low frequency content words, generic terms and components of metanarrative statements. For the right versus left task the number of informative lexical features was too small to support any specific inferences. An enriched feature set, including values derived from Quantitative Production Analysis (QPA) may shed further light on this little understood distinction. (C) 2013 Elsevier Ltd. All rights reserved.