Reduced syntactic complexity is the predominant linguistic impairment unique to the nonfluent agrammatic variant of PPA (naPPA). Reliable, objective and automatic methods for assessing syntactic complexity are currently lacking. Here, we introduce an automatic method for quantifying syntactic complexity in speech. We show that this method differentiates naPPA from other subtypes and monitors disease progression and severity. We analyzed speech samples of picture descriptions that were collected from PPA patients and healthy controls (HC) at the University of Pennsylvania during 2000-2021, including a subset of patients with longitudinal data. Ten syntactic features were automatically extracted from the speech samples and combined into a composite score, using principal component analysis, to represent syntactic complexity. Regression models compared syntactic complexity between PPA phenotypes, covarying for age, sex, education and MMSE. Linear mixed-effects models tested change in syntactic complexity over time. The correlation of syntactic complexity with neurofilament light chain (NfL) levels in the cerebrospinal fluid (CSF) was tested in naPPA and svPPA patients. We examined a total of 136 participants: HC (n=36, 31% males, age 69±8y, MMSE 29±1), naPPA (n=31, 51% males, age 71±8y, MMSE 24±5), lvPPA (n=32, 50% males, age 68±9y, MMSE 22±6) and svPPA (n=37, 49% males, age 64±7y, MMSE 22±7). Z-scored syntactic complexity scores showed significant group differences, with naPPA scoring lower than HC (β=1.08, p =.004), lvPPA (β=1.19, p =.001) and svPPA (β=.93, p =.008) [Fig-1]. Sensitivity to disease progression and severity was found only within the naPPA group. The syntactic score of naPPA significantly decreased over time (-.5 per year, p =.005), significantly different from the 0 decrease in lvPPA ( p =.001) or the .1 decrease in svPPA ( p =.001) [Fig-2]. naPPA patients had a significant inverse correlation of syntactic complexity score and CSF NfL levels (r=-.58, p =.03), while svPPA patients did not (r=.39, p =.9) [Fig-3]. We introduce a novel measure of syntactic complexity which can be extracted automatically from brief naturalistic speech samples. We validated this measure by relating it to naPPA and its biological underpinning. We provided objective evaluation of syntax, which has great potential as a non-invasive biomarker in the dementia research and clinic.
BackgroundReduced syntactic complexity is unique to the nonfluent-agrammatic variant of PPA (naPPA) compared to the semantic variant (svPPA) or the logopenic variant (lvPPA). However, there are no widely agreed, objective methods for quantifying syntactic complexity. Current methods either consider various syntactic features in isolation, or they use indirect measures as a proxy for syntactic complexity. Here, we propose a computational method for quantifying syntactic complexity in natural speech, which directly considers multiple syntactic features and combines them into a composite score. We test the sensitivity of the composite score in monitoring disease progression in naPPA in comparison to other PPA phenotypes. We examined cerebrospinal fluid (CSF) for additional biological validation. We examined two alternative indirect metrics that have been associated with syntactic complexity - namely, syntactic frequency and sentence length - and compared them with our novel score.MethodsSpeech samples of picture descriptions were collected from people with naPPA (n = 35, 50% males, age 70.0 +/- 8), svPPA (n = 37, 49% males, age 64.2 +/- 7) and lvPPA (n = 33, 49% males, age 67.6 +/- 9) and from healthy controls (HC; n = 36, 31% males, age 68.6 +/- 8). Nine syntactic features were automatically identified and tallied from the transcripts. Using principal component analysis (PCA), we calculated a syntactic complexity score that reflects the shared variability across the extracted syntactic features. After verifying group differences, we tested change over time using mixed effect models on a subset of the participants with follow-up recordings (N = 49). For biological validation, we examined the association of syntactic complexity with CSF concentration of neurofilament light chain (NfL).ResultsnaPPA scored lower than HC (beta = 1.19, CI = [0.5, 1.9]) and the other PPA groups (lvPPA: beta = 1.22, CI = [0.5, 1.9]; svPPA: beta = 1.00, CI = [0.3, 1.7]). Only in naPPA did syntactic complexity decrease over time (beta = -0.045, CI = [-0.08, -0.01]) and lower scores associate with increased CSF NfL concentration (beta = -0.44, CI = [-0.87, -0.003]). Sentence length also showed a longitudinal decline, but to a smaller extent (beta = -0.03, CI = [-0.06, -0.0009]). Syntactic frequency showed neither a longitudinal decline nor an association with NfL.DiscussionOur novel syntactic score, derived automatically from recorded naturalistic speech, may directly capture agrammatism and can be used as a clinical outcome assessment tool to capture worsening symptoms in naPPA. Sentence length can be a useful proxy for syntactic complexity.
BACKGROUND:Sentiment in the speech of people with schizophrenia spectrum disorder (SSD) may reflect psychosis severity. Previous research examines speech from semi-structured interviews or self-narrative prompts, where differences in measured sentiment may be driven by differences in life experiences. We measured sentiment in speech evoked from standardised stimuli among participants with a psychotic disorder. METHODS:Two cohorts (N = 97) participated in this study. Symptom domains were assessed using the Brief Psychiatric Rating Scale and were represented as Anxious Depression, Hostile Suspiciousness, Thought Disturbance, and Withdrawal Retardation. Participant speech during picture description tasks was quantified for sentiments: Valence, Arousal, Dominance, Happiness, Sadness, Anger, Fear, Disgust, and Surprise. Correlations between clinical and sentiment measures were conducted separately for the two cohorts and two timepoints in Cohort 1. Within-participant longitudinal relationships were examined with linear mixed models. RESULTS:Several replicable relationships between sentiment and symptom severity were found: two replicable findings among Cohorts 1 and 2 and three replicable findings across Cohort 1 timepoints. Five findings were also generalised to within-participant longitudinal relationships. CONCLUSIONS:Sentiment measures were related to the four symptom domains in the context of standardised stimuli, suggesting a disruption in emotion processing among people with a psychotic disorder.
Singleton mentions, i.e.~entities mentioned only once in a text, are important to how humans understand discourse from a theoretical perspective. However previous attempts to incorporate their detection in end-to-end neural coreference resolution for English have been hampered by the lack of singleton mention spans in the OntoNotes benchmark. This paper addresses this limitation by combining predicted mentions from existing nested NER systems and features derived from OntoNotes syntax trees. With this approach, we create a near approximation of the OntoNotes dataset with all singleton mentions, achieving ~94% recall on a sample of gold singletons. We then propose a two-step neural mention and coreference resolution system, named SPLICE, and compare its performance to the end-to-end approach in two scenarios: the OntoNotes test set and the out-of-domain (OOD) OntoGUM corpus. Results indicate that reconstructed singleton training yields results comparable to end-to-end systems for OntoNotes, while improving OOD stability (+1.1 avg. F1). We conduct error analysis for mention detection and delve into its impact on coreference clustering, revealing that precision improvements deliver more substantial benefits than increases in recall for resolving coreference chains.
Purpose: Multiple methods have been suggested for quantifying syntactic complexity in speech. We compared eight automated syntactic complexity metrics to determine which best captured verified syntactic differences between old and young adults. Method: We used natural speech samples produced in a picture description task by younger ( n = 76, ages 18–22 years) and older ( n = 36, ages 53–89 years) healthy participants, manually transcribed and segmented into sentences. We manually verified that older participants produced fewer complex structures. We developed a metric of syntactic complexity using automatically extracted syntactic structures as features in a multidimensional metric. We compared our metric to seven other metrics: Yngve score, Frazier score, Frazier–Roark score, developmental level, syntactic frequency, mean dependency distance, and sentence length. We examined the success of each metric in identifying the age group using logistic regression models. We repeated the analysis with automatic transcription and segmentation using an automatic speech recognition (ASR) system. Results: Our multidimensional metric was successful in predicting age group (area under the curve [AUC] = 0.87), and it performed better than the other metrics. High AUCs were also achieved by the Yngve score (0.84) and sentence length (0.84). However, in a fully automated pipeline with ASR, the performance of these two metrics dropped (to 0.73 and 0.46, respectively), while the performance of the multidimensional metric remained relatively high (0.81). Conclusions: Syntactic complexity in spontaneous speech can be quantified by directly assessing syntactic structures and considering them in a multivariable manner. It can be derived automatically, saving considerable time and effort compared to manually analyzing large-scale corpora, while maintaining high face validity and robustness. Supplemental Material: https://doi.org/10.23641/asha.24964179
This article describes the MyST corpus developed as part of the My Science Tutor project – one of the largest collections of children's conversational speech comprising approximately 400 hours, spanning some 230K utterances across about 10.5K virtual tutor sessions by around 1.3K third, fourth and fifth grade students. 100K of all utterances have been transcribed thus far. The corpus is freely available (https://myst.cemantix.org) for non-commercial use using a creative commons license. It is also available for commercial use (https://boulderlearning.com/resources/myst-corpus/). To date, ten organizations have licensed the corpus for commercial use, and approximately 40 university and other not-for-profit research groups have downloaded the corpus. It is our hope that the corpus can be used to improve automatic speech recognition algorithms, build and evaluate conversational AI agents for education, and together help accelerate development of multimodal applications to improve children's excitement and learning about science, and help them learn remotely.
The aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, deliver datasets encoded according to these standards, and developing methods for evaluating models carrying out this type of interpretation. Such expansion of the scope of anaphora resolution requires a comparable expansion of the scope of the scorers used to evaluate this work. In this paper, we introduce an extended version of the Reference Coreference Scorer (Pradhan et al., 2014) that can be used to evaluate the extended range of anaphoric interpretation included in the current Universal Anaphora proposal. The UA scorer supports the evaluation of identity anaphora resolution and of bridging reference resolution, for which scorers already existed but not integrated in a single package. It also supports the evaluation of split antecedent anaphora and discourse deixis, for which no tools existed. The proposed approach to the evaluation of split antecedent anaphora is entirely novel; the proposed approach to the evaluation of discourse deixis leverages the encoding of discourse deixis proposed in Universal Anaphora to enable the use for discourse deixis of the same metrics already used for identity anaphora. The scorer was tested in the recent CODI- CRAC 2021 Shared Task on Anaphora Resolution in Dialogues.
Background and Hypothesis Quantitative acoustic and textual measures derived from speech ("speech features") may provide valuable biomarkers for psychiatric disorders, particularly schizophrenia spectrum disorders (SSD). We sought to identify cross-diagnostic latent factors for speech disturbance with relevance for SSD and computational modeling. Study Design Clinical ratings for speech disturbance were generated across 14 items for a cross-diagnostic sample (N = 334), including SSD (n = 90). Speech features were quantified using an automated pipeline for brief recorded samples of free speech. Factor models for the clinical ratings were generated using exploratory factor analysis, then tested with confirmatory factor analysis in the cross-diagnostic and SSD groups. The relationships between factor scores and computational speech features were examined for 202 of the participants. Study Results We found a 3-factor model with a good fit in the cross-diagnostic group and an acceptable fit for the SSD subsample. The model identifies an impaired expressivity factor and 2 interrelated disorganized factors for inefficient and incoherent speech. Incoherent speech was specific to psychosis groups, while inefficient speech and impaired expressivity showed intermediate effects in people with nonpsychotic disorders. Each of the 3 factors had significant and distinct relationships with speech features, which differed for the cross-diagnostic vs SSD groups. Conclusions We report a cross-diagnostic 3-factor model for speech disturbance which is supported by good statistical measures, intuitive, applicable to SSD, and relatable to linguistic theories. It provides a valuable framework for understanding speech disturbance and appropriate targets for modeling with quantitative speech features.
Previous attempts to incorporate a mention detection step into end-to-end neural coreference resolution for English have been hampered by the lack of singleton mention span data as well as other entity information. This paper presents a coreference model that learns singletons as well as features such as entity type and information status via a multi-task learning-based approach. This approach achieves new state-of-the-art scores on the OntoGUM benchmark (+2.7 points) and increases robustness on multiple out-of-domain datasets (+2.3 points on average), likely due to greater generalizability for mention detection and utilization of more data from singletons when compared to only coreferent mention pair matching.
Most existing proposals about anaphoric zero pronoun (AZP) resolution regard full mention coreference and AZP resolution as two independent tasks, even though the two tasks are clearly related. The main issues that need tackling to develop a joint model for zero and non-zero mentions are the difference between the two types of arguments (zero pronouns, being null, provide no nominal information) and the lack of annotated datasets of a suitable size in which both types of arguments are annotated for languages other than Chinese and Japanese. In this paper, we introduce two architectures for jointly resolving AZPs and non-AZPs, and evaluate them on Arabic, a language for which, as far as we know, there has been no prior work on joint resolution. Doing this also required creating a new version of the Arabic subset of the standard coreference resolution dataset used for the CoNLL-2012 shared task (Pradhan et al.,2012) in which both zeros and non-zeros are included in a single dataset.
Much of our understanding of psychosis behavioral phenotypes occurs through the medium of speech – either through the content of what our patients report, or through how it is spoken. Speech is considered the observable surface phenomena that reveal, in part, the shrouded thoughts of the inner "mind." We will first review historical and current understandings of speech and language disturbance in psychosis, then examine data for latent factors, and finally relate these clinical phenomena to distinct computational linguistic features.
Graphical representations of speech generate powerful computational measures related to psychosis. Previous studies have mostly relied on structural relations between words as the basis of graph formation, i.e., connecting each word to the next in a sequence of words. Here, we introduced a method of graph formation grounded in semantic relationships by identifying elements that act upon each other (action relation) and the contents of those actions (predication relation). Speech from picture descriptions and open-ended narrative tasks were collected from a cross-diagnostic group of healthy volunteers and people with psychotic or non-psychotic disorders. Recordings were transcribed and underwent automated language processing, including semantic role labeling to identify action and predication relations. Structural and semantic graph features were computed using static and dynamic (moving-window) techniques. Compared to structural graphs, semantic graphs were more strongly correlated with dimensional psychosis symptoms. Dynamic features also outperformed static features, and samples from picture descriptions yielded larger effect sizes than narrative responses for psychosis diagnoses and symptom dimensions. Overall, semantic graphs captured unique and clinically meaningful information about psychosis and related symptom dimensions. These features, particularly when derived from semi-structured tasks using dynamic measurement, are meaningful additions to the repertoire of computational linguistic methods in psychiatry.
This paper describes the evolution of the Prop-Bank approach to semantic role labeling over the last two decades.During this time the Prop-Bank frame files have been expanded to include non-verbal predicates such as adjectives, prepositions and multi-word expressions.The number of domains, genres and languages that have been PropBanked has also expanded greatly, creating an opportunity for much more challenging and robust testing of the generalization capabilities of PropBank semantic role labeling systems.We also describe the substantial effort that has gone into ensuring the consistency and reliability of the various annotated datasets and resources, to better support the training and evaluation of such systems.
SOTA coreference resolution produces increasingly impressive scores on the OntoNotes benchmark. However lack of comparable data following the same scheme for more genres makes it difficult to evaluate generalizability to open domain data. This paper provides a dataset and comprehensive evaluation showing that the latest neural LM based end-to-end systems degrade very substantially out of domain. We make an OntoNotes-like coreference dataset called OntoGUM publicly available, converted from GUM, an English corpus covering 12 genres, using deterministic rules, which we evaluate. Thanks to the rich syntactic and discourse annotations in GUM, we are able to create the largest human-annotated coreference corpus following the OntoNotes guidelines, and the first to be evaluated for consistency with the OntoNotes scheme. Out-of-domain evaluation across 12 genres shows nearly 15-20% degradation for both deterministic and deep learning systems, indicating a lack of generalizability or covert overfitting in existing coreference resolution models.
This article discusses the requirements of a formal specification for the annotation of temporal information in clinical narratives. We discuss the implementation and extension of ISO-TimeML for annotating a corpus of clinical notes, known as the THYME corpus. To reflect the information task and the heavily inference-based reasoning demands in the domain, a new annotation guideline has been developed, “the THYME Guidelines to ISO-TimeML (THYME-TimeML)”. To clarify what relations merit annotation, we distinguish between linguistically-derived and inferentially-derived temporal orderings in the text. We also apply a top performing TempEval 2013 system against this new resource to measure the difficulty of adapting systems to the clinical domain. The corpus is available to the community and has been proposed for use in a SemEval 2015 task.
SOTA coreference resolution produces increasingly impressive scores on the OntoNotes benchmark. However lack of comparable data following the same scheme for more genres makes it difficult to evaluate generalizability to open domain data. Zhu et al. (2021) introduced the creation of the OntoGUM corpus for evaluating geralizability of the latest neural LM-based end-to-end systems. This paper covers details of the mapping process which is a set of deterministic rules applied to the rich syntactic and discourse annotations manually annotated in the GUM corpus. Out-of-domain evaluation across 12 genres shows nearly 15-20% degradation for both deterministic and deep learning systems, indicating a lack of generalizability or covert overfitting in existing coreference resolution models.
James H. Martin合作论文数Department of Computer Science and ; Center for Spoken Language Research and;University of Colorado;Institute of Cognitive Science 11