INTRODUCTION:The Toolkit to Examine Lifelike Language (TELL) is a web-based application providing speech biomarkers of neurodegeneration. After deployment of TELL v.1.0 in over 20 sites, we now introduce TELL v.2.0. METHODS:First, we describe the app's usability features, including functions for collecting and processing data onsite, offline, and via videoconference. Second, we summarize its clinical survey, tapping on relevant habits (e.g., smoking, sleep) alongside linguistic predictors of performance (language history, use, proficiency, and difficulties). Third, we detail TELL's speech-based assessments, each combining strategic tasks and features capturing diagnostically relevant domains (motor function, semantic memory, episodic memory, and emotional processing). Fourth, we specify the app's new data analysis, visualization, and download options. Finally, we list core challenges and opportunities for development. RESULTS:Overall, TELL v.2.0 offers scalable, objective, and multidimensional insights for the field. CONCLUSION:Through its technical and scientific breakthroughs, this tool can enhance disease detection, phenotyping, and monitoring.
Automated speech and language analysis (ASLA) is a promising approach for capturing early markers of neurodegenerative diseases. However, its potential remains underexploited in research and translational settings, partly due to the lack of a unified tool for data collection, encryption, processing, download, and visualization. Here we introduce the Toolkit to Examine Lifelike Language (TELL) v.1.0.0, a web-based app designed to bridge such a gap. First, we outline general aspects of its development. Second, we list the steps to access and use the app. Third, we specify its data collection protocol, including a linguistic profile survey and 11 audio recording tasks. Fourth, we describe the outputs the app generates for researchers (downloadable files) and for clinicians (real-time metrics). Fifth, we survey published findings obtained through its tasks and metrics. Sixth, we refer to TELL’s current limitations and prospects for expansion. Overall, with its current and planned features, TELL aims to facilitate ASLA for research and clinical aims in the neurodegeneration arena. A demo version can be accessed here: https://demo.sci.tellapp.org/ .
Maximum extractable value (MEV) has been extensively studied. In most papers, the researchers have worked with the Ethereum blockchain almost exclusively. Even though, Ethereum and other blockchains have dynamic gas prices this is not the case for all blockchains; many of them have fixed gas prices. Extending the research to other blockchains with fixed gas price could broaden the scope of the existing studies on MEV. To our knowledge, there is not a vast understanding of MEV in fixed gas price blockchains. Therefore, we propose to study Terra Classic as an example to understand how MEV activities affect blockchains with fixed gas price. We first analysed the data from Terra Classic before the UST de-peg event in May 2022 and described the nature of the exploited arbitrage opportunities. We found more than 188K successful arbitrages, and most of them used UST as the initial token. The capital to perform the arbitrage was less than 1K UST in 50% of the cases, and 80% of the arbitrages had less than four swaps. Then, we explored the characteristics that attribute to higher MEV. We found that searchers who use more complex mechanisms, i.e. different contracts and accounts, made higher profits. Finally, we concluded that the most profitable searchers used a strategy of running bots in a multi-instance environment, i.e. running bots with different virtual machines. We measured the importance of the geographic distribution of the virtual machines that run the bots. We found that having good geographic coverage makes the difference between winning or losing the arbitrage opportunities. That is because, unlike MEV extraction in Ethereum, bots in fixed gas price blockchains are not battling a gas war; they are fighting in a latency war.
Introduction Automated speech analysis has emerged as a scalable, cost-effective tool to identify persons with Alzheimer's disease dementia (ADD). Yet, most research is undermined by low interpretability and specificity. Methods Combining statistical and machine learning analyses of natural speech data, we aimed to discriminate ADD patients from healthy controls (HCs) based on automated measures of domains typically affected in ADD: semantic granularity (coarseness of concepts) and ongoing semantic variability (conceptual closeness of successive words). To test for specificity, we replicated the analyses on Parkinson's disease (PD) patients. Results Relative to controls, ADD (but not PD) patients exhibited significant differences in both measures. Also, these features robustly discriminated between ADD patients and HC, while yielding near-chance classification between PD patients and HCs. Discussion Automated discourse-level semantic analyses can reveal objective, interpretable, and specific markers of ADD, bridging well-established neuropsychological targets with digital assessment tools.
Serotonergic psychedelics are being studied as novel treatments for mental health disorders and as facilitators of improved well-being, mental function and creativity. Recent studies have found mixed results concerning the effects of low doses of psychedelics (“microdosing”) on these domains. However, microdosing is generally investigated using instruments designed to assess larger doses of psychedelics, which might lack sensitivity and specificity for this purpose. Following a double-blind and placebo-controlled experimental design, we explored natural language as a resource to identify speech produced under the acute effects of psilocybin microdoses, focusing on variables known to be affected by higher doses: verbosity, semantic variability and sentiment scores. Except for semantic variability, these metrics presented significant differences between a typical active microdose of 0.5 g of psilocybin mushrooms and an inactive placebo condition. Moreover, machine learning classifiers trained using these metrics were capable of distinguishing between conditions with high accuracy (AUC≈0.8). Our results constitute first proof that low doses of serotonergic psychedelics can be identified from unconstrained natural speech, with potential for widely applicable, affordable, and ecologically valid monitoring of microdosing schedules.
Serotonergic psychedelics have been suggested to mirror certain aspects of psychosis, and, more generally, elicit a state of consciousness underpinned by increased entropy of on-going neural activity. We investigated the hypothesis that language produced under the effects of lysergic acid diethylamide (LSD) should exhibit increased entropy and reduced semantic coherence. Computational analysis of interviews conducted at two different time points after 75 μg of intravenous LSD verified this prediction. Non-semantic analysis of speech organization revealed increased verbosity and a reduced lexicon, changes that are more similar to those observed during manic psychoses than in schizophrenia, which was confirmed by direct comparison with reference samples. Importantly, features related to language organization allowed machine learning classifiers to identify speech under LSD with accuracy comparable to that obtained by examining semantic content. These results constitute a quantitative and objective characterization of disorganized natural speech as a landmark feature of the psychedelic state.
While technology has dramatically changed medical practice, various aspects of mental health practice and diagnosis remain almost unchanged across decades. Here we argue that artificial intelligence — with its capacity to learn and infer from data the workings of the human mind — may rapidly change this scenario. However, this process will not happen without friction and will promote an explicit reflection of the overarching goals and foundational aspects of mental health. We suggest that the converse relation is also very likely to happen. The application of artificial intelligence to a field that relates to the foundations of what makes us human — our volition, our thoughts, our pains and pleasures — may shift artificial intelligence back to its earliest days, when it was mostly conceived of as a laboratory to explore the limits and possibilities of human intelligence.
Widespread commercialization of cannabis has led to the introduction of brand names based on users’ subjective experience of psychological effects and flavors, but this process has occurred in the absence of agreed standards. The objective of this work was to leverage information extracted from large databases to evaluate the consistency and validity of these subjective reports, and to determine their correlation with the reported cultivars and with estimates of their chemical composition (delta-9-THC, CBD, terpenes). We analyzed a large publicly available dataset extracted from Leafly.com where users freely reported their experiences with cannabis cultivars, including different subjective effects and flavour associations. This analysis was complemented with information on the chemical composition of a subset of the cultivars extracted from Psilabs.org . The structure of this dataset was investigated using network analysis applied to the pairwise similarities between reported subjective effects and/or chemical compositions. Random forest classifiers were used to evaluate whether reports of flavours and subjective effects could identify the labelled species cultivar. We applied Natural Language Processing (NLP) tools to free narratives written by the users to validate the subjective effect and flavour tags. Finally, we explored the relationship between terpenoid content, cannabinoid composition and subjective reports in a subset of the cultivars. Machine learning classifiers distinguished between species tags given by “Cannabis sativa” and “Cannabis indica” based on the reported flavours: = 0.828 ± 0.002 (p < 0.001); and effects: = 0.9965 ± 0.0002 (p < 0.001). A significant relationship between terpene and cannabinoid content was suggested by positive correlations between subjective effect and flavour tags (p < 0.05, False-Discovery-rate (FDR)-corrected); these correlations clustered the reported effects into three groups that represented unpleasant, stimulant and soothing effects. The use of predefined tags was validated by applying latent semantic analysis tools to unstructured written reviews, also providing breed-specific topics consistent with their purported subjective effects. Terpene profiles matched the perceptual characterizations made by the users, particularly for the terpene-flavours graph (Q = 0.324). Our work represents the first data-driven synthesis of self-reported and chemical information in a large number of cannabis cultivars. Since terpene content is robustly inherited and less influenced by environmental factors, flavour perception could represent a reliable marker to indirectly characterize the psychoactive effects of cannabis. Our novel methodology helps meet demands for reliable cultivar characterization in the context of an ever-growing market for medicinal and recreational cannabis.
The menstrual cycle affects many aspects of female physiology, from the immune system to behavioral and emotional regulation. It is unclear however if these physiological changes are reflected in everyday, naturalistic language production, and moreover whether these putative effects can be consistently quantified. Using a novel approach based on social networks, we characterized linguistic expression differences in female and male volunteers over the course of several months, while having no physiological or reported information of the female participants’ menstrual cycles. We used a simple algorithm to quantify the linguistic affect intensity of 418 (184 females and 234 males) subjects using their social networks production and found a 7-day modulatory cycle of affect intensity that corresponds to labor-week fluctuations, with no significant difference by biological sex, and a 28-day cycle over which females are significantly different than males. Our results are consistent with the hypothesis that the menstrual cycle modulates affective features of naturalistic linguistic production.
Background Commercially available cannabis strains have multiplied in recent years as a consequence of regional changes in legislation for medicinal and recreational use. Lack of a standardized system to label plants and seeds hinders the consistent identification of particular strains with their elicited psychoactive effects. The objective of this work was to leverage information extracted from large databases to improve the identification and characterization of cannabis strains.Methods We analyzed a large publicly available dataset where users freely reported their experiences with cannabis strains, including different subjective effects and flavour associations. This analysis was complemented with information on the chemical composition of a subset of the strains. Both supervised and unsupervised machine learning algorithms were applied to classify strains based on self-reported and objective features.Results Metrics of strain similarity based on self-reported effect and flavour tags allowed machine learning classification into three major clusters corresponding to Cannabis sativa , Cannabis indica , and hybrids. Synergy between terpene and cannabinoid content was suggested by significative correlations between psychoactive effect and flavour tags. The use of predefined tags was validated by applying semantic analysis tools to unstructured written reviews, also providing breed-specific topics consistent with their purported medicinal and subjective effects. While cannabinoid content was variable even within individual strains, terpene profiles matched the perceptual characterizations made by the users and could be used to predict associations between different psychoactive effects.Conclusions Our work represents the first data-driven synthesis of self-reported and chemical information in a large number of cannabis strains. Since terpene content is robustly inherited and less influenced by environmental factors, flavour perception could represent a reliable marker to predict the psychoactive effects of cannabis. Our novel methodology contributes to meet the demands for reliable strain classification and characterization in the context of an ever-growing market for medicinal and recreational cannabis.* THC : Tetrahydrocannabinol CBD : Cannabidiol AUC : Area under the receiver operating characteristic curve LSA : Latent Semantic Analysis SVD : Singular Value Decomposition PCA : Principal Component Analysis FDR : False Discovery Rate.
Background: Natural speech analytics has seen some improvements over recent years, and this has opened a window for objective and quantitative diagnosis in psychiatry. Here, we used a machine learning algorithm applied to natural speech to ask whether language properties measured before psilocybin for treatment-resistant can predict for which patients it will be effective and for which it will not. Methods: A baseline autobiographical memory interview was conducted and transcribed. Patients with treatment-resistant depression received 2 doses of psilocybin, 10 mg and 25 mg, 7 days apart. Psychological support was provided before, during and after all dosing sessions. Quantitative speech measures were applied to the interview data from 17 patients and 18 untreated age-matched healthy control subjects. A machine learning algorithm was used to classify between controls and patients and predict treatment response. Results: Speech analytics and machine learning successfully differentiated depressed patients from healthy controls and identified treatment responders from non-responders with a significant level of 85% of accuracy (75% precision). Conclusions: Automatic natural language analysis was used to predict effective response to treatment with psilocybin, suggesting that these tools offer a highly cost-effective facility for screening individuals for treatment suitability and sensitivity. Limitations: The sample size was small and replication is required to strengthen inferences on these results.
Language offers a privileged view into the mind; it is the basis by which we infer others’ thoughts. Subtle language disturbance is evident in schizophrenia prior to psychosis onset, including decreases in coherence and complexity, as measured using clinical ratings in familial and clinical high-risk (CHR) cohorts. Bearden et al previously used manual linguistic analysis of baseline speech transcripts in CHR to show that illogical and referential thinking, and poverty of content, predict later psychosis onset. Then, Bedi et al used automated natural language processing (NLP) of CHR transcripts to show that decreased semantic coherence and reduction in syntactic complexity predicted psychosis onset. To determine validity and reproducibility, we have applied automated NLP methods, with machine learning, to Bearden’s original CHR transcripts to identify a language profile predictive of psychosis. Participants in the Bearden UCLA cohort include 59 CHR, of whom 19 developed psychosis (CHR+) within 2 years, whereas 40 did not (CHR-), as well as 16 recent-onset psychosis and 21 healthy individuals, similar in demographics; speech was elicited using Caplan’s “Story Game. Participants in the Bedi NYC cohort include 34 CHR (29 CHR+), with speech elicited using open-ended interview. Speech was audiotaped, transcribed, de-identified and then subjected to latent semantic analysis to determine coherence and part-of-speech tagging to characterize syntactic structure and complexity. A machine-learning speech classifier of psychosis onset was derived from the UCLA CHR cohort, and then applied both to the NYC CHR cohort and to the UCLA psychosis/control comparison, with convex hull (three-dimension depiction of model) and receiver operating characteristics analyses. Correlational analyses with demographics, symptoms and manual linguistic features were also done. A four-factor model language classifier derived from the UCLA CHR cohort that comprised three semantic coherence variables and one syntax (usage of possessive pronouns) predicted psychosis t with accuracy of 83% (intra-protocol) for UCLA CHR, 79% (cross-protocol) for NYC CHR, and 72% for discriminating psychosis from normal speech (UCLA psychosis/control). Convex hulls were defined as the smallest space containing all datapoints within a set for CHR- or healthy controls: these convex hulls showed substantial overlap, with CHR+ and psychosis speech datapoints largely outside these convex hulls. Coherence was associated with age, but speech variables did not vary by gender, race, or socioeconomic status in this study. While automated text features were unrelated to prodromal symptom severity, they were highly correlated with manual text features (r = 0.7, p < .000001). In this small preliminary study, we identified and cross-validated a robust language classifier of psychosis risk that comprised measures of semantic coherence (flow of meaning in language) and syntactic usage (usage of possessive pronouns). This classifier had utility in discriminating speech in individuals with recent-onset psychosis from the norm. It demonstrated concurrent validity in that it was highly correlated with manual linguistic features previously identified by Bearden et al, important as automated methods are fast and inexpensive. Automated language features were unrelated to sex, ethnicity or social class in these small samples, and semantic coherence increased with age, consistent with prior studies of normal language development. Of interest, overlapping convex hulls could be defined for groups of individuals without psychosis (UCLA CHR-, NYC CHR- and UCLA healthy), suggesting a constrained hull of normal language in respect to syntax and semantics, from which pre-psychosis and psychosis speech deviates. The RDoC linguistic corpus-based variables of semantic coherence and syntactic structure hold promise as biomarkers of psychosis risk and expression, with initial validation and reproducibility. Next steps in biomarker development include larger multisite studies with standardization of protocols for speech elicitation, test-retest, and attention to traction/feasibility, acceptability, cost, and utility. Mechanistic studies can also yield neural and physiological correlates of abnormal semantic coherence and syntax.
Language and speech are the primary source of data for psychiatrists to diagnose and treat mental disorders. In psychosis, the very structure of language can be disturbed, including semantic coherence (e.g., derailment and tangentiality) and syntactic complexity (e.g., concreteness). Subtle disturbances in language are evident in schizophrenia even prior to first psychosis onset, during prodromal stages. Using computer‐based natural language processing analyses, we previously showed that, among English‐speaking clinical (e.g., ultra) high‐risk youths, baseline reduction in semantic coherence (the flow of meaning in speech) and in syntactic complexity could predict subsequent psychosis onset with high accuracy. Herein, we aimed to cross‐validate these automated linguistic analytic methods in a second larger risk cohort, also English‐speaking, and to discriminate speech in psychosis from normal speech. We identified an automated machine‐learning speech classifier – comprising decreased semantic coherence, greater variance in that coherence, and reduced usage of possessive pronouns – that had an 83% accuracy in predicting psychosis onset (intra‐protocol), a cross‐validated accuracy of 79% of psychosis onset prediction in the original risk cohort (cross‐protocol), and a 72% accuracy in discriminating the speech of recent‐onset psychosis patients from that of healthy individuals. The classifier was highly correlated with previously identified manual linguistic predictors. Our findings support the utility and validity of automated natural language processing methods to characterize disturbances in semantics and syntax across stages of psychotic disorder. The next steps will be to apply these methods in larger risk cohorts to further test reproducibility, also in languages other than English, and identify sources of variability. This technology has the potential to improve prediction of psychosis outcome among at‐risk youths and identify linguistic targets for remediation and preventive intervention. More broadly, automated linguistic analysis can be a powerful tool for diagnosis and treatment across neuropsychiatry.
BACKGROUND The Menstrual cycle affects many aspects of female physiology, from the immune system to behavioral and emotional regulation. It is unclear however if these physiological changes are reflected in everyday, naturalistic language production, and moreover whether these putative effects can be consistently quantified. OBJECTIVE Using a novel approach based on social networks data, we characterize linguistic expression differences in men and women participants consistent with the expected influence of the menstrual cycle on affect intensity. METHODS We use a simple algorithm to quantify the affect intensity of 418 (184 females and 234 males) subjects using their social networks production. RESULTS We find a 7-day modulatory cycle of affect intensity that corresponds to labor-week fluctuations, with no significant difference by gender, and a 28-day cycle over which females are significantly stronger in females, consistent with the expected effect of the menstrual cycle on affect. CONCLUSIONS Our result supports the hypothesis that language reflects affect modulation during the menstrual cycle.
Psychiatry is an area of medicine that strongly bases its diagnoses on the psychiatrists subjective appreciation.The task of diagnosis loosely resembles the common pipelines used in supervised learning schema.Therefore, we propose to augment the psychiatrists diagnosis toolbox with an artificial intelligence system based on natural language processing and machine learning algorithms.This approach has been validated in many works in which the performance of the diagnosis has been increased with the use of automatic classification.
The massive availability of digital repositories of human thought opens radical novel way of studying the human mind. Natural language processing tools and computational models have evolved such that many mental conditions are predicted by analysing speech. Transcription of interviews and discourses are analyzed using syntactic, grammatical or sentiment analysis to infer the mental state. Here we set to investigate if classification of Bipolar and control subjects is possible. We develop the Emotion Intensity Index based on the Dictionary of Affect, and find that subjects categories are distinguishable. Using classical classification techniques we get more than 75\% of labeling performance. These results sumed to previous studies show that current automated speech analysis is capable of identifying altered mental states towards a quantitative psychiatry.
Inner concepts are much richer than the words that describe them. Our general objective is to inquire what are the best procedures to communicate conceptual knowledge. We construct a simplified and controlled setup emulating important variables of pedagogy amenable to quantitative analysis. To this aim, we designed a game inspired in Chinese Whispers, to investigate which attributes of a description affect its capacity to faithfully convey an image. This is a two player game: an emitter and a receiver. The emitter was shown a simple geometric figure and was asked to describe it in words. He was informed that this description would be passed to the receiver who had to replicate the drawing from this description. We capitalized on vast data obtained from an android app to quantify the effect of different aspects of a description on communication precision. We show that descriptions more effectively communicate an image when they are coherent and when they are procedural. Instead, the creativity, the use of metaphors and the use of mathematical concepts do not affect its fidelity.
We investigate the dynamics of semantic organization using social media, a collective expression of human thought. We propose a novel, time-dependent semantic similarity measure (TSS), based on the social network Twitter. We show that TSS is consistent with static measures of similarity but provides high temporal resolution for the identification of real-world events and induced changes in the distributed structure of semantic relationships across the entire lexicon. Using TSS, we measured the evolution of a concept and its movement along the semantic neighborhood, driven by specific news/events. Finally, we showed that particular events may trigger a temporary reorganization of elements in the semantic network.
BACKGROUND/OBJECTIVES:Psychiatry lacks the objective clinical tests routinely used in other specializations. Novel computerized methods to characterize complex behaviors such as speech could be used to identify and predict psychiatric illness in individuals.AIMS:In this proof-of-principle study, our aim was to test automated speech analyses combined with Machine Learning to predict later psychosis onset in youths at clinical high-risk (CHR) for psychosis.METHODS:Thirty-four CHR youths (11 females) had baseline interviews and were assessed quarterly for up to 2.5 years; five transitioned to psychosis. Using automated analysis, transcripts of interviews were evaluated for semantic and syntactic features predicting later psychosis onset. Speech features were fed into a convex hull classification algorithm with leave-one-subject-out cross-validation to assess their predictive value for psychosis outcome. The canonical correlation between the speech features and prodromal symptom ratings was computed.RESULTS:Derived speech features included a Latent Semantic Analysis measure of semantic coherence and two syntactic markers of speech complexity: maximum phrase length and use of determiners (e.g., which). These speech features predicted later psychosis development with 100% accuracy, outperforming classification from clinical interviews. Speech features were significantly correlated with prodromal symptoms.CONCLUSIONS:Findings support the utility of automated speech analysis to measure subtle, clinically relevant mental state changes in emergent psychosis. Recent developments in computer science, including natural language processing, could provide the foundation for future development of objective clinical tests for psychiatry.