The COVID-19 pandemic has adverse consequences on human psychology and behavior long after initial recovery from the virus. These COVID-19 health sequelae, if undetected and left untreated, may lead to more enduring mental health problems, and put vulnerable individuals at risk of developing more serious psychopathologies. Therefore, an early distinction of such vulnerable individuals from those who are more resilient is important to undertake timely preventive interventions. The main aim of this article is to present a comprehensive multimodal conceptual approach for addressing these potential psychological and behavioral mental health changes using state-of-the-art tools and means of artificial intelligence (AI). Mental health COVID-19 recovery programs at post-COVID clinics based on AI prediction and prevention strategies may significantly improve the global mental health of ex-COVID-19 patients. Most COVID-19 recovery programs currently involve specialists such as pulmonologists, cardiologists, and neurologists, but there is a lack of psychiatrist care. The focus of this article is on new tools which can enhance the current limited psychiatrist resources and capabilities in coping with the upcoming challenges related to widespread mental health disorders. Patients affected by COVID-19 are more vulnerable to psychological and behavioral changes than non-COVID populations and therefore they deserve careful clinical psychological screening in post-COVID clinics. However, despite significant advances in research, the pace of progress in prevention of psychiatric disorders in these patients is still insufficient. Current approaches for the diagnosis of psychiatric disorders largely rely on clinical rating scales, as well as self-rating questionnaires that are inadequate for comprehensive assessment of ex-COVID-19 patients’ susceptibility to mental health deterioration. These limitations can presumably be overcome by applying state-of-the-art AI-based tools in diagnosis, prevention, and treatment of psychiatric disorders in acute phase of disease to prevent more chronic psychiatric consequences.
Comprehensive multimodal psychophysiological measurements and smart data analysis based on wearable and low-cost technologies could enhance traditional air traffic controller (ATC) selection process. Many recent studies in neuro-cognitive science and stress resilience illustrated effectiveness of these multimodal measurements and appropriate metrics in comprehensive assessment of ATCs' mental states, such as cognitive workload, cognitive decline, attention deficit, fatigue, emotional and behavioural problems, etc. Accordingly, this article is focused on innovation efforts in ATC selection protocols based on a set of comprehensive stimuli and corresponding multimodal psychophysiological measurements. The concept of enhancement of ATC selection process presented in this article includes complex physiological, oculometric and speech measurements and appropriate metrics. From these multimodal measurements during specific stimulation tasks, which include different versions of acoustic startle stimuli, airblasts, semantically relevant aversive images and sounds, different versions of Stroop tests, visual tracking test, a complex set of multimodal-multidimensional features is computed as predictors of ATC candidates' future performance, like: stress resilience, workload capacity, attention, visual performance, working memory etc. Such cost-effective, more objective, non-invasive preliminary measurements, lasting no longer than 45 minutes may have good discriminative power and might be used in ATC selection processes as enhancement of current selection procedures. Comprehensive analysis of presented multimodal features during different experimental conditions might also be very useful in selection processes of other stressful professional jobs, like first responders, pilots, astronauts etc.
This paper presents a dataset for multimodal classification of cognitive load recorded on a sample of students. The cognitive load was induced by way of performing basic arithmetic tasks, while the multimodal aspect of the dataset comes in the form of both speech and physiological responses to those tasks. The goal of the dataset was two-fold: firstly to provide an alternative to existing cognitive load focused datasets, usually based around Stroop tasks or working memory tasks; and secondly to implement the cognitive load tasks in a way that would make the responses appropriate for both speech and physiological response analysis, ultimately making it multimodal. The paper also presents preliminary classification benchmarks, in which SVM classifiers were trained and evaluated solely on either speech or physiological signals and on combinations of the two. The multimodal nature of the classifiers may provide improvements on results on this inherently challenging machine learning problem because it provides more data about both the intra-participant and inter-participant differences in how cognitive load manifests itself in affective responses.
The human voice is the most frequently used mode of communication among people. It carries both linguistic and paralinguistic information. For an emotion classification task, it is important to process paralinguistic information because it describes the current affective state of a speaker. This affective information can be used for health care purposes, customer service enhancement and in the entertainment industry. Previous research in the field mostly relied on handcrafted features that are derived from speech signals and thus used for the construction of mainly statistical models. Today, by using new technologies, it is possible to design models that can both extract features and perform classification. This preliminary research explores the performance of a model that comprises a convolutional neural network for feature extraction and a deep neural network that performs emotion classification. The convolutional neural network consists of three convolutional layers that filter input spectrograms in time and frequency dimensions and two dense layers forming the deep part of the model. The unified neural network is trained and tested spectrograms of speech utterances from the Berlin database of emotional speech.
Cognitive load classification has seen a boost in popularity lately among the speech analysis community. A number of handmade feature based methods and purely machine learning based methods were presented in the last few years, all trained on a small number of established datasets. This paper presents results of several machine learning methods used on an original dataset of voice samples from a preliminary pilot study into effects of cognitive load. Basic arithmetic problems were presented to the participants with instructions to answer them verbally. Acoustic voice features were extracted from the recorded utterances and modelled using methods like Support Vector Machines and Neural Networks. The accuracies of classification are presented over several conditions for a binary classification task (low cognitive load vs. high cognitive load). The viability of the basic arithmetic task as a dataset for cognitive load classification is discussed. Lessons learned during the analysis are also discussed and present a basis for a stronger experiment design using basic arithmetic tasks in the future.
This paper presents an extensive statistical analysis of the acoustic startle response of two vocal parameters: fundamental frequency (F0) and root-mean-square energy (E), as well as of the orbicularis oculi (eyeblink) surface electromyography (sEMG). An experiment was conducted in which fourteen participants were exposed to acoustic startle stimuli of varying parameters, i.e., intensity level, duration, rise time, and spectral type, during periods of sustained phonation. Voice recordings of the phonations were taken alongside several physiological signals, of which only the sEMG was analyzed in this paper. Response features (peak value, peak time, latency, rise time, fall time, and duration) were extracted on F0, E and sEMG data, and statistical analysis was conducted using linear mixed effects models to show the response behavior with the varying stimuli. The results for vocal F0 and E data were congruent with sEMG data and earlier work in the field. The results demonstrated that vocal analysis can be used as a feasible alternative to the sEMG eyeblink analysis of acoustic startle responses.
This paper presents a comparison of two methods for measuring the acoustic startle response: orbicularis oculi (eyeblink) electromyogram (EMG), which is the conventional measure, and voice fundamental frequency (F0) variations as a consequence of laryngeal muscle innervations. A comparative analysis of the two approaches was performed using statistical methods, as well as system identification modeling using ARX, ARMAX, Output Error and Box-Jenkins models. For this purpose, an experiment was designed in which fourteen participants sustained constant phonation and acoustic startle stimuli of varying parameters were delivered at random time points during the phonation. Physiological signals, including eyeblink EMG, were acquired in parallel to voice recording. The comparative analysis showed that by increasing intensity of acoustic stimulus, response peak amplitudes of both: F0 variations and rectified and smoothed EMG responses increased as well. Therefore, voice F0 may be useful for startle response analysis when eyeblink EMG measurement is unavailable or impractical.