Experiences of positive and negative affect can impact social motivation, including desire and effort towards social connection. Although people with schizophrenia report lower social motivation and more negative affect in daily life than matched controls, they also experience pleasure from social interactions. In the current study, we were interested in examining the extent to which positive and negative affect related to recent social interactions predicted social goal progress and motivation in daily life. Within the context of a digital health intervention (Motivation and Skills Support, or MASS), 30 people with a schizophrenia spectrum disorder worked towards a social goal and completed ecological momentary assessments twice a day for 60 days on affect experienced from recent social interactions, progress on one's goal, and different facets of social motivation. Results showed that participants reported more positive than negative affect in response to social interactions. However, only positive (not negative) affect predicted greater social goal progress and social motivation towards one's goal across the intervention. These results suggest that facilitating positive affect experiences during social interactions, over and above reducing negative affect, may be important for improving social motivation in this population.
Speech emotion recognition (SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach that exploits prior knowledge of testing environments to maximize SER performance under noisy conditions. To address this task, we propose a text-guided, environment-aware training where an SER model is trained with contaminated speech samples and their paired noise description. We use a pre-trained text encoder to extract the text-based environment embedding and then fuse it to a transformer-based SER model during training and inference. We demonstrate the effectiveness of our approach through our experiment with the MSP-Podcast corpus and real-world additive noise samples collected from the Freesound and DEMAND repositories. Our experiment indicates that the text-based environment descriptions processed by a large language model (LLM) produce representations that improve the noise-robustness of the SER system. With a contrastive learning (CL)-based representation, our proposed method can be improved by jointly fine-tuning the text encoder with the emotion recognition model. Under the -5dB signal-to-noise ratio (SNR) level, fine-tuning the text encoder improves our CL-based representation method by 76.4% (arousal), 100.0% (dominance), and 27.7% (valence).
OBJECTIVE:The COVID-19 pandemic has elicited wide-scale general psychological distress; however, longitudinal investigations are required to identify the critical resources that support individuals' adaptation to this type of unique situation over time. Hardiness, a cognitive trait that facilitates adaptation in the context of adversity and possible posttraumatic growth, may be particularly influential on mental health recovery during health disasters when other resources are not available or effective. METHOD:We tested the hypothesis that greater psychological hardiness prior to the pandemic would predict lower traumatic stress symptoms (TSSs) and loneliness early into the pandemic and decreases in TSSs and loneliness between early 2020 and late 2021. Predominantly ethnic minority (77% Latina/o/x or Asian American) female young adults (N = 80; Mage = 25 years; 88% female) attending a minority-serving public university completed a measure of hardiness in January 2020 as well as measures of pandemic-related TSSs and loneliness in April 2020, October 2020, and December 2021. RESULTS:Latent growth curve analyses indicated that hardiness was associated with lower initial loneliness as well as decreases in TSSs and loneliness over time. CONCLUSIONS:Consistent with previous research on adaptation to other potentially traumatic stressors, the current findings suggest that psychological hardiness may play a critical protective role during a global health disaster, both in terms of initial distress and changes in distress over time. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
People with schizophrenia often experience impaired social functioning and low satisfaction with relationships. Existing measures of social impairment in schizophrenia primarily assess pleasure derived from social interactions, rather than examining impairments in social motivation, including effortful behavior. We conducted a validation study of a recently developed self-report measure of social effort in 31 participants with schizophrenia spectrum disorders, including tests of association with standard assessments of social functioning and behavior in daily life using ecological momentary assessment (EMA). We also assessed predictive validity of the scale, measuring the extent to which social effort at baseline predicted changes in social functioning over a 60-day smartphone-based social intervention. Higher social effort was associated with greater social functioning and lower negative symptom severity at baseline. Baseline social effort did not predict changes in social functioning over the intervention period, nor was it related to EMA-reported social experiences. Thus, tendencies toward social effort exertion may capture meaningful variance in gold-standard assessments of social functioning and negative symptoms, but may not track with social experiences in daily life. Further research should examine whether social effort is sensitive to change, and evaluate the utility of targeting social effort in evidence-based interventions for improving the social functioning in schizophrenia.
Speech emotion recognition (SER) system can exploit an Speech enhancement (SE) model to increase its noise robustness by suppressing the background noise. However, SE could also suppress emotionally discriminative features, affecting the emotion prediction. We propose an alternative framework, Keep or Delete (KoD), to keep the information of the original speech while minimizing the influence of background noise. We train a frame reliability predictor that determines clean frames to keep, discarding the noisy frames. We expand this framework by replacing the dropped frames with those extracted from the enhanced speech to keep the lexical information. We refer to this implementation as Keep or Substitute (KoS). Our experiment shows that the KoD model improves the SER results under noisy conditions without fine-tuning the whole model. Also, the KoS framework performs better than enhancing all the frames, indicating the importance of avoiding speech distortion.
An appealing approach for speech emotion recognition (SER) is to pre-train a large speech representation model, such as Wav2Vec2.0 or HuBERT. However, this large model should be adapted to different environments when deployed in real-world applications. This approach demands additional training time and stored parameters for each target environment. This paper proposes a computation and memory-efficient adaptation method. The approach trains skip connection adapters that generate environmental representations from the convolutional encoder, and denoise the self-supervised speech representations. Our experiments with the clean and contaminated versions of the MSP-Podcast corpus show that our adapter-based approach not only improves the performance of the original fine-tuned SER model, but also reduces the computation and memory requirements. For each environment, the approach requires 59.16% less adaptation time and only 0.98% of the parameters of the transformer encoder.
Background Exposure to natural vegetation (ie, “greenspace”) is related to beneficial outcomes, including higher positive and lower negative affect, in individuals with and those without mental health concerns. Researchers have yet to examine dynamic associations between greenspace exposure and affect within individuals over time. Smartphone-based ecological momentary assessment (EMA) and passive sensors (eg, GPS, microphone) allow for frequent sampling of data that may reveal potential moment-to-moment mechanisms through which greenspace exposure impacts mental health. Objective In this study, we examined associations between greenspace exposure and affect (both self-reported and inferred through speech) in people with and those without schizophrenia spectrum disorder (SSD) at the daily level using smartphones. Methods Twenty people with SSD and 14 healthy controls reported on their current affect 3 times per day over 7 days using smartphone-based EMA. Affect expressed through speech was labeled from ambient audio data collected via the phone’s microphone using Linguistic Inquiry and Word Count (LIWC). Greenspace exposure, defined as the normalized difference vegetation index (NDVI), was quantified based on continuous geo-location data collected from the phone’s GPS. Results Overall, people with SSD used significantly more positive affect words (P=.04) and fewer anger words (P=.04) than controls. Groups did not significantly differ in mean EMA-reported positive or negative affect, LIWC total word count, or NDVI exposure. Greater greenspace exposure showed small to moderate associations with lower EMA-reported negative affect across groups. In controls, greenspace exposure on a given day was associated with significantly lower EMA-reported anxiety on that day (b=–0.40, P=.03, 95% CI –0.76 to –0.04) but significantly higher use of negative affect words (b=0.66, P<.001, 95% CI 0.29-1.04). There were no significant associations between greenspace exposure and affect at the daily level among participants with SSD. Conclusions Our findings speak to the utility of passive and active smartphone assessments for identifying potential mechanisms through which greenspace exposure influences mental health. We identified preliminary evidence that greenspace exposure could be associated with improved mental health by reducing experiences of negative affect. Future directions will focus on furthering our understanding of the relationship between greenspace exposure and affect on individuals with and those without SSD.
Studies have shown high performance in the speech emotion recognition (SER) task by fine-tuning a self-supervised speech representation model. Although this model can provide emotionally discriminative embedding in clean conditions, adapting it to a noisy target environment is still required when deployed on real-world applications. For adaptation, it is essential to balance between acquiring new knowledge from noisy speech and keeping the previous knowledge acquired during the pre-training and fine-tuning of the model. Therefore, we propose a contrastive teacher-student learning framework to retrain a self-supervised speech representation model for noisy SER. To keep the knowledge of the original model, we minimize the root mean square error between the clean embeddings from the original SER model and the noisy embeddings from the retrained model. To acquire the discriminative knowledge in the target noisy condition, we also minimize the InfoNCE loss by selecting the corresponding clean embedding as a positive sample and other noisy embeddings with different emotional labels as negative samples. Our experiment with the clean and noisy version of the MSP-Podcast corpus demonstrates that the contrastive teacher-student learning framework can significantly improve the performance of the model only trained with the clean speech in the target noisy condition for all the emotional attributes.
A speech emotion recognition (SER) system deployed on a real-world application can encounter speech contaminated with unconstrained background noise. To deal with this issue, a speech enhancement (SE) module can be attached to the SER system to compensate for the environmental difference of an input. Although the SE module can improve the quality and intelligibility of a given speech, there is a risk of affecting discriminative acoustic features for SER that are resilient to environmental differences. Exploring this idea, we propose to enhance only weak features that degrade the emotion recognition performance. Our model first identifies weak feature sets by using multiple models trained with one acoustic feature at a time using clean speech. After training the single-feature models, we rank each speech feature by measuring three criteria: performance, robustness, and a joint rank ranking that combines performance and robustness. We group the weak features by cumulatively incrementing the features from the bottom to the top of each rank. Once the weak feature set is defined, we only enhance those weak features, keeping the resilient features unchanged. We implement these ideas with the low-level descriptors (LLDs). We show that directly enhancing the weak LLDs leads to better performance than extracting LLDs from an enhanced speech signal. Our experiment with clean and noisy versions of the MSP-Podcast corpus shows that the proposed approach yields a 17.7% (arousal), 21.2% (dominance), and 3.3% (valence) performance gains over a system that enhances all the LLDs for the 10dB signal-to-noise ratio (SNR) condition.
Impaired social functioning contributes to reduced quality of life and is associated with poor physical and psychological well-being in schizophrenia, and thus is a key psychosocial treatment target. Low social motivation contributes to impaired social functioning, but is typically examined using self-report or clinical ratings, which are prone to recall biases and do not adequately capture the dynamic nature of social motivation in daily life. In the current study, we examined the utility of global positioning system (GPS)-based mobility data for capturing social motivation and behavior in people with schizophrenia. Thirty-one participants with schizophrenia engaged in a 60-day mobile intervention designed to increase social motivation and functioning. We examined associations between twice daily self-reports of social motivation and behavior (e.g., number of social interactions) collected via Ecological Momentary Assessment (EMA) and passively collected daily GPS mobility metrics (e.g., number of hours spent at home) in 26 of these participants. Findings suggested that greater mobility on a given day was associated with more EMA-reported social interactions on that day for four out of five examined mobility metrics: number of hours spent at home, number of locations visited, probability of being stationary, and likelihood of following one's typical routine. In addition, greater baseline social functioning was associated with less daily time spent at home and lower probability of following a daily routine during the intervention. GPS-based mobility thus corresponds with social behavior in daily life, suggesting that more social interactions may occur at times of greater mobility in people with schizophrenia, while subjective reports of social interest and motivation are less associated with mobility for this population.
Impaired social functioning contributes to reduced quality of life and is associated with poor physical and psychological well-being in schizophrenia, and thus is a key psychosocial treatment target. Low social motivation contributes to impaired social functioning, but is typically examined using self-report or clinical ratings, which are prone to recall biases and do not adequately capture the dynamic nature of social motivation in daily life. In the current study, we examined the utility of global positioning system (GPS)-based mobility data for capturing social motivation and behavior in people with schizophrenia. Thirty-one participants with schizophrenia engaged in a 60-day mobile intervention designed to increase social motivation and functioning. We examined associations between twice daily self-reports of social motivation and behavior (e.g., number of social interactions) collected via Ecological Momentary Assessment (EMA) and passively collected daily GPS mobility metrics (e.g., number of hours spent at home) in 26 of these participants. Findings suggested that greater mobility on a given day was associated with more EMA-reported social interactions on that day for four out of five examined mobility metrics: number of hours spent at home, number of locations visited, probability of being stationary, and likelihood of following one’s typical routine. In addition, greater baseline social functioning was associated with less daily time spent at home and lower probability of following a daily routine during the intervention. GPS-based mobility thus corresponds with social behavior in daily life, suggesting that more social interactions may occur at times of greater mobility in people with schizophrenia, while subjective reports of social interest and motivation are less associated with mobility for this population.
Speech emotion recognition (SER) system deployed in real-world applications often encounters noisy speech. While most noise compensation techniques consider all acoustic features to have equal impact on the SER model, some acoustic features may be more sensitive to noisy conditions. This paper investigates the noise robustness of each feature in the acoustic feature set. We focus on low-level descriptors (LLDs) commonly used in SER systems. We firstly train SER models with clean speech by only using a single LLD. Then, we rank each LLD with respect to the absolute performance on a development set contaminated with noise, and the relative performance decrease from the results from the models trained with the clean set. Our experiment shows that using all the LLDs leads to worse performance than training the system with a single robust LLD. We propose to select a group of robust features according to their performance and robustness in noisy condition. Without using any compensation method, our feature selection methods improve the performance by 24.4% (arousal), 23.9% (dominance), and 43.2% (valence) in the 10dB noisy condition. Moreover, even though the selection is conducted with the 10dB condition, our selection methods also yield performance improvements in unseen noisy recording conditions.
Digital mental health interventions, such as those provided by smartphone applications (apps), show promise as cost-effective approaches to increasing access to evidence-based psychosocial interventions for psychosis. Although it is well known that limited financial resources can reduce the benefits of digital approaches to mental healthcare, the extent to which cognitive functioning in this population could impact capacity to engage in and benefit from these interventions is less studied. In the current study we examined the extent to which cognitive functioning (premorbid cognitive abilities and social cognition) were related to treatment engagement and outcome in a standalone digital intervention for social functioning. Premorbid cognitive abilities generally showed no association with aggregated treatment engagement markers, including proportion of notifications responded to and degree of interest in working on app content, though there was a small positive association with improvements in social functioning. Social cognition, as measured using facial affect recognition ability, was unrelated to treatment engagement or outcome. These preliminary findings suggest that cognitive functioning is generally not associated with engagement or outcomes in a standalone digital intervention designed for and with people with schizophrenia spectrum disorders.
We examined relationships between general and specific anxiety symptoms and time perspective among adolescents and how these relationships varied by gender. Time perspective was conceptualized as a multidimensional construct and assessed with the following dimensions: time frequency, time attitudes, time orientation, and time relation. Multiple regression analyses indicated that participants (N = 771; Mage = 15.82, SD = 1.23; 54% female) with more anxiety (a) felt less positively about the present; (b) felt more negatively about the past, the present, and the future; and (c) thought the past was more important than the present and the future. Findings for specific anxiety subtypes were generally similar. Interactions between time perspective and gender indicated that more frequent thoughts about the past and the future may be associated with greater anxiety, especially for females. Results highlight the temporal qualities of anxiety and provide support for time perspective as a potential factor for understanding and supporting adolescents with anxiety.
Growing evidence suggests that psilocybin, the active ingredient in hallucinogenic mushrooms, can rapidly and durably improve symptoms of depression, leading to recent breakthrough status designation by the FDA and legalization for mental health treatment in some jurisdictions. Depression in bipolar disorder is associated with significant morbidity and has few effective treatments. However, there is little available scientific data on the risk of psilocybin use in people with bipolar disorder. Individuals with bipolar disorder have been excluded from modern clinical trials, out of understandable concerns of activating mania or worsening the illness course. As psilocybin becomes more available, people with these disorders will likely seek psilocybin treatment for depression and have likely already been doing so in unregulated settings. Our goal here is to summarize the known risks of psilocybin use (and similar substances) in bipolar disorder and to systematically evaluate examples of published case history data, in order to critically evaluate the relative risk of psilocybin as a treatment for bipolar depression. We found 17 cases suggesting that there is potential risk for activating a manic episode, thereby warranting caution. Nonetheless, the relative lack of systematic data or common case examples indicating risk appears to show that a cautious trial, using modern trial methods focusing on appropriate ‘set’ and ‘setting’, targeted at those lowest at risk for mania in the bipolar spectrum (e.g., bipolar 2 disorder), is very much needed, especially given the degree to which depression impacts this population.
Transfer learning is a promising approach to increase performance for many speech-based systems, including voice activity detection (VAD). Domain adaptation, a subfield of transfer learning, often improves model conditioning in the presence of a mismatch between train-test conditions. This study proposes a formulation for VAD based on the teacher-student training, where the teacher model, trained with clean data, transfers knowledge to the student model trained with a noisy, paired version of the corpus resembling the test conditions. The models leverage temporal information using recurrent neural networks (RNN), implemented with either bidirectional long short term memory (BLSTM) or the modern, continuous-state Hopfield network. We provide evidence that in-domain noise emulation for domain adaptation is viable under unconstrained audio channel conditions for VAD "in the wild." Our application domain is in healthcare, where multimodal sensors, including microphones, from portable devices are used to automatically predict social isolation in patients affected by schizophrenia. We empirically show positive results for domain emulation when the training conditions are similar to the target domain. We also show that the Hopfield network outperforms our best BLSTM for VAD on real-world benchmarks.
The serious mental illness (SMI) phenotype is marked by several different symptom domains and biomedical challenges. The nature of SMI renders in-person assessment challenging, due to problems in event recall, response biases, lack of experience in real-world functional domains, and difficulties identifying informants. Digital strategies offer a promising alternative to in-person assessments and allow for remote delivery of cognitive and social cognitive assessments in addition to continuous momentary assessment of activities, moods, symptoms, expressions, experiences, and psychophysiological variables. Remote assessments of mood, emotion, behavior, cognition, and self-assessment have been successfully collected across various SMI conditions. Both active (paging and triggered observations of facial and vocal expressions) and passive (global positioning, actigraphy) methods have been deployed remotely, similarly to in-person assessments previously conducted in the laboratory. Advanced strategies in data analysis are used to examine this information and to guide the development of newer advances in assessment of phenotypic variation in SMI.