Abstract AVATAR therapy is an effective relational therapy for persistent distressing auditory verbal hallucinations (voices). A digital representation of the embodied persecutory voice (avatar) is created and used in a series of dialogues in which the voice hearer is supported to be more assertive and the avatar concedes power. In the first mediation analysis of AVATAR therapy examining the role of power-related constructs, we investigate whether treatment effects on total severity, frequency, and distress of voices are mediated by changes in beliefs about voices and the self, voice relationship appraisals and anxiety. Mediation effects were evaluated in relation to decomposing treatment offer and treatment receipt effects using both Intention to treat (ITT) and Complier Average Causal Effect (CACE) analyses. One hundred and fifty participants from AVATAR1, a randomised control trial (RCT) comparing AVATAR therapy to Supportive Counselling took part in this study, with their baseline and end of treatment (12 weeks) data used. As hypothesised, across both ITT and CACE analyses, reductions in perceived voice omnipotence and increased assertiveness in relation to voices emerged as consistent mediators of AVATAR therapy on reductions in overall severity, frequency and distress of auditory hallucinations compared to SC, whereas voice malevolence, perceived power differential, self-esteem and anxiety did not. Exploratory analysis also indicated that increases in acceptance and autonomy in relation to voices mediated the impact of AVATAR therapy on overall voice severity and distress. This mediation analysis refines our understanding of AVATAR therapy and highlights agency, voice omnipotence and acceptance as intervention targets.
AVATAR therapy involves facilitated dialogs between a voice hearer and a digital embodiment of their distressing voice ("the avatar"). We conducted a multi-site single-blind randomized controlled trial to evaluate the efficacy of brief (AV-BRF) and extended (AV-EXT) forms of AVATAR therapy, compared with treatment as usual (TAU) alone (AVATAR2). This study reports the data from experience sampling method (ESM) assessments conducted at baseline, end of therapy (16 weeks), and follow-up (28 weeks). The research questions focused on whether those in the AV-BRF or AV-EXT arms experienced less voice-related distress, anxiety, and beliefs as measured by ESM, compared to TAU. Separate mixed-effects models were fitted for each research question. The final sample (n = 200) completed approximately 40% of questionnaires across all timepoints. Participants who received AV-EXT therapy, but not AV-BRF, reported reduced momentary voice-related distress at 16 (P = .022) and 28 weeks (p = .029). Appraisals of voice control were also reduced in the AV-EXT arm at 16 weeks when the voice was present (P = .002) or not (P = .008). Voice power appraisals were reduced (P < .035) in both arms when the voice was "not present but on my mind" at all timepoints. There were no changes in the frequency of voice hearing, appraisals of voice intent, or assertive responding. These findings from everyday life, reported for the first time, provide evidence of the impact on the primary AVATAR therapy treatment targets, including appraisals of voice power and control. The weight of evidence favors the AV-EXT protocol in the further development and implementation of AVATAR therapy.
Speaker embeddings are widely used in speaker verification systems and other applications where it is useful to characterise the voice of a speaker with a fixed-length vector. These embeddings tend to be treated as "black box" encodings, and how they relate to conventional acoustic and phonetic dimensions of voices has not been widely studied. In this paper we investigate how state-of-the-art speaker embedding systems represent the acoustic characteristics of speakers as described by conventional acoustic descriptors, age, and gender. Using a large corpus of 10,000 speakers and three embedding systems we show that a small set of 9 acoustic parameters chosen to be "interpretable" predict embeddings about the same as 7 principal components, corresponding to over 50
INTRODUCTION:Around 70% of people with psychosis experience auditory verbal hallucinations (AVHs), which can cause distress and impair the social functioning of the individual. AVATAR therapy works by facilitating a 'face-to-face' dialogue between the person and a digital representation (avatar) of their persecutory voice. Although there is cumulative evidence of this way of working with voices, enhancing the therapeutic focus on improved confidence and sense of control of the voices in social situations represents a promising way to boost generalisation of therapy gains into social contexts. We aim to enhance AVATAR therapy by incorporating immersive Virtual Reality (VR) social environments aiming to help the person to deal better with their voices in daily situations. METHODS AND ANALYSIS:A randomised controlled feasibility trial will be conducted. 40 patients aged 18 or above who are at early stages of psychosis (first episode of psychosis in the last five years) and report distressing and interfering voices will be recruited. Participants will be randomised to receive either a novel, enhanced version of AVATAR therapy (AVATAR_VRSocial) in addition to usual care or usual care alone. Assessor-blinded assessments will be conducted at baseline, 3 months (post-intervention) and 6 months (follow--up). Key therapeutic targets of AVATAR_VRSocial will be those established by the previous evidence of this approach (ie, power and control, self-esteem and future focus), while introducing exposure and management of distressing voices during social interactions. Analyses will focus on feasibility outcomes (recruitment, retention and completion rates) and preliminary estimates of intervention effects. Qualitative interviews will be carried out with participants allocated to AVATAR_VRSocial to gain a comprehensive understanding of participants' views on the acceptability of the intervention and research procedures. Thematic analysis of the qualitative interviews will assess the acceptability of the intervention, trial procedures and the new VR technology and software involved. ETHICS AND DISSEMINATION:The study has received ethical approval from the Ethics Commission at the Faculty of Psychology (Ruhr-Universität Bochum), and there is an independent Trial Steering Committee and Lived Experience Advisory Panel also supporting it. Findings will be disseminated through peer--reviewed publications, conference presentations and science dissemination events. TRIAL REGISTRATION NUMBER:ISRCTN35980117.
AVATAR therapy is an innovative form of relational therapy for the treatment of distressing auditory verbal hallucinations, or voice-hearing, targeted at reducing voice-related distress. AVATAR therapy involves the creation of a digital simulation of a single voice, termed an 'avatar', which is used in a series of three-way therapeutic dialogues. This paper presents the AVATAR Therapy Dialogues Corpus, a specialised corpus containing orthographic transcriptions of AVATAR therapy sessions. We offer an overview of the corpus contents, and a detailed discussion of the design and construction of the corpus. We describe the processes and specialised tools created, transcription conventions, and mark-up designed to capture para-linguistic and non-speech features which may have clinical relevance. Finally, we discuss the potential of the corpus to provide a genuine innovation in clinical care, offering clinicians a data stream that could augment their understanding of patient experiences.
Background AVATAR therapy, a digitally supported intervention, utilises avatars to promote recovery in people who experience distressing auditory hallucinations. This approach was recently evaluated in a multicentre randomised controlled trial comparing brief (AV-BRF) and extended (AV-EXT) forms of therapy with treatment as usual (TAU). There was evidence for the effectiveness of therapy, particularly for AV-EXT. However, value for money needs to be assessed. Aims To compare separately the cost utility of the brief and extended forms of AVATAR therapy with TAU. Method In a three-arm randomised controlled trial the use of health services was measured, and costs (2021/2022; pounds sterling) calculated from a health and social care perspective over a 28-week follow-up period. Quality-adjusted life years (QALYs; derived from the 5-level version of the EuroQol 5-Dimension questionnaire) were combined with costs. Results AV-BRF resulted in extra costs of £319 (95% CI, −£1558 to £2496), and AV-EXT in lower costs of £1965 (95% CI, −£1912 to £1519), compared with TAU. Over the follow-up, AV-BRF resulted in 0.0159 (95% CI, −0.0103 to 0.0422) and AV-EXT in 0.0173 (95% CI, −0.0049 to 0.0395) more QALYs than TAU. The cost per QALY for AV-BRF compared with TAU was £20 016, while AV-EXT dominated TAU (lower costs and more QALYs). Conclusions Neither version of AVATAR had a substantial impact on QALYs. However, AV-EXT did result in reduced care costs − albeit not statistically significant − and was potentially cost-effective compared with TAU. AV-BRF had an incremental cost-effectiveness ratio that indicated lower potential cost-effectiveness. These findings are uncertain, but could still inform decision-making regarding interventions in this field.
AVATAR therapy (AT) works by facilitating a 'face-to-face' dialog between the person and a digital representation (avatar) of their persecutory voice. Although there is cumulative evidence of this way of working with voices, enhancing the therapeutic focus on improved confidence and a sense of control of the voices in social situations represents a promising way to boost the generalization of therapy gains into social contexts. This paper presents a descriptive clinical case example of AVATAR_VRSocial therapy, a new augmented version of AT incorporating immersive Virtual Reality to help the person deal better with their voices in daily situations. "Laura" is a woman who was hearing a very distressing, threatening voice. She felt anxious and distressed when anticipating hearing it and would engage in safety-seeking behaviors to prevent hearing the voice. Laura was supported to stand up to her avatar and regain power over it by using assertive responses, both in active avatar dialog and when exposed to the avatar voice in VR scenarios, which turned into reduced distress when hearing the voice in her everyday life. Laura's dialog with her avatar evolved into a more explicit exploration of the meaning and the purpose of the voice in relation to previous trauma and personal relationships. The additional work in VR appeared to facilitate exposure to social situations while hearing the distressing voice, without performing seeking-safety behaviors, and to allow for practicing strategies to reduce the voice's interference, which evolved from the dialogic sessions with the personalized avatar.
Using avatars to explore communication accommodation through online interactions has become increasingly popular. In this study, an avatar was adopted as a medium and employed by an experimenter who communicated with participants through the avatar using real-time voice conversion, which differs from studies utilizing chatbots with text-to-speech functions. Through the avatar, the experimenter engaged participants in a word-guessing game, during which the avatar's facial expressions and pitch were manipulated. The results indicate that changes in the avatar's pitch significantly influenced participants' pitch height and range, demonstrating accommodation to the avatar's pitch. However, this accommodation was observed to be non-mutual. In contrast, neither changes in the avatar's facial expressions nor the interaction between pitch and facial expressions had a significant effect on participants' pitch. Overall, the successful replication of pitch accommodation observed in face-to-face interactions highlights the feasibility of using avatars to investigate communication accommodation.
As users are only too aware, contemporary large vocabulary speech recognition systems do not respond to speech in the same way as humans. The dictation systems that are in use today are very sensitive to disfluencies, restarts, background noise and change of speaker or voice quality. Furthermore the recognition mistakes they make seem to be very different to the ones that humans make even when listening in poor environments. There is no doubt that recognition systems will only become more comfortable to use when they act more like a human listener. This should mean that scientific knowledge about how humans process speech is relevant and important in the design of these systems. Unlike the situation in the early days of the field, it is now the case that scientific research into the human processing of language has diverged from research into systems. We now have separate and independent fields of ‘psycholinguistics’ and ‘spoken language engineering’. This article explores the relationship between the engineering and cognitive science communities within the relatively well-defined sub-field of spoken word recognition. That is we shall be mainly concerned with the processes by which word sequences are recovered from acoustic input. The article is in three parts: the roots of the divergence between engineering and cognitive science accounts of word recognition are explored in the first part. Differences in motivation, methodology and culture are all seen to play a part and are explored in a historical context. The second part of the article discusses the potential benefits of a re-convergence of the two scientific fields and argues that the time is ripe for progress now. Engineering systems are stable and successful enough to be worth interpreting in cognitive terms, while they are sophisticated enough to allow useful comparisons with humans to be undertaken. The final part of the article proposes some elements of a joint research programme which could act as a stimulus for the two communities to work together. Highlighted are the cognitive accounts of priming phenomena which relate to recent engineering work in adaptation, and cognitive accounts of morphological processing which relate to engineering problems of vocabulary selection and use. Other possibilities relate to phonetic reduction phenomena at the low end, and semantic grouping or phrasing at the high end of both human and machine recognition.
A speech intelligibility prediction model for hearing impaired listeners would be useful in the development of better signal enhancement methods and for the fitting of hearing aids. Most current prediction models use only information from a pure-tone audiogram to characterise impaired listeners, although evidence suggests that listeners vary in ways not captured by pure-tone thresholds. In this paper we evaluate a model in which each listener is described by three factors: average pure-tone thresholds, sensitivity to phonetic distortion and sensitivity to word likelihood. We build and evaluate the model using the corpus collected by the second Clarity Prediction Challenge, which contains over 13,000 intelligibility judgments by 31 hearing impaired listeners. We describe how the factors were estimated and test their independence. We show that incorporating the listener-dependent factors into an existing intelligibility metric can improve the accuracy of prediction on held-out test data with a 9.8% relative improvement in prediction error.
Distressing voices are a core symptom of psychosis, for which existing treatments are currently suboptimal; as such, new effective treatments for distressing voices are needed. AVATAR therapy involves voice-hearers engaging in a series of facilitated dialogues with a digital embodiment of the distressing voice. This randomized phase 2/3 trial assesses the efficacy of two forms of AVATAR therapy, AVATAR-Brief (AV-BRF) and AVATAR-Extended (AV-EXT), both combined with treatment as usual (TAU) compared to TAU alone, and conducted an intention-to-treat analysis. We recruited 345 participants with psychosis; data were available for 300 participants (86.9%) at 16 weeks and 298 (86.4%) at 28 weeks. The primary outcome was voice-related distress at both time points, while voice severity and voice frequency were key secondary outcomes. Voice-related distress improved, compared with TAU, in both forms at 16 weeks but not at 28 weeks. Distress at 16 weeks was as follows: AV-BRF, effect -1.05 points, 96.5% confidence interval (CI) = -2.110 to 0, P = 0.035, Cohen's d = 0.38 (CI = 0 to 0.767); AV-EXT -1.60 points, 96.5% CI = -3.133 to -0.058, P = 0.029, Cohen's d = 0.58 (CI = 0.021 to 1.139). Distress at 28 weeks was: AV-BRF, -0.62 points, 96.5% CI = -1.912 to 0.679, P = 0.316, Cohen's d = 0.22 (CI = -0.247 to 0.695); AV-EXT -1.06 points, 96.5% CI = -2.700 to 0.586, P = 0.175, Cohen's d = 0.38 (CI = -0.213 to 0.981). Voice severity improved in both forms, compared with TAU, at 16 weeks but not at 28 weeks whereas frequency was reduced in AV-EXT but not in AV-BRF at both time points. There were no related serious adverse events. These findings provide partial support for our primary hypotheses. AV-EXT met our threshold for a clinically significant change, suggesting that future work should be primarily guided by this protocol. ISRCTN registration: ISRCTN55682735 .
Aim There is growing interest in tailoring psychological interventions for distressing voices and a need for reliable tools to assess phenomenological features which might influence treatment response. This study examines the reliability and internal consistency of the Voice Characterisation Checklist (VoCC), a novel 10-item tool which assesses degree of voice characterisation, identified as relevant to a new wave of relational approaches. Methods The sample comprised participants experiencing distressing voices, recruited at baseline on the AVATAR2 trial between January 2021 and July 2022 ( n = 170). Inter-rater reliability (IRR) and internal consistency analyses (Cronbach’s alpha) were conducted. Results The majority of participants reported some degree of voice personification (94%) with high endorsement of voices as distinct auditory experiences (87%) with basic attributes of gender and age (82%). While most identified a voice intention (75%) and personality (76%), attribution of mental states (35%) to the voice (‘What are they thinking?’) and a known historical relationship (36%) were less common. The internal consistency of the VoCC was acceptable (10 items, α = 0.71). IRR analysis indicated acceptable to excellent reliability at the item-level for 9/10 items and moderate agreement between raters’ global (binary) classification of more vs. less highly characterised voices, κ = 0.549 (95% CI, 0.240–0.859), p < 0.05. Conclusion The VoCC is a reliable and internally consistent tool for assessing voice characterisation and will be used to test whether voice characterisation moderates treatment outcome to AVATAR therapy. There is potential wider utility within clinical trials of other relational therapies as well as routine clinical practice.
In this paper we evaluate the hypothesis that automated methods for diagnosis of voice disorders from speech recordings would benefit from contextual information found in continuous speech. Rather than basing a diagnosis on how disorders affect the average acoustic properties of the speech signal, the idea is to exploit the possibility that different disorders will cause different acoustic changes within different phonetic contexts. Any differences in the pattern of effects across contexts would then provide additional information for discrimination of pathologies. We evaluate this approach using two complementary studies: the first uses a short phrase which is automatically annotated using a phonetic transcription, the second uses a long reading passage which is automatically annotated from text. The first study uses a single sentence recorded from 597 speakers in the Saarbrucken Voice Database to discriminate structural from neurogenic disorders. The results show that discrimination performance for these broad pathology classes improves from 59% to 67% unweighted average recall when classifiers are trained for each phone-label and the results fused. Although the phonetic contexts improved discrimination, the overall sensitivity and specificity of the method seems insufficient for clinical application. We hypothesise that this is because of the limited contexts in the speech audio and the heterogeneous nature of the disorders. In the second study we address these issues by processing recordings of a long reading passage obtained from clinical recordings of 60 speakers with either Spasmodic Dysphonia or Vocal fold Paralysis. We show that discrimination performance increases from 80% to 87% unweighted average recall if classifiers are trained for each phone-labelled region and predictions fused. We also show that the sensitivity and specificity of a diagnostic test with this performance is similar to other diagnostic procedures in clinical use. In conclusion, the studies confirm that the exploitation of contextual differences in the way disorders affect speech improves automated diagnostic performance, and that automated methods for phonetic annotation of reading passages are robust enough to extract useful diagnostic information.