OBJECTIVE:This study applied a machine-learning-based skill assessment system to investigate the association between supportive counseling skills (empathy, open questions, and reflections) and treatment outcomes. We hypothesized that higher empathy and higher use of open questions and reflections would be associated with greater symptom reduction. METHOD:We used a data set with 2,974 sessions, 610 clients, and 48 therapists collected from a university counseling center, which included 845,953 rated therapist statements. Client outcome was routinely monitored by the Counseling Center Assessment of Psychological Symptoms Instruments. Therapists' skills were measured via computer by a bidirectional-long-short-term-memory-based system that rated use of supportive counseling skills. We used multilevel modeling to separate the between-therapist and the within-therapist associations of the skills and outcome. RESULTS:Use of open questions and reflections was associated with client symptom reduction between therapists but not within therapists. We did not find significant associations between therapist empathy and client symptom reduction but found that empathy was negatively associated with clients' baseline symptom level within therapists. CONCLUSIONS:Therapist exploration of clients' experience and expression of understanding may be important skills that are associated with clients' better outcomes. This study highlights the importance of support counseling skills, as well as the potential of machine-learning-based measures in psychotherapy research. We discuss the limitations of the study, including the limitations related to the speaker recognition system and potential reasons for the lack of association between empathy and client outcome. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
The current study sought to advance our understanding of the connections between stress, perceived control, affect, and physiology in daily life. To achieve this goal, we integrated hourly ambulatory physiological and experiential data from young adult participants who experienced work or academic stressors over the course of a day. Participants wore a cardiovascular monitor that recorded heart rate data continuously for 8 h while hourly random Ecological Momentary Assessment (EMA) data were collected in personally relevant settings via mobile phones to learn about stress, perceived control, and affect. The current findings provide a critical advance by demonstrating clear evidence for moderation by perceived control, wherein affective wellbeing was strongly associated with heart rate when one experienced a stressor outside their control. The innovative approach utilized in the current study in real-world settings provides further support for the value of integrating individuals' self-report and physiological experiences (e.g., the role of perceived control), as the information gained can provide critical insights into stress physiology (e.g., heart rate) and wellbeing (e.g., negative affect) connections. The present study thus provides a critical advance to the literature by connecting the literature on daily affect, perceived control, and physiological data streams. This innovation is particularly noteworthy given the general paucity of work that employs ambulatory assessments of physiological responses to daily life.
OBJECTIVE:The purpose of this study was to examine how often clients report discussing cultural identities during counseling sessions; the extent to which discussion of cultural identities during treatment varies across therapists; whether identifying as BIPOC (Black, Indigenous, and people of color) predicts clients' discussion of cultural identities in sessions; and whether differences in the frequency of cultural conversations (i.e., dialogue that focuses on client cultural identities) across client groups depend on the therapist. METHODS:This study examined variation in reports of engagement in cultural conversations during sessions (N=10,731) with 1,997 clients and 72 therapists from a university counseling center. Data were analyzed by using Bayesian multilevel models. RESULTS:Overall, clients reported having cultural conversations in 48.4% of sessions. Cultural conversations were much more likely to occur in sessions with BIPOC clients than with White clients: 66.2% of sessions with BIPOC clients involved conversations about cultural identities, compared with only 39.8% of sessions with White clients. Of note, the magnitude of this difference varied by therapist. CONCLUSION:Cultural conversations were more likely to occur in treatment with BIPOC clients than with White clients, and the presence of cultural conversations in treatment varied by therapist.
We developed an asynchronous online cognitive behavioral therapy (CBT) training tool that provides artificial intelligence- (AI-) enabled feedback to learners across eight CBT skills. We sought to evaluate the technical reliability and to ascertain how practitioners would use the tool to inform product iteration and future deployment. We conducted a single-arm 2-week field trial among behavioral health practitioners who treat outpatients with psychosis. Practitioners (N = 21) were invited to use the AI-enabled CBT training tool over a 2-week (15 days, inclusive) period. To enable naturalistic observation, no adjustments were made to their workloads nor were prescriptions on use provided. We conducted daily assessments and collected backend analytics for all users. At end point, we assessed acceptability, appropriateness, feasibility of implementation, perceived usability, satisfaction, and perceived impact of training. We observed four types of technical issues: broken links, intermittent issues receiving AI-enabled feedback, video replay errors, and an HTML error. Participants averaged 6.57 logins over the 2 weeks, with more than half engaging daily. Most participants (44.7%) engaged for < 30-min increments. Usability scores exceeded industry standard and satisfaction scores indicated good promotion of the tool. All participants endorsed high feasibility, acceptability, and appropriateness. Twelve participants (57%) used the AI-enabled feedback feature; those who did tended to report improved satisfaction, feasibility, and perceived impact of the training. The training tool was used by practitioners in a routine care setting, met or exceeded conventional implementation benchmarks, and may support skill improvement; however, data suggest that practitioners may need support or accountability to fully leverage the training tool. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
The accessibility of training and fidelity assessment is critical to implementing and sustaining empirically supported psychotherapies like cognitive behavioral therapy for psychosis (CBTp). We describe the development of an online CBTp training tool that incorporates behavioral rehearsal tasks to enable deliberate practice of cognitive and behavioral techniques for psychosis. The development process consisted of designing content, inclusive of didactics, client profiles, and learner prompts; constructing standardized performance tasks and metrics; collecting responses to learner prompts; establishing intraclass correlation (ICC) of responses among trained raters; and training a transformer-based machine learning (ML) model to meet or surpass human ICC. Authenticity ratings of each simulated client surpassed benchmarks. CBTp trainers (n = 12), clinicians (n = 78), and nonclinicians (n = 119) generated 3,958 unique verbal responses to 28 unique prompts (7 skills × 4 simulated clients), of which the coding team rated 1,961. Human ICC across all skills was high (mean ICC = 0.77). On average, there was a high correlation between ML and human ratings of fidelity (rs = .74). Similarly, the average percentage of human agreement was high at 96% (range = 87%-102%), where values greater than 100 indicate that the ML model agreed with a human rater more than two human raters agreed with each other. Results suggest that it is possible to reliably measure discrete CBTp skills in response to simulated client vignettes while capturing expected variation in skill utilization across participants. These findings pave the way for a standardized, asynchronous training that incorporates automated feedback on learners' rehearsal of CBTp skills. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Natural language processing (NLP) is a subfield of machine learning that may facilitate the evaluation of therapist–client interactions and provide feedback to therapists on client outcomes on a large scale. However, there have been limited studies applying NLP models to client-outcome prediction that have (a) used transcripts of therapist–client interactions as direct predictors of client-symptom improvement, (b) accounted for contextual linguistic complexities, and (c) used best practices in classical training and test splits in model development. Using 2,630 session recordings from 795 clients and 56 therapists, we developed NLP models that directly predicted client symptoms of a given session based on session recordings of the previous session (Spearman’s ρ = .32, p < .001). Our results highlight the potential for NLP models to be implemented in outcome-monitoring systems to improve quality of care. We discuss implications for future research and applications.
Objective: Counselor assessment of suicide risk is one key component of crisis counseling, and standards require risk assessment in every crisis counseling conversation. Efforts to increase risk assessment frequency are limited by quality improvement tools that rely on human evaluation of con- versations, which is labor intensive, slow, and impossible to scale. Advances in machine learning (ML) have made pos- sible the development of tools that can automatically and immediately detect the presence of risk assessment in crisis counseling conversations. Methods: To train models, a coding team labeled every statement in 476 crisis counseling calls (193,257 statements) for a core element of risk assessment. The authors then fine-tuned a transformer-based ML model with the labeled data, utilizing separate training, validation, and test data sets. Results: Generally, the evaluated ML model was highly consistent with human raters. For detecting any risk as- sessment, ML model agreement with human ratings was 98% of human interrater agreement. Across specific labels, average F1 (the harmonic mean of precision and recall) was 0.86 at the call level and 0.66 at the statement level and often varied as a result of a low base rate for some risk labels. Conclusions: ML models can reliably detect the presence of suicide risk assessment in crisis counseling conversations, presenting an opportunity to scale quality improvement efforts.
Recent scholarship has highlighted the value of therapists adopting a multicultural orientation (MCO) within psychotherapy. A newly developed performance-based measure of MCO capacities exists (MCO-performance task [MCO-PT]) in which therapists respond to video-based vignettes of clients sharing culturally relevant information in therapy. The MCO-PT provides scores related to the three aspects of MCO: cultural humility (i.e., adoption of a nonsuperior and other-oriented stance toward clients), cultural opportunities (i.e., seizing or making moments in session to ask about clients' cultural identities), and cultural comfort (i.e., therapists' comfort in cultural conversations). Although a promising measure, the MCO-PT relies on labor-intensive human coding. The present study evaluated the ability to automate the scoring of the MCO-PT transcripts using modern machine learning and natural language processing methods. We included a sample of 100 participants (n = 613 MCO-PT responses). Results indicated that machine learning models were able to achieve near-human reliability on the average across all domains (Spearman's ρ = .75, p < .0001) and opportunity (ρ = .81, p < .0001). Performance was less robust for cultural humility (ρ = .46, p < .001) and was poorest for cultural comfort (ρ = .41, p < .001). This suggests that we may be on the cusp of being able to develop machine learning-based training paradigms that could allow therapists opportunities for feedback and deliberate practice of some key therapist behaviors, including aspects of MCO. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Importance Use of asynchronous text-based counseling is rapidly growing as an easy-to-access approach to behavioral health care. Similar to in-person treatment, it is challenging to reliably assess as measures of process and content do not scale. Objective To use machine learning to evaluate clinical content and client-reported outcomes in a large sample of text-based counseling episodes of care. Design, Setting, and Participants In this quality improvement study, participants received text-based counseling between 2014 and 2019; data analysis was conducted from September 22, 2022, to November 28, 2023. The deidentified content of messages was retained as a part of ongoing quality assurance. Treatment was asynchronous text-based counseling via an online and mobile therapy app (Talkspace). Therapists were licensed to provide mental health treatment and were either independent contractors or employees of the product company. Participants were self-referred via online sign-up and received services via their insurance or self-pay and were assigned a diagnosis from their health care professional. Exposure All clients received counseling services from a licensed mental health clinician. Main Outcomes and Measures The primary outcomes were client engagement in counseling (number of weeks), treatment satisfaction, and changes in client symptoms, measured via the 8-item version of Patient Health Questionnaire (PHQ-8). A previously trained, transformer-based, deep learning model automatically categorized messages into types of therapist interventions and summaries of clinical content. Results The total sample included 166 644 clients treated by 4973 therapists (20 600 274 messages). Participating clients were predominantly female (75.23%), aged 26 to 35 years (55.4%), single (37.88%), earned a bachelor’s degree (59.13%), and were White (61.8%). There was substantial variability in intervention use and treatment content across therapists. A series of mixed-effects regressions indicated that collectively, interventions and clinical content were associated with key outcomes: engagement (multiple R = 0.43), satisfaction (multiple R = 0.46), and change in PHQ-8 score (multiple R = 0.13). Conclusions and Relevance This quality improvement study found associations between therapist interventions, clinical content, and client-reported outcomes. Consistent with traditional forms of counseling, higher amounts of supportive counseling were associated with improved outcomes. These findings suggest that machine learning–based evaluations of content may increase the scale and specificity of psychotherapy research.
Background The opioid epidemic has resulted in expanded substance use treatment services and strained the clinical workforce serving people with opioid use disorder. Focusing on evidence-based counseling practices like motivational interviewing may be of interest to counselors and their supervisors, but time-intensive adherence tasks like recording and feedback are aspirational in busy community-based opioid treatment programs. The need to improve and systematize clinical training and supervision might be addressed by the growing field of machine learning and natural language-based technology, which can promote counseling skill via self- and supervisor-monitoring of counseling session recordings. Methods Counselors in an opioid treatment program were provided with an opportunity to use an artificial intelligence based, HIPAA compliant recording and supervision platform (Lyssn.io) to record counseling sessions. We then conducted four focus groups—two with counselors and two with supervisors—to understand the integration of technology with practice and supervision. Questions centered on the acceptability of the clinical supervision software and its potential in an OTP setting; we conducted a thematic coding of the responses. Results The clinical supervision software was experienced by counselors and clinical supervisors as beneficial to counselor training, professional development, and clinical supervision. Focus group participants reported that the clinical supervision software could help counselors learn and improve motivational interviewing skills. Counselors said that using the technology highlights the value of counseling encounters (versus paperwork). Clinical supervisors noted that the clinical supervision software could help meet national clinical supervision guidelines and local requirements. Counselors and clinical supervisors alike talked about some of the potential challenges of requiring session recording. Conclusions Implementing evidence-based counseling practices can help the population served in OTPs; another benefit of focusing on clinical skills is to emphasize and hold up counselors’ roles as worthy. Machine learning technology can have a positive impact on clinical practices among counselors and clinical supervisors in opioid treatment programs, settings whose clinical workforce continues to be challenged by the opioid epidemic. Using technology to focus on clinical skill building may enhance counselors’ and clinical supervisors’ overall experiences in their places of work.
Mental health researchers have focused on promoting culturally sensitive clinical care (Herman et al., 2007; Whaley & Davis, 2007), emphasizing the need to understand how biases may impact client well-being. Clients report that their therapists commit racial microaggressions-subtle, sometimes unintentional, racial slights-during treatment (Owen et al., 2014). Yet, existing studies often rely on retrospective evaluations of clients and cannot establish the causal impact of varying ambiguity of microaggressions on clients. This study uses an experimental analogue design to examine offensiveness, emotional reactions, and evaluations of the interaction across three distinct levels of microaggression statements: subtle, moderate, and overt. We recruited 158 adult African American participants and randomly assigned them to watch a brief counseling vignette. We found significant differences between the control and three microaggression statements on all outcome variables. We did not find significant differences between the microaggression conditions. This study, in conjunction with previous correlational research, highlights the detrimental impact of microaggressions within psychotherapy, regardless of racially explicit content. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Researchers have historically focused on understanding therapist multicultural competency and orientation through client self-report measures and behavioral coding. While client perceptions of therapist cultural competency and multicultural orientation and behavioral coding are important, reliance on these methods limits therapists receiving systematic, scalable feedback on cultural opportunities within sessions. Prior research demonstrating the feasibility of automatically identifying topics of conversation in psychotherapy suggests that natural language processing (NLP) models could be trained to automatically identify when clients and therapists are talking about cultural concerns and could inform training and provision of rapid feedback to therapists. Utilizing 103,170 labeled talk turns from 188 psychotherapy sessions, we developed NLP models that recognized the discussion of cultural topics in psychotherapy (F - 1 = 70.0; Spearman's rho = 0.78, p < .001). We discuss implications for research and practice and applications for future NLP-based feedback tools.
Objective: This paper highlights the facilitation of dyadic synchrony as a core psychotherapist skill that occurs at the non-verbal level and underlies many other therapeutic methods. We define dyadic synchrony, differentiate it from similar constructs, and provide an excerpt illustrating dyadic synchrony in a psychotherapy session.Method: We then present a systematic review of 17 studies that have examined the associations between dyadic synchrony and psychotherapy outcomes. We also conduct a meta-analysis of 8 studies that examined whether there is more synchrony between clients and therapists than would be expected by chance.Results: Weighted box score analysis revealed that the overall association of synchrony and proximal as well as distal outcomes was neutral to mildly positive. The results of the meta-analysis indicated that real client-therapist dyad pairs exhibited synchronized behavioral patterns to a much greater extent than a sample of randomly paired people who did not actually speak.Conclusion: Our discussion revolves around how synchrony can be facilitated in a beneficial way, as well as situations in which it may not be beneficial. We conclude with training implications and therapeutic practices.
Supportive counseling skills like empathy and active listening are critical ingredients of all psychotherapies, but most research relies on client or therapist reports of the treatment process. This study utilized machine-learning models trained to evaluate counseling skills to evaluate supportive skill use in 3,917 session recordings. We analyzed overall skill use and variation in practice patterns using a series of mixed effects models. On average, therapists scored moderately high on observer-rated empathy (i.e., 3.8 out of 5), 3.3% of the therapists' utterances in a session were open questions, and 12.9% of their utterances were reflections. However, there were substantial differences in skill use across therapists as well as across clients within-therapist caseloads. These findings highlight the substantial variability in the process of counseling that clients may experience when they access psychotherapy. We discuss findings in the context of both the need for therapists to be responsive and flexible with their clients, but also potential costs related to the lack of a more uniform experience of care. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Abstract When humans interact, they tend to coordinate or synchronize their physiology and behavior spontaneously and mutually. This chapter highlights the facilitation of dyadic synchrony as a core psychotherapist skill that occurs at the nonverbal level and underlies many other therapeutic methods. The chapter defines dyadic synchrony, differentiates it from similar constructs, and provides an excerpt illustrating dyadic synchrony in a psychotherapy session. It also conducts a systematic review of studies that have examined the associations between dyadic synchrony and psychotherapy outcomes and conducts a meta-analysis that examines whether there is more synchrony between clients and therapists than would be expected by chance. The chapter includes situations in which synchrony may not be beneficial. It concludes with diversity considerations, training implications, and therapeutic practices.
Introduction With the increasing utilization of text-based suicide crisis counseling, new means of identifying at risk clients must be explored. Natural language processing (NLP) holds promise for evaluating the content of crisis counseling; here we use a data-driven approach to evaluate NLP methods in identifying client suicide risk. Methods De-identified crisis counseling data from a regional text-based crisis encounter and mobile tipline application were used to evaluate two modeling approaches in classifying client suicide risk levels. A manual evaluation of model errors and system behavior was conducted. Results The neural model outperformed a term frequency-inverse document frequency (tf-idf) model in the false-negative rate. While 75% of the neural model’s false negative encounters had some discussion of suicidality, 62.5% saw a resolution of the client’s initial concerns. Similarly, the neural model detected signals of suicidality in 60.6% of false-positive encounters. Discussion The neural model demonstrated greater sensitivity in the detection of client suicide risk. A manual assessment of errors and model performance reflected these same findings, detecting higher levels of risk in many of the false-positive encounters and lower levels of risk in many of the false negatives. NLP-based models can detect the suicide risk of text-based crisis encounters from the encounter’s content.
Ensuring the effectiveness of text-based crisis counseling requires observing ongoing conversations and providing feedback, both labor-intensive tasks. Automatic analysis of conversations-at the full chat and utterance levels-may help support counselors and provide better care. While some session-level training data (e.g., rating of patient risk) is often available from counselors, labeling utterances requires expensive post hoc annotation. But the latter can not only provide insights about conversation dynamics, but can also serve to support quality assurance efforts for counselors. In this paper, we examine if inexpensive-and potentially noisy-session-level annotation can help improve label utterances. To this end, we propose a logic-based indirect supervision approach that exploits declaratively stated structural dependencies between both levels of annotation to improve utterance modeling. We show that adding these rules gives an improvement of 3.5% f-score over a strong multi-task baseline for utterance-level predictions. We demonstrate via ablation studies how indirect supervision via logic rules also improves the consistency and robustness of the system.
Psychotherapy can be an emotionally laden conversation, where both verbal and non-verbal interventions may impact the therapeutic process. Prior research has postulated mixed results in how clients emotionally react following a silence after the therapist is finished talking, potentially due to studying a limited range of silences with primarily qualitative and self-report methodologies. A quantitative exploration may illuminate new findings. Utilizing research and automatic data processing from the field of linguistics, we analysed the full range of silence lengths (0.2 to 24.01 seconds), and measures of emotional expression - vocally encoded arousal and emotional valence from the works spoken - of 84 audio recordings Motivational Interviewing sessions. We hypothesized that both the level and the variance of client emotional expression would change as a function of silence length, however, due to the mixed results in the literature the direction of emotional change was unclear. We conducted a multilevel linear regression to examine how the level of client emotional expression changed across silence length, and an ANOVA to examine the variability of client emotional expression across silence lengths. Results indicated in both analyses that as silence length increased, emotional expression largely remained the same. Broadly, we demonstrated a weak connection between silence length and emotional expression, indicating no persuasive evidence that silence leads to client emotional processing and expression.