Sycophantic AI distorts social judgments and behaviors.
Empathy plays a central role in human emotional relationships. Empathic accuracy, the ability to accurately appraise another person's emotional state, varies by informational modality and, in humans, is often intertwined with emotional and motivational processes. This study examines whether state-of-the-art Large Language Models (LLMs)—GPT-4, Claude, and Gemini—demonstrate accuracy in emotional-appraisal, and how their accuracy compares to that of humans when presented with only the semantic content (transcripts of recorded videos) detailing ecological, complex autobiographical emotional narratives. We compared the emotional-appraisal of LLMs to the empathic accuracy of human participants (N = 127, randomly sampled students, both in-lab and online) who either read the same transcripts or watched the original videos, which enabled them to use facial and bodily expressions, as well as paralinguistic cues, in addition to semantics. LLMs were able to infer emotional states from semantic content alone with a precision that is equal to or surpasses human performance. This was true both generally and when analyzing positive and negative emotions separately. Theoretically, these findings suggest that semantic information alone can support accuracy in emotional appraisal, though humans may not fully leverage this potential. Practical implications are discussed regarding the use of LLMs in introspective and emotional contexts, while raising critical concerns about privacy, ethical risks, and the potential reshaping of emotional understanding, intimacy, and human connection in an increasingly AI-mediated world.
People devalue empathic responses believed to be edited by artificial intelligence (AI). I propose that this reflects the role of empathy as a predictive social signal: genuine human empathy updates expectations about closeness and future commitment, a function that is weakened when empathy becomes cheap, automated, or outsourced.
Understanding how parents perceive their children's abilities is crucial for family dynamics and intervention strategies, particularly in autism, where accurate parental assessment of social-cognitive capabilities can influence support approaches and developmental outcomes. This study introduces ToM2, a novel measure examining parents' ability to predict their autistic child's Theory of Mind (ToM) performance, representing a form of mentalization that requires parents to evaluate how their child understands others' mental states. We recruited 54 parent-child dyads (43 included in final analyses) from families with children diagnosed with autism (ages 42-70 months). Children completed a six-task ToM scale, while parents predicted their child's responses to each task. ToM2 accuracy was calculated based on the match between parental predictions and child performance. We examined the relationships between ToM2 accuracy and family accommodation for restricted and repetitive behaviors, autism symptom severity, and parental broader autism phenotype characteristics using logistic mixed-effects modeling. Results revealed that parents with higher levels of family accommodation demonstrated significantly lower ToM2 accuracy (p = 0.030), suggesting that higher accommodation is associated with reduced accuracy in perceiving social-cognitive abilities, consistent with bidirectional parent-child interaction patterns. Greater autism symptom severity showed a trend toward reduced ToM2 accuracy (p = 0.051), possibly suggesting that more pronounced autism characteristics may present greater challenges for parental mentalization. Parental broader autism phenotype was not associated with ToM2 accuracy. These findings suggest that ToM2 represents a useful framework for parental mentalization in autism and may inform family-centered interventions targeting both accommodation behaviors and parental perception accuracy.
Greater numbers of people are turning to artificial intelligence (AI) for empathy and emotional support. Here, we review and synthesize recent empirical work on how people perceive empathy from AI versus from humans. Growing evidence points to two dueling effects and a paradox: AI produces language that is rated higher in empathy than language written by humans, but when people perceive text as coming from AI versus a human, they rate it as less empathic. However, despite sometimes rating AI more empathic or having to wait for human empathy, people still show a preference for human empathy. This emerging literature carries significant implications for fundamental research on empathy and for public discourse as the use of AI for emotional support continues to grow.
In the past several years we have seen much interest in how people seek out empathy from artificial intelligence (AI) chatbots, especially Large Language Models like ChatGPT. In this short paper we describe a meta-analysis of several recently-published papers and preprints on this topic. This brief meta-analysis supports a more in-depth narrative review of the literature (Ong et al., 2026). Here, we sought to answer two research questions: Research Question 1: How do people perceive AI-generated empathic responses compared to human-written empathic responses? Research Question 2: How do people perceive empathic responses which are labeled as generated by AI compared to written by humans?
The impact of digitally mediated social interaction on understanding others and sharing their emotions has not been thoroughly investigated. We examined how live, video-mediated interaction, as opposed to watching a prerecorded video, affects behavioral, neural, and physiological aspects of empathy for pain. Thirty-five observers watched targets undergoing painful electric stimulation in an electroencephalogram study. We hypothesized that reduced temporal presence, the immediacy or delay with which information is transferred during social interactions, would result in diminished behavioral and electrophysiological empathic responses. However, observer’s behavioral empathic responses were not diminished with reduced temporal presence. On a neural level, midfrontal theta was sensitive to the other’s pain intensity, and we observed significant physiological coupling between participants. Conversely, Mu suppression was not modulated by pain intensity. Importantly, neural and physiological indices of empathy were independent of temporal presence. However, exploratory analyses indicated a latency effect of temporal presence on pain-related theta activity with an earlier theta increase in interactions with high temporal presence. The results suggest that the temporal presence of individuals may not be necessary for empathy towards another’s pain. Future studies may investigate more naturalistic social interactions and include motivational aspects of empathy. We discuss implications of these findings for debates on social presence and on second-person neuroscience.
Over the past decades, we have seen a fundamentally new form of social interaction become more and more common in our everyday lives: mediated social interactions, in which we do not speak to a person face-to-face, but the interaction takes place via a video call or some other form of technology. While most people would agree that a mediated interaction is somehow different from a face-to-face interaction, we currently lack a framework for conceptualizing what exactly is different. This is especially important when studying how mediated interactions shape our understanding of others' mental states and its neural basis. To address this gap, we propose a framework that conceptualizes the impact of social presence on social interactions and their neural underpinnings along a continuum of two key dimensions: sensory and temporal presence. Whereas sensory presence refers to the amount of sensory information available, temporal presence refers to the immediacy or temporal delay with which information is available. We argue that a systematic manipulation of both dimensions is essential for our understanding of the impact of social presence on social cognition and affect. Our conceptual framework will lay ground for integrating and advancing research on this topic in social neuroscience and beyond.
This research examines how visual and auditory channels affect empathic interactions. Going beyond past work, which focuses mainly on the person experiencing empathy (perceivers), we explore the experiences of empathy recipients (targets) and the shared dyadic experience. Across three studies (N = 710), targets shared emotional experiences with perceivers through online video calls, with cameras either on or off. Hearing the target was sufficient for perceivers to accurately identify a target’s emotions and to be considered attentive and empathic. However, visual information was key to creating shared positive emotions, we-ness, and mutual care and appeared to enhance prosocial behavior. Moreover, perceivers' vocal responsiveness was related to targets' empathy perceptions only when visual information was present. These findings reveal the different roles of visual and auditory information in empathic interaction, suggesting that while hearing each other might suffice for information sharing, optimal empathic interpersonal connection requires seeing and being seen.
Social communication difficulties in autism were traditionally attributed to deficits in empathy, that is, understanding others’ mental states and responding to these with a similar or appropriate emotion, within autistic individuals. The double empathy problem theory proposes that difficulties with cross-neurotype empathy are bidirectional. This study examines predictions from the double empathy problem theory, mainly whether autistic and non-autistic individuals differ in their empathy towards autistic versus non-autistic social targets, using an empathic accuracy paradigm alongside self-report measures of empathy and empathic interest. A novel empathic accuracy stimulus featuring video-recorded autobiographical stories from five autistic and five non-autistic adult storytellers was used. Participants [141 autistic, 94 non-autistic; mean age: 49.09 years (SD = 16.41)] were recruited through the Cambridge Autism Research Database. Each participant viewed two stories (one autistic, one non-autistic storyteller) in a randomized design, continuously rated storytellers’ emotional valence, and globally rated the degree to which they believed the storytellers felt each of 12 discrete emotions. Accuracy was computed as concordance with storytellers’ own ratings. Participants also self-reported empathy and empathic interest toward the storyteller in the target videos. Mixed linear models examined the effects of rater neurotype, target neurotype, and their interaction. Participants’ text responses were qualitatively analyzed using an inductive, data-driven approach to identify patterns within the data, which converged into subthemes and themes. No significant main effects emerged for rater or target neurotype on continuous valence or specific emotions’ empathic accuracy. A trend-level interaction (p=.059, d = 0.13) suggested that autistic raters showed relatively higher continuous valence empathic accuracy toward autistic targets than non-autistic raters. Non-autistic raters reported significantly higher self-reported empathy (p<.001, d=-0.67) and empathic interest (p<.001, d=-0.43) regardless of target neurotype. Qualitative analysis revealed that autistic participants described difficulties in identifying emotions and in performing the EA task, and engaged in metacognitive introspection, while non-autistic participants focused more on task design feedback. The use of video-based experimental design may not fully capture the complexity of spontaneous social encounters. The self-selected sample of autistic participants may not represent the full autism spectrum or include individuals requiring greater support, potentially affecting generalizability. Findings suggest partial support for the double empathy problem theory and potential underestimation of autistic participants of their own empathic abilities. Individual variability patterns caution against treating neurotypes as homogeneous categories that alone determine double empathy processes. Future research should examine real-world interactions and include measures of autistic traits across participants.
People increasingly turn to LLMs for emotional support and understanding how such AI support differs from humans is crucial for evaluating the consequences of this use. We focused on a highly researched emotion regulation strategy, cognitive reappraisal, a strategy for alleviating negative emotions through reinterpretation. We compared trained humans and GPT-4 in their reappraisal ability with vignettes (Study 1, N = 868) and in real-time interactions (Study 2, N = 386). GPT-4 was consistently rated as more effective, empathic, and novel by objective raters and recipients of the reappraisals. Incentivizing humans increased time spent on the reappraisals but did not close the gap (Study 3, N = 1477). Labeling reappraisals as AI reduced participants' evaluation of their effectiveness, though GPT-4 was still considered more effective (Study 4, N = 496). Language analysis revealed differences in vocabulary associated with quality, suggesting that specific language use is what drives AI superiority.
Background and Aims Empathy improves clinical outcomes, patient satisfaction, and adherence to treatment. Few studies have explored the real-world use of large language models in conveying empathy. We compared the empathy in emergency department (ED) discharge letters written by GPT-4 and physicians. Methods We conducted a retrospective, blinded, comparative study in a tertiary ED. All patients discharged for one 8-h shift were included. For each patient, we compared the original ED discharge letter to a GPT-4 generated letter. GPT-4 generated the letters using ED notes, excluding the original discharge letter. Seventeen evaluators (seven physicians, five nurses, five patients) compared the letters side by side. They were blinded to the source. Evaluators first chose between the AI and human letters. Then they rated each letter for empathy, overall quality, clarity of summary, and clarity of recommendations using a 5-point Likert scale. Results Evaluators preferred GPT-4 over physician letters in 83.7% of comparisons (1009 vs. 197; p < 0.001). GPT-4 letters received higher scores for empathy (median 4.0 vs. 3.0; p < 0.001), overall quality, and clarity of summary across all evaluator groups. Among patients, no significant difference was found in the clarity of recommendations ( p = 0.771). Qualitative analysis showed that GPT-4's empathetic expressions, though sometimes generic, were perceived as effective. Conclusion GPT-4 shows strong potential in generating empathetic ED discharge letters. These letters are preferred by healthcare professionals and patients. GPT-4 offers a promising tool to reduce the workload of ED physicians. Further research is necessary to explore patient perceptions and best practices for integrating AI with physicians in clinical practice.
The impact of digitally mediated social interaction on understanding others and sharing their emotions has not been thoroughly investigated. We examined how live, video-mediated interaction as opposed to watching a prerecorded video, affects behavioral, neural, and physiological aspects of empathy for pain. Thirty-five observers watched targets undergoing painful electric stimulation in an electroencephalogram study. We hypothesized that reduced temporal presence would result in diminished behavioral and electrophysiological empathic responses. However, observer’s behavioral empathic responses were not diminished with reduced temporal presence. On a neural level, mid-frontal theta was sensitive to the other’s pain intensity, and we observed significant physiological coupling between participants. Mu suppression, on the other hand, was not modulated by pain intensity. Importantly, neural and physiological indices of empathy were independent of temporal presence. However, exploratory analyses indicated a latency effect of temporal presence on pain-related theta activity with an earlier theta increase in interactions with high temporal presence. The results suggest that the temporal presence of individuals may not be necessary for empathy towards another’s pain. Future studies may investigate more naturalistic social interactions and include motivational aspects of empathy. We discuss implications of these findings for debates on social presence and on second-person neuroscience. ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft, KR3691/12-1
Complex post-traumatic stress disorder (CPTSD) stemming from childhood sexual abuse (CSA) is characterized by profound interpersonal difficulties in later life. Despite the crucial role empathy plays in social functioning, the specific deficits in emotional and cognitive empathy in CPTSD remain understudied. This study employs a rich, naturalistic empathy task to investigate empathic responses among women with CPTSD following childhood sexual abuse. Female participants with CPTSD following CSA and female controls viewed autobiographical videos containing emotional pain, while providing dynamic ratings of their own distress and post-viewing assessments of their own and the targets' emotions. The results showed that the CPTSD group had higher levels of baseline and sustained anxiety compared to the controls, which is consistent with their heightened distress and hyperarousal clinical profiles. Emotional empathy, operationalized as synchrony between participants' and targets' distress ratings, was significantly lower in the CPTSD group, indicating diminished alignment with others' emotional experiences. Cognitive deficits were evident in the systematic underestimation of targets' anger. This study significantly contributes to the understanding of empathy deficits in CPTSD following CSA and potentially informs therapeutic strategies targeted for this population. Specifically, the findings suggest that interventions aimed at improving emotional attunement and fostering the recognition and expression of anger may enhance social functioning and therapeutic outcomes for women with complex trauma stemming from childhood sexual abuse.
Three preregistered studies (N = 533) investigated the relationship between intellectual humility (IH) and cognitive and emotional empathy. Study 1 (n = 212) revealed a positive association between IH and empathic accuracy (EA), especially toward the outgroup. Study 2 (n = 112) replicated the significant association between IH and EA. Study 3 (n = 209) employed a manipulation to enhance IH to demonstrate causality. We found evidence for an indirect effect, wherein the manipulation increased state IH, which was associated with greater EA. A mini meta-analysis revealed that, on average, individuals with higher levels of IH exhibit increased EA, showing a greater understanding of others' emotional states. Moreover, IH predicts empathic resilience-buffering against personal distress while maintaining or increasing empathic concern for others. These findings highlight the positive influence of IH on empathy, emphasizing its potential for fostering deeper connections and better understanding in social interactions.
Artificial Intelligence (AI), and specifically large language models, demonstrate remarkable social-emotional abilities, which may improve human-AI interactions and AI’s emotional support capabilities. However, it remains unclear whether empathy, encompassing understanding, ’feeling with’, and caring, is perceived differently when attributed to AI versus humans. We conducted nine studies (N = 6,282) where AI-generated empathic responses to participants’ emotional situations were labeled as either provided by humans or AI. Human-attributed responses were rated as more empathic and supportive, and elicited more positive and fewer negative emotions, compared to AI-attributed ones. Moreover, participants own uninstructed belief that AI aided human-attributed responses, reduced perceived empathy and support. These effects replicated across varying response lengths, delays, iterations and LLMs, being primarily driven by responses emphasizing emotional sharing and care. Additionally, people consistently chose human interaction over AI when seeking emotional engagement. These findings advance our general understanding of empathy, and specifically human–AI empathic interactions.
Poor sleep is pervasive in modern society. Poor sleep is associated with major physical and mental health consequences, as well as with impaired cognitive function. Less is known about the relationship between sleep and emotional and interpersonal behavior. In this work, we investigate whether poor sleep impairs empathy, an important building block of human interaction and prosocial behavior. We aimed to capture the effects of poor sleep on the various aspects of empathy: trait and state, affect and cognition. Study 1 (n = 155) assessed daily habitual sleep over several days, and global sleep quality in the past month. Participants who reported worse sleep quality exhibited lower empathic caring and perspective-taking traits. Study 2 (n = 347) induced a one-night disruption of sleep continuity to test a causal relationship between sleep and empathy. Participants in the sleep disrupted condition had to briefly wake up five times over the night, whereas the sleep-rested controls slept normally. In the next morning, participants' empathy and prosocial intentions were assessed. Participants in the sleep disruption condition exhibited lower empathic sensitivity and less prosocial decision-making than sleep-rested controls. The main contribution of this work is in providing a robust demonstration of the multi-faceted detrimental effects of poor sleep on trait and state empathy. Our findings demonstrate that poor sleep causally impairs empathic response to the suffering of others. These findings highlight the need for greater public attention to adequate sleep, which may impact empathy on a societal level.
Accurately understanding others’ emotional states is fundamental to effective social functioning. While extensive research exists on how humans recognize different emotions, little is known about how people assess emotional intensity. Through a preliminary survey and seven multi-site studies (n = 2,866), we demonstrate that despite believing they gauge emotions accurately, systematic discrepancies emerge: individuals tend to rate others' emotions as more intense than those individuals rate themselves, particularly for negative emotions. This bias persists across text-based interactions, recorded videos, and live conversations, with both strangers and romantic partners. Interestingly, while people report preferring accurate judgments of their own emotional intensity, the discrepancy may serve adaptive functions, predicting higher empathic responses with strangers and greater relationship satisfaction in romantic relationships. These findings advance understanding of discrepancies in interpersonal emotional perception, highlighting their potential adaptive roles and providing insight into how they shape our social world and relationship outcomes.