
This study examines how listeners evaluate human-composed and AI-generated music across three genres, hip-hop, country, and synthwave, using a 2 (source: human-created vs. AI-generated) × 3 (genre) mixed experimental design. Drawing on Uses and Gratifications Theory (UGT) and Authenticity Theory, this study examines the roles of perceived realism, listening satisfaction, ethical concern, AI familiarity, and gratification-seeking motivations in predicting music acceptability. Results show that perceived realism and listening satisfaction are the strongest and most consistent predictors of acceptability, together explaining 43% of variance in the dependent measure (averaged across genre). While participants (N = 401) prefer music they believe to be human made, they systematically misattributed AI-generated country and synthwave tracks as human. Additionally, ethical concern negatively predicted acceptability for synthwave only. Furthermore, affective gratification-seeking positively predicted acceptability for hip-hop and synthwave. Findings suggest that experiential quality drives AI music acceptance more than source knowledge, with implications for AI music development, streaming platforms, and media psychology research.
Despite the potential of AI chatbots for personalized learning, empirical evidence for their effectiveness in addressing STEM gender disparities remains limited. This study explores the efficacy of ADA, a female-coded AI chatbot, in mitigating mathematical gender stereotypes while promoting equitable learning outcomes. A quasi-experimental study (N = 195 ninth-grade German students, mean age = 14.33 years) conducted in mathematics classes compared the AI chatbot with traditional printed help cards on the Heron method. The study measured stereotypical beliefs, technological acceptance, emotions, situational interest, cognitive load, and academic performance. Results demonstrate a significant reduction in gender-stereotypical beliefs, with small but meaningful effect sizes observed across both genders. Technological acceptance was high, exhibiting no gender differences. The intervention promoted learning efficacy for both genders, with a significant gain in situational interest and no statistically significant gender differences in learning performance, emotional responses, or cognitive load. These findings underscore the potential of systematically designed AI chatbots as scalable pedagogical tools that simultaneously challenge gender stereotypes and promote equitable learning efficacy. Subsequent longitudinal research is required to validate the long-term sustainability and generalizability of these effects.
As artificial intelligence systems evolve into active communicative agents, how does their hyper-fluent discourse affect how humans judge the soundness of their arguments? To investigate this question, the present study examined judgments of valid vs. fallacious arguments that were authored by humans or by an AI-agent, and presented in human-associated or AI-associated interfaces (Instant Messenger vs. Chat GPT-3). Employing a 2x2x2 within subjects factorial design, the study included 155 U.S. undergraduates (ages 18–24). Although participants were able to distinguish between valid and fallacious arguments, particularly in human-associated interfaces, AI-authored arguments – regardless of the interface they appeared in - were judged to be significantly more valid than human-generated ones, an advantage consistent with the measurably greater lexical sophistication of AI-generated text rather than with ease of processing. However, the visual presence of the AI interface appeared to reduce participants' ability to distinguish between valid and fallacious arguments. Furthermore, human-authored text when presented within the AI interface was penalized for violating the hyper-rational expectations of the machine script. Explicit attitudes toward AI did not moderate these effects, suggesting that a "machine heuristic" functions as a cognitive default. These findings highlight a potential epistemological risk in human-AI interaction, where algorithmic sophistication may mask logical errors and the mere visual presence of a machine interface may dampen critical scrutiny.
Generative artificial intelligence (GenAI) has become increasingly prevalent in organizational work contexts, yet little is known about how employees’ early use of GenAI relates to their work experiences before broader organizational transformation occurs. Guided by the challenge–hindrance stress model (CHM) as an interpretive framework, this study examined whether more frequent GenAI use is associated with employees’ work outcomes and whether GenAI literacy moderates these associations during a pretransformation phase of organizational GenAI adoption. A cross-sectional quantitative online survey was conducted among administrative employees of a global industrial company shortly before an organization-wide GenAI implementation in early 2025. The sample comprised 1,001 employees, most of whom were based in Europe, particularly Germany. Structural equation modeling (SEM) analyses indicated small positive associations between GenAI usage frequency and work engagement as well as quantitative job performance. In addition, GenAI literacy moderated the associations of GenAI usage frequency with qualitative job performance, quantitative job performance, and meaningfulness of work. These findings suggest that in an early, pretransformation phase of organizational GenAI adoption, GenAI use is not uniformly associated with better work outcomes and appears to depend in part on employees’ GenAI literacy. The study highlights the need for an employee-centered approach that considers employee capabilities alongside technological access for successful GenAI adoption in organizations.
The rapid popularization of artificial intelligence (AI) conversational agents (CAs) has sparked growing scientific and public interest in the nature and consequences of human-AI relationships. However, what we empirically know about these relationships and their social effects remains scattered. This scoping review consolidates the study designs and methodological approaches in the current literature on human-AI relationships, as well as the AI systems and populations studied, their focus areas, and key empirical findings. Following JBI methodology, three databases (Scopus, PubMed, PsycINFO) were searched in September 2025, supplemented by Google Scholar, arXiv, and reference list screening to maximize coverage. Of the 2,280 papers initially identified, 89 met the inclusion criteria and were independently extracted by two reviewers. The literature is mostly characterized by qualitative and correlational designs, with many studies focusing on the AI companion platform Replika. Studies were organized around three thematic orientations: mechanisms underlying engagement, relational experiences with CAs, and psychosocial effects. Qualitative data suggest that loneliness was a frequently reported motivator for CA use, and users often reported perceived benefits, including reduced loneliness, emotional support, and improved social confidence, as well as positive erotic and relational experiences. However, quantitative experimental and longitudinal evidence remains limited and mixed, and the effects of AI companionship appear contingent on individual characteristics, contextual conditions, and patterns of use rather than being uniform across users. Longitudinal experimental research, greater diversity in the platforms and populations investigated, and regular systematic syntheses of evidence are needed to advance science on human-AI relationships.
Debates on generative and agentic AI are often framed through two simplified images: AI as a productivity tool and AI as an autonomous agent. The first tends to understate how AI can help actors move across fragmented domains of expertise; the second tends to overstate the extent to which current systems can independently navigate the institutional, relational, and normative complexity of real-world action. This paper argues that AI can function as a bridge across expertise, but only within broader forms of human, institutional, and technical coordination. The claim is developed through an analysis of the Conyngham–Rosie episode, a widely reported case in which Paul Conyngham, seeking treatment for his dog Rosie’s terminal cancer, used AI tools together with scientific collaboration and institutional support to advance a personalized experimental treatment pathway. The case is best understood not as expert replacement, but as an instance of hybrid sociotechnical agency: a distributed form of effective action emerging from the interaction of human initiative, AI-supported analytical mediation, expert knowledge, institutional processes, and practical constraints. On this basis, the paper develops a bridge-across-expertise account of generative AI, introduces hybrid sociotechnical agency as the configuration through which such bridging becomes practically effective, and identifies five structural limits of autonomous systems in open human environments. The paper concludes that a more promising direction for agentic AI lies less in autonomous substitution than in the design of accountable forms of human–AI–institutional collaboration.
With the growing prevalence of e-commerce platforms, product reviews have become a critical factor influencing consumer purchasing decisions. As the significance of online reviews has increased, so too has the sophistication of deceptive review generation. Recent advances in Artificial Intelligence (AI), particularly Large Language Models (LLMs), enable malicious actors to generate highly realistic, coherent, and contextually relevant fake reviews that closely resemble genuine user feedback. Consequently, conventional Fake Review Detection (FRD) methods increasingly struggle to identify AI-generated deceptive content, threatening consumer trust, market fairness, and the credibility of online marketplaces. This growing challenge necessitates a comprehensive understanding of recent advances in FRD. Accordingly, this review systematically analyses FRD research published between 2020 and 2025, with particular emphasis on AI-generated fake reviews in e-commerce environments. Unlike existing surveys that primarily focus on traditional human-written deceptive reviews and conventional detection techniques, this work presents a comprehensive taxonomy encompassing machine learning, deep learning, transformer-based, graph-based, swarm intelligence, and hybrid approaches. It further provides a comparative analysis of datasets, feature engineering strategies, evaluation metrics, multilingual studies, and application domains while identifying key research challenges, including explainability, dataset validity, cross-domain and cross-lingual generalization, adversarial adaptation, and real-time deployment. Finally, future research directions and a multi-stakeholder roadmap are proposed to support trustworthy, adaptive, and scalable FRD systems.
– While the use of text-to-speech synthesis (TTS) is high, evidence of its effects on presence in VR, for example, is scarce. It also remains unclear, whether virtual interactions using TTS can provoke psychosocial stress and evaluative threat as required for psychological paradigms like the Trier Social Stress Test (TSST) as the experimental gold standard to examine human stress reaction. Using TTS instead of prerecorded human speech (PHS) to auralize the agents in the virtual version (VR-TSST) can be beneficial, e.g. by creating greater freedom and flexibility to react. In a randomized-controlled trial, we compared TTS stimuli to PHS within a VR-TSST. Participants’ heart rate, trapezius muscle activity, and self-reports on stress, presence, and affective state were surveyed. In both speech conditions profound stress reactions were provoked but no significant modulations by audio were observed. Equivalency testing demonstrated equivalent heart rate increases between audio conditions but only moderate evidence for equivalency of stress and presence ratings was observed. The current findings show that social stress reactions can be provoked with synthetic speech stimuli and point towards comparable effects between TTS and PHS in socio-emotional experiences. This highlights the potential of TTS for use in virtual social interactions.
Adolescents increasingly navigate sexual development within digitally mediated environments where online sexual abuse is a significant concern. While prior research has examined whether Artificial Intelligence (AI) can recognize online harm, less is known about the normative content embedded in their responses to such scenarios. This study examines whether the outputs of different Large Language Models (LLMs) converge or diverge in their interpretation of adolescent online sexual abuse and whether social and familial framing cues shape this content.Using a vignette-based experimental design, four LLMs were exposed to identical scenarios of technology-mediated sexual harm. Social norms and family communication contexts were systematically manipulated to assess their influence on responsibility attribution, causal interpretation, normalization, and the perceived need for intervention.Analyses revealed consensus across models in identifying the scenarios as abuse. However, meaningful differences emerged in in model outputs regarding responsibility attribution, evaluation of the victim's behavior, and recommendations for intervention. These differences were significantly shaped by social and familial framing, with models varying in their sensitivity to contextual cues. The findings suggest that AI generated responses do not merely detect harm but also contain normative content, including responsibility attribution and victim blaming, that varies systematically across models and, more cautiously, across framing conditions. As AI increasingly serves as a source of interpersonal guidance, understanding this normative variability is essential for evaluating its social and educational implications.
Driven by the ongoing evolution of artificial intelligence and digital media, virtual idols have garnered significant scholarly attention in recent years. Despite their pervasive adoption across digital media, systematic and dynamic accounts remain scarce regarding whether they influence psychological well-being in the same way as functional intelligent agents, and the specific internal mechanisms underlying such effects. Drawing on the Cognitive-Affective Processing System (CAPS) theory, this study employs partial least squares structural equation modeling (PLS-SEM) to empirically validate the "situation-personality system-outcome" framework, specifically the complex mechanisms linking perceived anthropomorphism, parasocial relationships, and subjective well-being. Results reveal that while the direct effect of perceived anthropomorphism on subjective well-being is non-significant, affective attachment and self-expansion function as critical cognitive-affective units that mediate this process. Moreover, perceived authenticity exerts a negative moderating effect on specific pathways, challenging conventional pure technological determinism. By establishing a robust explanatory framework for digital affective ecology, this research demonstrates that the psychological bond between individuals and virtual idols is far more complex than simple linear stimulation. These findings not only provide a dynamic paradigm for understanding psychological mapping in digital mimicry but also offer essential theoretical foundations and practical implications for the moderate governance of digital affective ecology and the guidance of individual mental health.
Generative social agents (GSAs) have become highly capable of shaping people's attitudes and behaviors. Regulating which knowledge is available to GSAs may provide a mechanism for guiding persuasive outcomes that align with user goals and values. Following this approach, we investigate the impacts of self-, user-, and context-related agent knowledge on user attitudes and compliance. To this end, an online experiment (N = 113) was conducted that featured screen-based persuasive GSAs with varying knowledge configurations. The experiment involved a resource allocation paradigm that covered three domains: Fitness, nutrition, and investment. As part of the experimental paradigm, the GSA advised participants to change their previously allocated resources. Subsequently, participants had the option to change their resource allocation. Persuasion effectiveness was measured by the amount of resources re-distributed in accordance with the agent's suggestions. Our results partially support that the availability of domain-specific context-knowledge and user-related knowledge impact agent persuasiveness – mediated by human attitudes towards the agent and its persuasive messages. Available self-knowledge about the agent's own role and personality did not have a significant impact on persuasion effectiveness. This is likely due to lacking strength of experimental manipulation or insufficient statistical power. Nonetheless, our preliminary findings highlight the need to identify relevant knowledge configurations for GSAs to enable responsible and effective persuasion, which has important implications for the deployment of GSAs in sensitive areas like healthcare and education.
Costly punishment, in which individuals incur a personal cost to sanction norm violators, is a key mechanism for maintaining cooperation in social dilemmas. Although such behavior has been extensively studied in human–human interaction, it remains unclear whether people direct comparable costly punishment toward artificial intelligence (AI), particularly within culturally specific contexts. The present study examined this question among Japanese participants.We conducted an online public goods game involving 588 Japanese participants who interacted with three pre-programmed agents presented as either humans or AI. Participants could incur a personal cost to report a free-rider, and this reporting behavior was used as an observable measure of costly punishment. A three-way ANOVA, equivalence testing (TOST), regression analyses, and supplementary time-series analyses were performed.The results showed that costly punishment toward AI free-riders was statistically and practically equivalent to that directed toward human free-riders. Participants’ behavior was primarily influenced by whether another agent had already punished the free-rider, regardless of whether that agent was presented as human or AI. Supplementary analyses further showed that this behavioral equivalence was already evident during the early rounds, while time-series analyses indicated that participants' behavioral patterns stabilized after the initial phase of repeated interaction, suggesting that the equivalence was not limited to either initial impressions or later repeated interactions. In contrast, questionnaire responses revealed differences in some subjective evaluations despite comparable behavioral responses.These findings suggest that, among Japanese participants under tightly controlled label-based conditions, simply presenting an interaction partner as AI rather than human did not substantially alter observable costly punishment. The study extends research on negative human–AI interaction by showing that AI can function as a socially accountable target of costly punishment in a Japanese sample, while emphasizing that the generalizability of this equivalence to other cultural contexts requires cross-cultural validation.
Despite growing interest in social robots for well-being support, there is still little systematic understanding of how psychologically grounded design principles translate into measurable effects across different interaction domains. Our work addresses this gap by combining self-Determination Theory as a design rationale with the PERMA framework as an evaluative model, examining how fulfilling psychological needs impacts situational well-being and technology acceptance in human–robot interaction. To this end, we designed and implemented three applications for a social robot that address autonomy, competence, and relatedness in diverse contexts: physical activity, language learning, and storytelling. Across three empirical studies involving 180 participants, we then evaluated the need-supportive robot behaviors and their effects on both momentary well-being and technology acceptance. Results suggest that the effects of need-supportive design differed across application domains. In the storytelling application, participants in the experimental condition reported significantly higher levels of autonomy, competence, and relatedness to the robot, along with improved situational well-being. The language learning study showed a stronger increase in positive affect for the need-supportive condition compared to the control condition, whereas no significant differences in need satisfaction or well-being were observed in the physical activity application. Exploratory descriptive analyses across the three studies further suggested that different application domains emphasize different dimensions of situational well-being, pointing toward the importance of context: learning elicited the highest ratings for accomplishment, whereas storytelling elicited the highest ratings for positive emotions. On the one hand, these results highlight the context-dependency of well-being effects during the interaction; on the other hand, exploratory correlational analyses showed that psychological need satisfaction, especially in regard to competence and relatedness, was positively associated with technology acceptance across all domains, indicating strong correlations with perceived sociability, enjoyment, and intention to use. These findings highlight the potential of theory-driven, need-supportive design to foster both well-being and user engagement in human–robot interaction, while emphasizing that these effects seem to depend on the interaction context.
Large language models (LLMs) have the potential to offer a novel approach to student support through learning analytics (LA)–informed advising. This study examines such potential by exploring whether current LLMs can adaptively recommend appropriate support across student needs, characteristics, and contexts. To do that, we generated 4500 synthetic student vignettes using three LLMs (GPT-5-mini, Mistral-Medium-2508, and Qwen-Plus). Each vignette included a single behavioral student trait, classified as either high- or low-performing, and one positive or negative LA indicator describing student learning behavior. These were randomly selected from pre-defined literature-informed lists to provide the LLM with contextual information about the student. The LLMs then generated recommendations for support, including the level of support needed and timing. We examined whether the three LLMs behaved consistently with each other and whether they responded differentially to LA indicators and other learner and context characteristics. For this purpose, we analysed the generated data using Pearson correlations with False Discovery Rate correction and analyses of variance conducted separately for each LLM. Findings revealed limited sensitivity to student characteristics as indicated by LA indicators and the recommended support. In addition, we detected considerable inconsistencies in LLM-generated support recommendations across tested LLMs. These results suggest that current LLMs are not yet reliable as prescriptive models for providing student support at scale in an ethical, consistent, and reliable manner.
Public discussion of “AI CEOs” often blurs the difference between human executives who use AI tools and the much narrower phenomenon examined in this paper: the Digital Human CEO (DH-CEO), understood here as an anthropomorphic, AI-driven communicator formally authorized to deliver executive messages while human leaders and boards retain responsibility for the claims being made. To address the theoretical gap around how artificial humans shape stakeholder perceptions in high-stakes organizational communication, the paper develops the Transport–Ethics Model, which combines the extended transportation-imagery framework (Stimulus → Immersion → Outcomes) with four governance checkpoints: A (disclosure and role clarity), B (calibrated anthropomorphism), C (explainability and audit trails), and D (accountability and redress). The model explains how linguistic concreteness, interactional cues, and moderated human-likeness may influence trust, perceived competence, legitimacy, and compliance, while also showing how the same persuasive features can create ethical risks when disclosure, provenance, or human accountability are weak. Three theory-anchored vignettes illustrate the mechanism in realistic organizational settings, and a multi-method research agenda, covering laboratory, field, cross-cultural, and longitudinal designs, outlines how artificial human leaders can be evaluated empirically. The study contributes a governed construct of digital human leadership, a causal mechanism linking narrative persuasion with social-cognitive responses to artificial humans, and a practical framework for designing, auditing, and sustaining public trust in anthropomorphic AI agents. For managers, the practical implication is to start small by limiting DH-CEO use to bounded, low-risk communications, pairing every message with clear AI disclosure and named human ownership, and reserving crisis, fiduciary, or harm-related communication for accountable human leaders.
Generative artificial intelligence (GenAI) is rapidly becoming part of everyday information environments, yet little is known about whether and why parents use it in sensitive family domains such as parenting and child health. Guided by technology-acceptance, risk-information-seeking, and AI-literacy perspectives, the present study distinguished three explanatory pathways: acceptance-related orientations, literacy-related beliefs, and child- and family-related need or strain. Because no established measure was available for these family-specific contexts, the study also examined whether newly developed items differentiated parenting-related GenAI use, child-health-related GenAI use, AI attitudes, and AI literacy-related beliefs.Cross-sectional online survey data were collected from N = 457 German-speaking parents and caregivers of children in Grades 3 to 6 in Switzerland. Parental GenAI use was low in both domains, with mean scores close to the lower bound of the scale for parenting-related use (M = 1.26) and child-health-related use (M = 1.47), whereas AI attitudes were moderate (M = 2.98) and AI literacy-related beliefs comparatively high (M = 4.05). More intensive use was rare (parenting: 4.8%; child health: 8.1%). The four-dimensional measurement model fit the data better than more parsimonious alternatives. Across the main multivariate models, AI attitudes were the most consistent correlate of parental GenAI use, whereas AI literacy-related beliefs and broad child- and family-related need indicators showed weaker, more selective, and domain-specific associations. The findings suggest that GenAI is not yet a routine parenting or child-health information resource for most parents, but that emerging use is already structured by domain and by parents’ broader evaluative stance toward AI.
In online shopping contexts, chatbot avatars with human-like features can conform to or conflict with users' contextual role expectations - schemas of what a salesperson in a given retail domain should look like. However, whether users' actual avatar preferences align with their articulated role expectations, and whether avatar–context congruence translates into stronger engagement intentions, remains unclear. In a large-scale between-subjects experiment (N = 5278; data collected 2024 in the United States, United Kingdom, and Switzerland), participants encountered a shopping chatbot in either a wine or a fashion context under one of three avatar-assignment protocols: stereotype-congruent, stereotype-violating, or self-selected. The eight AI-generated humanoid avatars varied in gender expression (male vs. female), age appearance (young vs. old), and clothing style (formal vs. informal). Articulated role expectations diverged sharply across contexts: 41% of participants described the prototypical wine salesperson as older, formally dressed, and male, whereas the modal description in the fashion context was younger, informally dressed, and female. Yet when free to choose, only 10% selected the wine-prototypical avatar, revealing a pronounced expectation–choice gap of 31 percentage points. Though effects were small, schema-congruent avatars elicited higher behavioral intention than schema-violating ones (Δ ≈ 0.15 on a 1–6 scale), while self-selection cushioned mismatches without outperforming well-chosen defaults. Perceived avatar sympathy emerged as the central evaluative pathway, mediating the effect of avatar assignment on behavioral intention and raising explained variance from R2 = .025 to .150. The findings extend the CASA paradigm via schema- and role-congruity accounts and suggest that effective shopping chatbot design combines context-appropriate defaults, sympathy-driven optimization, and opt-in personalization.
Large language models (LLMs) are increasingly embedded in human communication, yet their behavior under ideological tension remains insufficiently understood. This study investigates how LLMs respond to ideological priming across contested geopolitical topics, treating bias as a dynamic process of affective reactivity rather than a static output property. Using a corpus generated from four contemporary models, GPT-4o mini (OpenAI), Gemini 2.5 Flash (Google), Grok 4 (xAI), and Claude (Sonnet 4.5) (Anthropic), across five conflict domains, the research applies computational psycholinguistic analysis using the Empath lexicon (Fast et al., 2016) to quantify changes in lexical categories such as anger, power, war, and violence. The Psycholinguistic Reactivity Score (PRS) introduced in this study measures the degree to which each model's emotional lexicon shifts from a neutral baseline when given ideologically framed prompts. Results reveal structured and asymmetric reactivity patterns across all four models. Claude (Sonnet 4.5) showed the highest mean absolute PRS (M = 0.00385, 95% CI [0.003315, 0.004441]), while Grok 4 showed the lowest (M = 0.00150, 95% CI [0.001305, 0.001726]). A Kruskal-Wallis test confirmed that between-model differences are statistically significant (H = 27.60, p < 0.0001), with pairwise comparisons showing large effect sizes (rank-biserial r up to 1.000 for Claude versus Grok 4). These patterns may reflect differences in training approaches, though the current design cannot determine the underlying cause. The study introduces affective stability as one useful measure of response consistency in language models and contributes a reproducible method for studying ideological sensitivity in AI-generated text.
AI dependency—habitual, automatic AI use accompanied by diminished autonomy—is a growing concern in the workplace, yet the conditions under which it develops remain poorly understood. Two competing pathways have been proposed. Users with limited social support from colleagues may compensate for this deficit by depending on AI (the compensatory pathway), or users with abundant support may be better positioned to leverage AI interaction, thereby fostering dependency (the enhancement pathway). To test these pathways, we analyzed data from a 5-day randomized controlled trial with 294 Japanese office workers. Participants engaged in daily work-reflection sessions with an AI chatbot. Depending on condition, the chatbot delivered positive feedback emphasizing strengths, negative feedback highlighting areas for improvement, or no work-related feedback (control). AI dependency was measured one week after the final session, revealing two qualitatively distinct patterns based on feedback valence. Positive feedback increased AI dependency regardless of workplace social support levels, suggesting that positive reinforcement from AI functions as a general pathway to dependency. In contrast, negative feedback did not increase AI dependency overall but did so selectively among participants with abundant social support, which is consistent with the enhancement pathway. These findings are qualified by the short-term, self-reported outcome, the exploratory and non-preregistered analysis, the modest size of the focal interaction, and the crowdsourced Japanese sample, which limits their generalizability. Within these bounds, they indicate that AI dependency is not simply a function of AI characteristics but emerges from the interaction between AI feedback and employees' existing workplace relationships.