Generative artificial intelligences, particularly Large Language Models (LLMs), increasingly influence human decision-making, making it essential to understand how cognitive biases are reproduced or amplified in these systems. Building on evidence of the human “addition bias” – a preference for additive over subtractive problem-solving strategies1 – this research compared humans with GPT-4 (Study 1) and GPT-4o (Study 2) in spatial and linguistic tasks. Study 1 comprised four experiments (1a, 1b, 2a, 2b) with 588 human participants and 680 GPT-4 outputs; Study 2 included two experiments (3a, 3b) with 751 human participants and 1,080 GPT-4o outputs. We manipulated (a) solution efficiency and (b) instruction valence. Across both studies, a general addition bias emerged, more pronounced in the LLMs than in humans. Humans made fewer additive choices when subtraction was more efficient than addition (compared to when both were equally efficient), whereas GPT-4’s output showed the opposite pattern. GPT-4o’s outputs aligned with those of humans in the linguistic task but showed no efficiency effect in the spatial task. Instruction valence did not reach statistical significance for either agent in the spatial task. In the linguistic task, positive valence (compared to neutral valence) led to more additive outputs in both GPT models, but only in Study 2 for humans. These findings indicate that addition bias has been transferred to LLMs, which can replicate and, depending on context, amplify this human bias. This emphasizes the importance of further theoretical and empirical work on the cognitive and data-driven mechanisms underlying addition bias in both humans and LLMs. Six experiments examined addition bias in humans and in outputs of large language models (GPT-4, GPT-4o). Both human and model-generated solutions showed a preference for additive over subtractive strategies, with a stronger bias in the model output.
Educational videos have become central to learning environments, particularly following the COVID-19 pandemic, underscoring the need for optimized design to enhance learning outcomes. Concurrently, research on film cognition has provided valuable insights into how viewers process dynamic visual narratives in Hollywood cinema. This paper explores the parallels between Hollywood cinema and educational videos to determine how film cognition research can inform educational video design. We analyze how cinematic techniques align with educational principles (i.e. signaling, spatial and temporal contiguity) proposed by Cognitive Theory of Multimedia Learning. Additionally, we examine social cognition aspects, drawing connections to the learning principles (i.e., embodiment and image principles) in educational contexts. We propose that educational videos can enhance learner engagement by adopting attention-guiding methods and social cues used in film. We also suggest that compiling a shared collection of example videos and reusable materials could help teachers and researchers reuse and adapt resources more easily.
Visual narrative comprehension is essential for navigating modern society, where information, rules, and news are frequently communicated through images, diagrams, and visual stories. Encoding a coherent narrative from disparate elements is critical for all age groups. Although recent studies report a significant rise in stress and anxiety levels, driven by factors such as the COVID-19 pandemic and recent geopolitical conflicts, the impact of stress on visual narrative comprehension remains largely underexplored. This study explored how acute stress affects narrative comprehension in younger (N = 203, 18-57 years; M = 23 years; Experiment 1) and older adults (N = 212, 60-85 years; M = 67 years; Experiment 2). Participants were assessed under both acute stress and neutral conditions. A tool for inducing acute stress online employed mathematical and logical tasks under time pressure, along with elements that simulate social stress. Participants were presented with pictorial stories consisting of three panels, with the second panel intentionally left blank. Their task was to comprehend the storyline despite the missing information. On the following page, an image representing a possible bridging event was shown, either depicting the correct or an incorrect inference. Participants were asked to judge whether or not the presented image accurately reflected the missing event in the story. Results revealed that acute stress negatively impacted narrative comprehension in younger adults, while the older adults' comprehension remained unaffected by acute stress. Similarly, younger adults demonstrated reduced confidence in their responses under stress, whereas older adults' confidence levels remained unaffected. These findings highlight the relationship between visual narrative comprehension, stress, and aging, suggesting that, with age and experience, comprehenders may develop more differentiated event schemas, which makes their comprehension processes more resilient to stress. Understanding how cognitive and perceptional processes function under stress is crucial for daily life across all age groups. Our research demonstrates that younger adults exhibit poorer visual narrative comprehension under acute stress, whereas older adults' performance remains stable. This finding suggests that older adults may employ more differentiated event schemas, which help maintain their narrative comprehension in the face of stress. Consequently, narrative comprehension appears to be more resilient compared to other fundamental cognitive skills. These insights could inform interventions and strategies to support cognitive health across different age groups.
The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether AI-generated catalogue descriptions can approximate human-written quality and how generative AI might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content. After a human-in-the-loop curation process, they were evaluated by archive and archaeology experts (consisting of students and professionals) and non-experts in a human-centred, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect AI-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify. This suggests that even further human review is necessary to ensure the accuracy and quality of the catalogue descriptions generated by the out-of-the-box model and selected by humans, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends on technical advancements, such as domain-specific fine-tuning, and on establishing trust among professionals. Both could be fostered through a transparent and explainable AI pipeline.
This preregistered study investigates the availability of research data in articles from four selected educational research journals over a five-year period. Among the journals examined are two national (German) and two international journals from the years 2018, 2020, and 2023. The study focuses on changes in data-sharing practices as responses to evolving research data policies and infrastructural developments. Findings indicate that the availability of research data has increased substantially over time, with more recent articles providing data more frequently compared to earlier publications. A key observation is the positive correlation between journal-level data transparency standards and actual data availability. This suggests that explicit editorial guidelines and the introduction of Open Science badges effectively incentivize transparency. In contrast, institutional research data policies were found to have no significant impact on data-sharing practices. Finally, the study found no statistically significant evidence that the use of persistent identifiers (e.g., DOI) improves long-term data accessibility, although the effect estimates suggest a possible positive association that warrants further investigation. Overall, the results highlight the importance of clear editorial policies, the introduction of Open Science badges as part of an ongoing effort to strengthen incentive structures, and continued meta-scientific research aimed at monitoring and optimizing research infrastructures. These combined efforts are essential for advancing the practice of open data access in educational research, ensuring that data becomes a lasting resource that supports transparency, replicability, and scientific progress.
While historically, the aim of propaganda was to convince the public of a certain agenda, many political commentators contend that modern forms of disinformation come with a different goal in mind: To confuse, rather than convince. It is believed that to secure this goal, informational spaces are diluted, rather than dominated, a strategy termed "zone-flooding." The present research provides a unified psychological account of confusion that draws on methods from Signal Detection Theory and metacognition, to examine the public's object-level ability to discern truth from falsehood and accurate metacognitive insight into that distinction. In two preregistered studies, including one Registered Report, we presented quota-matched samples (Study 1, Germany: n = 1488; Study 2, USA: n = 1891) with true-only, false-only ("classical propaganda") or noisy ("zone-flooding") information about climate change in the form of tweets. Across both studies, across various operationalizations of zone-flooding (e.g., targeting the political left vs. right; different numbers of tweets), and against various baselines (politically equated vs. ecologically valid tweets), we find no empirical evidence for the notion that zone-flooding causes confusion in this initial test. The present work lays the theoretical basis for a psychology of zone-flooding and public confusion within one coherent theoretical and methodological framework.
Construal-level theory (CLT) proposes that psychological distance influences the level of abstraction at which something is mentally construed: Things perceived as less probable (likelihood) or further away from the here (spatial distance), now (temporal distance), or self (social distance) are thought about more abstractly. In this international multilab study, we tested four basic hypotheses derived from core assumptions of CLT and explore potential moderators and boundary conditions of the effects. Participants ( N = 11,775) from 27 countries and regions were randomly assigned to one of four experimental protocols focused on different types of psychological distance (temporal, spatial, social, or likelihood), and each experiment manipulated psychological distance (close vs. distant). The protocols for temporal distance ( n = 2,941) and spatial distance ( n = 2,973) were direct replications of Liberman and Trope (Study 1) and Fujita et al. (Study 1), respectively. The remaining two protocols were paradigmatic replications, applying to social distance ( n = 2,926) and likelihood ( n = 2,936). The effects of psychological distance on construal level for the four present studies were as follows (positive effects are consistent with hypotheses): temporal, d = 0.08, 95% confidence interval [CI] = [0.003, 0.16] (effect in original study: d = 0.92); spatial, d = 0.04, 95% CI = [−0.03, 0.11] (effect in original study: d = 0.55); social, d = −0.27, 95% CI = [−0.34, −0.19]; and likelihood, d = 0.03, 95% CI = [−0.05, 0.11]. Pretests indicated that valence and abstraction were confounded in response options on the outcome measure. Controlling for this confound eliminated the hypothesis-inconsistent effect of social distance, d = 0.006, 95% CI = [−0.05, 0.07]. These findings provide limited evidence for the predictions of the theory and present a critical challenge for CLT.
LLM agents can now pass as human participants, threatening the validity of online social science. We urge a shift from ad-hoc checks to multi-layered, adaptive defenses, borrowing from internet anti-bot practice, and call for cooperation across researchers, platforms, and institutions, to guard against this challenge. LLM agents can now pass as human participants, threatening the validity of online social science. We urge a shift from ad-hoc checks to multi-layered, adaptive defenses, borrowing from internet anti-bot practice, and call for cooperation across researchers, platforms, and institutions, to guard against this challenge.
Narrative comprehension involves deriving meaning from information and constructing a mental representation of a narrative's components. It conveys knowledge, information, and explanations, playing a crucial role in everyday life. Inferencing – is one of the key processes in narrative comprehension. While cognitive changes associated with aging are well-documented, research on narrative comprehension and inferencing in older adults remains inconclusive, highlighting the need for further study. Despite being ever-present in modern life, the role of stress in this process also remains largely unexplored. This online study involved 298 participants, who were randomly assigned to either a stress group, which underwent stress induction, or a control group. Participants viewed picture-based narratives and were tasked with filling in missing elements by describing the middle part of the story in their own words. This approach differs from previous studies that primarily assessed inferences using comprehension questions or true-false statements, enabling us to capture the inference generation process as it unfolds. The findings show that narrative comprehension, including inference generation, remains stable across age and education groups and is unaffected by acute stress. This highlights the resilience of these abilities, even under challenging conditions. The implications for cognitive theories and directions for future research are discussed.
Communicating the scientific consensus on climate change acts as a gateway to increasing climate beliefs, a process described by the Gateway Belief Model (GBM). Although the effectiveness of this approach is well-documented, critical gaps in the literature persist. Few studies have scrutinized the risks of low-consensus messages and their correctability, the potential for consensus messaging to have an impact beyond self-reported measures, or the cognitive mechanisms driving belief updates. This study addresses these three critical areas in a German sample (N = 941). First, we demonstrate the powerful adverse effect of a low-consensus message on the perceived scientific consensus (PSC) and show that this effect can be corrected through subsequent accurate messaging. Second, we introduce learning as a novel cognitive outcome, finding that high-consensus messaging can act as a gateway to knowledge acquisition. Third, we explore initial confidence as a key psychological moderator, revealing that it influences belief updating under certain conditions.
Human behavior is frequently constrained by the behavior of other agents (increasingly artificial agents like machines and software). In a cooperative setting, each individual needs to understand the partners’ intentions and corresponding actions to plan their actions adequately. Misunderstandings have adverse effects and diminish the efficiency of cooperation. We created a game-based experiment to study the effects of understanding a partner’s actions in a cooperative setting. In this game, the players depend on each other’s ability to make good decisions to succeed. In two studies (N = 87, N = 281), we collected data on the understanding of an artificial agent, operationalized as the ability to predict its actions and the skill at the task itself. The participants improve at predicting the agent’s actions and at their subtasks in the game over time. Following a misunderstanding, the participants’ performance at their subtask was worse, as measured by their actions’ quality, speed, and efficiency. Results of Study 2 suggest that the improvements in predicting the agent’s actions are likely the result of an improved understanding of the game rather than an improved understanding of the agent.
Robots are increasingly present in our society. Their successful integration depends, however, on understanding and fostering pro-social behavior towards robots, in this case, helping. To better understand people’s reported willingness to help robots across different contexts (delivery, medical, service, and security), we conducted two preregistered studies on a German-speaking population (N = 415, and N = 542, representative of age and gender). We assessed attitudes, knowledge about robots, and anthropomorphism and investigated their effect on reported willingness to help. Results show that positive attitudes significantly predicted reported higher willingness to help. Contrary to our hypothesis, having more knowledge about robots increased reported willingness to help. Additionally, we found no effect of anthropomorphism, neither in the form of robot appearance nor as participant’s own view about robots, on reported willingness to help. Furthermore, results point to a context-dependency for willingness to help with participants preferring to help robots in a medical context compared to a security one, for example. Our findings thus highlight the relevance of knowledge and attitudes in understanding helping behavior towards robots. Additionally, our results raise questions about the relevance of anthropomorphism in pro-sociality towards robots.
Dynamic events can be perceived at different levels of granularity, with coarser contexts (e.g., “morning routine”) requiring more conceptual integration than finer contexts (e.g., “having breakfast”). We investigated whether this increased integration yields more abstract mental representations (i.e., fewer perceptual details) across three experiments. In these experiments, participants encoded events (presented via text or video) in coarse- or fine-grained contexts and then matched these events to corresponding stimuli. Our results indicated that coarse context consistently took longer to process, which is consistent with higher integration demands. Participants were also faster when encoding and test modalities matched (suggesting modality-specific processing) and when the test modality was video (reflecting slower reading times in text). Crucially, we found no interaction between context grain and encoding or test modalities. Contrary to the expectation that coarser contexts would produce more amodal mental representations, participants showed no coarse-grain cross-modal advantage in recognizing events. We propose that the contents of everyday events might be abstracted similarly within mental representations for both fine and coarse contexts. Furthermore, our results provide evidence for amodal representations when the environment is unpredictable. That is, participants transformed their mental representations into both verbally and visually compatible formats, regardless of the original encoding modality.
Online videos have become a central tool in modern education. Alongside this shift, Artificial Intelligence (AI) is reshaping personalized learning experiences, with generative large language models like ChatGPT offering new ways to tailor information to individual learners. Based on the Cognitive Theory of Multimedia Learning (CTML), which proposes two principles that relate to the interaction with the learning material (segmenting and generative activity), we conducted two experiments in which participants were asked to pause an educational video at times of comprehension difficulty. In Experiment 1 (N = 101), we examined whether GPT-generated summaries -introduced at self-paced pause points-result in better learning compared to video transcripts. In Experiment 2 (N = 215), we compared the role of GPT-generated summaries and GPT-generated reflective prompts. Those elicited open-ended answers from the participants. We measured retention and transfer learning, as well as mental effort, and perceived task difficulty. Contrary to our expectations, we observed no differences between AI summaries and transcripts in terms of retention and transfer outcomes. Participants showed a learning effect indicating more correct answers after watching the video, but this effect did not differ between conditions. We can especially note that the motivation to engage in the material, as well as the difficulty and length of the video, may have affected the results. As research investigating the role of AI in educational settings is still new, future research can delve into finding the optimal conditions under which AI can benefit learning outcomes.
Background: Understanding narratives is essential for societal participation. However, insufficient literacy or age-related cognitive changes can limit narrative comprehension and create participation barriers.Aims: This study investigates the potential of pictorial narratives to convey information beyond text and break down barriers to comprehension.Sample: A representative adult sample (N = 1487).Methods: The experimental study assessed the influences of age and education on the comprehension of textual and pictorial narratives. Participants were tested on the generation of bridging inferences, a central aspect of narrative comprehension. The narratives used were based on the "Multilingual Assessment Instrument for Narratives" (MAIN), and comprehension was measured through a task where participants identified correct and false inference statements related to missing parts of the narratives.Results: Narrative comprehension was higher within higher educated groups and stable across the measured adult age span. Comprehension was generally better for pictorial narratives than textual ones across all education and age groups. Frequentist and Bayesian analyses supported these findings, showing significant effects of education and narrative modality on comprehension but no interaction between these factors.Conclusions: The findings lay the groundwork for more effectively addressing underprivileged groups and refining theories on narrative comprehension. The results suggest that pictorial narratives could be a valuable approach to enhance comprehension and participation for individuals with literacy challenges or cognitive changes due to aging. This study also emphasizes the role of education in narrative comprehension and suggests stable comprehension abilities across the adult age span.
Humans extend social perceptions not only to other people but also to non-human entities such as pets, toys, and robots. This study investigates how individuals differentiate between various relationship partners—a familiar person, a professional person, a pet, a cuddly toy, and a social robot—across three interaction contexts: caregiving, conversation, and leisure. Using the Repertory Grid Method, 103 participants generated 811 construct pairs, which were categorized into seven psychological dimensions: Verbal Communication, Assistance and Competences, Liveness and Humanity, Emotional and Empathic Ability, Autonomy and Voluntariness, Trust and Closeness, and Physical Activity and Responsiveness. Cluster analyses revealed that in Verbal Communication and Assistance and Competences, robots were perceived similarly to human partners, indicating functional comparability. In contrast, for Liveness and Humanity and Emotional and Empathic Ability, humans clustered with pets—distinct from robots and cuddly toys—highlighting robots’ lack of perceived emotional richness and animacy. Interestingly, in Autonomy and Voluntariness and Trust and Closeness, robots were grouped with professional humans, while familiar persons, pets, and cuddly toys formed a separate cluster, suggesting that robots are seen as formal, emotionally distant partners. These findings indicate that while robots may match human partners in communicative and task-oriented domains, they are not regarded as emotionally intimate or fully animate beings. Instead, they occupy a hybrid role—competent yet impersonal—situated between tools and social agents. The study contributes to a more nuanced understanding of human-robot relationships by identifying the psychological dimensions that shape perceptions of sociality, animacy, and relational closeness with non-human partners.
Large language models (LLMs) increasingly mimic human cognition in various language-based tasks. However, their capacity for metacognition-particularly in predicting memory performance-remains unexplored. Here, we introduce a cross-agent prediction model to assess whether ChatGPT-based LLMs align with human judgments of learning (JOL), a metacognitive measure where individuals predict their own future memory performance. We tested humans and LLMs on pairs of sentences, one of which was a garden-path sentence-a sentence that initially misleads the reader toward an incorrect interpretation before requiring reanalysis. By manipulating contextual fit (fitting vs. unfitting sentences), we probed how intrinsic cues (i.e., relatedness) affect both LLM and human JOL. Our results revealed that while human JOL reliably predicted actual memory performance, none of the tested LLMs (GPT-3.5-turbo, GPT-4-turbo, and GPT-4o) demonstrated comparable predictive accuracy. This discrepancy emerged regardless of whether sentences appeared in fitting or unfitting contexts. These findings indicate that, despite LLMs' demonstrated capacity to model human cognition at the object-level, they struggle at the meta-level, failing to capture the variability in individual memory predictions. By identifying this shortcoming, our study underscores the need for further refinements in LLMs' self-monitoring abilities, which could enhance their utility in educational settings, personalized learning, and human-AI interactions. Strengthening LLMs' metacognitive performance may reduce the reliance on human oversight, paving the way for more autonomous and seamless integration of AI into tasks requiring deeper cognitive awareness.