The Programme for International Student Assessment (PISA) 2022 assessed creative thinking (CT) skills for the first time. While this assessment offers a unique snapshot of the creative abilities of 15-year-olds worldwide, it also raises questions about the influence of the low-stakes school context in which it was administered, where academic domains are more routinely assessed than creativity. This study examined differences in careless responding across PISA domains and within CT subdomains and facets. Within CT, it also tested whether cognitive and behavioral engagement and self-reported disengagement due to PISA’s low-stakes nature predicted item-level performance. It further investigated whether school-level testing practices moderated the link between disengagement and performance. A secondary analysis of the PISA 2022 data was conducted on 142,564 students who completed at least one CT item. Results showed that careless responding rates were lower for CT items than for core academic domains. Within CT, careless responding also varied across subdomains and facets, with the lowest rates observed for the Social Problem Solving (SoPS) domain and the Generate Creative Ideas (GCI) facet. Performance on CT items was positively associated with both cognitive and behavioral engagement, and negatively associated with disengagement. Finally, among students in schools that routinely use tests for promotion or retention, the negative impact of disengagement on CT performance was weaker. These findings suggest that students took CT tasks at least as seriously as traditional academic ones. We discuss how, despite their potentially more engaging nature, CT assessments are not immune to the same engagement-related issues that affect core domains in low-stakes contexts.
The serial order effect (SOE), i.e., the tendency for the creative quality of ideas to increase with increasing response position, has been widely documented in divergent thinking (DT) research. However, its magnitude is hardly reported, and its sensitivity to methodological effects (e.g., DT task used, scoring method) has rarely been investigated. More critically, the stability of inter-individual differences in SOE has not been investigated to date. Is the SOE a sizable, dependable (i.e., immune to methodological variations), and importantly, a trait-like psychological phenomenon? This study examined SOE strength and inter-individual consistency across three Alternate Uses Task (AUT) prompts and three scoring approaches: subjective ratings, frequency-based scoring, and semantic similarity. While the SOE emerged across all scoring methods and functional forms examined (linear and quadratic), effect sizes were consistently small in magnitude (explaining at most 3.5% of variance) and varied significantly by stimulus and scoring method. Crucially, evidence for stable individual differences in SOE slopes was absent (ICCs < 0.31 across AUT prompts), raising a potential “reliability paradox”: the SOE appears robust at the group level (i.e., reproducible, though small), yet unreliable in terms of inter-individual stability. These findings underscore the need for systematic evaluation of SOE effect sizes, greater attention to task and scoring differences, and closer investigations clarifying whether the SOE reflects a stable trait interpretable at the individual level, or an aggregate phenomenon with limited practical significance.
The Programme for International Student Assessment (PISA) is one of the most costly, influential, and wide-reaching research efforts in educational psychology. But are the measures they administer as valid as they should be? As part of its 2022 cycle, PISA included a focus on creative thinking, incorporating scales in the student and context questionnaires designed to measure various creativity-related constructs. However, information on the development of these items is limited, and the extent to which these items validly capture their intended constructs has not been systematically evaluated. If these measures lack validity, the field risks interpreting current findings and secondary analyses of PISA data in ways that may not be fully psychologically meaningful, possibly leading to incorrect recommendations for educational practice. This issue is addressed by employing a multistage expert judgment procedure (involving a total of N = 11 experts) to assess the content validity of 127 creativity-related items identified from an initial pool of 936 items from PISA 2022 student and context questionnaires. After an extensive pilot phase (Study 1), experts categorized items according to predefined construct and domain definitions, evaluated item quality, and rated construct coverage (Study 2). Findings revealed substantial misalignment between the constructs assigned by the Organisation for Economic Co-operation and Development (OECD) and those identified by the expert panel (an average overlap of 41%). The results ultimately offer a detailed item-content map as a ground for revised measurement of creativity-relevant constructs, which are intended to support future analyses grounded in psychologically and educationally valid measures.
Students often encounter difficulties when learning and processing fractions. In fraction comparison tasks, they tend to rely on a mixture of successful strategies, e.g., fraction magnitude processing or benchmarking, and erroneous strategies, e.g., natural number-based reasoning or gap thinking. Reinhold et al. (2023) used a theory-driven approach to classify students into distinct profiles depending on their strategy choice based on their performance on a comparison task with 24 single-digit fractions. The authors identified single strategy profiles (e.g., typical natural number bias) and composite profiles (e.g., benchmarking or typical bias). The current study aimed to replicate this study (RO1) and extend its design by incorporating multi-digit fractions (i.e., the denominator has at least two digits; RO2) and additional biased comparison strategies, specifically gap thinking (RO3). A set of 101 fraction comparison tasks was administered to 285 fifth- and sixth-grade students in Flanders, Belgium, controlling for benchmarking to 1/2, numerical distance, natural number-based reasoning and gap thinking. A Bayesian classification approach based on students’ performance (accuracy, reaction time, and individual distance effect) replicated the distinct profiles identified by Reinhold et al. (2023), not only for single- (RO1) but also for multi-digit fractions (RO2). Furthermore, we identified additional single strategy and composite profiles (e.g., applying benchmarking where possible, otherwise gap thinking), both based on accuracy and reaction time data (RO3). This study enhances our understanding of individual differences in fraction processing and emphasizes the importance of controlling for gap thinking.
Within the space of dialogue systems for language learning, the rapid advance of conversational artificial intelligence, powered among others by large language models, is currently driving innovations that enable learners of a second or foreign language (L2) to practice dialogic interaction at their own pace, including through functional tasks. Such language practice is particularly relevant when learners have limited opportunities to interact with more proficient speakers of that L2. However, to ensure a meaningful contribution of dialogue systems to L2 development, technology-mediated practice of dialogic interaction needs to be adapted to the needs and proficiency of learners. This requires accurate and transparent assessment of L2 performance that is both driven by theories about L2 acquisition and practically feasible with state-of-the-art technologies. The end goal is that task-based dialogue practice can be scaffolded through individualized feedback and other forms of learning support. This study models task performance in Language Hero, a game-based spoken dialogue system designed for Dutch-speaking learners who want to practice French as a L2. Using explanatory item response analysis, we explored to what extent the completion of functional spoken tasks can be predicted from learning process data, more specifically from fully-automated measures of L2 task performance. Data were drawn from 263 participants who completed a total of 739 tasks in the system, comprising 22,074 spoken responses. The results indicate that 12 fully-automated measures of previous task performance, including complexity, accuracy, fluency, and functional adequacy, as well as two measures of hint use, significantly predicted future task completion. A multilevel model with fixed and random effects accounted for 45% of the variance in task completion. This study demonstrates the potential of data-driven learner models for micro-adaptivity in dialogic technology-mediated practice while simultaneously highlighting the need to include complementary predictors as well as human evaluation.
Interest in understanding creativity through Programme for International Student Assessment (PISA) data is on the rise, yet researchers face methodological challenges in synthesizing findings across various constructs, measures, and datasets. Meta-analysis-a valuable methodology for synthesizing quantitative data-remains underutilized in creativity research involving large-scale assessments like PISA. This paper provides guidelines for applying meta-analytic techniques to PISA creative thinking assessment data to help researchers address these challenges. It introduces meta-analysis by outlining its definition and advantages, followed by key steps and methodological considerations for synthesizing bivariate and multivariate relationships within PISA. Finally, the paper discusses techniques for managing the computational complexity of meta-analyzing PISA data. Ultimately, these guidelines aim to support researchers in effectively synthesizing PISA data to advance the study of creativity.
Researchers and educators interested in creative writing need a reliable and efficient tool to score the creativity of narratives, such as short stories. Typically, human raters manually assess narrative creativity, but such subjective scoring is limited by labor costs and rater disagreement. Large language models (LLMs) have shown remarkable success on creativity tasks, yet they have not been applied to scoring narratives, including multilingual stories. In the present study, we aimed to test whether narrative originality-a component of creativity-could be automatically scored by LLMs, further evaluating whether a single LLM could predict human originality ratings across multiple languages. We trained three different LLMs to predict the originality of short stories written in 11 languages. Our first monolingual model, trained only on English stories, robustly predicted human originality ratings (r = .81). This same model-trained and tested on multilingual stories translated into English-strongly predicted originality ratings of multilingual narratives (r >= .73). Finally, a multilingual model trained on the same stories, in their original language, reliably predicted human originality scores across all languages (r >= .72). We thus demonstrate that LLMs can successfully score narrative creativity in 11 different languages, surpassing the performance of the best previous automated scoring techniques (e.g., semantic distance). This work represents the first effective, accessible, and reliable solution for the automated scoring of creativity in multilingual narratives.
Background The potential of adaptive feedback in digital educational games remains largely unexplored. Fractions are a suitable topic for investigating the effectiveness of adaptive feedback, as the complexity of this domain highlights the need for adequate feedback. Objectives This study examines the effectiveness of explanatory adaptive feedback in a digital educational game to address two particular misconceptions regarding fractions (i.e., Natural Number Bias and Unit of Reference). Methods A total of 197 4th graders were randomly assigned to two different conditions, each playing a different version of a digital educational game: one with corrective feedback and one with explanatory adaptive feedback. During gameplay, we collected log data of students' item-wise correctness and misconception errors. Results Explanatory item response analyses indicated that correctness improved in both game versions, with a more pronounced increase for the game with explanatory adaptive feedback compared to the game with corrective feedback. However, no decrease in misconception errors was observed in either game version. Moreover, neither the type of misconception nor students' prior fraction knowledge were moderating factors. These results suggest that adaptive feedback can support students in learning fractions; however, to reduce misconception errors concrete feedback should be optimised.
OBJECTIVE:This meta-analysis explores the relationship between Big Five personality traits and flow. It also examines the moderating roles of demographic factors (i.e., gender and age), cultural differences, contextual variations, flow dimensions, and the instruments used to assess personality and flow. METHOD:A systematic search was conducted across ProQuest, Scopus, and Web of Science, identifying 24 eligible studies reporting associations between Big Five traits and flow. A total of 352 effect sizes were analyzed using a three-level random-effects model. Moderator analyses examined the influence of demographic, cultural, contextual, and methodological factors. RESULTS:Results reveal a medium-sized positive association between Conscientiousness and flow (r = 0.33), while Extraversion (r = 0.25), Openness (r = 0.18), and Agreeableness (r = 0.16) show smaller positive relationships. Neuroticism has a small negative relationship with flow (r = -0.16). Significant moderating effects were identified for culture, with stronger correlations in Eastern cultures for Extraversion, Openness, and Agreeableness. CONCLUSIONS:These findings emphasize the importance of considering personality traits when studying flow. Future research should expand cross-cultural studies, explore flow across a broader range of contexts, incorporate multimodal measurement techniques, and develop interventions that enhance flow experiences by aligning them with individuals' personality profiles and contextual characteristics.
BackgroundAugmented reality (AR) is receiving increasing interest as a tool to create an interactive and motivating learning environment. Yet, it is unclear how instructional support affects performance in AR.ObjectivesThis study sought to explore how varying the instructional support in AR can affect performance-related behaviours of students with low cognitive abilities during assembly work.MethodsA total of 90 Belgian secondary school students repeatedly executed four different realistic assembly tasks. Three levels of instructional support (low, medium, and high) in AR as well as a control condition with paper instructions with a high level of detail were systematically varied across tasks and participants.Results and ConclusionsMultilevel regression analyses showed that AR instructions yielded lower assembly times and a lower perceived physical effort than paper instructions. Additionally, participants perceived tasks as less complex when given AR instructions with a high or a medium level of detail than when given a low level of detail. No effects of instructional support were established for other performance-related behaviours, namely necessary assistance, error-making, cognitive load, competence frustration, and stress. Effect sizes were small, at least among the instructional support conditions studied, yielding a limited base for adaptivity. Presumably, tailoring the instructional support in AR is only beneficial for highly complex tasks. The results might be useful for the design and implementation of AR in educational settings. What is currently known about the subject matterAugmented reality is knowing increased use in education.Little is known about how to design effective instructional support in augmented reality.This study represents an initial insight into how to personalize instructional support.The impact of instructional support on augmented reality performance of students with low cognitive abilities was investigated.The study suggests that varying instructional support may lead to differences in performance outcomes.Implications for practitionersHow instructional support is constructed may affect augmented reality learning.The results may inform the design of effective augmented reality learning environments for students with special needs.How individual, contextual, and task-specific characteristics moderate the effectiveness of instructional support should be further investigated.
Single-case experimental designs (SCEDs) may offer a reliable and internally valid way to evaluate technology-enhanced learning (TEL). A systematic review was conducted to provide an overview of what, why and how SCEDs are used to evaluate TEL. Accordingly, 136 studies from nine databases fulfilling the inclusion criteria were included. The results showed that most of the studies were conducted in the field of special education focusing on evaluating the effectiveness of computer-assisted instructions, video prompts and mobile devices to improve language and communication, socio-emotional, skills and mental health. The research objective of most studies was to evaluate the effects of the intervention; often no specific justification for using SCED was provided. Additionally, multiple baseline and phase designs were the most common SCED types, with most measurements in the intervention phase. Frequent data collection methods were observation, tests, questionnaires and task analysis, whereas, visual and descriptive analysis were common methods for data analysis. Nearly half of the studies did not acknowledge any limitations, while a few mentioned generalization and small sample size as limitations. The review provides valuable insights into utilizing SCEDs to advance TEL evaluation methodology and concludes with a reflection on further opportunities that SCEDs can offer for evaluating TEL.
The objective of this study is to explore the relationship between personality and peer-rated team role behavior on the one hand and team role behavior and verbal behavior on the other hand. To achieve this, different data types were collected in fifteen professional teams of four members (N = 60) from various private and public organizations in Flanders, Belgium. Participants’ personalities were assessed using a workplace-contextualized personality questionnaire based on the Big Five, including domains and facets. Typical team role behavior was assessed by the team members using the Team Role Experience and Orientation peer rating system. Verbal interactions of nine of the teams (n = 36) were recorded in an educational lab setting, where participants performed several collaborative problem-solving tasks as part of a training. To process these audio data, a coding scheme for collaborative problem solving and linguistic inquiry and word count were used. We identified robust links and logical correlation patterns between personality traits and typical team role behaviors, complementing prior research that only focused on self-reported team behavior. For instance, a relatively strong correlation was found between Altruism and the Team builder role. Next, the study reveals that role taking within teams is associated with specific verbal interaction patterns. For example, members identified as Organizers were more engaged in responding to others’ ideas and monitoring execution.
The Kaufman Domains of Creativity Scale (K-DOCS), a self-report measure designed to capture creative behaviors across various domains, has been utilized and validated across different cultural contexts. The present study sought to assess the psychometric properties of an Arabic version of the K-DOCS. Using exploratory graph analysis followed by confirmatory factor analysis and item response theory analysis, the factor structure of the K-DOCS was assessed. Additionally, the criterion validity of the K-DOCS was assessed in relation to measures of openness and emotional intelligence. Beyond validation, the study examined the network structure of the K-DOCS domains to understand their interconnections and investigated potential domain network differences based on gender, age, and academic major. Data were collected among 2,594 Egyptian university students. The results suggest that the K-DOCS has a five-factor structure broadly consistent with the theoretical factor structure and demonstrates acceptable criterion validity. The results further reveal that the K-DOCS domains cluster together into a single interconnected community, with significant differences in domain connectivity based on gender and age, but not on academic major. The implications of these results for the conceptualization and measurement of creativity are discussed.
This study aims to investigate the interplay between personality traits and verbal behavior within the context of computer-supported collaborative problem solving (CPS).To address this, audio data were collected from nine professional teams during CPS tasks as part of a training.Each team comprised four members, originating from various private and public organizations in Flanders, Belgium.Personality assessments were conducted using the Business Attitudes Questionnaire.Using the audio data, measures of verbal interactions were processed using (a) content analysis based on a coding scheme for computer-supported CPS and (b) linguistic inquiry and word count.The results of this research provide first insights into the influence of personality traits on team's verbal interactions in CPS processes.
Single-case experimental designs (SCEDs) may offer a reliable and internally valid way to evaluate technology-enhanced learning (TEL). A systematic review was conducted to provide an overview of what, why and how SCEDs are used to evaluate TEL. Accordingly, 136 studies from nine databases fulfilling the inclusion criteria were included. The results showed that most of the studies were conducted in the field of special education focusing on evaluating the effectiveness of computer-assisted instructions, video prompts and mobile devices to improve language and communication, socio-emotional, skills and mental health. The research objective of most studies was to evaluate the effects of the intervention; often no specific justification for using SCED was provided. Additionally, multiple baseline and phase designs were the most common SCED types, with most measurements in the intervention phase. Frequent data collection methods were observation, tests, questionnaires and task analysis, whereas, visual and descriptive analysis were common methods for data analysis. Nearly half of the studies did not acknowledge any limitations, while a few mentioned generalization and small sample size as limitations. The review provides valuable insights into utilizing SCEDs to advance TEL evaluation methodology and concludes with a reflection on further opportunities that SCEDs can offer for evaluating TEL.Practitioner notesWhat is already known about this topicWhat this paper addsImplications for practice and/or policy SCEDs use multiple measurements to study a single participant over multiple conditions, in the absence and presence of an intervention SCEDs can be rigorous designs for evaluating behaviour change caused by any intervention, including for testing technology-based interventions. Reveals patterns, trends and gaps in the use of SCED for TEL. Identifies the study disciplines, EdTech tools and outcome variables studied using SCEDs. Provides a comprehensive understanding of how SCEDs are used to evaluate TEL by shedding light on methodological techniques. Enriches insights about justifications and limitations of using SCEDs for TEL. Informs about the use of the rigorous method, SCED, for evaluation of technology-driven interventions across various disciplines. Contributes therefore to the quality of an evidence base, which provides policymakers, and different stakeholders a consolidated resource to design, implement and decide about TEL.
Achieving creativity in the real-world depends on multiple individual and environmental factors. Among them, divergent thinking (DT) has long been considered a key ingredient of creativity and an essential criterion for predicting real-life creative outcomes. However, the link between DT and creative achievement (CA) has yielded heterogeneous results, as outlined by a prior meta-analysis on the DT-CA link published in 2008. Given several limitations of this meta-analysis and the large body of relevant studies that have been published since then, the present article aimed to offer an updated and methodologically rigorous meta-analytical examination of the DT-CA link. A total of 766 effect sizes from 70 studies encompassing 14,901 subjects were analyzed using a meta-analytic three-level model. The results showed that DT was positively, albeit weakly, linked to CA, with only 3% of shared variance. Moderator analyses indicated that this link was robust to variations in DT and CA measures used, gender, educational level, measurement interval between DT and CA, and country of study, but differed by DT task modality, CA domain, and intellectual giftedness. Specifically, the strength of the DT-CA link was significantly larger for (a) verbal DT tasks, (b) CA in the performance domain, and (c) gifted subjects. A significant interaction effect was also found between CA domain and intellectual giftedness, with the DT-CA link being strongest among gifted subjects in the performance domain. Implications of these results for the study and measurement of creativity are discussed.
Meta-analysis is often recognized as the highest level of evidence due to its notable advantages. Therefore, ensuring the precision of its findings is of utmost importance. Insufficient reporting in primary studies poses challenges for meta-analysts, hindering study identification, effect size estimation, and meta-regression analyses. This manuscript provides concise guidelines for the comprehensive reporting of qualitative and quantitative aspects in primary studies. Adhering to these guidelines may help researchers enhance the quality of their studies and increase their eligibility for inclusion in future research syntheses, thereby enhancing research synthesis quality. Recommendations include incorporating relevant terms in titles and abstracts to facilitate study retrieval and reporting sufficient data for effect size calculation. Additionally, a new checklist is introduced to help applied researchers thoroughly report various aspects of their studies.
Society is largely shaped by creativity, making it critical to understand why, despite minimal mean gender differences in creative ability, substantial differences exist in the creative achievement of men and women. Although the greater male variability hypothesis (GMVH) in creativity has been proposed to explain women's underrepresentation as eminent creators, studies examining the GMVH are sparse and limited. This systematic review and meta-analysis were conducted to examine whether the GMVH in creativity can adequately explain the gender gap in creative achievement. We examined the GMVH in creativity, along with mean gender differences, in a range of indicators of creativity and across different sample characteristics and measurement approaches. Effect sizes (k = 1,003) were calculated using information retrieved from 194 studies (N = 68,525). Data were analyzed using three-level meta-analysis and metaregression and publication bias was evaluated using Egger's regression test and contour-enhanced funnel plots. Results revealed minimal gender differences overall, with a slight mean advantage for females (g = -0.10, 95% CI [-0.13, -0.06]) and a trivial variability advantage for males (lnVR = 0.02, 95% CI [0.004, 0.04]) in creative ability scores. However, the magnitude of the effect sizes was moderated by creative domain, task type, scoring type, and study region for mean differences and by country-level gender egalitarianism values for variability. Taken together, gender differences in the mean and variability of creative ability scores are minimal and inconsistent across different contexts, suggesting that the GMVH may not provide much explanatory power for the gender gap in creative achievement.
Reading is a fundamental skill to acquire during children's school career. The present meta-analysis examined research on the effectiveness of digital technologies to foster early reading skills during Tier -1 interventions (ie, high-quality core reading instruction which is intended to promote learning for all children). Unlike previous meta-analyses, this meta-analysis investigated the effectiveness in a broad way, taking into account cognitive versus non-cognitive learning outcomes, near versus far transfer outcomes and immediate versus delayed outcomes. Furthermore, different study characteristics were taken into account including participant characteristics, the targeted reading subskills, duration of intervention, type of technology and the level of integration. A total of 568 effect sizes from 72 studies encompassing 60,890 participants were analysed using a meta-analytic three -level model. A Hedges'g effect size of 0.37 was obtained, suggesting that using digital technologies generally have a positive, albeit small, effect compared to traditional teaching methods. Moderator analyses indicated that this effect was robust to cognitive and non-cognitive outcomes, near and far transfer outcomes, and immediate and delayed outcomes, but differed by participants' age and study quality. Recommendations are formulated to push forward research on how digital interventions can be effectively implemented in the classroom.