
People tend to avert their eyes or blink to reduce distraction, but visual decoupling can also occur when cognition turns inward, for instance, during memory retrieval. Using mobile eye tracking, a latent trial-level measure of this shift is introduced - the Visual Decoupling Index (VDI) - derived from blinks, fixations, and gaze aversion. In Experiment 1, VDI was assessed during creative verbal generation in a natural face-to-face setting. A Sentence Continuation Task (SEN) required switching between listening (externally directed cognition) and generating creative continuations (internally directed cognition). VDI increased reliably during generation vs. listening, indicating robust visual decoupling during internally directed cognition. Experiment 2 added two control tasks - Number Sequence Continuation (NUM) and Animal Visual Description (VIS). VDI differed by task such that most decoupling occurred during the creative task (SEN > VIS > NUM). Participant-level creativity scores (human ratings and GPT-4 rankings) converged strongly and related to divergent thinking, but neither predicted mean VDI. Big-five personality traits and creative achievement showed no reliable associations with VDI. Together, ecological eye-tracking, psychometric modeling, and human/AI scoring of creative output yield a principled index of internal attention: a signature of decoupling that is pronounced during creative cognition but not specifically related to creative quality.
Strong and growing interest in Malevolent Creativity has created a need for valid and reliable measures of, among other things, malevolent creative ideation. The Malevolent Creativity Behavior Scale (MCBS) was created in response to weaknesses identified in earlier studies of malevolent creativity. However, concerns over the face validity of the MCBS, coupled with an increasing popularity of the scale in empirical studies, prompted an examination of the construct and concurrent (i.e. predictive) validity of the MCBS. The results of this study suggest that the MCBS has poor validity: it measures neither creative nor malevolent ideation sufficiently well to warrant its use in empirical studies of malevolent creativity. These findings caution against the uncritical use of the MCBS and underscore the need for new instruments, and approaches, that better measure the intersection of novelty, effectiveness and harmful intent.
It is argued in this Commentary that Green et al. proposal for completely separate or orthogonal definitions of creative processes and creative products is problematic regarding products. On conceptual grounds, drawing on recent philosophical analyses, it is argued that for a product to be classed as creative, it must be produced in the right sort of way, i.e. by an agential creative process. It is suggested that Green et al.'s process definition could be used in a nonorthogonal definition of creative products as a necessary process underlying creative products. It is also noted that the proposal for nonorthogonal definitions of creative products, where product creativity is process-dependent, make the creative process definition primary in defining creativity. The arguments here also support the use of process measures in empirical studies to back up judgments regarding the creative nature of any putative creative products.
Cognitive fixation mechanisms, such as functional fixedness, mental set (including Einstellung), and design fixation, have long been studied as mental barriers to individual problem solving and creativity. While the literatures on each of these mechanisms draw from similar theoretical traditions and have a shared language focused on cognitive fixation, each literature developed with its own set of foci, methodological paradigms, and findings, resulting in an empirically rich, but disconnected, broader literature on cognitive fixation and creativity. This narrative review organizes the literatures on these cognitive fixation mechanisms and their relationships with creativity, summarizing their conceptual and methodological underpinnings, key findings, and limitations. More importantly, this review integrates these literatures by mapping common and distinct features across fixation types. The results demonstrate how fixation mechanisms can be viewed as a family of interrelated, yet distinct, cognitive phenomena that influence creativity and the thinking processes underlying creativity. Finally, this review highlights theoretical and practical implications and provides an agenda for future research on cognitive fixation mechanisms and creativity.
Creative thinking is emphasized by educators to prepare students for a rapidly changing world. Recent studies have suggested that creative thinking may not develop uniformly across individuals but instead cluster into distinct classes. Identifying these classes is critical for effectively supporting students' creative skill development and tailoring educational strategies. This study investigates latent classes of creative thinking skills among 15-year-old students using data from the PISA 2022 Creative Thinking Assessment and applies a mixture item response model (MixIRT). The results revealed the three distinct latent classes: Class 3 (high creativity) was more likely to have high creative thinking skills and exhibited high proficiency in mathematics, reading, and science than the other classes. Members of Class 2 (medium creativity) demonstrated a medium level of creative thinking skills. Finally, members of Class 1 (low creativity) tended to have low creative thinking skills and displayed the lowest proficiency in mathematics, reading, and science. These findings highlight important distinctions among students and suggest opportunities for targeted educational interventions. Developing creativity training programs, especially for students in the low creativity class, could enhance their cognitive skills and promote the development of creative thinking, better preparing them to succeed in a complex world.
Creativity is widely recognized as a core skill for the twenty-first century, yet its early developmental pathways remain insufficiently understood. This longitudinal study examined trajectories of creativity in a sample of 230 children aged 4-7, assessed across five waves from pre-kindergarten to first grade. Using group-based trajectory modeling, we identified three distinct patterns: "Accelerated Growth" (22.3%), "Gradual Growth" (72.6%), and "Early Peak and Recovery" (5.1%). Higher quality of fantasy during pretend play quality predicted membership in the Accelerated Growth pattern (OR = 4.91), while higher receptive language predicted Early Peak and Recovery (OR = 1.17). Demographic factors such as gender and school-level socioeconomic vulnerability were not significant predictors. Trajectory membership was associated with developmental outcomes: children in higher-growth groups demonstrated superior expressive and receptive language skills, as well as higher divergent thinking scores, by the first semester of 1st grade. Socioemotional differences were limited, with fewer peer problems observed in the Accelerated Growth group. Findings underscore the importance of promoting high-quality fantasy play and language development in early childhood education. Limitations include small subgroup sizes and reliance on visual-narrative creativity measures. Future research should incorporate multidimensional assessments and explore diverse populations to advance the understanding of creativity development.
The Runco Ideational Behavior Scale (RIBS) is a widely used instrument for assessing students' creative ideation; however, a reliable measure of ideational behavior has yet to be established in the Chinese context. This study aimed to validate a Chinese version of the RIBS with a sample of 1,331 students, comprising 719 fourth-graders and 612 eighth-graders. Exploratory factor analysis showed that the original 23-item scale was not suitable for the Chinese sample, which resulted in a shorter 20-item version. Confirmatory factor analysis revealed a two-factor structure representing two types of creative ideation: constructive, goal-directed ideational behavior and intrusive, disorganized ideational flow. This adapted version demonstrated satisfactory internal consistency reliability and composite reliability. The scale also exhibited reasonable factorial, convergent, and discriminant validity. Furthermore, the measurement invariance across gender and grade level confirmed the applicability of the RIBS for both boys and girls, as well as primary and middle school students. This study provides robust evidence for the utility of the 20-item RIBS in the Chinese context, facilitating the assessment of students' creative potential through the measurement of their ideational behavior.
This exploratory study investigates the role of executive control in metaphor comprehension. Using event-related potentials (ERPs) and standardized low-resolution brain electromagnetic tomography analysis (sLORETA), we examined the neural mechanisms of individuals with high and low executive control in understanding metaphors with varying levels of familiarity. We found that metaphor comprehension is constrained by both executive control and familiarity. The results indicated that both the left and right hemispheres were involved in metaphor comprehension, but as familiarity decreased, more activations in the right hemisphere were observed. Individuals with low executive control elicited more negative amplitudes of N400 and P600, compared to those with high executive control, and their performance was influenced by familiarity. For individuals with high executive control, more activation of the middle temporal gyrus was found when processing unfamiliar metaphors. These findings suggest that executive control plays a critical role in metaphor comprehension by suppressing irrelevant information. Readers with low executive control consume more cognitive resources to understand metaphors. In contrast, those with high executive control may exhibit greater neural flexibility to allocate cognitive resources more effectively for metaphor comprehension. Furthermore, the right hemisphere may act as a potential reserve resource, contributing to the comprehension of unfamiliar metaphors.
This second-order meta-analysis examined 52 first-order creativity meta-analyses (164 effect sizes) which met inclusion and exclusion criteria. These meta-analyses provided data from 2609 studies (1,248,416 research participants) and 16,153 primary effect sizes. The original meta-analyses all examined creativity, but each had a specific focus. Even with that large range of foci, certain questions could be addressed with a second-order meta-analysis. More specifically, a comprehensive statistical analysis provided an effect size for all previous meta-analyses (a) where creativity had been used as a predictor and (b) where creativity was the outcome. This moderator was statistically significant, F(1, 16.3) = 9.43, p = .007. The average second-order effect size was .12 when creativity was the criterion and .29 when it was a predictor. Four categories were created for the type of creativity variable used in first-order meta-analyses: creative outcomes, creative process, divergent thinking, and overall creativity. In the first model, where creative process and divergent thinking were analyzed separately, creative process had an effect size of .27, while divergent thinking had an effect size of .14. Different correlates of creativity were also analyzed. Of these, interventions and educational programs yielded the largest effect size (r = .20). Implications and limitations are discussed.
Using data from the Programme for International Student Assessment (PISA) 2022, this study examines the impact of digital capital on middle school students' creative thinking. The sample includes 12,333 15-year-old students from Hong Kong, Macao, and Taiwan in China. Digital capital is measured across three dimensions: usage frequency, usage ability, and usage interest. The results show that digital capital has an overall positive effect on creative thinking. Further analysis reveals clear dimensional heterogeneity: usage frequency shows a negative or non-linear association, whereas usage ability and usage interest have significant positive effects. These findings suggest that creative development depends not simply on more digital participation, but on the quality and motivational orientation of digital engagement. Further analyses show that the positive effect varies by gender, parental education, and school location. Mechanism analysis indicates that digital capital promotes creative thinking through self-efficacy, assertiveness, empathy, and curiosity. These findings deepen understanding of how digital capital shapes creative development and provide implications for fostering creative thinking in the digital era.
This study introduces the concept of creative opportunities as a new aspect of work design and provides evidence for its validity. The concept of creative opportunities describes the perceived degree to which the content and organization of the employees' work tasks, activities, relationships, and responsibilities within an organizational or work-related context (i.e. work design) affords the employees possibilities for creative action. In contrast to creative requirements, creative opportunities focus more on the possibilities for creative action in opposition to expectations and duties to be creative. Creative opportunities are expected to be related to employee creativity as well as to well-being indicators such as meaningfulness, thriving, and life satisfaction. A cross-sectional survey study (N = 403) shows evidence for the incremental validity of creative opportunities in predicting these outcomes beyond creative requirements and work autonomy.
The structure of semantic networks reflects how knowledge is organized in memory and shapes creative thinking by influencing how concepts are accessed and combined. However, most existing findings are cross-sectional, leaving it unclear how changes in knowledge structures shape corresponding changes in semantic networks and creative development over time. This longitudinal study tested whether optimizing children's knowledge structures through thinking-based teaching fosters dynamic changes in semantic networks and creative thinking. A total of 108 Chinese elementary students participated (44 girls and 64 boys; M = 9.85 years, SD = 0.50) in the intervention and pre-post changes in creativity and semantic network measures were examined using paired-sample t-tests and regression analyses. Results indicated that students receiving the intervention showed significant gains in fluency, flexibility, and originality compared with controls (t = 2.48, p = .015), and a follow-up study confirmed longitudinal improvements in semantic network efficiency and creativity (t = 3.90, p < .001). Regression analyses indicated that baseline cognitive flexibility and originality jointly predicted gains in originality (F(2, 42) = 6.46, p = .0036, R2 = .235). Overall, findings provide longitudinal, intervention-based evidence that optimizing knowledge structures reorganizes semantic networks and supports improvements in creative thinking.
This systematic review examined the characteristics and psychometric properties of creativity assessments for PK-12 students. Using the PRISMA framework, we identified 169 unique studies published between 2002 and 2023, which encompassed 288 assessments meeting the inclusion criteria. Four coders independently analyzed all studies using a comprehensive system that included participant demographics, facets of creativity measured, measurement approaches, and reported score reliability and validity evidence. Several areas for growth were identified. Very few studies worldwide reported race/ethnicity (10%); in the U.S. race was reported in 48% of studies. Only 9.46% of studies reported students' gifted/talented status. Disability status was reported in just seven studies (4.14%), five of which were U.S.-based. Pre-school students were included in only 3.8% of studies. Few assessments captured the creative process (5.28%) or environmental influences (5.69%). Although usefulness is a core component of creativity, it was assessed in just seven studies (1.59%), mostly within creative problem-solving. Fewer than half the studies (40.10%) reported validity evidence. Together, this review highlighted promising developments and provides guidance for researchers and practitioners seeking to assess creativity equitably and comprehensively in educational settings.
Our study analyzes "creative inflection points"-key moments when design direction fundamentally shifts-to compare human and generative AI (GenAI) contributions to innovation. Through three design challenges (visual communication, product ideation, speculative interaction), we find that GenAI performs well in low-constraint settings, offering high productivity and visual novelty. However, its outputs often lack semantic depth and contextual grounding. AI breakthroughs typically emerge as unanchored metaphorical leaps, while human shifts are driven by cognitive insights or strategic reframing. Rather than replacing human creativity, GenAI serves as a catalyst for ideation. We advocate for hybrid design workflows that leverage AI's divergent exploration and human contextual reasoning, advancing new paradigms of human-AI co-creation.
Creative thinking is a primary driver of innovation in science, technology, engineering, and math (STEM), allowing students and practitioners to generate novel hypotheses, flexibly connect information from diverse sources, and solve ill-defined problems. To foster creativity in STEM education, there is a crucial need for assessment tools for measuring STEM creativity that educators and researchers can apply to test how different teaching approaches impact scientific creativity in undergraduate education. In this work, we introduce the Scientific Creative Thinking Test (SCTT). The SCTT includes three subtests that assess cognitive skills important for STEM creativity: generating hypotheses, research questions, and experimental designs. In five studies with young adults, we demonstrate the reliability and validity of the SCTT - including test-retest reliability and convergent validity with measures of creativity and academic achievement - as well as measurement invariance across race/ethnicity and gender. In addition, we present a method for automatically scoring SCTT responses, training the large language model Llama 2 to produce originality scores that closely align with human ratings - demonstrating STEM-specific, automated creativity assessment for the first time. The full SCTT, along with the code to automatically score it, are available on a repository in the Open Science Framework.
Are individual-level factors necessary for creativity to occur in the workplace? Using a novel statistical approach, Necessary Condition Analysis, we explored empirically whether individual factors were critical to creativity in the workplace, drawing on Amabile's componential theory, using a sample of 1384 workers in France. We examined three types of drivers: intrinsic motivation, job self-efficacy and creative self-efficacy. We examined four known creativity-relevant processes: openness to experience, creative personality, creative personal identity, as well as creative process engagement. Crucially, results showed that creative self-efficacy was a critical driver for creativity: an employee will not be able to achieve high creative performance if he or she does not have strong confidence about his or her creative capacities, regardless of other factors. However, neither job self-efficacy nor intrinsic motivation proved to be necessary for creativity: their absence can be compensated by other factors. We observed that all creativity-relevant processes appeared necessary for creativity, even though they each showed small effect sizes. This study and its findings highlight the need to distinguish between what makes a variable "important" versus "necessary," in the field of creativity and innovation.
Creativity assessment at scale is difficult because expert ratings are resource-intensive and hard to use in dynamic settings. Large language models (LLMs) offer potential for automated assessment, yet their validity for evaluating multimodal creative artifacts in authentic educational contexts remains unexplored. Here we evaluated whether LLMs can assess human creativity in Physics Playground, an educational physics game where students design playable levels. We compared rubric-guided and rubric-free prompting approaches across 421 student-created levels, tested three multimodal input configurations, and examined reliability and model capacity effects using GPT and Gemini model families. Rubric-guided prompting yielded strong agreement with human expert ratings compared to rubric-free approaches (rs ranged from .61 to .81). Multimodal inputs combining images with structured data significantly enhanced validity compared to text-only methods. These effects were consistent across GPT-4o and Gemini 2.5 Flash. Also, single model calls achieved comparable reliability to averaged responses. Model capacity substantially influenced performance, with larger, high-capacity models (e.g., GPT-4o) consistently outperforming smaller, low-capacity variants (e.g., GPT-4o-mini). Theoretically, these findings extend creativity assessment to multimodal artifacts in authentic contexts. Practically, embedding assessment in learning games enables them to foster creativity and support STEM learning and AI literacy.
Divergent thinking (DT) ability is a fundamental aspect of creativity, but its assessment remains challenging by the reliance on effortful human ratings and persistent uncertainty regarding how to aggregate scores across a variable number of responses. Recent work demonstrated that automated scoring based on large language models (LLMs) can predict human creativity ratings. Other research has evaluated the psychometric quality of different response aggregation methods for human ratings by comparing their convergent construct validity with respect to external criteria. The present study integrates these two lines of work and investigates the construct validity evidence of automated creativity scorings derived from five LLMs (CLAUS, Ocsai, GPT-4, Llama 3.3, and Claude 3.5) under different aggregation methods. Instead of just relating LLM-based ratings to human ratings, we compared the validity evidence between rater-based and LLM-based scores which opens up the possibility that automated scoring could prove even more valid. Analyses were based on data from 300 participants who completed five Alternate Uses Tasks. Findings showed that general-purpose LLMs yielded equal or even slightly higher construct validity evidence compared to human ratings, especially when using max-3 scoring. These results suggest that automated DT scoring can serve as a psychometrically sound alternative to rater-based scoring.
Although objective, automated measures of verbal creativity have shown promise by correlating well with subjective creativity ratings, objective measures of musical creativity remain underexplored. This study investigated whether melodic complexity - specifically the Lempel-Ziv complexity of pitch sequences - is related to perceived musical creativity in improvisations. We further investigated whether this relationship is affected by participants' openness to experience and musical expertise. In an online study, 495 participants with at least five years of musical experience rated improvisations on liking, originality, and appropriateness. They also completed the openness to experience scale of the Big Five Inventory and provided information about their musical expertise. Results revealed inverted U-shaped relationships between Lempel-Ziv complexity and ratings, with improvisations of moderate complexity receiving the highest evaluations. However, openness to experience did not significantly affect this relationship. Participants with greater musical expertise tended to rate improvisations lower overall, while they rated improvisations made by professionals higher than those by amateurs on liking and appropriateness, highlighting the role of expertise in evaluation. These findings demonstrate individual differences in the assessment of musical creativity. While Lempel-Ziv complexity is a relatively simple metric, the results suggest its potential for informing the development of quantitative measures of musical creativity.