The PISA 2022 Creative Thinking Assessment measures the ability to generate diverse ideas, that is, flexibility. However, since test-takers only provide up to three responses, scoring flexibility is challenging. Employed scoring systems rely on dimensions such as linguistic diversity or category systems to quantify how dissimilar a test-taker's responses are. In this work, I will revisit the flexibility scoring in the PISA 2022 Creative Thinking Assessment from the perspective of general diversity measure frameworks, which distinguish variety, balance, and disparity, in the hope of identifying ways to gain more information on test-takers' performances. An empirical illustration demonstrates the potential of diversity scores to better differentiate between different response patterns.
Creative thinking is a primary driver of innovation in science, technology, engineering, and math (STEM), allowing students and practitioners to generate novel hypotheses, flexibly connect information from diverse sources, and solve ill-defined problems. To foster creativity in STEM education, there is a crucial need for assessment tools for measuring STEM creativity that educators and researchers can apply to test how different teaching approaches impact scientific creativity in undergraduate education. In this work, we introduce the Scientific Creative Thinking Test (SCTT). The SCTT includes three subtests that assess cognitive skills important for STEM creativity: generating hypotheses, research questions, and experimental designs. In five studies with young adults, we demonstrate the reliability and validity of the SCTT - including test-retest reliability and convergent validity with measures of creativity and academic achievement - as well as measurement invariance across race/ethnicity and gender. In addition, we present a method for automatically scoring SCTT responses, training the large language model Llama 2 to produce originality scores that closely align with human ratings - demonstrating STEM-specific, automated creativity assessment for the first time. The full SCTT, along with the code to automatically score it, are available on a repository in the Open Science Framework.
Abstract Background Creativity is essential for engineering design, yet its assessment remains challenging due to the resource‐intensive nature of traditional evaluation methods. Purpose/Hypothesis(es) This study investigates the potential of automatic item generation (AIG) using large language models (LLMs) to create psychometrically sound items for the design problem task (DPT), which measures creative thinking in engineering. Design/Method We developed and validated engineering design problems across three domains: ability difference and limitations (e.g., assisting people with learning impairments), transportation and mobility (e.g., reducing traffic congestion in mega cities), and social environments and systems (e.g., improving access to clean water in remote areas). The study comprised three phases with samples matched on race and ethnicity: (1) content validation with a diverse sample of 40 engineers evaluating item clarity and validity; (2) item administration to 462 engineering students; and (3) response evaluation by 65 expert raters assessing originality and effectiveness. Results Results demonstrated that LLM‐generated items achieved comparable or higher content validity rates than expert‐written items (43% vs. 20% success). Bayesian confirmatory factor analysis supported a unidimensional model for fluency, originality, and effectiveness scores, with excellent reliability estimates (.92–.95). While fluency showed minimal correlation with originality ( r = −.11) and effectiveness ( r = −.04), originality and effectiveness were strongly positively correlated ( r = .73). Conclusions The present research advances our understanding of automated assessment generation in engineering education, provides empirical evidence for the psychometric properties of AI‐generated engineering creativity tasks, and offers a scalable approach for measuring creative thinking in engineering classrooms.
Even in the time of artificial intelligence, creative thinking is considered an important 21st century skill. Nevertheless, our understanding of how contextual factors such as socio-economic status (SES) and gender affect creativity is still limited – especially from an international perspective. In the current study, we thus examined the impact of gender and SES on creative thinking across 62 countries, using data from the 2022 Programme for International Student Assessment (PISA). Creative thinking, alongside mathematics and reading, was analyzed using a two-stage meta-analytic approach, integrating effect sizes from country-specific samples with a total sample size of N = 493,660. Our results revealed consistent gender disparities, with females outperforming males in creative thinking and reading, a trend robust across countries but with considerable variability. Gender disparities were less pronounced in the mathematical domain. Moreover, SES was found to be a strong predictor of creative thinking, mathematics, and reading, with higher SES associated with better performance across all domains. There was no substantial interaction effect between gender and SES for creative thinking and reading, suggesting that SES advantages are consistent across genders. Our analyses indicated substantial heterogeneity between countries, emphasizing the need for context-specific educational policies. These findings highlight the pervasive influence of gender and SES on fundamental educational outcomes and hence stress the necessity of tailored interventions to address these disparities around the globe.
The Programme for International Student Assessment (PISA) is one of the most costly, influential, and wide-reaching research efforts in educational psychology. But are the measures they administer as valid as they should be? As part of its 2022 cycle, PISA included a focus on creative thinking, incorporating scales in the student and context questionnaires designed to measure various creativity-related constructs. However, information on the development of these items is limited, and the extent to which these items validly capture their intended constructs has not been systematically evaluated. If these measures lack validity, the field risks interpreting current findings and secondary analyses of PISA data in ways that may not be fully psychologically meaningful, possibly leading to incorrect recommendations for educational practice. This issue is addressed by employing a multistage expert judgment procedure (involving a total of N = 11 experts) to assess the content validity of 127 creativity-related items identified from an initial pool of 936 items from PISA 2022 student and context questionnaires. After an extensive pilot phase (Study 1), experts categorized items according to predefined construct and domain definitions, evaluated item quality, and rated construct coverage (Study 2). Findings revealed substantial misalignment between the constructs assigned by the Organisation for Economic Co-operation and Development (OECD) and those identified by the expert panel (an average overlap of 41%). The results ultimately offer a detailed item-content map as a ground for revised measurement of creativity-relevant constructs, which are intended to support future analyses grounded in psychologically and educationally valid measures.
A fundamental question in creativity research is whether chance models of creativity align with the idea that cognitive ability can explain individual differences in creative cognition. Using two datasets ( NDataset1 = 462 and NDataset2 = 331) we extended previous work on the equal odds baseline (i.e., a chance model of creativity) by using a latent variable analytic approach to model latent residual factors rectified for fluency. We compared three measurement models: the EOB latent variable model, a residualized model, and a ratio score model, examining their reliability and associations with cognitive abilities (i.e., fluid intelligence and working memory capacity). We found that when a chance model of creativity is deemed to be an appropriate fit for a given dataset, as evidenced by Dataset 2 in our study, a cognitive interpretation of originality factor η is warranted (i.e., around 21% of variance can be explained by working memory capacity). Our work, thus, refines discussions about the compatibility of chance models with cognitive explanations of creative cognition.
Recent research has investigated the relationship between creativity and its defining components, which, according to the standard definition of creativity, are novelty and value. Previous work has shown that nonlinear relationships between creativity and its components, as well as interactions among the components, are necessary to explain overall creativity. Building on and extending previous research, we propose to adapt response surface analysis as a tool for creativity researchers to explore the multidimensional relationship between creativity and its constituents. This idea is explored empirically in two large data sets. Data Set 1 includes responses from a creative verb association task (N = 4,468), while Data Set 2 includes ideas from a real-life creative problem-solving task (N = 662). We employed response surface analysis to creativity measures mapping onto the standard two-component definition of creativity (Data Sets 1 and 2) and a three-component extension (Data Set 2). Consistent with the standard two-component definition, we found that creativity ratings were highest when novelty and value were both of maximally high quality. Furthermore, our investigation of a three-component definition revealed positive linear effects of novelty and feasibility on creativity ratings, whereas value did not incrementally contribute to the prediction of creativity beyond novelty and feasibility. We discuss the empirical findings as well as potential future applications of response surface analysis to understand the relationship between creativity and its components.
Automated scoring with machine learning has received considerable attention, especially for the Alternative Uses Test (AUT) of divergent thinking. In particular, supervised learning on word embeddings and approaches based on large language models predict human ratings with sufficient accuracy for many practical purposes. However, since automated systems will potentially be used without a ``human in the loop'', the robustness of any model is crucial, not only for the validity of basic research, but also for applications. We investigate the potential of adversarial examples (AEs) as robustness checks. AEs are synthetic responses obtained by subtle permutations of original responses that are assigned a substantially different score by the automated system. Specifically, we propose to employ synonyms to obtain semantically close synthetic responses. A synthetic response is an AE, when the deviation of the automated scores of the synthetic response and the original response is at least as large as the prediction error of the automated system. A random forest trained on GloVe word embeddings predicts human ratings of 2,690 distinct original AUT answers (stimuli: brick, paperclip) and for 42% AEs can be found, including both single and multi-word responses. On average, the AEs have slightly higher automated scores than the originals. While a small fraction of the AEs are invalid AUT responses, pronounced large deviations indicate that the prediction model is non-robust. Retraining including the AEs in the training data improves the robustness. We recommend to use AEs routinely to assess the robustness of automated scoring systems.
Aesthetics is crucial for product design and success, yet there are no differentiated and validated evaluation tools for the assessment of interactive products (e.g. IT products, household appliances, home entertainment devices or power tools). This contrasts with the significant efforts invested in the development and optimization of such products. To close this gap, we developed and validated the Product Aesthetics Inventory (PAI) and its short version PAI-S through three studies (with a total of > 7,000 respondents). In a pre-study with design experts (N = 6) and product users (N = 4), we developed the initial item set. This set was then used by 6,002 participants in an online survey to evaluate household appliances. In Study 1, data from 3,000 participants helped determine the number of factors for product aesthetics using factor analysis. Study 2 verified this structure with the remaining 3,002 participants, confirming construct validity with established aesthetic scales, intentions and general judgements. Both studies demonstrated strict measurement invariance across four household appliance categories, underscoring the psychometric quality of the scales. Study 3 (N = 1,028) extended these findings to power tools, office equipment and consumer electronics, affirming reliability, validity and measurement invariance for most PAI dimensions. The final PAI has 32 items measuring eight dimensions: visual aesthetics, operating elements, brand logo, feedback sounds, operating sounds, haptics, interaction aesthetics and impression. An overall product aesthetics factor can be assessed with the PAI and the 8-item short version PAI-S. This article also presents advances in scaling approaches for the PAI and PAI-S, along with practical benchmarks for three broad product categories, offering valuable guidance for the application and interpretation of product aesthetics assessments.
The alternate uses task (AUT) is the most popular measure when it comes to the assessment of creative potential. Since their implementation, AUT responses have been rated by humans, which is a laborious task and requires considerable resources. Large language models (LLMs) have shown promising performance in automatically scoring AUT responses in English as well as in other languages, but it is not clear which method works best for German data. Therefore, we investigated the performance of different LLMs for the automated scoring of German AUT responses. We compiled German data across five research groups including ~50,000 responses for 15 different alternate uses objects from eight lab and online survey studies (including ~2300 participants) to examine generalizability across datasets and assessment conditions. Following a pre-registered analysis plan, we compared the performance of two fine-tuned, multilingual LLM-based approaches [Cross-Lingual Alternate Uses Scoring (CLAUS) and the Open Creativity Scoring with Artificial Intelligence (OCSAI)] with the Generative Pre-trained Transformer (GPT-4) in scoring (a) the original German AUT responses and (b) the responses translated to English. We found that the LLM-based scorings were substantially correlated with human ratings, with higher relationships for OCSAI followed by GPT-4 and CLAUS. Response translation, however, had no consistent positive effect. We discuss the generalizability of the results across different items and studies and derive recommendations and future directions.
Researchers and educators interested in creative writing need a reliable and efficient tool to score the creativity of narratives, such as short stories. Typically, human raters manually assess narrative creativity, but such subjective scoring is limited by labor costs and rater disagreement. Large language models (LLMs) have shown remarkable success on creativity tasks, yet they have not been applied to scoring narratives, including multilingual stories. In the present study, we aimed to test whether narrative originality-a component of creativity-could be automatically scored by LLMs, further evaluating whether a single LLM could predict human originality ratings across multiple languages. We trained three different LLMs to predict the originality of short stories written in 11 languages. Our first monolingual model, trained only on English stories, robustly predicted human originality ratings (r = .81). This same model-trained and tested on multilingual stories translated into English-strongly predicted originality ratings of multilingual narratives (r >= .73). Finally, a multilingual model trained on the same stories, in their original language, reliably predicted human originality scores across all languages (r >= .72). We thus demonstrate that LLMs can successfully score narrative creativity in 11 different languages, surpassing the performance of the best previous automated scoring techniques (e.g., semantic distance). This work represents the first effective, accessible, and reliable solution for the automated scoring of creativity in multilingual narratives.
Divergent thinking (DT) ability is widely regarded as a central cognitive capacity underlying creativity, but its assessment is challenged by the fact that DT tasks yield a variable number of responses. Various approaches for the scoring of DT tasks have been proposed, which differ in how responses are evaluated and aggregated within a task. The present study aimed to identify methods that maximize psychometric quality while also reducing the confounding effect of DT fluency. We compared traditional scoring approaches (summative and average scoring) to more recent methods such as snapshot as well as top- and max-scoring. We further explored the moderating role of task complexity as well as metacognitive abilities. A sample of 300 participants was recruited via Prolific. Reliability evidence was assessed in terms of internal consistency, concurrent criterion validity in terms of correlations with real-life creative behavior, creative self-beliefs, and openness. Findings confirm that alternative aggregation methods reduce the confounding effect of DT fluency. Reliability tends to increase as a function of the number of included responses with three responses as a minimal requirement for decent reliability evidence. Convergent validity was highest for snapshot as well as max-scoring when using a medium number of three ideas.
Statistical modeling of scientific productivity and impact provides insights into bibliometric measures used also to quantify differences between individual scholars. The Q model decomposes the log-transformed impact of a published paper into a researcher capacity parameter and a random luck parameter. These two parameters are then modeled together with the log-transformed number of published papers (i.e., an indicator of productivity) by means of a trivariate normal distribution. In this work we propose a formulation of the Q model that can be estimated as a structural equation model. The Q model as a structural equation model allows to quantify the reliability of researchers’ Q parameter estimates, it can be extended to incorporate person covariates, and multivariate extensions of the Q model could also be estimated. We empirically illustrate our approach to estimate the Q model and also provide openly available code for R and Mplus.
In recent years, the importance of mobile devices has increased for education in general and more specifically for science and mathematics education. In the classroom, approaches for teaching with mobile devices include using student-owned devices (“bring your own device”; BYOD approach) or using school-owned devices from central pools (POOL approach). While many studies point out features of mobile learning and BYOD that are conducive to learning, a research gap can be identified in the analysis of effects of mobile device access concepts on teaching–learning processes. Thus, this study aimed to empirically compare BYOD and POOL approaches in terms of learning performance and cognitive performance (subject knowledge development, cognitive load, concentration performance). Furthermore, the analyses included specific characteristics and preconditions (gender, socioeconomic status, fear of missing out, problematic smartphone use). A quasi-experimental study (two groups) was conducted in year 8 and 9 physics classes (N = 339 students) in which smartphones are used for different purposes. The present data show no group differences between the BYOD and the POOL approach in the group of learners with respect to subject knowledge development, cognitive load, and concentration performance. However, individual findings in subsamples indicate that the POOL approach may be beneficial for certain learners (e.g., learners with low fear of missing out or learners tending toward problematic smartphone use). For school practice, these results indicate that organizational, economic, and ecological aspects appear to be the main factors in deciding about the mobile device access concept.
Various bibliometric indicators have been used to assess the researchers’ impact, but composites of such indicators, namely a metric that combines various individual indicators to describe a complex construct, have received a strong critique thus far. We employ concepts from psychometrics to revisit a composite proposed by Ioannidis et al. (2020) that aimed to represent researcher impact. Based on a selected sample of highly cited researchers, our proof-of-concept study presents a psychometrically principled composite formation. Specifically, by relying on the congeneric measurement model (and related models) rooted in classical test theory, we found that one of the proposed indicators clearly violated the congeneric model’s fundamental assumption of unidimensionality, and two other indicators were excluded for redundancy. The resulting composite based on only three bibliometric indicators was found to display excellent reliability. Importantly, the reliability approached that of the composite based on five indicators, and it was clearly better than the original six-indicator composite. Further, we found rather homogeneous effective weights (i.e., relative contributions of each indicator to composite variance) for simple sum scores, and these weights were close to those calculated using an algorithm for equally effective weights. While the congeneric measurement model also showed strong measurement invariance across sexes, this model’s loadings and intercepts were not measurement invariant across scientific fields and academic age groups. Notably, we found that various derived composites correlate positively with academic age, hinting at a lack of fairness of the composites.
While a rich methodology for analyzing response patterns for accuracy and time-on-task is at hand via Item Response Theory (IRT), tests with time cutoffs are so far harder to handle. Given that this test mode is widely applied, especially in the context of paper-and-pencil testing, there is a lack of psychometric techniques for a relevant number of tests. In this context, the original work of Rasch and his Rasch Poisson Counts model indeed offers an approach for this scenario that is adequate to solve the problem but which leads to model violations in many cases. Recent developments in statistical modeling - the so-called Conway Maxwell Poisson Counts Model (CMPCM) - can solve the problem of under- and overdispersion. We apply this model to the norm data of the ELFE II reading comprehension test and analyze patterns of over- and underdispersion with regard to speededness and mode effects. CMPCM with subtest-specific dispersion was adequate to model the raw test data, with underdispersion occurring mainly in highly speeded subtests with low difficulty and overdispersion in less speeded subtests with high difficulty. Thus, the CMPCM could contribute to psychometric methodology to appropriately model tests with time cutoffs on the subtest level.
Divergent thinking (DT) tasks are among the most established approaches to assess creative potential. Although DT assessments are widely used, there exist many variants on how DT tasks can be administered and scored. We present findings from a preregistered, systematic review of DT assessment methods aiming to determine the prevalence of various DT assessment conditions as well as to identify recent trends in the field. We searched two electronic databases for studies that have investigated creativity with DT. We then screened a total of 2066 publications published between 1957 and 2022 and identified 451 eligible studies within 396 articles. The employed data coding system discerned more than 110 conditions and options that establish the specific administration and scoring of DT tasks. Amongst others, we found that the Alternate Uses Task is the most used DT task, task time is commonly set between two and three minutes, and responses are often scored by human raters. While traditional task instructions emphasized idea fluency, more recent studies often encourage creativity and originality. Trends in scoring include the assessment of response quality (i.e., originality/creativity) and response aggregation methods that account for the confounding effect of fluency (e.g., average or subset scoring such as top- and max-scoring) and generally an increasing instruction-scoring-fit. Overall, numerous studies lacked information regarding the precise procedures for both administration and scoring. In sum, the review identifies established practices and trends but also highlights substantial heterogeneity and underreporting in DT assessments that poses a risk to reproducibility in creativity research.
The term "creative" is commonly used in everyday language and in academic discourse to discuss the nature of artistic and innovative productions. This usage inherently implies the existence of a variable of creativity that allows different creative works to be compared. The standard definition of creativity asserts that a production must possess both value and novelty in order to be considered truly creative. However, previous psychometric studies aimed at establishing the existence of such a creativity variable based on these two dimensions have produced results that seem to demonstrate their independence or even negative association, based on a weak to negative correlation between value and novelty. These widely replicated empirical results seem to call into question the notion of a single creativity variable associated with productions, leading to a paradoxical use of the term "creative" to describe the object produced. In our study, we aimed to reproduce these results while addressing methodological errors made in previous efforts to establish construct validity. This work led us to test the existence of a common cause for the observed variations in novelty and value. The higher order factor we obtain in our analysis encompasses subtle differences from the conventional creativity axis and interacts negatively with novelty, while correlating positively with value.
In psychology and education, tests (e.g., reading tests) and self-reports (e.g., clinical questionnaires) generate counts, but corresponding Item Response Theory (IRT) methods are underdeveloped compared to binary data. Recent advances include the Two-Parameter Conway-Maxwell-Poisson model (2PCMPM), generalizing Rasch's Poisson Counts Model, with item-specific difficulty, discrimination, and dispersion parameters. Explaining differences in model parameters informs item construction and selection but has received little attention. We introduce two 2PCMPM-based explanatory count IRT models: The Distributional Regression Test Model for item covariates, and the Count Latent Regression Model for (categorical) person covariates. Estimation methods are provided and satisfactory statistical properties are observed in simulations. Two examples illustrate how the models help understand tests and underlying constructs.
Aesthetics is essential to the design of products. Nevertheless, aesthetic quality is often assessed with inaccurate, ad hoc scales. Therefore, we have developed the Product Aesthetics Inventory (PAI) and its short version, the PAI-S. A Pre-Study using face-to-face interviews (N = 6 design experts, N = 4 product users) served as basis for the development of test items. The resulting item battery was then used by N = 6,002 persons in an online survey to evaluate various types of household appliances. In Study 1, n = 3,000 of those participants' data were used to determine the number of product aesthetics factors and to select the optimal items by combining exploratory graph analysis (EGA) and confirmatory factor analysis (CFA) approaches. We found an eight-factor structure consisting of the dimensions Visual Aesthetics, Operating Elements, Logo, Feedback Sounds, Operating Sounds, Haptic, Interaction Aesthetics, and Impression. In Study 2, we confirmed this structure by a CFA with the remaining n = 3,002 participants, resulting in excellent model fit. Further, construct validity was confirmed by correlations with other established aesthetic scales, intentions, and general judgments. In both studies, strict measurement invariance was achieved across four household appliances, which further highlights the psychometric quality of the newly developed scales. In Study 3 (N = 1,028), we demonstrated that most of the PAI factors as well as reliability, validity, and measurement invariance findings generalized to power tools, office equipment, and home entertainment devices. Finally, we provide recommendations for the usage of the PAI and implications for further research.