Recent research has investigated the relationship between creativity and its defining components, which, according to the standard definition of creativity, are novelty and value. Previous work has shown that nonlinear relationships between creativity and its components, as well as interactions among the components, are necessary to explain overall creativity. Building on and extending previous research, we propose to adapt response surface analysis as a tool for creativity researchers to explore the multidimensional relationship between creativity and its constituents. This idea is explored empirically in two large data sets. Data Set 1 includes responses from a creative verb association task (N = 4,468), while Data Set 2 includes ideas from a real-life creative problem-solving task (N = 662). We employed response surface analysis to creativity measures mapping onto the standard two-component definition of creativity (Data Sets 1 and 2) and a three-component extension (Data Set 2). Consistent with the standard two-component definition, we found that creativity ratings were highest when novelty and value were both of maximally high quality. Furthermore, our investigation of a three-component definition revealed positive linear effects of novelty and feasibility on creativity ratings, whereas value did not incrementally contribute to the prediction of creativity beyond novelty and feasibility. We discuss the empirical findings as well as potential future applications of response surface analysis to understand the relationship between creativity and its components.
The present study investigated the extent to which linguistic features of children's stories (analysed using automated techniques), predicted human-rated Creative Expressiveness and Logic scores (both assessed with the Consensual Assessment Technique). A sample of 160 children (Mage = 8.99 years, SD = 0.3) wrote stories based on three pictures. Eleven linguistic characteristics were measured: Length, Grammar, Originality, Controlled Lexical Diversity, Uncontrolled Lexical Diversity, Divergent Semantic Integration (DSI), Referential Cohesion, Narrativity, Syntactic Simplicity, Word Concreteness and Deep Cohesion. The results showed that 51 % of the variance in Creative Expressiveness was explained by Length, DSI, Originality, Grammar, and Controlled Lexical Diversity (sr2 = 0.01 to 0.14). In comparison, 28 % of the variance in Logic scores was accounted for by DSI, Grammar, Controlled Lexical Diversity, Syntactic Simplicity, and Narrativity (sr2 = 0.01 to 0.06). These findings offer insights for educational practices by identifying the linguistic characteristics relevant to children's creative writing as opposed to logical narration.
We describe the process undertaken by a six-person faculty committee at Washington and Lee University to develop an open-ended empirically informed student evaluation of teaching (SET) and a process to guide interpretation of SET results. Our work focused on (1) Identifying empirically based principles and resources to guide SET development; (2) Developing and pilot-testing a new SET instrument; and (3) Creating a process for faculty and department heads to summarize SET responses and use them in formative and summative assessment. Importantly, our SET instrument was created to elicit shoulds (characteristics that have been empirically associated with positive learning outcomes that students are able to validly assess) and to avoid eliciting should nots (characteristics that have not been reliably associated with positive learning or that students are not able to validly assess). Pilot testing of our SET (N = 99 student participants) evaluated the following seven areas of teaching effectiveness: setting clear expectations, creating a welcoming environment, providing encouragement and challenge, actively engaging students in learning, explaining the purpose of activities and assignments, clarifying the relevance of material beyond the classroom, and providing actionable feedback on student work. It is our hope that this summary of our process, from articulating guiding principles to bringing the SPoT (Student Perceptions of Teaching) instrument and accompanying materials before the Faculty for approval, provides guidance for other institutions seeking to create their own SET instrument and process. Our committee emphasizes the necessity of using SETs in concert with multiple additional methods of assessing teaching effectiveness within a holistic framework.
We introduce thematic profile analysis, a computational method that combines topic modeling with large language models to quantify the semantic distance between an individual's creative ideas and themes from a corpus of creative ideas. If individuals generated creative uses for a brick, a thematic profile is the collection of semantic distances between each individual's creative idea (e.g., foot scrubber) and each theme (e.g., wall-garden-path). Analyzing 15,678 ideas from 3,213 participants, we found that an individual's thematic profile strongly predicted human creativity ratings (R = .50-.79) across diverse creativity tasks. Thematic profiles highlight which themes are most likely to enhance creative performance and which themes will diminish it. Thematic profiles enable idea-level diagnostics of why a particular creative idea was rated low or high in creativity, adding value beyond receiving a single creativity score. As automated creativity scoring becomes commonplace due to large language model scoring, it will be increasingly important to develop tools, like thematic profile analysis, to allow researchers to peer into the black box and examine influential determinants of creativity scores.
Creativity is increasingly recognized as a core competency for the 21st century, making its development a priority in education, research, and industry. To effectively cultivate creativity, researchers and educators need reliable and accessible assessment tools. Recent software developments have significantly enhanced the administration and scoring of creativity measures; however, existing software often requires expertise in experiment design and computer programming, limiting its accessibility to many educators and researchers. In the current work, we introduce CAP—the Creativity Assessment Platform—a free web application for building creativity assessments, collecting data, and automatically scoring responses (cap.ist.psu.edu). CAP allows users to create custom creativity assessments in ten languages using a simple, point-and-click interface, selecting from tasks such as the Short Story Task, Drawing Task, and Scientific Creative Thinking Test. Users can automatically score task responses using machine learning models trained to match human creativity ratings—with multilingual capabilities, including the new Cross-Lingual Alternate Uses Scoring (CLAUS), a large language model achieving strong prediction of human creativity ratings in ten languages. CAP also provides a centralized dashboard to monitor data collection, score assessments, and automatically generate text for a Methods section based on the study’s tasks, metrics, and instructions—with a single click—promoting transparency and reproducibility in creativity assessment. Designed for ease of use, CAP aims to democratize creativity measurement for researchers, educators, and everyone in between.
Researchers and educators interested in creative writing need a reliable and efficient tool to score the creativity of narratives, such as short stories. Typically, human raters manually assess narrative creativity, but such subjective scoring is limited by labor costs and rater disagreement. Large language models (LLMs) have shown remarkable success on creativity tasks, yet they have not been applied to scoring narratives, including multilingual stories. In the present study, we aimed to test whether narrative originality-a component of creativity-could be automatically scored by LLMs, further evaluating whether a single LLM could predict human originality ratings across multiple languages. We trained three different LLMs to predict the originality of short stories written in 11 languages. Our first monolingual model, trained only on English stories, robustly predicted human originality ratings (r = .81). This same model-trained and tested on multilingual stories translated into English-strongly predicted originality ratings of multilingual narratives (r >= .73). Finally, a multilingual model trained on the same stories, in their original language, reliably predicted human originality scores across all languages (r >= .72). We thus demonstrate that LLMs can successfully score narrative creativity in 11 different languages, surpassing the performance of the best previous automated scoring techniques (e.g., semantic distance). This work represents the first effective, accessible, and reliable solution for the automated scoring of creativity in multilingual narratives.
This study explores what care-experienced young people want from mental health services. Six care-experienced young people were interviewed, and an interpretative phenomenological analysis applied. Three key themes emerged demonstrating that the way support is delivered, the people who deliver it, and the environment of mental health services are all important to care-experienced young people. Along with these findings, this study demonstrates that engaging vulnerable young people in research and service design is beneficial.
Poetry is one of the most creative expressions of language, but how we evaluate the creativity of a poem is not properly characterized.The present study investigated the role of various subjective qualitiesclarity, aesthetic appeal, felt valence, arousal, and surprisein predicting the creativity judgment of English poems.Participants (N=129) were presented with a broad range of English poems; they rated each poem on six characteristics: clarity, aesthetic appeal, felt valence, felt arousal, surprise and overall creativity.Linear multilevel analysis showed that aesthetic appeal was the strongest predictor of poetic creativity, followed by surprise and felt valence.Multilevel mediation analysis indicated significant mediation by surprise and felt valence on the relationship between aesthetic appeal and creativity at both within and between-participant levels.Further, expertise in English literature was found to significantly moderate the effects of all three predictors on the evaluation of creativity.The study simultaneously captured the surprise-evoking line(s).Using the semantic distance computing approach, we have shown the objective validation of the subjectively chosen line(s) of surprise.Altogether, our findings suggest a parsimonious model of evaluation of creativity of poems and its interaction with expertise.
The current engineering education places a heavy emphasis on fostering creative and innovative problem-solving skills of students. One of the crucial factors of the creative design process that is linked to students' creative outputs is the design prompt they receive. More specifically, it was found that slight variations in the design prompt, or whether the design prompt is easy to understand can significantly influence the students' performance. However, the construction of design prompts, which is a crucial component of a design project, has received relatively little attention when compared with research on other steps of the design process. Therefore, the current work seeks to contribute to this gap in research by exploring 45 design prompts at their objective and text level using linguistic characteristics to investigate what could influence students' perception of its readability and creativity. The results of this study show that at the text level, the overall objective only differed significantly in terms of word prevalence. The three objective groups studied here varied minimally in terms of their understandability and creativity perception. Finally, it was found that both the objective and the linguistic characteristics were unable to predict the percentage of people in each answer category for understandability and creativity. These results contribute towards understanding the construction of design prompts, as well as pointing out additional areas of future research in this area.
Creativity research often relies on human raters to judge the novelty of participants’ responses on open-ended tasks, such as the Alternate Uses Task (AUT). Albeit useful, manual ratings are subjective and labor intensive. To address these limitations, researchers increasingly use automatic scoring methods based on a natural language processing technique for quantifying the semantic distance between words. However, many methodological choices remain open on how to obtain semantic distance scores for ideas, which can significantly impact reliability and validity. In this project, we propose a new semantic distance-based method, maximum associative distance (MAD), for assessing response novelty in AUT. Within a response, MAD uses the semantic distance of the word that is maximally remote from the prompt word to reflect response novelty. We compare the results from MAD with other competing semantic distance-based methods, including element-wise-multiplication—a commonly used compositional model—across three published datasets including a total of 447 participants. We found MAD to be more strongly correlated with human creativity ratings than the competing methods. In addition, MAD scores reliably predict external measures such as openness to experience. We further explored how idea elaboration affects the performance of various scoring methods and found that MAD is closely aligned with human raters in processing multi-word responses. The MAD method thus improves the psychometrics of semantic distance for automatic creativity assessment, and it provides clues about what human raters find creative about ideas.
Semantic distance scoring provides an attractive alternative to other scoring approaches for responses in creative thinking tasks. In addition, evidence in support of semantic distance scoring has increased over the last few years. In one recent approach, it has been proposed to combine multiple semantic spaces to better balance the idiosyncratic influences of each space. Thereby, final semantic distance scores for each response are represented by a composite or factor score. However, semantic spaces are not necessarily equally weighted in mean scores, and the usage of factor scores requires high levels of factor determinacy (i.e., the correlation between estimates and true factor scores). Hence, in this work, we examined the weighting underlying mean scores, mean scores of standardized variables, factor loadings, weights that maximize reliability, and equally effective weights on common verbal creative thinking tasks. Both empirical and simulated factor determinacy, as well as Gilmer-Feldt's composite reliability, were mostly good to excellent (i.e., > .80) across two task types (Alternate Uses and Creative Word Association), eight samples of data, and all weighting approaches. Person-level validity findings were further highly comparable across weighting approaches. Observed nuances and challenges of different weightings and the question of using composites vs. factor scores are thoroughly provided.
Abstract: Semantic distance scoring provides an attractive alternative to other scoring approaches for responses in creative thinking tasks. In addition, evidence in support of semantic distance scoring has increased over the last few years. In one recent approach, it has been proposed to combine multiple semantic spaces to better balance the idiosyncratic influences of each space. Thereby, final semantic distance scores for each response are represented by a composite or factor score. However, semantic spaces are not necessarily equally weighted in mean scores, and the usage of factor scores requires high levels of factor determinacy (i.e., the correlation between estimates and true factor scores). Hence, in this work, we examined the weighting underlying mean scores, mean scores of standardized variables, factor loadings, weights that maximize reliability, and equally effective weights on common verbal creative thinking tasks. Both empirical and simulated factor determinacy, as well as Gilmer-Feldt’s composite reliability, were mostly good to excellent (i.e., > .80) across two task types (Alternate Uses and Creative Word Association), eight samples of data, and all weighting approaches. Person-level validity findings were further highly comparable across weighting approaches. Observed nuances and challenges of different weightings and the question of using composites vs. factor scores are thoroughly provided.
Creativity research commonly involves recruiting human raters to judge the originality of responses to divergent thinking tasks, such as the alternate uses task (AUT). These manual scoring practices have benefited the field, but they also have limitations, including labor-intensiveness and subjectivity, which can adversely impact the reliability and validity of assessments. To address these challenges, researchers are increasingly employing automatic scoring approaches, such as distributional models of semantic distance. However, semantic distance has primarily been studied in English-speaking samples, with very little research in the many other languages of the world. In a multilab study (N = 6,522 participants), we aimed to validate semantic distance on the AUT in 12 languages: Arabic, Chinese, Dutch, English, Farsi, French, German, Hebrew, Italian, Polish, Russian, and Spanish. We gathered AUT responses and human creativity ratings (N = 107,672 responses), as well as criterion measures for validation (e.g., creative achievement). We compared two deep learning-based semantic models-multilingual bidirectional encoder representations from transformers and cross-lingual language model RoBERTa-to compute semantic distance and validate this automated metric with human ratings and criterion measures. We found that the top-performing model for each language correlated positively with human creativity ratings, with correlations ranging from medium to large across languages. Regarding criterion validity, semantic distance showed small-to-moderate effect sizes (comparable to human ratings) for openness, creative behavior/achievement, and creative self-concept. We provide open access to our multilingual dataset for future algorithmic development, along with Python code to compute semantic distance in 12 languages.
Psychological consultation is a key means of informing care and practice with psychological theory and evidence. The current paper sought to investigate what elements of psychological consultation are useful for social workers when consulting on high-risk youth, due to the current gap in the literature. Seven social workers shared their experiences during one-to-one interviews. The data was analysed through thematic analysis and the emerging themes were organised into three categories: Helpful elements, such as a safe space, independent expertise, and a shared understanding; Unhelpful elements, including consultee anxiety and the unheard young person; A Mediating element in the form of feasible recommendations. The implications of these findings are discussed, as well as the limitations of this paper and recommendations for future research.
We developed a novel conceptualization of one component of creativity in narratives by integrating creativity theory and distributional semantics theory. We termed the new construct divergent semantic integration (DSI), defined as the extent to which a narrative connects divergent ideas. Across nine studies, 27 different narrative prompts, and over 3500 short narratives, we compared six models of DSI that varied in their computational architecture. The best-performing model employed Bidirectional Encoder Representations from Transformers (BERT), which generates context-dependent numerical representations of words (i.e., embeddings). BERT DSI scores demonstrated impressive predictive power, explaining up to 72% of the variance in human creativity ratings, even approaching human inter-rater reliability for some tasks. BERT DSI scores showed equivalently high predictive power for expert and nonexpert human ratings of creativity in narratives. Critically, DSI scores generalized across ethnicity and English language proficiency, including individuals identifying as Hispanic and L2 English speakers. The integration of creativity and distributional semantics theory has substantial potential to generate novel hypotheses about creativity and novel operationalizations of its underlying processes and components. To facilitate new discoveries across diverse disciplines, we provide a tutorial with code (osf.io/ath2s) on how to compute DSI and a web app ( osf.io/ath2s ) to freely retrieve DSI scores.
Semantic distance is increasingly used for automated scoring of originality on divergent thinking tasks, such as the Alternate Uses Task (AUT). Despite some psychometric support for semantic distance - including positive correlations with human creativity ratings - additional work is needed to optimize its reliability and validity, including identifying maximally reliable items (objects) for AUT administration. We identify a set of 13 AUT items based on a systematic item-selection strategy (belt, brick, broom, bucket, candle, clock, comb, knife, lamp, pencil, pillow, purse, sock). This item-set resulted in acceptable reliability estimates and was found to be moderately related to both human creativity ratings and a creative personality factor (Study 1). These results replicated in a new sample of Participants (Study 2). We conclude with the following recommendations for reliable and valid assessment of AUT originality using semantic distance: 1) make choices based on theoretical/practical considerations, 2) administer (some or all of) the 13 items from this study; 3) if other items must be used, avoid compound words as AUT items (e.g., guitar string); 4) include as many AUT items as time permits; 5) instruct participants to "be creative"; and 6) address fluency confounds that conflate idea quantity and quality (e.g., via max scoring).
Abstract The chapter authors describe Scotland as at a tipping point for change with its residential care provision. They detail recent legislative initiatives and an independent care review. Renewed emphasis on the critical importance of relationships and youth rights is also noted. The chapter concludes with the matrix used throughout this volume, which includes the current policy context, key trends and initiatives, characteristics of children and youth served, preparation of residential care personnel, promising programmatic innovations, and present strengths and challenges. Scotland has a relatively low rate of utilization of residential child and youth care among all forms of out-of-home care.
The ability to search memory for diverse contextual usages of words is proposed to play a role in creative idea generation. However, it is not yet known whether searching for more diverse contextual usages of words enhances the novelty of ideas during idea generation. In three experiments, participants were given a noun as a creativity prompt (e.g., chamber) and generated words creatively linked to the prompt (e.g., secret, relic). We directly manipulated the semantic diversity of creativity prompts to induce differential levels of semantic context search. Semantic diversity refers to the diversity of contexts in which a word is used (e.g., chamber orchestra vs. judge's chamber). Semantic context search involves searching semantic memory for different contextual usages of a prompt word. Supporting the facilitative role of semantic context search, participants given prompts high in semantic diversity generated responses higher in novelty, compared to participants given prompts low in semantic diversity (Experiments 1 and 3). Capitalizing on recent advances in distributional semantics, our novel experimental approach led to a critical new insight-semantic context search can enhance the novelty of ideas.
Narrative text permeates our lives from job applications to journalistic stories to works of fiction. Developing automated metrics that capture creativity in narrative text has potentially far reaching implications. Human ratings of creativity in narrative text are labor-intensive, subjective, and difficult to replicate. Across 27 different story prompts and over 3,500 short stories, we used distributional semantic modeling to automate the assessment of creativity in narrative texts. We tested a new metric to capture one key component of creativity in writing – a writer’s ability to connect divergent ideas. We termed this metric, word-to-word semantic diversity (w2w SemDiv). We compared six models of w2w SemDiv that varied in their computational architecture. The best performing model employed Bidirectional Encoder Representations Transformer (BERT), which generates context-dependent numerical representations of words (i.e., embeddings). The BERT w2w SemDiv scores demonstrated impressive predictive power, explaining up to 72% of the variance in human creativity ratings, even exceeding human inter-rater reliability for some tasks. In addition, w2w SemDiv scores generalized across Ethnicity and English language proficiency, including individuals identifying as Hispanic and L2 English speakers. We provide a tutorial with R code (osf.io/ath2s) on how to compute w2w SemDiv. This code is incorporated into an online web app (semdis.wlu.psu.edu) where researchers and educators can upload a data file with stories and freely retrieve w2w SemDiv scores.
While idea diversity has long been considered a critical facet of creativity, little is known about how strongly people consider it during idea generation and evaluation. A semantic distance-based assessment of idea novelty and diversity was introduced to answer this question and circumvent issues of subjectivity and labor intensiveness in human coding. In response to a given noun (e.g., ship), participants were asked to generate creative words (e.g., bottle, pirate). We manipulated participants' awareness and explicit focus on idea diversity during idea generation (Experiments 1 and 2) and idea evaluation (Experiments 3 and 4). Results indicated a substantial underutilization of idea diversity during idea generation and evaluation. Overcoming idea diversity neglect during idea generation required explicit instruction to generate ideas from divergent contexts. In contrast, idea diversity neglect was easily overcome during idea evaluation by drawing participants' attention to it before idea evaluation. We propose a new approach to assessment where groups of ideas are assessed throughout the creative process instead of solely focusing on single ideas.