In recent decades, there have been considerable advances in the area of bilingual cognition, but research on cognition across dialects remains limited. Here, we consider how the intuitive link between these experiences of linguistic diversity could drive empirical and theoretical work on the cognitive basis of dialect production. We begin with descriptive evidence to motivate contrasting hypotheses: the Difference Threshold Hypothesis, which argues for a qualitative difference between bilingual and bidialectal cognition; and the Linguistic Continuum Hypothesis, which argues that bidialectalism and bilingualism represent the same phenomenon across different degrees of linguistic overlap. To differentiate these hypotheses at the level of individual cognition, we argue that (1) research must expand beyond the lexical level to consider phonological and morphosyntactic differences across codes and (2) extra consideration is warranted for cases of bilingualism with apparent dense code-switching, which more closely approximates “bidialectal” practices.
Listeners use their knowledge about a talker to guide speech perception. In the present study, we manipulated both the familiarity and predictability of this talker-specific knowledge. Twentytwo listeners from western Pennsylvania completed an audio-visual lexical decision task while EEG was recorded. Critically, listeners were introduced to the talkers beforehand, with short videos establishing the talker's U.S. English accent identity: Mainstream, more similar to the participants themselves; Southern, a relatively less familiar variety; and Unpredictable, which switched between Mainstream and Southern accents. During the task, listeners watched these talkers produce tokens with accents that either aligned with or violated their accent identity. Behaviorally, alignment between talker and token accent influenced performance on Southern-accented tokens, while accuracy on Mainstream-accented tokens was nearly at ceiling regardless of talker accent identity. Neurally, accent familiarity had the largest effect pre-speech (in response to the visual presentation of the talker), while both predictability and familiarity impacted processing of the speech signal itself. Overall, our results suggest that listeners use information about a talker's accent, even when it is unfamiliar or unpredictable, to alter their listening strategies for more successful speech recognition.
The influence of a talker's first language (L1) on the production of their second (L2) can complicate the task of word recognition. Nevertheless, previous research has shown that listeners demonstrate remarkable speed and plasticity in adapting to L2-accented talkers. Listeners are also able to generalize this learning to novel L2-accented talkers. There are two competing explanations for these talker-independent adaptation effects: one argues that increasing variation during exposure facilitates generalization to a novel test talker (exposure-to-variability hypothesis), while the other argues that increasing acoustic-phonetic similarity between the exposure and test talkers facilitates generalization (similarity-based hypothesis). We conducted three experiments to compare the effects of exposure variability and exposure-test similarity on generalization. During exposure, monolingual English participants performed an auditory lexical decision task with multisyllabic English tokens spoken by three Spanish-accented talkers. After exposure, we tested performance on a novel Spanish-accented talker with a cross-modal matching task (Experiments 1 and 2) or a primed cross-modal lexical decision task (Experiment 3). Exposure variability increased both identity and competitor priming, consistent with a relaxation of phonetic categorization criteria. Exposure-test similarity decreased competitor priming without changing identity priming, while exposure uniformity increased identity priming without changing competitor priming, providing some evidence of phonetic recalibration in these conditions. Our results provide equivocal support for the similarity-based hypothesis and against the exposure-to-variability hypothesis.
Creativity is a key 21st-century skill and a consistent predictor of academic learning outcomes. Despite decades of research on creativity and learning, little is known about the cognitive mechanisms underlying their relationship. In two studies, we examined whether creativity supports associative learning through associative thinking—the ability to generate novel word associations—an ability central to creativity which has not been previously tied to associative learning. In Study 1, we found that students who generated more novel word associations learned more words on a foreign language learning test 24 h later. In Study 2, we replicated and extended the effect to naturalistic creativity tasks (i.e., writing short stories and sketching line drawings), finding associative thinking mediated the relationship between creativity and associative learning. Importantly, both studies controlled for general intelligence. Our findings suggest that creativity’s contribution to learning operates partly through a shared cognitive capacity for making new connections.
A growing literature uses event-related potentials (ERPs) to investigate novel metaphor processing as a window into creative processes like conceptual expansion. Modulations of the N400 generally indicate that while novel metaphors are initially processed as semantic anomalies, after a connection is found relating the concepts, they pattern more with literal sentences. Existing research largely focuses on monolinguals, but less is known about novel metaphor processing in bilinguals' second language (L2). Here, we combine robust single-trial ERPs and behavioral measures to investigate how L2 English users process full-sentence novel metaphors. We compare our results to a previous study with English monolinguals using the same experimental design to test three competing hypotheses: L2 conceptual expansion will be more effortful than, more efficient than, or similar to L1. Group differences suggest more effortful processing for L2 English users than monolinguals. Behaviorally, L2 users show more sentence evaluation errors than monolinguals, particularly for anomalous sentences. ERP results in L2 users reveal an N400 semantic anomaly effect at the sentence-final position, with no significant differences between metaphorical and literal or metaphorical and anomalous sentences. Monolinguals show a graded N400 effect, with significant differences between literal and anomalous as well as metaphorical and anomalous sentences. By comparing L2 users' results with monolingual English users and using naturalistic full-sentence structures, our findings contribute to the emergent literature on L2 novel metaphor processing and conceptual expansion while also unraveling the cognitive challenges associated with incremental processing and integration of L2 metaphorical sentences.
The rise of artificial intelligence has prompted increased scrutiny of systemic biases in automatic speech recognition technologies. One focal topic of discussion has been the degraded performance for speakers of African American and Southern U.S. English. This study aims to contribute to the research on bias in voice-AI by investigating speech recognition performance for Appalachian English, an often-stigmatized variety in American society. Participants were recruited from Southern Appalachia (Eastern Tennessee), with a non-Southern Appalachian (Central Pennsylvania) sample included as a reference group. The participants read aloud a version of the "Goldilocks and the Three Bears" fairy tale and the Rainbow Passage, and the recordings were processed using Dartmouth Linguistic Automation (DARLA). The authors conducted two sets of analyses on the vowel phonemes. The first analysis assessed DARLA's effectiveness in recognizing vowels. The system returned higher phoneme error rates for Southern Appalachian speech compared to the non-Southern dataset. Next, they conducted a detailed error analysis on the misrecognized input-output phoneme pairs. The results suggested dialect bias in the system, with 50.2% of the errors in the Southern dataset attributed to participation in the Southern Vowel Shift. These findings underscore the importance of integrating sociolectal variation into the acoustic model to mitigate dialect bias for currently underserved users.
The present study examined how Arabic-Hebrew-English trilinguals process double and triple cognate words in their third language (L3) across three different experiments. Utilizing the same set of critical cognate items, trilinguals completed a semantic relatedness task, a lexical decision task, or a sentence reading eye-tracking task. The results revealed a significant cognate facilitation effect in the semantic relatedness task, with no consistent differences in the magnitude of facilitation across double and triple cognates, suggesting that both L1 and L2 are activated during L3 processing. In contrast, no cognate facilitation was observed in the lexical decision or the sentence reading tasks. These results demonstrate that the cognate facilitation is task-dependent, varying with the degree to which meaning is activated, sentential context is available, and orthographic cues are involved. Critically, the study extends findings of phonologically mediated cross-language activation from bilinguals to trilinguals. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
This study examined the integration of face cue and native- and nonnative-accented English speech by manipulating the face cue predictability of a speaker's accent (predictable: one accent (American-accent or Chinese-accented), not predictable: two accents (American-accented and Chinese-accented). Monolingual listeners were first familiarized with each speaker's number and type of accent(s). Then, they completed an EEG-recorded auditory go/no-go animal decision task where the no-go items were critical words and nonwords. Listeners saw a face cue (speaker image) before speech onset and concurrently with the speech. Pre-speech ERP results revealed that listeners processed face cues differently based on face cue predictability. Post-speech ERP analyses revealed N400 lexicality effects for native-accented speech, and face cue predictability effects for nonnative-accented speech. No N400 effects were found for an audio-only experiment. This indicates that the monolingual listeners integrate face cues and auditory cues during real-time nonnative-accented speech processing, but not during native-accented speech processing.
Research shows that nonnative accents differing from a listener's own can impede comprehension, as described by the Interlanguage Speech Intelligibility Benefit (ISIB). While extensively studied in nonnative contexts, native regional varieties have been less frequently studied, with mixed findings. This study examined native listeners' real-time sentence processing of geographically distant Spanish varieties. Mexican Spanish speakers listened to accents that matched (Mexican) or mismatched (Peninsular Spain, Puerto Rico) their own, along with nonnative English-accented Spanish. Behavioral results showed high comprehension across all varieties. ERP findings revealed semantic violation N400 effects for the Mexican and familiar Peninsular Spain but not for the less-familiar Puerto Rican accent. An N400 and late negativity appeared for nonnative English-accented Spanish. Results indicate that less-familiar native language varieties challenge, while familiar accents facilitate, lexico-semantic access during real-time sentence processing. Findings support a generalized intra-language processing benefit for regional varieties beyond matched speech, further refining the ISIB hypothesis.
Creativity is increasingly recognized as a core competency for the 21st century, making its development a priority in education, research, and industry. To effectively cultivate creativity, researchers and educators need reliable and accessible assessment tools. Recent software developments have significantly enhanced the administration and scoring of creativity measures; however, existing software often requires expertise in experiment design and computer programming, limiting its accessibility to many educators and researchers. In the current work, we introduce CAP—the Creativity Assessment Platform—a free web application for building creativity assessments, collecting data, and automatically scoring responses (cap.ist.psu.edu). CAP allows users to create custom creativity assessments in ten languages using a simple, point-and-click interface, selecting from tasks such as the Short Story Task, Drawing Task, and Scientific Creative Thinking Test. Users can automatically score task responses using machine learning models trained to match human creativity ratings—with multilingual capabilities, including the new Cross-Lingual Alternate Uses Scoring (CLAUS), a large language model achieving strong prediction of human creativity ratings in ten languages. CAP also provides a centralized dashboard to monitor data collection, score assessments, and automatically generate text for a Methods section based on the study’s tasks, metrics, and instructions—with a single click—promoting transparency and reproducibility in creativity assessment. Designed for ease of use, CAP aims to democratize creativity measurement for researchers, educators, and everyone in between.
Researchers and educators interested in creative writing need a reliable and efficient tool to score the creativity of narratives, such as short stories. Typically, human raters manually assess narrative creativity, but such subjective scoring is limited by labor costs and rater disagreement. Large language models (LLMs) have shown remarkable success on creativity tasks, yet they have not been applied to scoring narratives, including multilingual stories. In the present study, we aimed to test whether narrative originality-a component of creativity-could be automatically scored by LLMs, further evaluating whether a single LLM could predict human originality ratings across multiple languages. We trained three different LLMs to predict the originality of short stories written in 11 languages. Our first monolingual model, trained only on English stories, robustly predicted human originality ratings (r = .81). This same model-trained and tested on multilingual stories translated into English-strongly predicted originality ratings of multilingual narratives (r >= .73). Finally, a multilingual model trained on the same stories, in their original language, reliably predicted human originality scores across all languages (r >= .72). We thus demonstrate that LLMs can successfully score narrative creativity in 11 different languages, surpassing the performance of the best previous automated scoring techniques (e.g., semantic distance). This work represents the first effective, accessible, and reliable solution for the automated scoring of creativity in multilingual narratives.
Creative thinking is a vital skill for engineers. Prior work suggests that social dynamics-such as critical feedback from a high-authority figure-can influence the ideation process. Yet little is known about the neurocognitive mechanisms through which feedback shapes creative thinking. In this study, engineering students completed a creative ideation task while EEG was recorded. Midway through the experiment, a professor gave the participant either supportive or unsupportive critical feedback on their performance. Supportive feedback was expected to positively influence creativity compared to unsupportive feedback: participants in the supportive feedback condition were predicted to show greater idea originality and fluency after receiving feedback, as well as a greater EEG power increase in the alpha frequency band (8-12 Hz) that is robustly associated with creativity. We found that after receiving feedback-whether supportive or unsupportive-participants produced fewer but more highly original responses and showed increased alpha power. These results indicate that feedback can cause engineers to generate fewer but more original ideas by driving alpha-band activity in the brain. In further analyses, we found decreased beta-band activity before feedback only in the unsupportive condition, possibly reflecting increased cognitive stress and internally directed attention required to adjust performance in the post-feedback phase.
AbstractOver the past decades, bilingualism researchers have come to a consensus around a fairly strong view of nonselectivity in bilingual speakers, often citing Van Hell and Dijkstra (2002) as a critical piece of support for this position. Given the study’s continuing relevance to bilingualism and its strong test of the influence of a bilingual’s second language on their first language, we conducted an approximate replication of the lexical decision experiments in the original study (Experiments 2 and 3) using the same tasks and—to the extent possible—the same stimuli. Unlike the original study, our replication was conducted online with Dutch–English bilinguals (rather than in a lab with Dutch–English–French trilinguals). Despite these differences, results overall closely replicated the pattern of cognate facilitation effects observed in the original study. We discuss the replication of outcomes and possible interpretations of subtle differences in outcomes and make recommendations for future extensions of this line of research.
We examined the neural correlates underlying the semantic processing of native- and nonnative-accented sentences, presented in quiet or embedded in multi-talker noise. Implementing a semantic violation paradigm, 36 English monolingual young adults listened to American-accented (native) and Chinese-accented (nonnative) English sentences with or without semantic anomalies, presented in quiet or embedded in multi-talker noise, while EEG was recorded. After hearing each sentence, participants verbally repeated the sentence, which was coded and scored as an offline comprehension accuracy measure. In line with earlier behavioral studies, the negative impact of background noise on sentence repetition accuracy was higher for nonnative-accented than for native-accented sentences. At the neural level, the N400 effect for semantic anomaly was larger for native-accented than for nonnative-accented sentences, and was also larger for sentences presented in quiet than in noise, indicating impaired lexical-semantic access when listening to nonnative-accented speech or sentences embedded in noise. No semantic N400 effect was observed for nonnative-accented sentences presented in noise. Furthermore, the frequency of neural oscillations in the alpha frequency band (an index of online cognitive listening effort) was higher when listening to sentences in noise versus in quiet, but no difference was observed across the accent conditions. Semantic anomalies presented in background noise also elicited higher theta activity, whereas processing nonnative-accented anomalies was associated with decreased theta activity. Taken together, we found that listening to nonnative accents or background noise is associated with processing challenges during online semantic access, leading to decreased comprehension accuracy. However, the underlying cognitive mechanism (e.g., associated listening efforts) might manifest differently across accented speech processing and speech in noise processing.
Over the past decades, bilingualism researchers have come to a consensus around a fairly strong view of nonselectivity in bilingual speakers, often citing Van Hell and Dijkstra (2002) as a critical piece of support for this position. Given the study's continuing relevance to bilingualism and its strong test of the influence of a bilingual's second language on their first language, we conducted an approximate replication of the lexical decision experiments in the original study (Experiments 2 and 3) using the same tasks and-to the extent possible-the same stimuli. Unlike the original study, our replication was conducted online with Dutch-English bilinguals (rather than in a lab with Dutch-English-French trilinguals). Despite these differences, results overall closely replicated the pattern of cognate facilitation effects observed in the original study. We discuss the replication of outcomes and possible interpretations of subtle differences in outcomes and make recommendations for future extensions of this line of research.
The past decades have seen an explosion of research using electrophysiological or neuroimaging techniques for studying the neurocognitive underpinnings of second language (L2) processing. Although this field has a shorter history than does research on language learning more generally, important insights into the neurocognitive basis of L2 processing have driven it to the center stage of language science. In this target article for Language Learning's 75th Jubilee volume, I illustrate the field's impressive achievements by selectively reviewing electrophysiological and neuroimaging research on L2 processing and bilingual brain organization. I also review changing perspectives in the field (including individual difference and experience-based perspectives, neural network approaches, neuroplasticity, and L2-learning related neural changes) and identified challenges, promises, and future directions (revisit native-speaker benchmark, increase linguistic diversity, enhance ecological validity, intensify research on child L2 learners' brain, adopt lifelong approach to L2 learning) that can lead to a better understanding of the neural underpinnings of L2 learning and processing.
Behavioral studies have established that cross-dialectal communication is typically harder than within-dialect communication: listeners make more mistakes and are slower to respond to less familiar accents. In this study, we use event-related potential (ERP) analysis of electroencephalography (EEG) to capture the neurocognitive correlates of these patterns. Speakers of Mainstream US English from Western Pennsylvania (N = 23) participated in an auditory go-no-go task where they heard both Southern- and Mainstream-accented US English real and nonsense words (and were asked to respond to real, animal words). We see effects of accent for real words on the P200 and N400, reflecting more effortful processing in both the acoustic-phonetic and lexical-semantic stages for Southern versus Mainstream accents. We see no differences in nonsense word processing, and together, these findings suggest that the difficulty in normalizing Southern-accented tokens at the acoustic-phonetic level disrupted lexico-semantic access later on. We are currently running the same study with Southern-accented listeners in Southwest Virginia, to see whether we see inverse results or whether listeners with substantial exposure to both dialects (which we expect to be true of our Southern listeners) show different response profiles.
Creativity research commonly involves recruiting human raters to judge the originality of responses to divergent thinking tasks, such as the alternate uses task (AUT). These manual scoring practices have benefited the field, but they also have limitations, including labor-intensiveness and subjectivity, which can adversely impact the reliability and validity of assessments. To address these challenges, researchers are increasingly employing automatic scoring approaches, such as distributional models of semantic distance. However, semantic distance has primarily been studied in English-speaking samples, with very little research in the many other languages of the world. In a multilab study (N = 6,522 participants), we aimed to validate semantic distance on the AUT in 12 languages: Arabic, Chinese, Dutch, English, Farsi, French, German, Hebrew, Italian, Polish, Russian, and Spanish. We gathered AUT responses and human creativity ratings (N = 107,672 responses), as well as criterion measures for validation (e.g., creative achievement). We compared two deep learning-based semantic models-multilingual bidirectional encoder representations from transformers and cross-lingual language model RoBERTa-to compute semantic distance and validate this automated metric with human ratings and criterion measures. We found that the top-performing model for each language correlated positively with human creativity ratings, with correlations ranging from medium to large across languages. Regarding criterion validity, semantic distance showed small-to-moderate effect sizes (comparable to human ratings) for openness, creative behavior/achievement, and creative self-concept. We provide open access to our multilingual dataset for future algorithmic development, along with Python code to compute semantic distance in 12 languages.
We developed a novel conceptualization of one component of creativity in narratives by integrating creativity theory and distributional semantics theory. We termed the new construct divergent semantic integration (DSI), defined as the extent to which a narrative connects divergent ideas. Across nine studies, 27 different narrative prompts, and over 3500 short narratives, we compared six models of DSI that varied in their computational architecture. The best-performing model employed Bidirectional Encoder Representations from Transformers (BERT), which generates context-dependent numerical representations of words (i.e., embeddings). BERT DSI scores demonstrated impressive predictive power, explaining up to 72% of the variance in human creativity ratings, even approaching human inter-rater reliability for some tasks. BERT DSI scores showed equivalently high predictive power for expert and nonexpert human ratings of creativity in narratives. Critically, DSI scores generalized across ethnicity and English language proficiency, including individuals identifying as Hispanic and L2 English speakers. The integration of creativity and distributional semantics theory has substantial potential to generate novel hypotheses about creativity and novel operationalizations of its underlying processes and components. To facilitate new discoveries across diverse disciplines, we provide a tutorial with code (osf.io/ath2s) on how to compute DSI and a web app ( osf.io/ath2s ) to freely retrieve DSI scores.
We examined the impact of images on novel word learning and consolidation, in a conceptual replication of Liu and Van Hell (2020). After participants had learned one set of novel words with definitions and images on Day 1 (remote words) and a different set on Day 2 (recent words), they judged the semantic relatedness of word pairs on Days 2 and 8 while event-related potentials (ERPs) were recorded. Day 2 ERPs showed that remote, but not recent, novel words elicited a late positive component. By Day 8, both remote and recent novel words elicited a late positive component. We observed no N400 on either day. Comparing these learners (definition-image group) with learners trained with definitions only (using data from Liu & Van Hell, 2020) revealed that the groups’ ERP patterns did not differ, but definition recall and relatedness judgment performances were higher for the definition-image group than for the definition-only group. Learning novel word meanings through definitions and images strengthened behavioral outcomes but did not affect ERP signatures of learning and consolidation.