
Collaborative problem-solving (CPS) is an essential 21st-century competency, yet its valid and ecologically sound assessment remains challenging due to inconsistent frameworks, poorly aligned task design, and over-reliance on behavior frequency counting in scoring. This study developed and provided empirical support for a theory-driven game-based CPS assessment embedded in a real-time strategy game, using a customized Collaborative Siege Task with asymmetric roles and ill-structured problems, to elicit more realistic collaborative interactions. Based on a validated three-facet CPS framework from Chinese cultural context, the task systematically triggers social and cognitive processes of collaboration across five progressive stages. By sampling from the eastern provinces of China, a pilot study with 36 participants confirmed task feasibility and indicator effectiveness, while a formal study with 298 valid participants examined psychometric properties using item response theory and confirmatory factor analysis. Findings supported the hierarchical higher-order factor structure, satisfactory item difficulty, discrimination, reliability, and criterion-related validity against peer assessment, reasoning ability, and personality. The assessment showed no sex, major, or role bias. This study provides a reliable gamified assessment tool for CPS measurement with demonstrated potential for scalability pending further development of automated scoring procedures, offering implications for large-scale evaluation and competency-oriented instruction.
Large language model (LLM)-based artificial intelligence is increasingly used in ethically consequential human decision-making, yet fully autonomous machine ethics remains unrealistic, motivating architectures that support rather than replace human ethical judgment. This study introduces the Cognitive–Reflective Equilibration Architecture (CREA), a cognitively grounded artificial moral advisor that operationalizes the Cognitive–Reflective Equilibration Model (CREM), in which reflective reasoning guides ethical judgment from intuitive cognition toward a more advanced equilibrium among competing values, drawing on Piaget and Rawls. CREA implements CREM’s 20-step process through four stage-aligned reasoning agents—Cognitive, Reflective, Equilibration, and Evaluation—coordinated via multi-LLM orchestration, in which auxiliary models independently explore principles, generate counterarguments, and score supporting and opposing considerations to externalize reflective deliberation. The architecture was empirically evaluated by comparing four configurations—single-agent, multi-agent, multi-LLM, and multi-LLM with knowledge- and reasoning-bank augmentation—across four indicators of advice quality using 500 matched execution units per configuration. All comparisons are system-internal: advice quality was scored by CREA’s own multi-LLM measurement pipeline rather than by human ethicists, so the findings reflect relative differences among architectures under LLM-based self-evaluation, not normative validity. Within that scope, distributing reflective reasoning across multiple models was associated with higher reason-giving (justifiability) and normative-alignment scores relative to simpler configurations. CREA therefore offers an empirically characterized, auditable advisor architecture whose potential to scaffold human ethical judgment remains a hypothesis for user-centered validation rather than a demonstrated outcome.
Recent studies suggest forms of unconscious processing in classical insight problem solving. In particular, the unconscious analytic thought (UAT) approach speculates that the creative act of restructuring such very demanding problems implies a form of high-level unconscious thought, mainly active during incubation. The present study aims to investigate whether pupillary dilation, typically associated with conscious cognitive effort, can also represent an indicator of unconscious cognitive effort (UAT). In the experimental condition, participants attempted to solve the problem and, during the incubation phase, performed a series of single-digit arithmetic calculations. In the control condition, participants only performed the arithmetic calculations. Results showed a significant interaction between group (solvers, non-solvers and control) and temporal course on pupillary dilation. In particular, solvers, conversely to the other two groups, showed high and constant pupillary dilation during all the incubation phase, suggesting the maintenance of constant cognitive effort at an unconscious layer not attributable to anything other than the processing of the insight problem solution. In the other two groups, instead, the pupil size—starting at the same level as solvers—progressively decreased during the execution of the arithmetic task. Our results are discussed in the light of the debate on dual process theories.
In an era of information explosion and overload, individuals face a constant influx of mixed-quality, contradictory data. Navigating this complex landscape requires critical thinking, which enables individuals to dynamically adjust their reasoning and refine judgments. Consequently, accurately measuring this dynamic ability is essential. However, existing assessments predominantly focus on single-round argumentation and fail to capture the iterative process, potentially yielding unrepresentative evaluations of real-world performance. To address this gap, this study develops the Iterative Argumentation Task (IAT), a novel assessment tool that presents participants with a sequence of interrelated passages simulating real-world information flow. By capturing the process of evaluation, position changes, and argument revisions, the IAT evaluates three ability dimensions: single-round argument analysis, inter-argument relationship analysis, and argument integration. We recruited 355 undergraduates to validate the IAT. Confirmatory factor analysis provided evidence for the proposed measurement structure. Furthermore, correlations with established critical thinking and reasoning tests demonstrated convergent and discriminant validity, respectively. Moreover, performance differences between the undergraduates and 33 debaters established known-groups validity. Finally, latent profile analysis identified four distinct profiles across the IAT dimensions. Together, these findings establish the IAT as a theoretically sound and practically effective tool for assessing critical thinking.
Academic procrastination is a widespread self-regulatory failure among university students and a costly one for both academic achievement and psychological well-being. In this study, procrastination was examined within the framework of non-cognitive individual differences. The model tested here posits that perfectionism and locus of control affect procrastination through negative automatic thoughts and general self-efficacy. These variable-centered analyses were then complemented with a person-centered approach in order to capture distinct risk profiles. Cross-sectional data were collected from 1003 Turkish university students (71.5% women; Mage = 21.93 years) and analyzed with structural equation modeling, latent profile analysis (LPA), and random forest regression. The structural model accounted for 62% of variance in procrastination (CFI = 0.958, RMSEA = 0.074). Based on these results, general self-efficacy was identified as a common pathway through which distal non-cognitive influences converge on procrastination. The direct effect of automatic thoughts was non-significant, whereas the indirect effect through self-efficacy was strong (indirect effect = 0.51, 95% CI [0.44, 0.58]). LPA identified three profiles: adaptive, average, and vulnerable; the vulnerable profile showed markedly higher procrastination than the adaptive profile (d = 1.21). Random forest analysis indicated no meaningful nonlinearity, with self-efficacy as the dominant predictor. These findings suggest that interventions may benefit from prioritizing general self-efficacy beliefs.
Background: Developmental topographical disorientation (DTD)-like navigational difficulties can impair everyday orientation and spatial information processing. Objective: We examined whether 8-week virtual reality (VR) orienteering training improves performance on a laboratory spatial memory task in college students with operationally defined DTD-like navigational difficulties and examined task-related cortical changes. Methods: In a randomized controlled trial, 96 students were assigned to VR orienteering, VR exercise, or control groups (n = 32/group). The exercise groups trained twice weekly for 45 min for 8 weeks at moderate intensity (64–76% HRmax). Before and after the intervention, all participants completed a delayed visuospatial recognition task designed to assess spatial-memory-related processing, while functional near-infrared spectroscopy measured hemodynamic changes in prefrontal and sensorimotor cortices. The post-test session was conducted 48 h after the final scheduled session, with no immediate post-exercise assessment. Results: VR orienteering produced greater improvement in spatial-memory-task accuracy and reaction time than the other groups. After BH-FDR correction across eight ROIs, training-specific reductions in task-related activation were observed in the right primary motor cortex and left frontopolar area. Within the VR orienteering group, pre-to-post change in left orbitofrontal activation was negatively associated with change in spatial-memory-task accuracy, Pearson r(30) = −0.819, 95% CI [−0.908, −0.658], p < 0.001, R2 = 0.670; the association remained significant after eight-ROI BH-FDR correction (adjusted p < 0.001). Conclusions: VR orienteering improved performance on the laboratory spatial memory task and was accompanied by reduced task-related cortical recruitment. Greater reductions in L-OFC activation were linked to greater improvements in accuracy, identifying L-OFC activation change as a neural correlate of training-related spatial-memory improvement.
In computerized assessments, response accuracy and response time together reflect how examinees engage with test items. Existing joint models typically assume a linear relationship between these two outcomes, yet empirical observations often reveal more complex patterns. This study proposes a Quadratic Bifactor Hierarchical Model (QBi-HM) that captures nonlinear speed–accuracy dependencies through a shared general factor with quadratic terms, while preserving separate ability and speed components. Simulation results across sample sizes (N = 1000, 1500, 3000) and test lengths (m = 30, 60) demonstrated that QBi-HM recovered parameters accurately under both linear and nonlinear conditions, whereas the conventional linear bifactor model produced biased estimates when nonlinearity was present. An empirical application to PISA 2012 computer-based mathematics data (N = 1527 examinees from four economies, m = 10 items) showed that QBi-HM achieved superior model fit (AIC = 41,825.190, BIC = 42,251.675, SABIC = 41,997.535) compared with Bi-HM (AIC = 41,969.750, BIC = 42,289.613, SABIC = 42,099.099) and revealed an inverted U-shaped speed-accuracy pattern. These findings may suggest that modeling nonlinear dependencies can improve the precision of ability estimation in large-scale assessments and inform item design by identifying task features that elicit distinct response processes. The QBi-HM framework offers a flexible tool for researchers and practitioners seeking to leverage response time data to better understand test-taking behavior.
Research exploring the relationship between emotional intelligence and workplace outcomes is predominantly variable-centered and treats the dimensions of emotional–social competencies as separate from one another, thereby losing sight of the intra-individual interplay of these competencies and the potential heterogeneity of their effects. This study examines teachers’ emotional–social intelligence using a person-centered approach. Using a sample of 587 Hungarian teachers from 26 public education institutions, we conducted latent profile analysis based on Bar-On’s five composites and then used unbiased three-step procedures (BCH; categorical BCH) to examine the antecedents of profile membership, as well as perceptions of school organizational climate and dominant conflict management style. Four distinct profiles emerged, primarily differing at the global level—a uniformly low profile (12%), a medium-low profile (25%), a dominant, balanced medium-high profile (53%), and a uniformly high profile (9%). Participants with higher profiles perceived the school’s organizational climate as more supportive (η2 = 0.07) and exhibited more adaptive conflict management patterns (less avoidance, more problem-solving; χ2(12) = 37.24, p < .001). A compositional (log-ratio) reanalysis of the complete conflict-mode profiles, reported as the primary conflict analysis, showed that the relative weight of problem-solving rises across profiles and plateaus between the modal and the highest profile, that avoidance declines with profile elevation, and that an apparent surplus of accommodating in the highest profile was an artifact of dominant-style coding rather than an indication of a broader repertoire. At the same time, profile membership was essentially independent of job position, career stage, and gender, and the relationships with outcomes were of modest to moderate strength. The results extend person-centered, profile-based intelligence research to the emotional–social domain and identify promising directions for research on teacher professional development.
Digital game-based learning is increasingly used to support children’s engagement, interaction, and social–emotional development. However, its contribution to emotional intelligence, particularly in inclusive classroom contexts, remains theoretically underdeveloped and methodologically fragmented. This systematic review and critical narrative synthesis examines how teacher- or facilitator-mediated digital games address children’s emotional intelligence and related social–emotional outcomes in inclusive educational contexts. Following PRISMA 2020 reporting guidance, the review examined peer-reviewed articles indexed in Web of Science and Scopus between 2015 and 2026. After screening 347 unique records and assessing 31 full-text reports, 14 studies were included in the final synthesis. The review focused on four interrelated areas: the emotional and social–emotional outcomes addressed in digital game-based studies; the roles of teachers or facilitators in structuring, mediating, and debriefing game-based activities; the inclusive classroom conditions that support equitable participation and emotional safety; and the risks, limitations, or implementation challenges associated with classroom use. The findings suggest that digital games may become educationally meaningful when they are linked to explicit emotional learning goals, embedded in structured classroom or school-based activities, supported by teacher or facilitator mediation, and followed by reflection or debriefing. Based on the synthesis, the review proposes a teacher-mediated inclusive digital game-based emotional learning framework consisting of seven dimensions: emotional learning target, game selection and pedagogical fit, teacher mediation, inclusive participation, classroom interaction, reflective debriefing, and risk monitoring.
Foreign language reading comprehension is a critical indicator of university students’ academic cognitive ability and global competence, yet effectively assessing their latent learning potential remains a major challenge. Therefore, from a neurocognitive perspective, this study combined functional near-infrared spectroscopy (fNIRS) with dynamic assessment (DA) to investigate the activation characteristics of the left frontal and temporal regions in non-English majors during reading tasks across two difficulty levels (CET-4/CET-6) and two assessment conditions (static/dynamic testing). The results showed that reading tasks activated the dorsolateral prefrontal cortex and middle temporal gyrus. Compared with CET-4, the more difficult CET-6 task showed a broader descriptive cortical activation pattern. Notably, mediated prompts were associated with different behavioral and activation-frequency patterns across task difficulty levels: in the CET-4 task, dynamic assessment showed higher normalized activation-frequency values than static assessment in areas corresponding to the middle and inferior frontal gyrus, accompanied by rising prompt scores; however, the CET-6 task exhibited the opposite pattern. These findings suggest that mediated prompts may be more effective in stimulating learning potential when task difficulty aligns with the learner’s zone of proximal development. Overall, by considering behavioral dynamic-assessment indicators in foreign-language reading together with fNIRS-based cortical activation patterns, this exploratory study offers preliminary insights into learning potential, personalized instructional interventions, and academic cognitive assessment.
Creativity has a dark side—malevolent creativity (MC), which is the application of original ideas to purposely harm others; i.e., creative problem-solving with malicious intent. Here, the influence of past MC behavior on MC ideation was examined while controlling for other individual-level factors relevant to the AMORAL model of MC. College undergraduates (N = 276 initially recruited; N = 248 final sample analyzed; 82% female) were surveyed on demographics, self-perception of problem-solving potential (SP-PSP), dark triad traits, and past MC behavior, and completed an adapted Malevolent Creativity Task (MCT), generating revenge ideas to an unfair scenario. Sequential regression models were used to predict MCT outcomes (fluency index, malevolence index, and composite score), incrementally controlling for age, sex, SP-PSP, dark triad traits, past MC behavior, and provocation level. In the final model, SP-PSP significantly predicted the MCT fluency index (R2 = 0.065). Past MC behavior significantly increased the R2 of the MCT malevolence index; the final model was significant (R2 = 0.084). Past MC behavior and narcissism significantly predicted the MCT composite score (R2 = 0.091). These findings align with and extend previous research linking past MC behavior to MC ideation by highlighting unique and incremental impacts of SP-PSP, narcissism, past MC behavior, and provocation on MC ideation.
Higher education is undergoing a rapid transformation driven by advancements in artificial intelligence (AI), creating new opportunities to support the development of higher-order cognitive skills through carefully designed learning environments. This study aimed to examine the effectiveness of an educational program that integrates AI-powered intelligent assistants into a structured e-learning environment to support the development of digital entrepreneurship and future-oriented problem-solving skills among graduate students. A quasi-experimental design was used with 82 graduate students, randomly assigned to two groups: an experimental group (n = 41) and a control group (n = 41). Both groups studied the same educational content and completed identical learning activities; however, the experimental group participated in a structured learning framework that included AI-powered intelligent assistants, guided learning activities, and instructor facilitation, while the control group completed the same activities without AI support. The educational program comprised five modules focusing on chatbot design, intelligent platform development, digital content production, educational data analysis, and future-oriented problem-solving. Learning outcomes were assessed using a Digital Entrepreneurship Product Rating Scale and a Future-Oriented Problem-Solving Scale, which measures skills in visualization, prediction, foresight, and planning. The results revealed statistically significant differences favoring the experimental group on both outcome scales, with large effect sizes. These findings suggest that integrating AI-powered intelligent assistants within a structured learning framework may support the development of digital entrepreneurship and future-oriented problem-solving skills among graduate students. This study contributes empirical evidence to the potential educational value of integrating AI-powered intelligent assistants with pedagogically designed learning activities to foster higher-level cognitive development in higher education.
With the widespread use of digital technologies, it has become increasingly important to understand how children adapt to digital environments. However, limited research has examined digital maturity among children beyond digital access and familiarity with technologies, particularly in non-Western contexts. Building on recent work on digital maturity, the present study aimed to validate the Digital Maturity Inventory (DIMI) among Chinese primary school students and to identify their latent profiles. A sample of 434 students (221 girls; Meanage = 10.68, SD = 0.86) completed the DIMI for validation analyses, and an additional sample was simultaneously recruited, resulting in a total sample of 1026 students (500 girls; Meanage = 10.38, SD = 1.09), for latent profile analysis. The findings demonstrated that the DIMI exhibited satisfactory psychometric properties in the Chinese context. Moreover, digital maturity, in addition to digital native, emerged as a unique and significant predictor of artificial intelligence (AI) literacy. Three distinct subgroups of students with heterogeneous patterns of digital maturity were also revealed: an unbalanced low group, an inconsistent medium group, and a relatively high group. These findings not only support the cross-cultural applicability of DIMI, but also demonstrate its relevance to children’s AI-related competencies. The implications of this study were also discussed.
Guaranteeing high-quality, equitable, and inclusive basic education requires upholding the right of all students to reach their full potential and ensuring their well-being. Students with High Intellectual Ability (HIA) represent a vulnerable group that is often under-identified in schools, a fact that hinders the personalization of their education and the ability to meet their specific needs. In 2019, as part of its plan to address the educational needs of this group, the Department of Education of the Basque Government prioritized the design and validation of a screening protocol to support the identification of students with HIA. This article presents a study on the validity and reliability of the High Intellectual Ability Screening Protocol (SACI, by its Spanish acronym), applicable in Primary Education. A two-phase controlled screening design was employed, involving the random assignment of participants and the use of specific statistical techniques to evaluate the psychometric soundness of the measure. The results provide highly promising evidence, supporting the conclusion that the SACI protocol is a valid and reliable screening instrument for use with Primary Education students in the Basque Autonomous Community.
The present study addressed a timely and important issue in cognitive and media psychology about the relationship between media multitasking and cognition. This study examined attention control, working memory, convergent associative thinking, and fluid reasoning, examining multiple functions within the same analytical framework and moving beyond the narrow focus on single cognitive domains that characterizes much prior research. The main aim was to deepen the understanding of the associations of media multitasking with these core domains of human cognition. A total of 214 healthy young adults (18–30 years) completed the Short Media Multitasking Questionnaire, Raven’s Progressive Matrices, the Digit Span Test, the Trail-Making Test, and the Remote Associates Task. Multivariate general linear models and subsequent regression analyses were performed while controlling for age and gender. The results did not reveal a significant overall association between media multitasking and the combined cognitive outcomes. At the univariate level, higher media multitasking scores were associated with lower performance on the Remote Associates Task, although the observed effect size was small. Given the cross-sectional design and reliance on self-reported media multitasking, longitudinal studies employing objective measures of media multitasking use are needed to clarify the robustness and directionality of these associations.
Service-learning interventions have been proposed as a context for examining trait emotional intelligence, a core socioemotional competence associated with cognitive and behavioral outcomes. This pilot quasi-experimental study examined whether participation in a PLACE model (Prepare, Link, Action, Celebrate, Effect) service-learning program was associated with differences in trait emotional intelligence; civic engagement attitudes and behaviors; academic motivation and perseverance; learning mindset; school belonging; dropout intentions; and grade point average among secondary school students. Participants were assigned to an intervention group (n = 30) or a control group (n = 27). Data were collected using validated instruments, including the Trait Emotional Intelligence Questionnaire, the Civic Engagement Scale, and the Students’ Engagement and Motivation Scale 3.0, along with a single-item measure of dropout intention and official grade point average records. Analyses of covariance, controlling for baseline scores, revealed no statistically significant group differences. Effect sizes showed a mixed pattern, with small effects favoring both groups depending on the outcome. These patterns should be interpreted cautiously, as the study was underpowered to detect small effects. Overall, the findings do not support intervention efficacy but may inform future research with larger samples.
This study examined whether a single total-score cutoff in the Mawhiba Multiple Cognitive Aptitude Test (MMCAT) may conceal meaningful cognitive profile differences between students across four domains: mental flexibility, verbal reasoning and reading comprehension, mathematical–spatial reasoning, and scientific–mechanical reasoning. Archival score data from 5146 school-nominated students across three assessment levels in Al-Ahsa, Saudi Arabia, were analyzed. Latent profile models were estimated separately for each assessment level. A three-profile solution was retained for Levels 1 and 2, and a four-profile solution for Level 3. The retained profiles differed mainly in the overall performance level rather than the profile shape. The operational cutoff of 1540 captured globally high-performing profiles at considerably higher rates than adjacent middle-range profiles, and the identified groups remained internally heterogeneous across all levels. The findings indicate that the current identification rule effectively captures broad, high-level performance, yet it does not represent all adjacent cognitive profiles equally. This has implications for identification policy and educational programming for gifted students.
One of the defining characteristics of the Aha! experience is the profound sense of certainty with which it arrives—you don’t just know, you know you know. A large body of research suggests this confidence is generally warranted; ideas that arrive with a feeling of insight are more likely to be accurate than those that do not. However, the insight–accuracy relationship has been examined almost exclusively in convergent thinking tasks, leaving the diagnostic value of insights in divergent contexts uncertain. The present research tests whether this relationship differs between problem-solving contexts. Study 1 re-analyses daily diary data from professional writers and physicists, used as proxies for divergent and convergent domains. Writers initially rated their Aha! ideas as highly creative but downgraded them at 6-month follow-up; physicists’ Aha! ideas, by contrast, retained or gained value over time. Study 2 used a within-subjects design with laboratory measures of convergent and divergent thinking. Insight intensity predicted accuracy on both tasks but predicted subjective creativity ratings far more strongly than objective ones on the divergent task, producing a systematic bias we term the Aha! glow. Implications are discussed, including a proposed adaptive, motivational function for this phenomenon.
Understanding how environmental contexts influence cognitive functioning is a key issue in cognitive science and applied ergonomics. This study examined whether everyday environmental settings modulate task performance and perceived cognitive workload during a card-matching memory task, using high-fidelity immersive Virtual Reality (VR) environments. Participants performed the task by sequentially uncovering cards and remembering their spatial locations to identify matching pairs in two virtual scenarios: a naturalistic environment (urban park) and an artificial setting (office). The number of moves and time of execution were collected as behavioral measures. Perceived cognitive workload was assessed after task completion in each condition using the Simulation Task Load Index (SIM-TLX). Behavioral performances did not differ between environments. However, participants reported lower perceived cognitive workload when performing the task in the naturalistic environment. This effect was not uniform across individuals: the reduction in perceived cognitive workload was primarily observed among participants exhibiting high levels of task efficiency. These findings suggest that environmental context interacts with individual cognitive efficiency by modulating the perceived cognitive workload without necessarily affecting objective performance outcomes. Overall, the results suggest that naturalistic immersive virtual contexts may mitigate perceived cognitive workload during memory-based tasks, highlighting the role of contextual factors in influencing cognitive efforts under ecologically valid conditions.
Metacognitive monitoring, students’ evaluations of their own performance, is central to self-regulated learning but often misaligns with actual performance. While situational factors have been studied extensively, less is known about how learning-related beliefs and motivational dispositions—hereafter termed learning-related attributes—relate to individual differences in metacognitive monitoring. This study examined how these learning attributes, alongside cognitive and demographic variables, were associated with students’ general confidence and resolution. Data were drawn from 3946 Chinese sixth-grade students who completed academic assessments in the verbal and mathematical domains, provided post-test performance estimates, and completed questionnaires measuring their learning attributes and demographic characteristics. We conducted multilevel analyses with a random-intercept–random-slope specification and identified several learning attributes as significant predictors of metacognitive monitoring, primarily in terms of general confidence. Adding learning-related attributes after the cognitive and demographic blocks increased marginal R2, primarily for general confidence; this order-dependent increment was not interpreted as evidence that the block is inherently more important. The findings extend existing accounts of metacognitive monitoring by suggesting that students’ general confidence is associated with non-cognitive learning attributes—and may thus be partly separable from cognitive ability—whereas their resolution is comparatively less strongly associated with these attributes. Theoretical as well as practical implications are discussed.