Modern large language models (LLMs) exhibit remarkable reasoning abilities, yet it remains unclear whether gains in reasoning ability are accompanied by corresponding improvements in metacognitive sensitivity, which refers to the ability to discriminate, on an item-by-item basis, between correct and incorrect inferences. Here we systematically evaluate this relationship in five experiments across 17 models spanning five LLM series. Although reasoning ability reliably scales with model size, metacognitive sensitivity does not; in some series, the smallest models even exhibit higher metacognitive sensitivity than their larger counterparts. This deviation from scaling law arises from the intermediate reasoning steps produced by large models: long and deterministic reasoning traces boost performance but hinder accurate self-evaluation. Furthermore, distilling the reasoning traces of large models into small models transfers the impairment in metacognitive sensitivity. These findings demonstrate that metacognitive sensitivity is not an automatic by-product of improved reasoning.
Background Despite demonstrating comparable early aptitude, girls often demonstrate lower long-term motivation in mathematics, manifesting an enduring gender paradox in STEM pathways. While prior research frequently attributes this gender divergence to differences in perceived competence, the role of domain-specific metacognitive bias—how accurately students monitor their own performance—remains largely unexplored. Integrating a metacognitive perspective with Expectancy-Value Theory, this study investigated whether early metacognitive bias initiate or/and sustain divergent motivational trajectories. We tracked 4,073 students in China from primary through early middle school, utilizing longitudinal growth modeling to examine the longitudinal interplay between metacognitive bias and a critical math motivation --math intrinsic value (interest) across genders. Results Analyses revealed a general downward trend in math interest for all students, but girls experienced a significantly steeper decline compared to boys. Crucially, the structural pathways initiating this decline were moderated by gender. For girls, early overconfidence provided a temporary boost to initial interest but ultimately predicted a significantly steeper subsequent motivational decline. Conversely, boys’ math interest remained largely independent of early metacognitive bias. Furthermore, the downstream pathway linking the math interest trajectory to later metacognitive bias was positively significant and gender-invariant, leading to sustainment of gender gap in motivational and metacognitive profile. Conclusions These findings suggest that gender disparities in STEM motivation may be better explained by domain-specific metacognitive biases than by global competence beliefs. For girls, early overconfidence presents a developmental trade-off that initiates a steeper motivational decline, whereas boys’ math interest remains anchored in actual achievement. Educational interventions should therefore move beyond merely enhancing self-evaluation accuracy and focus on decoupling intrinsic motivation from competence-related calibration to build long-term motivational resilience.
Testing studied information not only facilitates long-term consolidation of tested content but also enhances subsequent acquisition of new knowledge. These dual aspects of test-enhanced learning have predominantly been examined using verbal materials. In the present study, we extended this investigation to low-level visual features, in particular visual salience (luminance and color saturation), to determine whether interim testing can retrospectively consolidate and prospectively boost learning of continuous visual information. Across two experiments, participants studied five lists of scene images with varying levels of visual salience. Participants in the interim test group were tested on each list, while those in the interim restudy group restudied the prior list after studying Lists 1-4 and were tested only on List 5. A cumulative final test was administered after either a brief delay of 30 s (Experiment 1) or an extended delay of 2 days (preregistered Experiment 2). Both experiments showed that interim testing prospectively enhanced the precision, absolute accuracy, and relative accuracy of visual memory in the List 5 interim test without altering memory bias. Similar retrospective effects of interim testing were observed on the cumulative test. These findings demonstrate the generalizability of test-enhanced learning to the precision and accuracy of continuous visual memory for subtle features and provide support for reset-of-encoding and motivation-based accounts of test-enhanced learning while challenging semantic explanations. In addition, design-related insights into the replicability of memory fading are provided, and practical implications are discussed. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Does overconfidence really confer adaptive benefits to children's learning? Through a tripartite investigation involving a preregistered replication (Study 1; N = 30, children aged 6-8 years), computational simulation (Study 2), and an experimental intervention (Study 3; N = 64, children aged 6-8 years), we first replicated previous findings that highly overconfident (HO) children exhibited less negative performance change across a memory task than their low-overconfidence (LO) counterparts. However, this pattern was driven by participant-selection bias and regression-to-the-mean effects rather than by adaptive benefits of childhood overconfidence. When experimentally manipulating children's overconfidence levels to eliminate these methodological drawbacks, the difference in performance changes between HO and LO children disappeared. These findings challenge an influential hypothesis about the adaptive nature of childhood overconfidence, underscore the risks of median-split designs with difference scores, highlight the necessity of causal experimental approaches in developmental research, and raise concerns about educational practices promoting positive illusions in children.
Previous studies on judgments of learning (JOLs) have primarily focused on how individuals integrate multiple cues to predict subsequent memory performance through experience-based and belief-based processes. However, few studies have examined the relative contributions of global prior beliefs about one's overall memory ability versus item-level processing experiences during JOL formation, possibly due to the absence of formal computational models that can characterize how these two sources of information are integrated. To address this issue, the Bayesian inference model for metamemory (BIM) introduced the parameter Pexp to quantify the reliance on item-level processing experience relative to global prior beliefs in forming JOLs. This study aimed to investigate whether the contribution of item-level processing experience versus global prior beliefs to JOLs (Pexp) is related to JOL resolution, defined as the accuracy in distinguishing between correctly and incorrectly remembered items. Experiments 1 and 2 employed longitudinal designs in which participants completed memory tasks twice, separated by either a one-month or a one-week interval. The results demonstrated individual consistency in Pexp over time, as well as a bidirectional relationship between Pexp and JOL resolution. Experiment 3 and a re-analysis of a previous meta-analytic dataset further revealed that the association between Pexp and JOL resolution persisted even in the presence of metamemory illusion, and was observed across datasets involving different material types and task formats. These findings highlight the complex interplay between JOL resolution and the process underlying JOL formation, and providing new insights into the cognitive mechanisms of metamemory.
Metacognitive monitoring, students’ evaluations of their own performance, is central to self-regulated learning but often misaligns with actual performance. While situational factors have been studied extensively, less is known about how learning-related beliefs and motivational dispositions—hereafter termed learning-related attributes—relate to individual differences in metacognitive monitoring. This study examined how these learning attributes, alongside cognitive and demographic variables, were associated with students’ general confidence and resolution. Data were drawn from 3946 Chinese sixth-grade students who completed academic assessments in the verbal and mathematical domains, provided post-test performance estimates, and completed questionnaires measuring their learning attributes and demographic characteristics. We conducted multilevel analyses with a random-intercept–random-slope specification and identified several learning attributes as significant predictors of metacognitive monitoring, primarily in terms of general confidence. Adding learning-related attributes after the cognitive and demographic blocks increased marginal R2, primarily for general confidence; this order-dependent increment was not interpreted as evidence that the block is inherently more important. The findings extend existing accounts of metacognitive monitoring by suggesting that students’ general confidence is associated with non-cognitive learning attributes—and may thus be partly separable from cognitive ability—whereas their resolution is comparatively less strongly associated with these attributes. Theoretical as well as practical implications are discussed.
Repeatedly asking individuals to answer the same question can increase their confidence in response accuracy, even when it does not improve accuracy per se. Fiechter and Kornell (2021) proposed two theoretical explanations to account for this repeated questioning effect on confidence, namely a fluency hypothesis and a response-repetition hypothesis. Their results supported the response-repetition hypothesis by showing a mediation effect of response repetition in the relationship between repeated questioning and confidence. Meanwhile, their results did not support the fluency hypothesis by showing no mediation effect of trial-completion times. However, fluency can be defined in multiple ways. To further test the fluency hypothesis, the current study used the same experimental procedure as in Fiechter and Kornell (2021), but employed first-key latency as a measure of retrieval fluency. The results showed a reliable mediation effect of first-key latency. Consistent with Fiechter and Kornell (2021), there was also a reliable mediation effect of response repetition. The documented findings support the fluency and response-repetition hypotheses to jointly account for the repeated questioning effect.
JOLs are widely used to measure metacognitive monitoring, yet their elicitation can reactively enhance memory—a phenomenon known as the positive reactivity effect. The enhanced engagement theory posits that JOLs improve memory by increasing attentional and cognitive engagement during encoding, but direct experimental evidence remains scarce. Across three experiments, we directly manipulated key components of learning engagement—attentional focus (via silent vs. aloud production), cognitive effort (via massed vs. spaced repetition), and motivational involvement (via standard vs. time-saving instructions)—while assessing their impact on the JOL reactivity effect in word recognition memory. Results consistently demonstrated robust positive reactivity effects, critically, the magnitude of these effects was significantly attenuated under high-engagement conditions (aloud reading, spaced learning, and heightened motivation). These converging findings provide the first direct, multi-method experimental support for the enhanced engagement theory, specifying that making JOLs benefit memory most when baseline engagement is low. The results delineate boundary conditions under which making JOLs yield beneficial effects and provide practical insights into leveraging JOLs to regulate engagement in real-world learning environments.
Taking photos is a ubiquitous everyday activity, woven into social life from classrooms and meetings to travel and social gatherings. Photography is known to impair memory for photographed experiences, a phenomenon termed the photo-taking impairment effect (PTIE), but prior work has focused solely on the photographer. The present research introduces and tests a social extension of this phenomenon: One person's photo-taking activity can impair a nearby companion's memory for the same event, which we term the companion-PTIE. Across four experiments spanning controlled laboratory tasks and a simulated art exhibition, we show that photo-taking reduces a nonphotographing companion's subsequent memory to a degree comparable to taking photos oneself. Mechanistically, we dissociate cognitive offloading from attentional disengagement. Instructing photographers to delete each photo immediately, thereby removing any expectation of later access, left both the self- and companion-PTIE unchanged, arguing against cognitive offloading as a primary driver. By contrast, preventing a nonphotographing companion from seeing the act of photo-taking with an opaque barrier abolished the companion-PTIE, implicating attentional disengagement as the causal pathway. Robust Bayesian meta-analyses estimated a reliable companion-PTIE with no meaningful difference in magnitude from the self-PTIE. These findings reframe the PTIE as a social phenomenon: The memory costs of ubiquitous photography extend beyond photographers to copresent others, revealing how everyday technologies reshape collective memory for shared experiences. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Previous research has extensively examined how various factors (e.g., font size, relatedness) influence Judgments of Learning (JOLs), yet comparatively little is known about whether phonological factors, such as rhyme, similarly influence metacognitive judgments. Across four experiments, the present study investigated how rhyme affects memory performance, JOLs, and metacognitive accuracy, and explored the underlying mechanisms involved. Experiments 1 and 2 revealed that rhyme yielded reliable mnemonic benefits only when study involved reading aloud, emphasizing the role of auditory encoding in enhancing memory. Moreover, participants consistently expected higher recall for rhyming than non-rhyming pairs. Experiment 3 confirmed that participants explicitly believed rhyme enhances memorability. Experiment 4 further showed that these beliefs partially mediated rhyme effects on JOLs and that learners updated these beliefs based on study experience. Importantly, despite context-dependent mnemonic benefits, rhyme reduced relative metacognitive accuracy, diminishing item-level discrimination between remembered and forgotten materials. Together, these findings suggest that phonological cues such as rhyme robustly shape metacognitive judgments while sometimes compromising relative accuracy. Practically, instructors and learners should use rhyme judiciously and combine it with more diagnostic cues to support both learning and monitoring accuracy.
An emerging body of studies has observed that soliciting judgments of learning (JOLs) reactively changes recall or recognition of specific study items, a phenomenon known as the reactivity effect of JOLs on memory. The current studies explored whether soliciting JOLs reactively affects continuous color memory and whether it affects memory accessibility or memory precision (or a combination of both). Experiment 1 employed a classical continuous color memory task in which participants studied animal images in different colors and then reconstructed the colors on a continuous matching wheel, and found that making JOLs reactively enhanced accessibility but impaired precision of memory. Experiment 2 replicated these dissociated reactivity effects and further found that articulatory suppression successfully eliminated these effects, suggesting that making JOLs reactively alters color memory through prompting individuals to favor the verbal-labeling strategy for encoding and retrieving visual information. The findings support the strategy-change theory to account for the JOL reactivity effect.
Metacognition, particularly the ability to monitor and regulate cognitive processes, plays a crucial role in effective learning [...]
Over the past few decades, Swahili-English and Lithuanian-English word pair databases have been extensively utilized in research on learning and memory. However, these normative databases are specifically designed for generating study stimuli in learning and memory research involving native (or fluent) English speakers. Consequently, they are not suitable for investigations that encompass populations whose first language is not English, such as Chinese individuals. Notably, native Chinese speakers constitute a substantial proportion, approximately 18%, of the global population. The current study aims to establish a new database of translation equivalences, specifically tailored to facilitate research on learning, memory, and metacognition among the Chinese population. We present a comprehensive set of normative measures for 200 Swahili-Chinese paired associates, including recall accuracy, recall latency, error patterns, confidence ratings, perceived learning difficulty, judgments of learning, and perceived learning interestingness for the entire word pairs. Additionally, we include word-likeness ratings and word length for the Swahili words, and concreteness ratings, familiarity ratings, word frequency, and number of strokes for the Chinese words. This diverse array of measures, gathered across a substantial number of Swahili-Chinese word pairs, is poised to effectively support future research seeking to investigate the intricate processes of learning, memory and metacognition within the Chinese population.
Effective learning involves not only the ability to quickly acquire knowledge and skills, but also the capacity to accurately monitor one's ongoing learning progress. The present research probed the relation between learning ability and monitoring accuracy. A meta-analysis (Study 1, N = 2,406) counterintuitively found that individuals with superior learning ability exhibited slightly poorer monitoring accuracy (measured as the resolution of judgments of learning). Study 2 reanalyzed the meta-analysis data and observed that expert learners remembered more items they erroneously believed they would not remember, and this underconfidence in expert learners led to a negative association between learning ability and monitoring accuracy. Studies 3 (N = 102, adults aged 18-23) and 4 (N = 481, adults aged 18-59) conceptually replicated the findings of Studies 1 and 2 in controlled experiments. These findings challenge the conventional wisdom that good learners are also good monitors, suggesting instead that expert learners are actually the ones with monitoring deficits.
An emerging body of studies has demonstrated that asking participants to make concurrent judgments of learning (JOLs) during learning can reactively change (typically enhance) their memory performance, a phenomenon known as the reactivity effect. The current study conducted the first exploration of individual differences in the JOL reactivity effect by employing a large-scale (N = 284 participants) approach. The reactivity effect was measured in a related word pair learning task, and each of four higher-order cognitive constructs, including working memory capacity (WMC), attentional control (AC), episodic memory (EM), and general fluid intelligence (gF), was assessed by multiple tasks. The results showed that making JOLs enhanced cued recall of related word pairs, reflecting an overall positive reactivity effect. WMC independently and positively predicted JOL reactivity and this prediction effect survived when controlling for the prediction effects of other cognitive constructs. After controlling for the effects of WMC, EM, and gF, AC negatively predicted JOL reactivity. Neither EM nor gF predicted reactivity. These findings lend support to the learning engagement and dual-task costs theories to jointly account for the JOL reactivity effect. Practical implications for guiding learning practices and for mitigating JOL reactivity in future metacognition research are discussed.
Retrieval practice is well-established as a powerful tool for reinforcing long-term learning. Most previous research has concentrated on the effectiveness of overt retrieval, involving recalling information from memory and generating overt responses by writing, typing, or speaking aloud the retrieved information. Here we ask whether covert retrieval, involving mentally retrieving information without producing overt responses, can enhance learning and consolidate long-term memory, and whether it does so as effectively as overt retrieval. The current meta-analysis integrated data from 2560 participants across 18 studies to investigate the magnitude, boundary conditions, and underlying mechanisms of the covert retrieval effect and the relative efficacy of overt and covert retrieval. The results showed that covert retrieval enhances learning to a small but significant extent (g = 0.23), and its effectiveness is moderated by several factors including provision of corrective feedback, control strategy, and retention interval. The results support the additional exposure and desirable difficulty theories to jointly account for the covert retrieval effect. The meta-analysis also found that overt retrieval is more effective than covert retrieval (g = 0.17), with the effect size of this additional benefit being moderated by the mode by which covert retrieval is performed. The results support the truncated search and desirable difficulty explanations of the relative benefit of overt compared to covert retrieval. Overall, the documented findings provide practical implications for optimizing learning and teaching practices and highlight several important directions for future research.
Recent studies established that engaging metacognitive monitoring via making judgments of learning (JOLs) can directly enhance young adults' recognition memory, a phenomenon termed the reactivity effect of JOLs. The present study explored the reactive influence of making JOLs on older adults' recognition memory and probed the potential age-related differences in this effect. In three experiments, participants were instructed to study four lists of words, with two lists studied with concurrent JOLs and the other two without, followed by a recognition test. The results provided strong evidence that making JOLs improves older adults' recognition performance (Experiments 1-3) through enhancing both recollection- and familiarity-based recognition (Experiment 3). But the positive reactivity effect on recognition memory for older adults was weaker than that for young adults (Experiments 2 and 3). To elucidate potential mechanisms underlying age-related differences in the reactivity effect, the present study also measured participants' learning engagement and cognitive abilities. The model results substantiated the mediating role of learning engagement, supporting the enhanced learning engagement theory, rather than the dual-task hypothesis, as an account for the reactivity effect on recognition memory. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
With the global aging of the population, the importance of understanding the characteristics and mechanisms of developmental changes in later life has grown. The present study explored age-related differences in the effect of emotion on judgments of learning (JOLs) in Chinese participants and delved deeper into the mechanisms underlying this effect. Experiment 1 observed that older participants showed a positivity effect on JOLs, whereas young participants demonstrated an emotional salience effect on JOLs, reflecting age-related differences in the effect of emotion on JOLs. To investigate the mechanisms underlying these age-related differences, Experiment 2 measured participants' metamemory beliefs about the effect of emotion on memory and found that older participants held a belief of the positivity effect, whereas young participants possessed a belief of the emotional salience effect. Experiment 3 collected data of beliefs and JOLs from the same participants and provided further evidence highlighting the contribution of metamemory beliefs to age-related differences in the effect of emotion on JOLs. These findings are essential for advancing the theoretical framework of metamemory and for extending lifespan theory of socioemotional selectivity. (PsycInfo Database Record (c) 2025 APA, all rights reserved).