The possibility of LLM self-awareness and even sentience is gaining increasing public attention and has major safety and policy implications, but the science of measuring them is still in a nascent state. Here we introduce a novel methodology for quantitatively evaluating metacognitive abilities in LLMs. Taking inspiration from research on metacognition in nonhuman animals, our approach eschews model self-reports and instead tests to what degree models can strategically deploy knowledge of internal states. Using two experimental paradigms, we demonstrate that frontier LLMs introduced since early 2024 show increasingly strong evidence of certain metacognitive abilities, specifically the ability to assess and utilize their own confidence in their ability to answer factual and reasoning questions correctly and the ability to anticipate what answers they would give and utilize that information appropriately. We buttress these behavioral findings with an analysis of the token probabilities returned by the models, which suggests the presence of an upstream internal signal that could provide the basis for metacognition. We further find that these abilities 1) are limited in resolution, 2) emerge in context-dependent manners, and 3) seem to be qualitatively different from those of humans. We also report intriguing differences across models of similar capabilities, suggesting that LLM post-training may have a role in developing metacognitive abilities.
It has been reported that LLMs can recognize their own writing. As this has potential implications for AI safety, yet is relatively understudied, we investigate the phenomenon, seeking to establish: whether it robustly occurs at the behavioral level, how the observed behavior is achieved, and whether it can be controlled. First, we find that the Llama3-8b–Instruct chat model - but not the base Llama3-8b model - can reliably distinguish its own outputs from those of humans, and present evidence that the chat model is likely using its experience with its own outputs, acquired during post-training, to succeed at the writing recognition task. Second, we identify a vector in the residual stream of the model that is differentially activated when the model makes a correct self-written-text recognition judgment, show that the vector activates in response to information relevant to self-authorship, present evidence that the vector is related to the concept of ``self'' in the model, and demonstrate that the vector is causally related to the model’s ability to perceive and assert self-authorship. Finally, we show that the vector can be used to control both the model’s behavior and its perception, steering the model to claim or disclaim authorship by applying the vector to the model’s output as it generates it, and steering the model to believe or disbelieve it wrote arbitrary texts by applying the vector to them as the model reads them.
Grouping items facilitates working memory performance by reducing memory load. Recent studies reported that within-object binding of multiple features from the same dimension may facilitate encoding and maintenance of visual working memory for item-specific information, such as color or size. Further, perceptual inter-item configuration strongly influences working memory performance for item-specific spatial information. Here, we test whether forming a configuration between two items enhances short-term memory for the relationship between features of those items. To do so, we asked subjects to perform a change detection task based on the relationship between two items while the strength of the inter-item configuration was manipulated. Critically, an absolute feature value for each item varied between sample and test periods. This design ensured that subjects extracted and maintained relation information between items independent of item-specific information. We found that when items were located side-by-side, making the formation of an inter-item configuration easy, subjects’ performances for remembering 1, 2, 3, and 4 relations (2, 4, 6 and 8 items present, respectively) were not different from remembering specific information for 1, 2, 3 and 4 items. When items within a pair were placed offset, thus making it more difficult to form an inter-item configuration, memory performance for relations was significantly worse than for items. Our findings suggest that encoding and maintenance of relation information are strongly influenced by perceptual factors; the presence of a strong inter-item configuration facilitates processing of relation working memory within those pairs of items. Meeting abstract presented at VSS 2012
Item-specific spatial information is essential for interacting with objects and for binding multiple features of an object together. Spatial relational information is necessary for implicit tasks such as recognizing objects or scenes from different views but also for explicit reasoning about space such as planning a route with a map and for other distinctively human traits such as tool construction. To better understand how the brain supports these two different kinds of information, we used functional MRI to directly contrast the neural encoding and maintenance of spatial relations with that for item locations in equivalent visual scenes. We found a double dissociation between the two: whereas item-specific processing implicates a frontoparietal attention network, including the superior frontal sulcus and intraparietal sulcus, relational processing preferentially recruits a cognitive control network, particularly lateral prefrontal cortex (PFC) and inferior parietal lobule. Moreover, pattern classification revealed that the actual meaning of the relation can be decoded within these same regions, most clearly in rostrolateral PFC, supporting a hierarchical, representational account of prefrontal organization.
Past research has probed the working memory representations of visual object features, but less is known about how visual relational information is stored in working memory. To investigate this, we employed a delayed recognition behavioral paradigm using visual object features that also afford relational comparisons between objects: specifically, the magnitude dimensions of size and luminance. In a series of experiments, we examined whether the working memory capacity for relations is similar to the capacity for objects and their features, whether memory for relational information is decomposed into magnitude and direction components, and whether relative magnitudes (eg, “X much bigger than”) are encoded similarly to absolute magnitudes (eg, “X big”). Results for object features reproduce earlier findings for nonscalar visual dimensions. Accuracy was equally high when subjects had to encode either size or luminance as when they had to encode both size and luminance, indicating that multiple features of an item can be remembered as well as a single feature of an item. Relational magnitude, however behaved differently. When subjects had to encode both the size differential and the luminance differential of two pairs of objects, accuracy was significantly lower than when they had to encode only the size or luminance differential of two pairs of objects. This suggests that visual comparative relations are not maintained in separate feature stores, nor are they automatically bound into an integrated “object-like” multidimensional relational representation. Rather, size and luminance relational representations compete with each other for a limited shared memory resource, and the capacity of this resource is similar to that found for objects. These behavioral results are consistent with fMRI results from our lab comparing the neural representations of item-specific and relational information along the dimensions of size and luminance.