Targeting the debate about the foundational nature of serial-context markers in verbal working memory, Attout et al. (2026a) recently concluded that temporal codes dominate spatial codes for serial order across a fronto-parietal network. Based on methodological and empirical arguments, in the current commentary we argue that this conclusion is not warranted from their findings. Overall, our aim is to caution against overstating the relevance of their findings to the debate at stake, and to inspire further theoretical and empirical work building on their study.
Recent advances in generative and embodied AI have been driven by large-scale predictive learning over multimodal data. However, the resulting systems remain largely based on passive training regimes where linguistic regularities create the scaffold onto which information from other modalities is attached. Conversely, neuroscience and cognitive science suggest that biological intelligence is organized in the opposite way, where grounded world models acquired through interaction with the environment provide the semantic scaffold to which language is attached. Here, we illustrate five examples of neural circuits supporting grounded world modelling, which underlie navigation in physical and conceptual spaces, affordance-based perception and interaction with objects, active perception and exploratory learning, allostatic control and emotion, and the distinction between self- and world-generated outcomes. These examples highlight several features largely missing from current embodied AI, including the role of intrinsic dynamics as a foundation for learning, the centrality of action in aligning these dynamics with the external world, the prominence of autonomous experience and open-ended learning over passive assimilation of externally provided data, and the fact that early predictive and control mechanisms scaffold higher cognitive abilities such as reasoning, conceptual navigation, planning, imagination, understanding others' minds, and communication. Finally, we discuss whether and how principles derived from biological systems may inform future embodied AI, including training regimes based on social interaction to construct world models that are not only grounded but also socially shared and aligned with human norms and values.
Humans conceptualize time in terms of space, allowing flexible time construals from various perspectives. We can travel internally through a timeline to remember the past and imagine the future (i.e., mental time travel) or watch from an external standpoint to have a panoramic view of history (i.e., mental time watching). However, the neural mechanisms that support these flexible temporal construals remain unclear. To investigate this, we asked participants to learn a fictional religious ritual of 15 events. During fMRI scanning, they were guided to consider the event series from either an internal or external perspective in different tasks. Behavioral results confirmed the success of our manipulation, showing the expected symbolic distance effect in the internal-perspective task and the reverse effect in the external-perspective task. We found that the activation level in the posterior parietal cortex correlated positively with sequential distance in the external-perspective task but negatively in the internal-perspective task. In contrast, the activation level in the anterior hippocampus positively correlated with sequential distance regardless of the observer’s perspectives. These results suggest that the hippocampus stores the memory of the event sequences allocentrically in a perspective-agnostic manner. Conversely, the posterior parietal cortex retrieves event sequences egocentrically from the optimal perspective for the current task context. Such complementary allocentric and egocentric representations support both the stability of memory storage and the flexibility of time construals.
Cognitive maps provide a powerful way to compress experience into low-dimensional structure that supports flexible learning, inference, and generalization across spatial and conceptual domains. Yet relational structure must ultimately be converted into action and decision-making. Here we advance a unifying account of how hippocampal–entorhinal cognitive maps become action-relevant, translating core principles from spatial navigation to nonspatial declarative knowledge. We first synthesize evidence that conceptual navigation engages complementary allocentric and egocentric reference frames. We then operationalize egocentric coding in conceptual spaces as an attentional “point of view,” and show how mental actions can be implemented as shifts of attention over relational structures. Finally, we outline how reference-frame transformations and state-updating mechanisms—well characterized in spatial navigation—can generalize to conceptual spaces, yielding representations that support both overt behavior and internal mental operations.
The hippocampal-entorhinal system represents relations between states in spatial and nonspatial cognitive maps. Critical to understanding how these memory representations are used for cognition is to determine whether the actions underlying state transitions are incorporated in entorhinal cognitive maps. Participants learned to transition between states using different actions, operationalized as mathematical operations. We found that the entorhinal cortex represented the afforded actions across the states. This action representation was not explained by other properties of the task space, such as link distance between the states or reaction times. Furthermore, gaze behavior reflected the direction of afforded actions in the horizontal axis, and the strength of this lateralization predicted both performance and entorhinal pattern similarities, suggesting a link between gaze behavior and neurocognitive mechanisms for navigating conceptual spaces. In sum, this study provides first evidence for the integration of action information into ocular and entorhinal representations of conceptual spaces, suggesting that these may not just map out experiences, but provide information about how to explore knowledge.
The limits of unconscious processing in the semantic domain are highly debated. While prior research presents polarized views on whether word representations are accessible when presented subliminally, this work proposes a more fine-grained investigation into which aspects of word meaning can be accessed unconsciously. Specifically, we explore the conditions under which high-level semantic information, such as metaphorical relations, can be processed subliminally. We rely on space-time and space-number metaphorical mappings, which have been observed in many experiments as well as in spontaneous behavior, such as the association of the past (or subtraction) and the future (or addition) with "left" and "right", respectively. We exploited the fact that some of these conceptual associations (i.e., sagittal or vertical mappings) are also present in language (e.g., "you have a bright future in front of you"; "taxes are going down"), whereas others (i.e., lateral mapping) are not (e.g., "you have a bright future on your right"; "taxes are going left"). In two experiments, space-time and space-number semantic priming consistent with canonical metaphorical mappings emerged when both prime (e.g., "right") and target (e.g., "tomorrow") were consciously perceived, confirming their conceptual association. However, with masked priming, only language-encoded associations were activated. These results suggest that consciousness is necessary to process even ubiquitous and overlearned metaphorical associations and that putatively unconscious semantic priming, when present, may be lexical in nature.
How early visual cortex (EVC) supports language and semantic processing in sighted and congenitally blind individuals remains debated. Some predict that blindness induces radical functional reorganization of EVC, whereas others suggest more modest scaling-up changes of a neurofunctional architecture present in sighted individuals. We recorded whole-head MEG in 19 early blind (EB) and 21 sighted adults (SC) listening to spoken adjectives belonging to different semantic categories (i.e., abstract, concrete visual, concrete non-visual) and performing lexical and semantic decision tasks. In the lexical task, lexical information (i.e., word vs pseudoword) could be reliably decoded in perisylvian language areas and the EVC of both groups, with group differences (EB > SC) localized to occipital areas between 0.4 and 0.8 s after word onset. During the semantic task, semantic category information could be decoded in a network overlapping with the canonical semantic system and encompassing the EVC of both groups, with group differences (EB > SC) localized to occipital areas between 0.9 and 1 s after stimulus onset. We further characterized EVC properties and showed that i) EVC semantic information decoding is task sensitive, with no above-chance decoding in EB or SC in the lexical task ii) in the blind EVC it was possible to decode abstract from concrete concepts but not visual from non-visual concepts, a profile similar to that of posterior cingulate areas, but different to anterior cingulate cortex which showed above chance classification for all semantic categories. These results suggest that deprived EVC is integrated into a distributed lexical-semantic network carrying behaviourally relevant semantic information, although to a different extent.
Grid cells in human entorhinal cortex encode spatial layouts for real-world navigation, yet their role in conceptual navigation remains unclear. Here we show that mentally transforming tones within a purely auditory pitch-duration space engages spatial circuits, and that such resources are causally necessary. In Experiment 1, participants trained for five days to navigate through a purely conceptual pitch-duration auditory space, then underwent fMRI on Day 6. We observed a six-fold modulation of entorhinal BOLD signals aligned to each participant's trajectory angles, similar to grid-cell firing in physical space. Stronger grid-like coding predicted larger training-related gains. In Experiment 2, a new cohort performed the same task under either a spatial or non-spatial interference load. Only the spatial condition selectively disrupted performance on trials requiring mental "movement," indicating a causal reliance on spatial resources. These findings provide evidence that auditory conceptual transformations recruit--and depend on--spatial grid-like computations in the entorhinal-hippocampal system, pointing to a domain-general role for spatial coding in organizing new knowledge along continuous dimensions. ### Competing Interest Statement The authors have declared no competing interest.
Structural representations in the entorhinal cortex are a crucial feature of cognitive maps. However, a complete understanding of task structure often requires knowledge of the actions available at each state (analogous to the moves of different pieces on a chessboard). Moreover, action representations are key to several models of the hippocampal-entorhinal system. In this study, we tested whether the entorhinal cortex represents the actions that allow transitioning between states of a conceptual space. In particular, participants learned to transition between numerically-labelled states using different mathematical operations. Behaviourally, participants learned and generalised action information across states. Action information was also reflected in eye movements, indicating the active processing of the possible actions for a given state during the trials. Importantly, we found that neural pattern similarity in the right entorhinal cortex scaled with the similarity of action affordances across the states. This affordance representation was not explained by other properties of the task space such as link distance between the states or action magnitude. In sum, this study provides first evidence for the integration of action information into ocular and entorhinal representations of conceptual spaces, suggesting that these may not just store experiences, but provide information about how to explore knowledge. ### Competing Interest Statement The authors have declared no competing interest.
Theories of Embodied and grounded cognition posit that knowledge retrieval is rooted in sensorimotor simulations of past experiences. Accordingly, individuals with diverse sensorimotor experiences may retrieve knowledge differently. Here, we asked whether and how congenital blind individuals remap the representation of vision-related verbs in the motor system. Participants memorized lists of phrases combining an object to an action-related ("to take a guitar"), vision-related ("to see a guitar"), or control verb ("to hear a guitar"). The lists were either learned with the hands at rest or behind their back. Results replicated previous findings showing that recall for action-related phrases was lower in both groups when they were learned with the hands behind the back. As expected, posture impacted the memory of vision-related phrases only in blind people, although in the opposite direction. These findings provide evidence for the sensorimotor grounding of knowledge and shed light on how blind individuals represent knowledge.
In the physical world, we efficiently organize objects in space to facilitate finding them when mostly needed. Does this tendency to order the external environment extend to conceptual information that is internally stored in memory? Here we report that when people were asked to retrieve different concepts from memory, they spontaneously directed their gaze to distinct and minimally overlapping portions of their visual space. This effect was replicated across two experiments, and was correlated with retrieval performance. Interestingly, this categorical partitioning, which differentiated the content of the retrieved information, was paired with a temporal partitioning, which reflected the temporal progression of the retrieval experience. These findings reveal a tendency of humans to separately organize their memories during retrieval, disentangling content (the “what”) and time (the “when”) via a spontaneous oculomotor behavior.
Serial order working memory, the ability to retain sequences of information, is crucial for language, action planning, and recall, yet its representational basis is debated. A prominent account explains serial order through spatial coding, whereby positions in a sequence are aligned along an internal spatial axis and accessed via shifts of spatial attention. So far spatial coding may not fully explain how order is maintained. A complementary possibility is that the brain forms a representation of order, which makes abstraction of the original spatial frame and is generalizable across contexts. Whether these two mechanisms coexist, and how they are instantiated in the brain, remains unclear. Here, we used fMRI to dissociate spatial and abstract representations of serial order. Participants memorized sequences presented left-to-right or right-to-left, reversing the mapping between ordinal position and space. Behavioral results showed that item retrieval was influenced by the spatial condition, with distinguishable spatial-order associations across encoding directions. MVPA revealed position-specific decoding in the left intraparietal sulcus (IPS) for left-to-right sequences and in the right IPS for right-to-left sequences, without generalization across directions, showing the spatial nature of these representations. By contrast, direction-invariant ordinal coding emerged in the supramarginal gyrus and precentral gyrus. In addition, activity in medial parietal cortex and hippocampus scaled with ordinal distance in the sequence. These results support a dual mechanism account: serial order is supported by a spatial code tied to experience, and by a more abstract ordinal code that extends beyond it to support serial order processes.
How does the brain access stored knowledge? It has been proposed that conceptual search engages neurocognitive processes similar to foraging in physical space. We tested this idea using intracranial EEG in patients performing a verbal fluency task, where they spontaneously explored their own knowledge of the world, sampling words from semantic memory. We found that hippocampal theta power increased during conceptual search and scaled with the semantic distance between successive words, paralleling dynamics observed in spatial navigation. Critically, people transitioned between conceptual clusters, resembling transitions between resource patches in foraging behavior. These shifts were marked by enhanced theta-gamma coupling, both within the hippocampus and between the hippocampus and lateral temporal cortex, a key hub of the semantic network. These findings support a mechanistic account of memory search grounded in navigation and foraging principles, suggesting that the hippocampus orchestrates local computations and long-range interactions to enable flexible retrieval of conceptual knowledge. ### Competing Interest Statement The authors have declared no competing interest. European Research Council, https://ror.org/0472cxd90, 804422, 101125658 Max Planck Society, https://ror.org/01hhn8329 The Kavli Foundation, https://ror.org/00kztt736