In studying primate vision, a large body of work focuses on the first feedforward sweep. During this initial time window, information is thought to pass through ventral stream regions in a stage-like fashion in an effort to extract high-level information from the retinal input. Consequently, electrophysiological analyses commonly focus on spatial response patterns, either by averaging data in time, or by applying decoders in a temporally local fashion. By analysing data recorded simultaneously across multiple arrays placed along the macaque ventral stream, we here show that this prior approach may be missing key aspects of information encoding. First, time-resolved, multivariate analyses of information transfer between V4 and IT reveal temporally and semantically varied information content as being exchanged within the first 100ms of processing. Second, by employing recurrent neural network (RNN) decoding techniques that extend across the temporal domain, we demonstrate that the neural pattern dynamics themselves carry categorical information far beyond the spatially encoded information available at any given time point. These findings challenge the prevailing view of a single, stage-like feedforward process and suggest that even the earliest parts of visual processing are better characterised as a spatiotemporally evolving process that encodes information in its dynamics rather than purely spatial response patterns.
Here we present the Active Visual Semantics (AVS) dataset, a large-scale collection of magnetoencephalography (MEG) and eye-tracking data recorded while five participants freely explored 4,080 natural scenes over 10 sessions each, yielding more than 200,000 fixation epochs in total. Unlike existing neuroimaging datasets that rely on passive viewing with enforced central fixation, AVS captures brain activity during active scene exploration, including self-generated saccades and fixations. A semantic captioning task on 25
The rapid advancement of artificial intelligence (AI) has led to its increased integration into domains involving physical agency. While traditional AI applications maintain a clear division of responsibility with the decision-action loop being fully controlled by the human user (e.g., AI-assisted medical diagnosis) or the AI-based system (e.g., autonomous driving), novel assistive technologies disrupt this boundary. Here, we examine the responsibility gap in the case of HANS (Human-AI Navigation System), a system designed to assist visually impaired users in grasping. HANS splits the perception-action loop across the system boundary, with the AI making continuous spatial decisions and guiding the user via tactile vibrations, and the user executing the physical movement. The collaborative tandem enhances the capability of each agent and restores autonomy of the visually impaired user at the cost of a blurred responsibility boundary.
Abstract During natural vision, the human brain constantly has to account for self-produced interruptions of visual input, through blinks and saccades, to maintain a stable perception of the world. It remains unclear whether these event-related neural responses are treated similarly due to a common underlying mechanism. To systematically investigate this functional connection, we recorded synchronized mobile EEG and eye-tracking data from freely moving participants as they (visually) explored a city center. Our results replicate prior findings showing that a substantial portion of the neural response is aligned to event onsets in comparison to offset, underscoring the relevance of investigating the beginning and end of these events. Using advanced deconvolution techniques, we disentangled overlapping, time-locked neural contributions associated with the onsets and offsets of saccades and blinks. The results demonstrate that the deconvoluted neural kernels are highly correlated when comparing saccade onset to blink onset, as well as saccade offset to blink offset, for both evoked and induced activity across electrodes. Together, our findings demonstrate that despite their distinct physical profiles, saccades and blinks share a common neural signature that governs both their onset and offset dynamics during unconstrained real-world behavior. Highlights Co-registration of portable EEG and Eye-Tracking yielded meaningful real world data. We studied blinks and saccades as natural interruptions of visual processes. Aligning EEG data to onset and offset underscored the relevance of both timepoints. We deconvoluted overlapping neural activity for onsets/offsets of saccades/blinks. ERPs and TFRs are highly correlated for onset/offset of blinks and saccades.
Before each of around 200,000 eye movements we make each day, the brain decides how long to fixate before shifting gaze to new information. Here we investigate this process using a large-scale scene-viewing experiment (4,080 natural scenes, five participants) that combines magnetoencephalography, eye tracking and a semantic captioning task. Using multivariate analysis of magnetoencephalography source-space patterns, behavioral analyses and artificial neural network (ANN) modeling, we show that longer fixations do not reflect prolonged visual processing but relate to downstream memory encoding. First, temporal variability of ventral stream representational dynamics did not explain variability in fixation duration. Second, fixation durations were anticorrelated with ANN-estimated patch classification difficulty. Third, fixation durations correlate positively with ANN-predicted patch memorability and caption-inclusion and co-occur with increased theta-gamma phase-amplitude coupling, particularly in frontal and hippocampal regions. These results indicate that eye-movement timing decisions are shaped by memory-encoding demands rather than by perceptual processing limits.
Abstract Balancing exploration and exploitation is a fundamental challenge for adaptive behavior, yet it remains unclear whether visual sampling and spatial locomotion reflect a single cross–domain trait or operate independently. We addressed this question by recording head–mounted eye–tracking and full-body motion tracking while 26 participants freely navigated “Westbrook”, a large–scale virtual city for a total of 150 min across five sessions. From the movement trajectories we derived three spatial descriptors: median walking speed, occupancy entropy, and the proportion of explorative route choices. From the gaze data, we computed 38 robust visual descriptors encompassing fixation dynamics, pupil size, saccadic amplitude, gaze–head alignment, and transition entropy. Principal–component analysis reduced the visual descriptors to three components that captured 58 % of variance, with the first component (PC1) reflecting “gaze dynamism” (frequent shifts, larger saccades, higher transition entropy). Canonical correlation analysis revealed a strong coupling between spatial and visual behaviours: the first pair of canonical variates correlated at r = 0.68 (cross–validated r = 0.45), driven primarily by the association of high walking speed and occupancy entropy with elevated gaze dynamism. In contrast, the proportion of explorative route choices contributed little to this coupling. These findings demonstrate that individual differences in low–level locomotor speed and spatial coverage co–vary with an exploratory visual style, supporting the existence of a domain–general “exploration” factor that shapes both how people move through, and attend to, complex environments.
Spatial learning emerges not only from static environmental cues but also from the social and semantic context embedded in our surroundings. This study investigates how human agents influence visual exploration and spatial knowledge acquisition in a controlled Virtual Reality (VR) environment, focusing on the role of contextual congruency. Participants freely explored a 1 km2 virtual city while their eye movements were recorded. Agents were visually identical across conditions but placed in locations that were either congruent, incongruent, or neutral with respect to the surrounding environment. Using Bayesian hierarchical modeling, we found that incongruent agents elicited longer fixations and higher gaze transition entropy (GTE), a measure of scanning variability. Crucially, GTE emerged as the strongest predictor of spatial recall accuracy. A counterfactual mediation analysis indicated a small but reliable pathway via GTE and, for incongruent agents, a larger direct component not captured by GTE. These findings suggest that human-contextual incongruence promotes more flexible and distributed visual exploration, thereby enhancing spatial learning. By showing that human agents shape not only where we look but how we explore and encode space, this study contributes to a growing understanding of how social meaning guides attention and supports navigation.
In the last five years, dynamic optical coherence tomography (DOCT) has become a rapidly advancing field of research, which extends OCT to include functional contrast. In conjunction with increasing imaging speed and resolution this technique has broadened the application areas of OCT, unveiling previously hidden structures perceptibly. However, owing to the rapid growth of the field and its parallel development by multiple groups for diverse applications and imaging setups, a wide range of feature extraction metrics and corresponding visualization approaches has emerged. Here, we present an overview of the historical development of DOCT and its predominant algorithms for data evaluation as well as visualization. Recently proposed improvements are addressed, and lastly we highlight some of the applications of DOCT.
IntroductionEarly-life dysbiosis is associated with increased risk of asthma development but the underlying mechanisms remain unclear. Although eosinophils have been reported in the developing lung, their contributions to alveolar morphogenesis and lung mechanics have not been functionally interrogated.MethodsMaternal exposure to antibiotics (ABX) was used to induce early-life offspring dysbiosis, and the effects on lung function and development was assessed. Similar measurements were made in mice lacking eosinophils due to genetic modification, or administration of IL-5 blocking agents.ResultsABX exposure between Embryonic Day 15 (E15) and post-natal day 28 (PN28), increased allergen-induced, and baseline airway hyperreactivity (AHR). Similar observations were made when maternal ABX exposure was limited to PN10 to PN20. Complete characterization of baseline lung mechanics demonstrated downward-shifted pulmonary PV loops, increased small airway resistance, decreased compliance, and reduced inspiratory capacity at weaning and 14 months of age. Consistent with observation of small airway dysfunction, offspring of ABX-exposed dams demonstrated significantly smaller alveoli at multiple stages of lung development. Examination of recruitment to developing lungs demonstrated an exaggerated recruitment of eosinophils at key developmental periods (PN14) in offspring of ABX-exposed dams. Mice with fewer eosinophils (through genetic knockout, or treatment with anti-IL-5) display altered patterns of lung mechanics opposite to that seen in offspring of ABX-exposed dams.DiscussionThese data underscore an underappreciated role of eosinophils in homeostatic lung development and suggest that early life modulation of pulmonary eosinophil activity has long-term effects on susceptibility to the development of chronic lung diseases such as asthma.
BACKGROUND:Approach-avoidance behaviors are fundamental mechanisms guiding our interactions with the environment, driven by the emotional valence of stimuli. While previous research has extensively explored behavioral aspects of the AAB, the neural dynamics underlying these processes remain insufficiently understood. OBJECTIVES:The present study employs electroencephalography (EEG) to systematically investigate the neural correlates of AAB in a non-clinical population, focusing on stimulus- and response-locked event-related potentials (ERPs). METHODS:Forty-three participants performed a classic Approach-Avoidance Task (AAT) while EEG activity was recorded. RESULTS:Behavioral results confirmed the AAB effect, with faster reaction times in congruent compared to incongruent trials, as well as for positive versus negative trials. ERP analyses revealed significant differences in the Valence factor, with early effects for stimulus-locked trials and late differences at the parietal-occipital region for response-locked trials. However, no significant effects were found for the Condition factor, suggesting that the neural mechanisms differentiating congruent and incongruent responses might not be optimally captured through EEG. Additionally, frontal alpha asymmetry (FAA) analyses showed no significant differences between conditions, aligning with the literature. CONCLUSIONS:These findings provide novel insights into the temporal and spatial characteristics of AAB-related neural activity, emphasizing the role of early visual processing and motor preparation in affect-driven decision-making. Future research should incorporate methodological approaches for assessing AAB in ecologically valid settings.
After the offset of complex visual stimuli, rich stimulus information remains briefly available to the observer, reflecting a rapidly decaying iconic memory trace. Here we found that even if the cues are presented in the final stage of the stimulus presentation, the reportable information already starts decaying. Using closely spaced readout cues and a theoretical model of information availability, we observed that a cue has to be presented around 10 to 30 milliseconds before stimulus offset to access the full sensory information. We suggest that this does not reflect an early loss in sensory encoding, but instead it is a consequence of a latency in the processing of the cue that postpones the readout of the sensory representation by 10 to 30 milliseconds. Our analysis also shows that spatial proximity of items in complex arrays impacts sensory representation during both perceptual encoding and initial memory decay. Overall, these results provide a theoretical and empirical characterization of the readout from visual representations and offer a detailed insight into the transition from perception into iconic memory.
Current research strives to investigate cognitive processes under natural conditions. Virtual reality and EEG are promising techniques combining naturalistic settings with close experimental control. However, many questions and technical challenges remain, e.g., are saccade onsets a suitable replacement of fixation onsets as key events in continuous gaze trajectories ( Amme et al., 2024), and consequently, can VR capture differences across different stimulus categories associated with varying saccade durations? To address both questions, we investigate the N170 face effect in humans (14 males, 19 females, zero diverse) using a free-viewing and free-movement immersive VR study that contained houses, various background stimuli, and, notably, static and moving pedestrians to study face perception under naturalistic conditions. Our results show that aligning trials to saccade onsets leads to more well-defined ERPs than fixation onsets, especially for the P100 component, demonstrating that saccade-onset ERPs are a better-suited analysis method for this type of experiment. Furthermore, we observe an evolution of category-based differences, i.e., face versus background saccade-onset ERPs, compatible with previous reports but extending in a large temporal window and including all electrode sites at different points in time. In summary, employing VR, EEG, and eye-tracking to investigate differences across fixation categories provides insights into the relevance of saccadic onsets as event triggers and enhances our understanding of cognitive processes in naturalistic settings.
We demonstrate a 3.2 MHz-OCT system for inter-volumetric dynamic optical coherence tomography of ex vivo porcine kidney tissue. Employing a home-built Fourier Domain mode locking (FDML) laser with a 1310 nm wavelength, the system achieved a lateral resolution of 3.48 mu m and a frame rate of 612 Hz. A motorized XYZ positioning stage enabled the precise acquisition of multiple volumes, which were seamlessly stitched together to generate a comprehensive dataset with a total area of 2.6 x 2.6 mm(2). Validations against histological sections confirmed the system's ability to visualize cellular tissue structures.
Grasping constitutes a critical challenge for visually impaired people. To address this problem, we developed a tactile bracelet that assists in grasping by guiding the user's hand to a target object using vibration commands. Here we demonstrate the fully automated system around the bracelet, which can confidently detect and track target and distractor objects and reliably guide the user's hand. We validate our approach in three tasks that resemble complex, everyday use cases. In a grasping task, the participants grasp varying target objects on a table, guided via the automated hand navigation system. In the multiple objects task, participants grasp objects from the same class, demonstrating our system's ability to track one specific object without targeting surrounding distractor objects. Finally, the participants grasp one specific target object by avoiding an obstacle along the way in the depth navigation task, showcasing the potential to utilize our system's depth estimations to navigate even complex scenarios. Additionally, we demonstrate that the system can aid users in the real world by testing it in a less structured environment with a blind participant. Overall, our results demonstrate that the system, by translating the AI-processed visual inputs into a reduced data rate of actionable signals, enables autonomous behavior in everyday environments, thus potentially increasing the quality of life of visually impaired people.
Large language models (LLMs) have revolutionized human-machine interaction, and have been extended by embedding diverse modalities such as images into a shared language space. Yet, neural decoding has remained constrained by static, non-interactive methods. We introduce CorText, a framework that integrates neural activity directly into the latent space of an LLM, enabling open-ended, natural language interaction with brain data. Trained on fMRI data recorded during viewing of natural scenes, CorText generates accurate image captions and can answer more detailed questions better than controls, while having access to neural data only. We showcase that CorText achieves zero-shot generalization beyond semantic categories seen during training. Furthermore, we present a counterfactual analysis that emulates in-silico cortical microstimulation. These advances mark a shift from passive decoding toward generative, flexible interfaces between brain activity and language.
The proxemics theory explains the consistent social boundaries surrounding individuals as reported (Hall in The Hidden Dimension, Doubleday, Garden City, 1966), yet little is known about the social boundaries surrounding pairs or groups of people. The current study explored interpersonal proxemics behavior in a virtual environment, focusing on distances maintained towards individual pedestrians, pairs, and groups. Using virtual reality to simulate a city center, participants freely navigated it while their movements and gazes were captured. Importantly, the city was populated by pedestrians in different social configurations. Eye movements identified interactions defined by gaze-onsets towards a pedestrian’s head. Our results indicate that participants approached individuals with a median distance of 3.18 unity units aligned with the social space boundary as reported (Hall in The Hidden Dimension, Doubleday, Garden City, 1966). Distances kept from pairs and groups were similarly centered within the social space, revealing no significant difference in approaching behavior across different social configurations. The consistency in approaching distances suggests that personal and social spaces are not substantially altered, irrespective of the social context.
Dynamic optical coherence tomography (DOCT) enhances conventional OCT by providing specific information related to flow dynamics, cell motility, and organelle metabolic activity. These biological phenomena can be detected with varying sensitivity depending on the OCT architecture parameters, including wavelength, numerical aperture, and implementation method (time domain or Fourier domain). Despite its potential, the field lacks standardization as various research groups have independently developed algorithms for specific applications. In this paper, we compare four widely used DOCT algorithms, each employing a distinct analytical approach: power spectral density moment analysis, frequency band visualization, logarithmic intensity variation evaluation, and motility-based analysis. These algorithms were originally optimized for different OCT technologies (full-field OCT, microscopic OCT, swept-source OCT, and spectral domain OCT), which vary in temporal and spatial resolution as well as susceptibility to motion artifacts. To conduct a fair evaluation, we perform comprehensive cross-wise comparisons using datasets acquired from each of these setups. Our findings reveal that each method exhibits unique advantages in specific imaging environments, thereby providing valuable guidance for algorithm selection based on particular application requirements.