
Visual search theories distinguish between attentional capture and disengagement, yet how multiple feature dimensions combine to delay disengagement remains unclear. This study investigated whether shared features (color, shape, object category) produce additive or super-additive costs in oculomotor disengagement. Forty-three participants performed a delayed-disengagement task where a central fixated stimulus shared zero to three features with a peripheral target. Saccadic latency (SL) served as the index of disengagement cost. Results showed that SL increased nonlinearly with the number of shared features, following a quadratic trend. A component decomposition revealed that this acceleration was driven primarily by color: Color overlap produced large main effects, whereas shape and object-category overlap produced costs primarily when combined with color. Our findings suggest that disengagement is not a simple summation of feature similarities. Instead, color may act as a high-priority gate: A color match may open a capacity-limited verification stage in which each additional shared feature disproportionately prolongs inspection. We interpret these findings in terms of a candidate hierarchical gating account, in which high-salience features such as color may govern entry into verification.
The gulf between mathematically motivated stimuli-such as sine waves and noise ensembles-and images drawn from natural scenes may not be as wide as it appears. For the former, the mathematical properties that support their utility as analytic probes are explicit. For naturalistic stimuli, this is not the case, but they are created by physical processes whose outcomes-at least in principle-heavily constrain their statistics. Images generated computationally occupy a middle ground: Their mathematical properties remain implicit in the rules that generate them, but these rules are simpler than the processes that generate natural images. It is suggested that a greater understanding of the statistics of natural images, while challenging, will enhance their utility as experimental tools and thereby lead to new insights into the design principles of sensory systems.
When searching for objects, our sensory resources are limited. According to the zoom lens model of attention, the visual system can flexibly adapt the size of the attentional spotlight applied to an array of items depending on task difficulty. When search tasks are relatively straightforward, attention can be distributed widely across space to process several items simultaneously. However, when discrimination becomes more difficult, the attentional spotlight constricts to concentrate resources to a limited area resulting in the processing of only one or very few items simultaneously at any one time. The present study aimed to empirically test the zoom lens account by assessing search performance while systematically altering the visible area in the search display (field of view [FOV]). We presented two search tasks of contrasting difficulty in a virtual reality environment with a head-contingent display. When the FOV became smaller, there were no differences in search times between the two search tasks, a result which seems to contradict the predictions of the zoom lens model. The zoom lens model predicts that restricted viewing in a demanding search task should exact a smaller cost because of the presence of the smaller attentional window. However, we did not find any evidence of such an effect. Results from an analysis of head and eye movements also suggested that when using both fixations and head movements to scan their environment, participants used head movements to reveal new areas of the visual scene in a manner largely disconnected from classic eye movement and attentional processes. Once the head came to rest, attentional shifts and eye movements were used to search within the newly revealed area.
Free-viewing perimetry is a continuous psychophysical paradigm that aims to assess visual field defects (VFD) while relying solely on an analysis of recorded gaze. Such an approach could benefit patients who cannot be assessed using standard automated perimetry. Previous research has demonstrated that gaze behavior during the free viewing of movie clips contains information on the integrity of the visual field. Yet, thus far, no single analytical approach could reconstruct all types of (simulated) VFD. In the current study, we take a distinctive analytical approach towards VFD reconstruction by explicitly taking into account events that were not gazed at, rather than primarily focussing on those that were. To demonstrate the potential of this approach, we re-analyzed previously recorded data from a gaze-contingent simulated VFD experiment. We compared a participant's gaze behavior to that of control participants while they watched different movie clips. In our analysis, we give specific weight to the events that were prioritized by the control participants, yet missed out upon by those with simulated defects. Our findings indicate that this approach allows us to reliably detect the presence of, and to varying degrees reconstruct five types of simulated visual field defects. We conclude that free-viewing perimetry is feasible and may form the foundation for an inclusive complement to standard automated perimetry.
This study investigated attentional templates and color tuning during efficient visual search. To evaluate optimal tuning and relational accounts, we used a paradigm where participants searched for a color singleton target while ignoring a cross-dimensional shape distractor. Crucially, the distractor's color varied along a precise continuum within the same dimension, spanning values toward and away from the nontarget background. We modeled response times (RTs) and oculomotor capture by fitting asymmetric Gaussian functions to the response curves. Supporting optimal tuning, oculomotor capture (Experiment 1) and manual responses under fixed fixation (Experiment 2) revealed peak interference systematically shifted away from the nontarget color to maximize target-nontarget contrast. Intriguingly, this peak shift was significantly more pronounced in early oculomotor capture than manual RTs, suggesting that optimal tuning biases peripheral attentional guidance (guiding template) more strongly than foveated target verification (target template). Furthermore, response functions exhibited profound asymmetry, displaying shallower slopes for colors extending along the target-defining vector. Whereas Experiment 3 confirmed highly efficient search slopes, Experiment 4 demonstrated that a feature-dissimilar distractor produced negligible capture, validating that selection was governed by a feature-based template. Together, these findings indicate that although attentional templates center on a specifically tuned, contrast-maximizing peak, their overall distribution remains bound to a relative, directional feature rule, consistent with the relational account. Data and scripts are available at https://doi.org/q9f4.
Attentional templates in visual working memory (VWM) are essential for guiding visual search. While much debate has focused on whether these templates are stored in discrete states or continuous resources, a question remains: to what extent can the priority of a template be strategically tuned to match environmental probabilities? We investigated this flexibility by manipulating the reliability of retro-cues (the probability that a cued item is the target) across three experiments combining memory recall and visual search. Crucially, we tested whether strategic prioritization persists under varying memory loads (two vs. three items). We observed that retro-cue validity affected both search efficiency and memory performance, and that these effects were larger under higher cue reliability. This reliability-dependent pattern was clearly present when memory load increased to three items (Experiment 2) and was also observed when item shape signaled the functional role of each representation (Experiment 3). Together, these results show that the influence of a cued representation varies systematically with cue reliability. This pattern is consistent with flexible prioritization but does not distinguish between a continuous adjustment of template strength and probabilistic switching between discrete representational states. Overall, our findings suggest that VWM prioritization is sensitive to expected relevance and adapts to memory load and task structure.
Systematic intra-saccadic displacements of the saccade target object induce oculomotor learning. This type of manipulation remains undetected because of saccadic suppression of displacement. The subtle intra-saccadic manipulation causes the target object to appear at a retinal position that is inconsistent with the computed displacement of visual space during the saccade, leading to the computation of a postdicted motor error and corresponding adaptive changes to the saccade amplitude and the perceived location of objects in space. Adaptive changes to the saccade amplitude or direction have also been observed following obvious intra-saccadic manipulations, but very little is known about the characteristics of this learning process or its underlying mechanisms. The current study therefore compares oculomotor learning in response to subtle and obvious changes of the stimulus display. By measuring their effect on object localization and saccade gain, we were able to show that two different kinds of sensory error can induce oculomotor learning: a global spatial error and a local feature error. While experiencing the former leads to the typical adaptation-induced shift in perceived object location, experiencing the latter does not. On top of that, we demonstrate that adaptive adjustments to the oculomotor behavior do not necessarily rely on those sensory errors. Instead, the overt oculomotor behavior minimizes a task error.
A complete empirical characterization of color discrimination in three dimensions has long remained out of reach. Classical studies, beginning with MacAdam's ellipses, provided local measurements in restricted chromatic planes, but a spatially dense and internally consistent mapping of discrimination structure across full color space has not yet been achieved. Here, we present such a systematic three-dimensional measurement of color discrimination in the standardized sRGB stimulus space. Eight observers measured discrimination regions at 35 reference colors distributed on a body-centered cubic lattice within the RGB cube. At each location, color differences were probed along seven orientations, yielding 14 directional extents. These measurements defined centrally symmetric convex regions that were fitted with minimum-volume ellipsoids, providing a compact description of local discrimination structure. Ellipsoids were represented as symmetric positive-definite matrices and analyzed using a Frobenius geometry, enabling normalization across observers and smooth interpolation to arbitrary locations. The resulting metric field is spatially smooth, highly structured, and remarkably consistent across observers up to an individual global scale factor. Grain size increases along the achromatic axis and exhibits systematic chromatic asymmetries. Comparison with CIEDE2000 reveals substantial agreement in overall scale variation but systematic differences in local anisotropy. Because the metric field is defined geometrically, it can be transformed into linear RGB or any other color representation through the corresponding coordinate transformations. Together, these data provide a coherent three-dimensional empirical mapping of color discrimination across a large ecologically relevant volume of color space and establish an empirical framework for perceptual color metrics.
Accurately perceiving the emotions of other people is an important task for the human visual system. Falsely judging a group of people to be happy when they are angry can have major negative consequences. While the effects of surrounding faces on the perception of a target face (crowding) and the perception of ensemble features of face groups (ensemble perception) have been extensively investigated separately, here we investigated how observers perceive the emotion of both individual faces and groups of faces with identical stimuli to better understand the relations of crowding and ensemble perception. In two experiments, we asked observers to report either the emotion of a central face surrounded by eight faces (central task) or the average emotion of all nine faces (averaging task). In both tasks, only one to two faces were strongly weighted compared with the other faces in the group, suggesting a similar number of sampled items across both tasks. However, we found a strong task- and location-dependent asymmetry: in the central task, the central face and the innermost face on the horizontal meridian (closest to fixation) were strongly weighted, whereas in the averaging task, mainly the two innermost faces on the horizontal meridian and in the upper visual field were strongly weighted. The results indicate that the perception of emotions in groups is strongly constrained by anisotropic weighting mechanisms due to discriminability; however, target discriminability alone did not predict weighting, which was modulated by task requirements. We suggest that the visual system optimizes perceptual selection and information integration according to task demands and visual uncertainty.
Many classic questions in vision science require presenting stimuli at precisely controlled locations in the visual field, a demand that has traditionally been met with specialized laboratory setups involving chinrests, high-resolution displays, and commercial eye trackers. Recent advances in portable mixed-reality hardware raise the possibility that some forms of visual psychophysics could be conducted on consumer devices outside these fixed laboratory environments. Here, we evaluate one such instrument, the Apple Vision Pro, by asking whether it can support psychophysical experiments on two well-studied aspects of peripheral vision: crowding and polar angle asymmetries. These phenomena provide useful benchmarks because both depend on precise control over stimulus position and their behavioral signatures have been well characterized. At the same time, the Apple Vision Pro poses important methodological challenges for vision research, including the inability to access continuous real-time eye-position data and the use of visionOS "points," rather than degrees of visual angle to specify stimulus size and location. We outline strategies for working within these constraints and argue that the usefulness of the platform is not all-or-none, but depends on the scientific question being asked. Specifically, the Apple Vision Pro may be well suited to paradigms in which fixation can be verified at trial onset, stimulus presentation is brief enough to preclude eye movements, and relative positioning of stimuli is sufficient for the comparison of interest, while remaining less suitable for paradigms requiring continuous gaze monitoring, precise specification of absolute eccentricity, or direct quantitative comparison with conventional lab-based setups.
We often consider a visual space to be cluttered if it contains a large number of items that are crowded and untidy. However, measuring clutter is complex, and clutter has been quantified using objective measures, such as mathematical models, most of which capture different aspects of clutter, and subjective measures, such as human judgments. In this study, using a well-controlled semi-realistic image database, we examined the contributions of visual factors, such as the number of items in a scene, their spatial organization, and the complexity of the surface where items are placed, to clutter by testing the predictive performance of clutter metrics and collecting human judgments of perceived clutter with respect to these factors. Our results revealed that, although all factors contribute to clutter in visual scenes, estimates by humans and metrics account only for a subset of these factors, suggesting that visual clutter is a multifaceted construct, and subjective and objective clutter should be considered together to account for several of its dimensions as a more integrative measure.
It remains unclear whether different material attributes, such as glossiness and aesthetics, rely on different stages of visual processing. This study investigated how "temporal preparation," defined here as participants' expectation of stimulus duration, modulates the neural dynamics of material perception, thereby helping reveal the underlying processing stages. Electroencephalography (EEG) was recorded while participants viewed images of various objects and made judgments of either glossiness or aesthetic value. Temporal preparation was manipulated by changing the ratio of short (100 ms) to long (1000 ms) stimulus presentations within each session. This manipulation was intended to alter participants' expectations of stimulus duration without directly constraining the actual presentation duration. Decoding analyses were used to classify binary perceptual judgments from the EEG signals. The results showed that expectations of shorter durations enhanced decoding accuracy for glossiness, whereas expectations of longer stimulus durations tended to enhance decoding accuracy for aesthetics. These findings suggest that temporal preparation influences neural representations differently depending on the material attribute. This approach provides a useful methodological framework for dissociating perceptual strategies and offers new insights into the temporal dynamics of different material impressions.
Predictive processing theories suggest that perception is shaped by prior knowledge, with long-term priors formed through lifelong experiences and short-term priors emerging from immediate contextual cues. Although prior research has explored the influence of expectations on action perception, the interaction between short-term and long-term priors remains understudied. This study examines how these priors interact in visual action perception by embedding animated human and robot agents within a probabilistic-cueing paradigm. Participants observed agents performing biological or mechanical motion, with cues providing either form-based or motion-based short-term priors. Behavioral results showed that long-term priors guided perceptual decisions across conditions when the target stimuli were congruent with them. When expectations were violated, performance patterns reflected a greater weighting of motion-based short-term cues, while form-based cues contributed minimally. These findings highlight the hierarchical nature of predictive processing in action perception, and the precision assigned to different priors is adjusted according to context, favoring motion-based information instead of form-based when resolving perceptual ambiguity.
Our daily visual environment abounds with motion. Accurately perceiving object speed is essential when avoiding collisions or intercepting moving objects. Accurately remembering object speed is equally essential, as moving objects may become temporarily occluded. Most studies of speed use texture motion stimuli (e.g., dot motion), which minimize spatial cues but do not reflect the natural spatiotemporal dynamics of moving objects. In contrast, object motion contains rich spatial cues that better capture real-world motion. Neural mechanisms underlying texture and object motion processing differ across the visual hierarchy. Might memory for speed also differ between texture and object motion? To explore this question, we tested human participants on speed recall for both texture and object motion stimuli across different speeds (2-32°/s) and delays (1-8 s). Both types of target stimuli moved along a circular trajectory (for 4-6 s) and were recalled via the method of adjustment. Replicating previous work, we found a strong decrease in speed recall performance with increasing target speed. Importantly, we also observed that recall was significantly worse for texture stimuli than for object motion stimuli, implying an advantage for spatiotemporally bound objects. For object motion, longer target and delay durations led to better and worse speed recall, respectively. Gaze was subtly biased toward the location of the object moving along its trajectory. For both types of stimuli, responses were biased toward the recall probe, and congruency between target and probe direction improved speed recall. Together, our findings provide a solid psychophysical basis for understanding human short-term memory for speed.
We investigated mental representations of real road scenes using a prediction task. On each trial, 48 licensed drivers watched a road scene for 2 seconds and immediately after selected what the scene would look like 2 seconds in the future in a five-alternative forced choice (5AFC) task. We factorially manipulated three visual properties of the preview to examine their effects on performance: motion, scene context, and stereoscopic depth. We developed a binocular dashcam video dataset of real road scenes simulating human interpupillary distance. We manipulated motion by displaying still images or video previews. Scene context was manipulated by showing videos recorded on urban roads, which are visually denser, or in highway roads, which are visually sparser. To investigate whether small disparity signals are used in this task, we manipulated stereoscopic depth by varying stimulus presentation: Each eye received video from the corresponding camera, only one eye received video, or both eyes received identical videos. Participants were able to correctly select the future appearance of the road to some extent, with better precision for video than stills and for urban scenes than highway scenes. Stereoscopic cues did not affect performance, indicating that participants likely relied more heavily on monocular depth cues in this task. These results suggest that mental representations of complex natural scenes are temporally imprecise and are consistent with, but do not uniquely require, projections of future states.
This article introduces a pyramid-based multiresolution Bayesian framework for high-resolution behavioral analysis from sparse data. By restricting covariance modeling to the coarsest layer of a parameter pyramid while using a difference pyramid for fine-scale refinement, the framework overcomes major computational and statistical challenges of traditional hierarchical Bayesian models with covariance (HBMc). The framework is evaluated against three central claims: (1) scalability through dramatically reduced computational cost, (2) precision under sparsity even with very few trials per block, and (3) improved interpretability via complementary information across resolution layers. Three Bayesian variants-Bayesian inference procedure (BIP; independent parameters), hierarchical Bayesian model with variance only (HBMv), and HBMc-were implemented in PyMC. The best-performing model, PyramidHBMc (HBMc at the top layer combined with HBMv refinement), consistently achieved the best performance (Watanabe-Akaike information criterion weight = 1.0), with the lowest root mean square error and standard deviation, reducing errors and variability by up to 74.1% and 78.5% relative to BIP across four datasets spanning one-dimensional temporal and two-dimensional spatial functions. These results directly support the three central claims and demonstrate the framework's broad applicability in perceptual and cognitive science.
Closed captions, originally for hearing-impaired audiences, are now increasingly used, adding an additional visual element to the multisensory experience of movie viewing. This study examines how gaze behavior adapts to the absence of auditory information by testing whether increased caption engagement reflects more frequent switching between closed captions and the visual scene, longer uninterrupted reading intervals, or both. Using eye-tracking, we analyzed saccades and fixations in designated areas of interest as participants viewed a 20-minute movie with alternating presence and absence of audio. Viewers integrated the two visual input streams by frequently shifting their gaze between the captions and the scene. The balance between reading captions and watching visuals shifts in the absence of sound by adjusting transition probabilities (i.e., increasing gaze switching between captions and scene, rather than changing the duration of reading intervals). This suggests that increased caption engagement in the absence of sound reflects more frequent sampling of captions, rather than longer uninterrupted reading. Modeling the probabilistic characteristics of fixations and saccades revealed more consistent and focused eye movement patterns on the caption text compared to the visual scene, highlighting distinct reading and gazing behaviours. Notably, these differences remained stable regardless of the presence or absence of sound. In addition to the availability of audio, we observed the presence of dialogue and the emotional content also influence caption engagement. These findings highlight how viewers adapt to missing audio through more frequent switching between text and scene while maintaining low-level properties of eye movements.
Autonomous vehicles (AVs) use sensors that exceed human capabilities in resolution and reaction speed, yet a catastrophic failure mode persists: the inability to detect massive, white obstacles (such as semi-trailers) against bright skies. Although engineering analyses attribute these accidents to sensor dynamic range or radar filtration errors, we argue they represent a fundamental failure of perceptual organization. We propose the term "white whale effect" to describe a form of machine visual blindness where high-luminance objects are misclassified as ambient background because of failures in lightness anchoring, figure-ground segregation, and global motion perception. We conclude that safer AV perception requires moving beyond pixel-level classification to architectures constrained by the principles of psychophysics.
In immersive virtual environments and other interactive settings the motion gain linking stimulus motion to self-movement is often manipulated to study perceptual performance. We show that this introduces a previously unidentified source of external noise we term gain-dependent noise. The noise arises from trial-to-trial variability in self-movement and can distort the shape of the psychometric function. This leads to overestimates of threshold and sensory measurement noise, and misinterpreted lapse rates when fitting standard sigmoidal functions like a cumulative Gaussian. Using a model observer with minimal assumptions, simulations show the situation is more critical when the coefficient of variation of self-movement (i.e., across-trial standard deviation/mean) is relatively high, and measurement noise associated with image-based and idiothetic signals is relatively low. We derive a new gain-dependent psychometric function that separates sensory measurement noise from gain-dependent noise, and provide a MATLAB fitting routine (fitgdpmf).
Human visual motion processing is thought to involve two stages: local motion processing and global pooling of local motion signals. We used an efficient equivalent noise paradigm to explore temporal aspects of motion processing, specifically manipulating the lifetime of moving elements and overall stimulus-duration, to examine (respectively) local and global processing. Participants judged whether the overall direction of a dot pattern was clockwise or anti-clockwise of a reference direction. An adaptive staircase determined a) the minimum direction offset required for reliable judgment of direction when dots moved in a similar direction and b) the maximum directional noise (range of directions) that still allowed participants to tell if a stimulus was moving either ±45° from the reference direction. The two sets of thresholds allowed us to infer both a) the precision of participants' judgment of the direction of a single dot (i.e., local noise) and b) the number of samples they were effectively averaging (i.e., global sampling). Dots moved at 10 deg/s and we tested various dot lifetimes (33.36, 66.72, 133.4, and 266.8 ms) with a stimulus-duration of 1,000 ms, and various stimulus-durations (62, 125, 250, 500, 1,000, and 2,000 ms) with a dot lifetime of 133.4 ms. Longer element-lifetimes increased local directional precision but did not influence global integration (number of samples effectively averaged). Conversely, longer exposures did not greatly influence local precision, but increased global averaging by an average of 13% for every doubling of stimulus-duration. Thus, element-lifetime and stimulus-duration have dissociable effects on motion perception improving local and global motion processing, respectively.