
The mechanisms of goal-driven attention are often characterized as serving the selection of task-relevant information. However, decades of research have now shown that selection is biased by many different sources of extra-target information. These sources are often characterized as being "task irrelevant." In this review, we systematically describe different sources that carry uncertainty-reducing signals, and we argue that these "irrelevant" sources are not only relevant but also necessary for efficient processing of our environment. We provide evidence that "irrelevant" information carries target-predictive signals integral to rapid goal-directed attentional processing. We argue that the extra-target sources of information are a form of cognitive shortcuts that contribute to greater efficiency in information selection in the natural world.
The primate retina transmits visual information to the brain through parallel pathways carried by diverse retinal ganglion cell (RGC) types, each of which encodes specific features in the visual environment. These parallel channels provide input to a variety of brain regions to support both conscious and reflexive visual functions. Although the midget, parasol, and small bistratified pathways have dominated classical models of primate vision, accumulating evidence points to a broader diversity of wide-field RGC types whose functions are less well understood. Recent breakthroughs in single-cell sequencing have provided new insights into the transcriptomic signatures of these wide-field cells, uncovering molecular markers that can be used to link sparse RGC types to their distinct morphologies and physiological functions. This review synthesizes current knowledge about the diversity of wide-field RGCs in the primate retina and highlights recent insights into cells that encode complex visual features, such as direction-selective ganglion cells. We emphasize how integrative methodological approaches are refining our understanding of the roles of different wide-field RGC types in primate vision.
Humans execute rapid saccadic eye movements to inspect objects of interest in high-acuity foveal vision. Even though each saccade entails a large-scale displacement of the retinal image, vision is continuous, and we easily keep track of the location and identity of relevant stimuli. Here, we highlight the contribution of predictive foveal processing to visual continuity. Specifically, we summarize evidence that foveal and even foveolar vision is nonuniform and modulated by attentional allocation. We then describe a set of psychophysical and neuroimaging studies demonstrating that defining features of the eye movement target are predicted in presaccadic foveal vision. We explore a parsimonious implementational mechanism and propose the contribution of foveal prediction to several seemingly unconnected phenomena that may be unified under a straightforward assumption: Peripheral saccade target features predictively alter feature tuning in high-acuity foveal vision, facilitating a smooth perceptual transition once the target shifts to the center of gaze.
Eye-hand coordination (EHC) is central to everyday behavior. This is often described as a sequential process in which the eyes move first to guide the hand. However, converging behavioral and neurophysiological evidence supports a fundamentally different view: Coordination arises from a distributed network that issues parallel commands and aligns effectors through structured interareal communication. Saccade and reach planning are typically initiated concurrently, with apparent timing differences driven largely by effector dynamics. Experimental dissociations reveal that coupling enhances performance but is not obligatory, particularly in bimanual or naturalistic contexts. Here we emphasize the posterior parietal cortex as a key hub integrating sensory and motor signals for planning and coordination through effector-specific subregions and their interareal interactions. Oscillatory dynamics, notably beta-band coherence, consistently associate with coordination between oculomotor and manual circuits, although whether they causally implement routing or instead index structured interactions remains unresolved. Together, these findings point to a distributed, intercommunicating network that flexibly aligns eye and hand control.
The visual system is attuned to the statistical regularities of the visual world, enabling rapid, almost reflexive, and highly accurate recognition of object identities and categories. This article presents recent advances across psychophysics, computational modeling, and cognitive neuroscience, which together suggest that the visual system is just as attuned to the laws of physics governing our physical world. We review psychophysical work showing that vision incorporates intuitive physics as a rapid, spontaneous, and stimulus-driven process. We review the computational framework of physics-based analysis by synthesis, which suggests that the visual perception of intuitive physics corresponds to building and manipulating approximate structure-preserving representations of physical scenes. We end with an outline of an integrative program of computational modeling work, along with neuroscientific and psychophysical studies, toward psychologically and neurally refined mechanistic accounts of the visual perception of intuitive physics.Updated on July 8, 2026.
In the retina, two-photon (2P) excitation induces fluorescence emission both from endogenous fluorophores, including vitamin A metabolites that sustain vision, and from exogenous dyes or fluorescent proteins expressed selectively in individual retinal cells. The simultaneous absorption of two near infrared photons generates 2P excited fluorescence and results in high signal-to-noise ratio images deep in the tissue, enabling acquisition of 3D volumes over the entire thickness of the retina in an intact eye. Advancements in noninvasive 2P-based techniques offer detailed characterization of retinal structure and function at subcellular levels. This capability is especially valuable for measuring transfection efficiency of newly developing gene-editing approaches, including clustered regularly interspaced short palindromic repeats-CRISPR-associated nuclease 9 (CRISPR-Cas9) systems and viral-vector-mediated therapies, designed to treat inherited eye diseases. Furthermore, the simultaneous absorption of two photons by visual pigments can directly initiate phototransduction cascades, providing unique insights into precise detection of photoreceptor sensitivity and unexplored mechanisms of visual perception. Two-photon processes enable real-time study of biochemical transformations in living retinal tissue, advancing the development of novel treatments and facilitating assessment of therapeutic interventions.
Visual inputs evoke dynamic changes in the responses of visual neurons that evolve over both space and time. All visual experiences must somehow stem from stimulus-evoked spiking patterns in visual brain regions, but determining which neurons and periods of activity causally give rise to perception is one of the grand challenges in neuroscience. Targeted manipulations of neuronal activity while subjects perform sensory tasks have been indispensable for probing computations underlying perception. However, limits on the spatiotemporal precision of neuronal perturbations have constrained the scope of inquiry. Recent advances in optogenetic stimulation approaches have opened the door to augmenting neuronal activity on the spatiotemporal scales of visual computations. Applications of patterned optogenetic stimulation in behaving subjects have revealed many important insights into how different aspects of visual neuronal responses contribute to perception and ultimately behavior. First, we survey optical tools for precise optogenetic perturbations in mice and monkeys. Next, we discuss how patterned optogenetic experiments in behaving subjects have been used to causally test the neural mechanisms underlying perception. Last, we highlight a few key areas that we believe will be important for continued progress in this emerging research area.
Studies of chromatic discrimination have historically varied in their experimental procedures and in the units in which thresholds are expressed, but by 1950, it was clear that thresholds depend both on differences in the ratios of cone excitations and on the properties of post-receptoral channels. Discrimination was found to be optimal at the chromaticity to which the observer was currently adapted. Consequently, the most famous dataset in this field, MacAdam's discrimination ellipses, cannot validly be used to construct a uniform color diagram or to estimate the number of discernible colors: MacAdam's single observer, Perley Nutting, would have been in different states of adaptation when matching different colors. Furthermore, a sensation of difference is not the same as a difference of sensations: Some color discriminations may depend on channels that extract differences between adjacent regions, and different pairs of chromaticities could then give rise to the same sensation of difference.
Deep neural networks (DNNs) offer highly promising neurocomputational models of the visual system, yet vast gaps remain between DNNs and human observers. By some accounts, DNNs are approaching near ceiling levels in their ability to predict human neural responses to clear real-world images. However, even modest diversions toward more ambiguous viewing conditions can readily expose the brittle and inflexible nature of these networks. Human vision remains robust when faced with noise, blur, occlusion, and other challenges, whereas DNNs trained to classify large image datasets typically lack such robustness. Here, we discuss the ecological vision hypothesis, proposing that the robustness of human vision is acquired via learning from prevalent encounters with challenging viewing conditions, such that DNNs trained with similar challenges should become more robust and human-aligned. In particular, the prevalence of blur in everyday vision may enhance sensitivity to global shape and attenuate reliance on local textural cues. We conjecture that providing DNNs with ecologically relevant information to learn 3D scene and shape properties will further advance DNN-to-human alignment.
Binocular integration is a well-established feature of neuronal processing in the primary visual cortex, where such integration is thought to first emerge. However, accumulating evidence demonstrates that subcortical retinorecipient nuclei possess sophisticated binocular processing capabilities, with important implications for cortical function, visual behavior, and non-imaging-forming physiology. This review synthesizes our current understanding of the circuit origins and functional relevance of binocular integration and modulation in the dorsal lateral geniculate nucleus, superior colliculus, and other subcortical targets. We describe how the definition of binocularity has evolved beyond simple ocular dominance to encompass diverse modes of neuronal modulation, including facilitation, summation, suppression, and emergent responses. We highlight the prevalence of multiple wiring motifs subserving binocular convergence, including direct retinal inputs, lateral circuits via local interneurons, feedback from cortical or subcortical sources, and indirect relays through intra- or interhemispheric connections. Recent anatomical and functional studies reveal substantial binocular integration despite apparent eye-specific input segregation, with region- and species-specific differences reflecting distinct ethological demands. Finally, elucidation of a critical role for subcortical binocular processing in prey capture and threat responses is expanding our cortex-centric view and revealing new complexity in the regional distribution of visual computations essential for survival behaviors.
Inferring 3D surface structure is one of the most fundamental functions of vision. There are many well-known depth cues, such as shading, texture, and highlights. However, how these cues are extracted from images-and what exactly they tell the brain about 3D shape-is not fully understood. Here, we describe how these seemingly distinct 3D shape cues could share a common currency for the first stages of shape estimation. The key insight is that when patterns such as shading or texture are projected from a 3D object into the 2D retinal image, they are spatially distorted, with profound consequences for local image statistics. The distortions create highly organized patterns of local image orientation (orientation fields) that are systematically related to specific 3D shape properties. Orientation fields can be reliably measured by filter populations and predict both successes and failures of human shape perception across diverse conditions.
Depth perception depends on two powerful cues, stereopsis and motion parallax, that share similar geometric foundations but operate under distinct constraints. When observer motion is self-generated and the scene is stationary, the two can produce comparable depth percepts. However, disparity generally supports finer discrimination and more accurate depth magnitude percepts. Real-world depth perception from both cues is complicated by eye movements and matching ambiguity, but motion parallax faces additional challenges due to head movements and the presence of object and scene motion. Empirical studies and cue combination models show that the interaction of the two cues is nonlinear, with stereopsis often dominating depth percepts when both cues are available. These issues are increasingly relevant to immersive displays, where head tracking restores natural parallax but optical and geometric mismatches distort the resulting perceived depth. Understanding how the visual system integrates disparity and parallax remains essential for both theory and application.
Underserved communities in the United States continue to face substantial barriers to accessing eye care, leading to preventable vision loss. This review synthesizes United States-based teleophthalmology programs designed to expand eye care access among under-resourced populations, including older adults, Indigenous communities, children, veterans, people experiencing homelessness, and patients in safety-net systems. Most programs used asynchronous nonmydriatic fundus photography captured by trained technicians in community or primary care settings, with remote specialist interpretation and frequent integration into electronic health records. Across studies, teleophthalmology increased diabetic eye screening rates by 15-40% and identified ocular disease in 20-25% of participants. Many initiatives partnered with community organizations to address cultural and logistical barriers unique to the populations they served. Despite these gains, challenges persist, particularly poor follow-up adherence, financial instability, limited technical capacity, and workflow integration. Emerging opportunities include artificial intelligence-assisted automation, sustainable reimbursement mechanisms, home-based screening, and virtual eye clinics to enhance equitable, population-level vision care.
The primate brain excels at transforming photons into knowledge. When light strikes the back of the eye, opsin molecules within rods and cones absorb photons, triggering a change in membrane potential. This energy transfer initiates a cascade of neural events that endows us with useful knowledge. This knowledge manifests as subjectively experienced perceptual interpretations and mostly pertains to the 3D structure of the visual environment and the affordances of the objects within the scene. However, some of this knowledge instead pertains to the quality of these interpretations and contributes to our sense of confidence in perceptual decisions. Because such confidence reflects knowledge about knowledge, psychologists consider this the domain of metacognition. Here, we examine what is known about the neuronal basis of perceptual decision confidence, with a focus on vision. We review the crucial computational processes and neural operations that underlie and constrain the transformation of photons into visual metacognition.
Over the past decade and a half, a new understanding has emerged of the role of vision during the critical period in the primary visual cortex. Rather than driving competition for cortical space, vision is now understood to inform the establishment of feature conjunctions that cannot be constructed intrinsically. Longitudinal imaging studies reveal that the establishment of these higher-order feature detectors is a remarkably dynamic process involving the gain and elimination of neurons from functional groups (e.g., binocular neurons with nonlinear response tuning). Experience exerts its influence selectively on this developing circuitry; some pathways require experience for normal development, while others appear to be intrinsically established. This difference drives the network dynamism that is exploited to construct novel cortical representations that best encode our local environment and inform our actions in it.
Information visualization is central to how humans communicate. Designers produce visualizations to represent information about the world, and observers construct interpretations based on the visual input as well as their heuristics, biases, prior knowledge, and beliefs. Several layers of processing go into the design and interpretation of visualizations. This review focuses on processes that observers use for interpretation: perceiving visual features and their interrelations, mapping those visual features onto the concepts they represent, and comprehending information about the world based on observations from visualizations. Observers are more effective at interpreting visualizations when the design is well-aligned with the way their perceptual and cognitive systems naturally construct interpretations. By understanding how these systems work, it is possible to design visualizations that play to their strengths and thereby facilitate visual communication.
Animal camouflage in the natural world has been studied for over a century, with early research often relying on descriptive accounts of patterning as perceived by human observers. Recent advances, however, have leveraged a deeper understanding of visual processing across a wide range of predators. This review examines literature illustrating how insights from vision science have enriched research on camouflage. We focus on three areas: color and texture, motion processing, and the perception of shape and depth. We discuss findings from vision research that show how animals seeking to remain undetected optimize their camouflage. We also explore how predator visual systems have evolved to break that camouflage. Last, we highlight gaps where vision science has yet to be applied to research on camouflage, with the hope of encouraging further interdisciplinary work.
Crowding is ubiquitous: When objects are surrounded by other elements, their perception may be impaired depending on factors such as the proximity of the surrounding elements and the grouping of elements and targets. Crowding research aims to identify these factors, for instance, which elements interfere with one another and how close they need to be to cause crowding. Traditionally, crowding was thought to occur only within narrow temporal and spatial limits around the target. Recent studies, however, reveal that crowding may result from both low- and high-level processes, such as perceptual grouping and timing, as well as the arrangement of complex visual stimuli. This review highlights these new insights, suggesting that overall organization, as well as both feedforward and feedback processes, plays a role. Crowding emerges as a highly complex and dynamic phenomenon, underscoring the need for a more integrated approach to fully capture its intricacies, which may carry broader implications not only for crowding but also for vision science as a whole.
Layer 6 corticothalamic (L6 CT) pyramidal neurons send feedback projections from the primary visual cortex to both first- and higher-order visual thalamic nuclei. These projections provide direct excitation and indirect inhibition through thalamic interneurons and neurons in the thalamic reticular nucleus. Although the diversity of L6 CT pathways has long been recognized, emerging evidence suggests multiple subnetworks with distinct connectivity, inputs, gene expression gradients, and intrinsic properties. Here, we review the structure and function of L6 CT circuits in development, plasticity, visual processing, and behavior, considering computational perspectives on their functional roles. We focus on recent research in mice, where a rich arsenal of genetic and viral tools has advanced the circuit-level understanding of the multifaceted roles of L6 CT feedback in shaping visual thalamic activity.
Visual attention prioritizes relevant stimuli in complex environments through top-down (goal-directed) and bottom-up (stimulus-driven) mechanisms within cortical networks. This review explores the neural mechanisms underlying visual attention, focusing on how attentional control is encoded and decoded from prefrontal signals in both spatial and temporal domains. Decoding methods enable real-time tracking of covert visual attention from prefrontal activity with high spatial and temporal resolution, as a neurophysiological proxy of the attentional spotlight. This research provides insights into stimulus selection mechanisms, proactive and reactive suppression of irrelevant stimuli, the rhythmic nature of attentional shifts and attentional saccades, the balance between focus and flexibility, and the variation of these processes along epochs of sustained attention. Additionally, the review highlights how recurrent neural networks in the prefrontal cortex contribute to supporting these attention dynamics. These findings collectively offer a comprehensive model of attention that integrates dynamic prioritization processes at short and longer timescales.