The visual system is attuned to the statistical regularities of the visual world, enabling rapid, almost reflexive, and highly accurate recognition of object identities and categories. This article presents recent advances across psychophysics, computational modeling, and cognitive neuroscience, which together suggest that the visual system is just as attuned to the laws of physics governing our physical world. We review psychophysical work showing that vision incorporates intuitive physics as a rapid, spontaneous, and stimulus-driven process. We review the computational framework of physics-based analysis by synthesis, which suggests that the visual perception of intuitive physics corresponds to building and manipulating approximate structure-preserving representations of physical scenes. We end with an outline of an integrative program of computational modeling work, along with neuroscientific and psychophysical studies, toward psychologically and neurally refined mechanistic accounts of the visual perception of intuitive physics.Updated on July 8, 2026.
Visual processing seems specialized for perceiving other agents,1,2 as in biological motion perception: displays of surprisingly few moving dots ("point-light walkers"; PLWs) give rise to rich percepts of locomoting agents.3 Despite hundreds of experiments over decades of research,4,5,6 a foundational question remains unanswered: how specific are such phenomena to biological motion? Addressing this question has been historically challenging, largely due to the absence of non-biological comparison stimuli-since the translation or rotation of rigid objects (as in "structure-from-motion" displays7,8) lacks the rich relative motion of the points that is characteristic of PLWs. Here, to fill this gap, we introduce the perception of rich behavior from dynamic point-light cloths (PLCs)-as when a ribbon (or a sheet on a clothesline) is waving in the wind. Across 13 preregistered experiments, while focusing on several of the most foundational properties of biological motion, we found broad similarities between PLWs and PLCs-in terms of both experimental results and phenomenological demonstrations: percepts from PLCs (1) arise spontaneously and robustly, even in dynamic noise; (2) depend on cohesive relative motion, since they disappear both in static displays and in dynamic displays with spatially scrambled points; (3) do not require consistent local motion, since they persist in "limited-lifetime" displays; and (4) extend to rich secondary properties such as a fabric's stiffness. These results collectively demonstrate that, even beyond agents and biology, the visual system extracts rich structure from surprisingly limited input.
We are intimately familiar with liquids in our visual experience, yet the computational basis of liquid perception remains underexplored. This is an important knowledge gap because liquids, with their mutable shapes and complex intrinsic dynamics, differ remarkably from the commonly studied categories in computational vision, such as rigid objects or non-rigid solids. To understand the computational basis of liquid perception, we implemented different models of this ability and tested them in a new behavioral study. The models realize two distinct theoretical possibilities for the visual perception of liquid viscosity. The first possibility, and the focus of most existing work, explains the representation of liquid viscosity as a consequence of high-level image and motion statistics discriminative of the gradations of this physical property. A second, much different possibility is that the perceptual representations of liquids functionally map the physical processes of how viscosity and external forces (e.g., gravity, rigid surfaces) shape the way liquids move. We task these models and humans in a new behavioral task: making similarity judgments of liquid viscosity across pairs of animations depicting qualitatively different scenarios - e.g., a metal ball falling into a liquid container vs. liquid pouring over a non-flat surface. We find that a new model, Ripple, which builds and manipulates physics-based representations of liquid viscosity from sensory inputs, explains substantial variance in human judgments beyond powerful, previously behaviorally validated, statistical representations of viscosity. Moreover, statistical representations of viscosity across vastly different model architectures - a task-specific DNN and a general video foundation model - converge with one another, while remaining equally differentiated from Ripple. These results suggest that liquid perception extends beyond image statistics to also involve simulation-based intuitive physics.
Diffusion-based image generation models such as DALL-E 3 and Stable Diffusion-XL demonstrate remarkable capabilities in generating images with realistic and unique compositions. Yet, these models are not robust in precisely reasoning about physical and spatial configurations of objects, especially when instructed with unconventional, thereby out-of-distribution descriptions, such as "a chair with five legs". In this paper, we propose a language agent with chain-of-3D-thoughts (L3GO), an inference-time approach that can reason about part-based 3D mesh generation of unconventional objects that current data-driven diffusion models struggle with. More concretely, we use large language models as agents to compose a desired object via trial-and-error within the 3D simulation environment. To facilitate our investigation, we develop a new benchmark, Unconventionally Feasible Objects (UFO), as well as SimpleBlenv, a wrapper environment built on top of Blender where language agents can build and compose atomic building blocks via API calls. Human and automatic GPT-4V evaluations show that our approach surpasses the standard GPT-4 and other language agents (e.g., ReAct and Reflexion) for 3D mesh generation on ShapeNet. Moreover, when tested on our UFO benchmark, our approach outperforms other state-of-the-art text-to-2D image and text-to-3D models based on human evaluation.
How do our goals continually impact perceptual processing? The answer could arise from a computational specification of perception in terms of visual tasks, or perhaps several mechanisms operating over specific contexts. Here, we suggest an alternative: adaptive computation, a new algorithmic account of attention that rations the general resource of perceptual computations according to their impact on decision making.
A key role for attention is to continually focus visual processing to satisfy our goals. How does this work in computational terms? Here we introduce adaptive computation-a new computational mechanism of human attention that bridges the momentary application of perceptual computations with their impact on decision outcomes. Adaptive computation is a dynamic algorithm that rations perceptual computations across objects on-the-fly, enabled by a novel and general formulation of task relevance. We evaluate adaptive computation in a case study of multiple object tracking (MOT)-a paradigmatic example of selection as a dynamic process, where observers track a set of target objects moving amid visually identical distractors. Adaptive computation explains the attentional dynamics of object selection with unprecedented depth. It not only recapitulates several classic features of MOT (e.g., trial-level tracking accuracy and localization error of targets), but also captures properties that have not previously been measured or modeled-including both the subsecond patterns of attentional deployment between objects, and the resulting sense of subjective effort. Critically, this approach captures such data within a framework that is in-principle domain-general, and, unlike past models, without using any MOT-specific heuristic components. Beyond this case study, we also look to the future, discussing how adaptive computation may apply more generally, providing a new type of mechanistic model for the dynamic operation of many forms of visual attention. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Computational explorations of human cognition have been especially successful when applied to visual perception. Existing models have primarily focused on rigid objects, emphasizing shape-preserving invariance to changes in viewpoint, lighting, object size, and scene context. Yet many objects in our everyday environments, such as cloths, are soft. This poses both quantitatively greater and qualitatively different challenges for models of perception, due to soft objects' dynamic and high-dimensional internal structure, as in the changing folds and wrinkles of a cloth waving in the wind. Soft object perception is also correspondingly rich, involving distinct properties such as stiffness. Here we explore the ability of different kinds of computational models to capture visual perception of the physical properties of cloths (e.g., their degrees of stiffness) undergoing different naturalistic transformations (e.g., falling vs. waving in the wind). Across visual matching tasks, both the successes and failures of human performance are well explained by Woven: a new model that incorporates physics-based simulations to infer probabilistic representations of cloths. Woven outperforms powerful, performance-equated alternatives, including its ablations and a deep neural network, and suggests that humanlike machine vision may also require representations that transcend image statistics, and involve intuitive physics.
Stimulus-driven, multiarea processing in the inferotemporal (IT) cortex is thought to be critical for transforming sensory inputs into useful representations of the world. What are the formats of these neural representations and how are they computed across the nodes of the IT networks? A growing literature in computational neuroscience focuses on the computational-level objective of acquiring high-level image statistics that supports useful distinctions, including between object identities or categories. Here, inspired by classic theories of vision, we suggest an alternative possibility. We show that inferring 3D objects may be a distinct computational-level objective of IT, implemented via an algorithm analogous to graphics-based generative models of how 3D scenes form and project to images, but in the reverse order. Using perception of bodies as a case study, we show that inverse graphics spontaneously emerges in inference networks trained to map images to 3D objects. Remarkably, this correspondence to the reverse of a graphics-based generative model also holds across the body processing network of the macaque IT cortex. Finally, inference networks recapitulate the feedforward progression across the stages of this IT network and do so better than the currently dominant vision models, including both supervised and unsupervised variants, none of which aligns with the reverse of graphics. This work suggests inverse graphics as a multiarea neural algorithm implemented within IT, and points to ways for replicating primate vision capabilities in machines.
Vision is widely understood as an inference problem. However, two contrasting conceptions of the inference process have each been influential in research on biological vision as well as the engineering of machine vision. The first emphasizes bottom-up signal flow, describing vision as a largely feedforward, discriminative inference process that filters and transforms the visual information to remove irrelevant variation and represent behaviorally relevant information in a format suitable for downstream functions of cognition and behavioral control. In this conception, vision is driven by the sensory data, and perception is direct because the processing proceeds from the data to the latent variables of interest. The notion of "inference" in this conception is that of the engineering literature on neural networks, where feedforward convolutional neural networks processing images are said to perform inference. The alternative conception is that of vision as an inference process in Helmholtz's sense, where the sensory evidence is evaluated in the context of a generative model of the causal processes giving rise to it. In this conception, vision inverts a generative model through an interrogation of the evidence in a process often thought to involve top-down predictions of sensory data to evaluate the likelihood of alternative hypotheses. The authors include scientists rooted in roughly equal numbers in each of the conceptions and motivated to overcome what might be a false dichotomy between them and engage the other perspective in the realm of theory and experiment. The primate brain employs an unknown algorithm that may combine the advantages of both conceptions. We explain and clarify the terminology, review the key empirical evidence, and propose an empirical research program that transcends the dichotomy and sets the stage for revealing the mysterious hybrid algorithm of primate vision.
Decades of research document an association between neurocognitive dysfunction and externalizing behaviors, including rule-breaking, aggression, and impulsivity. However, there has been very little work that examines how multiple neurocognitive functions co-occur within individuals and which combinations of neurocognitive functions are most relevant for externalizing behaviors. Moreover, Latent Profile Analysis (LPA), a widely used method for grouping individuals in person-centered analysis, often struggles to balance the tradeoff between good model fit (splitting participants into many latent profiles) and model interpretability (using only a few, highly distinct latent profiles). To address these problems, we implemented a non-parametric Bayesian form of LPA based on the Dirichlet process mixture model (DPM-LPA) and used it to study the relationship between neurocognitive functioning and externalizing behaviors in adolescents participating in the Adolescent Brain Cognitive Development Study. First, we found that DPM-LPA outperformed conventional LPA, revealing more distinct profiles and classifying participants with higher certainty. Second, latent profiles extracted from DPM-LPA were differentially related to externalizing behaviors: profiles with deficits in working memory, inhibition, and/or language abilities were robustly related to different expressions of externalizing. Together, these findings represent a step towards addressing the challenge of finding novel ways to use neurocognitive data to better describe the individual. By precisely identifying and specifying the variation in neurocognitive and behavioral patterns this work offers an innovative empirical foundation for the development of assessments and interventions that address these costly behaviors.
Ensuring robustness of image classifiers against adversarial attacks and spurious correlation has been challenging. One of the most effective methods for adversarial robustness is a type of data augmentation that uses adversarial examples during training. Here, inspired by computational models of human vision, we explore a synthesis of this approach by leveraging a structured prior over image formation: the 3D geometry of objects and how it projects to images. We combine adversarial training with a weight initialization that implicitly encodes such a prior about 3D objects via 3D reconstruction pre-training. We evaluate our approach using two different datasets and compare it to alternative pre-training protocols that do not encode a prior about 3D shape. To systematically explore the effect of 3D pre-training, we introduce a novel dataset called Geon3D, which consists of simple shapes that nevertheless capture variation in multiple distinct dimensions of geometry. We find that while 3D reconstruction pre-training does not improve robustness for the simplest dataset setting, we consider (Geon3D on a clean background) that it improves upon adversarial training in more realistic (Geon3D with textured background and ShapeNet) conditions. We also find that 3D pre-training coupled with adversarial training improves the robustness to spurious correlations between shape and background textures. Furthermore, we show that the benefit of using 3D-based pre-training outperforms 2D-based pre-training on ShapeNet. We hope that these results encourage further investigation of the benefits of structured, 3D-based models of vision for adversarial robustness.