A key to interacting with the physical world is the ability to visually infer object properties, like elasticity, allowing us to anticipate object behaviour. Such perceptual inferences continue to challenge artificial intelligence systems, highlighting the complexity of the underlying computations. How does the human brain solve this task? Here, we propose a resource-rational model based on learned statistics of object motion to explain how humans judge elasticity. We created 100 000 physics-based simulations of bouncing cubes with different elasticities and found that even tiny changes in initial conditions (e.g. orientation) yield starkly different trajectories. Yet, across these simulations, we identified 23 motion features that capture natural variations in elasticity. Although a weighted combination of these features reliably predicts physical elasticity, surprisingly, humans do not seem to employ cue combination when judging elasticity. Instead, they switch between different cues. A series of experiments designed to carefully tease apart several competing heuristics suggests that observers switch between different computationally efficient yet informative cues depending on the information available in the stimulus.
Extensive prior work has identified regions of the human brain associated with visual perception of objects (lateral occipital complex [LOC]) and their physical properties and interactions ("frontoparietal physics network" [FPN]). However, this work has nearly exclusively tested the response of these regions to rigid objects. Deformable or nonsolid substances, or "stuff," including liquids such as water or honey and granular materials such as sand or snow, are of similar importance in everyday life but have different physical properties and invite different actions. Little is known about the brain basis of stuff perception. Here, we scan participants with functional MRI (fMRI) while they view videos of rigid and non-rigid objects ("things") and liquid and granular substances (stuff). We find double dissociations between the processing of things and stuff within both the ventral and dorsal visual pathways. These findings suggest that distinct mental algorithms are engaged when we perceive things and stuff, as they are in artificial physics engines.
Imagine pouring a box of granola into a bowl. Are you considering hundreds of individual chunks or the motion of the group as a whole? Human perceptual limits suggest we cannot be representing the individuals, implying we simulate ensembles of objects. If true, we would need to represent group physical properties beyond individual aggregates, similar to perceiving ensemble properties like color, size, or facial expression. Here we investigate whether people do hold ensemble representations of mass, using tasks in which participants watch a video of a single marble or set of marbles falling onto an elastic cloth and judge the individual or average mass. We find first that people better judge average masses than individual masses, then find evidence that the better ensemble judgments are not just due to aggregating information from individual marbles. Together, this supports the concept of ensemble perception in intuitive physics, extending our understanding of how people represent and simulate sets of objects.
The ability to build and reason about models of the world is essential for situated language understanding. But evaluating world modeling capabilities in modern AI systems-especially those based on language models-has proven challenging, in large part because of the difficulty of disentangling conceptual knowledge about the world from knowledge of surface co-occurrence statistics. This paper presents Elements of World Knowledge (EWOK), a framework for evaluating language models' understanding of the conceptual knowledge underlying world modeling. EWOK targets specific concepts from multiple knowledge domains known to be important for world modeling in humans, from social interactions (help, deceive) to spatial relations (left, right). Objects, agents, and locations in the items can be flexibly filled in, enabling easy generation of multiple controlled datasets. We then introduce EWOK-CORE-1.0, a dataset of 4,374 items covering 11 world knowledge domains. We evaluate 20 open-weights large language models (1.3B-70B parameters) and compare them with human performance. All tested models perform worse than humans, with results varying drastically across domains. Performance on social interactions and social properties was highest and performance on physical relations and spatial relations was lowest. Overall, this dataset highlights simple cases where even large models struggle and presents rich avenues for targeted research on LLM world modeling capabilities.
In a seminal paper published two decades ago, Adelson (2001) noted that "Our world contains both things and stuff, but things tend to get the attention." This remains the case today in the field of cognitive neuroscience. The large number of publications using fMRI to explore the lateral occipital complex (LOC) have focused almost exclusively on the role of this region in extracting the 3D shape of Things, without asking whether this region may also respond to Stuff with no fixed shape like honey, sand, or water. Similarly, investigations of the "physics network" previously implicated in visual intuitive physics (Fischer et al, 2016) have to date tested only Things, even though the physics of Stuff plays a comparable role in everyday life. Here, we asked whether LOC and the physics network are engaged when observing Stuff. We created 120 photorealistic short movie clips of four different computer-simulated substances—liquids and granular Stuff, and non-rigid and rigid Things—interacting with other objects, e.g., colliding with obstacles. The four types of videos, as well as scrambled versions of each, were presented in a blocked fMRI design while subjects (N=6) performed an orthogonal color-change detection task. Independently-localized LOC and the physics network showed higher activation for all materials than for scrambled controls (p < .05), whereas the opposite pattern was found for V1 (p < .05). Most importantly, we find that the physics network responded more to rigid and non-rigid Things than liquid and granular Stuff (p < .05), whereas LOC responded at least as strongly to Stuff as to Things. These findings suggest that the physics network may be more engaged in the physics of Things than Stuff, whereas LOC is not restricted to extracting the fixed 3D shape of Things but is equally engaged by Stuff with dynamically changing shapes.
Selecting suitable grasps on three-dimensional objects is a challenging visuomotor computation, which involves combining information about an object (e.g., its shape, size, and mass) with information about the actor's body (e.g., the optimal grasp aperture and hand posture for comfortable manipulation). Here, we used functional magnetic resonance imaging to investigate brain networks associated with these distinct aspects during grasp planning and execution. Human participants of either sex viewed and then executed preselected grasps on L-shaped objects made of wood and/or brass. By leveraging a computational approach that accurately predicts human grasp locations, we selected grasp points that disentangled the role of multiple grasp-relevant factors, that is, grasp axis, grasp size, and object mass. Representational Similarity Analysis revealed that grasp axis was encoded along dorsal-stream regions during grasp planning. Grasp size was first encoded in ventral stream areas during grasp planning then in premotor regions during grasp execution. Object mass was encoded in ventral stream and (pre)motor regions only during grasp execution. Premotor regions further encoded visual predictions of grasp comfort, whereas the ventral stream encoded grasp comfort during execution, suggesting its involvement in haptic evaluation. These shifts in neural representations thus capture the sensorimotor transformations that allow humans to grasp objects.SIGNIFICANCE STATEMENT Grasping requires integrating object properties with constraints on hand and arm postures. Using a computational approach that accurately predicts human grasp locations by combining such constraints, we selected grasps on objects that disentangled the relative contributions of object mass, grasp size, and grasp axis during grasp planning and execution in a neuroimaging study. Our findings reveal a greater role of dorsal-stream visuomotor areas during grasp planning, and, surprisingly, increasing ventral stream engagement during execution. We propose that during planning, visuomotor representations initially encode grasp axis and size. Perceptual representations of object material properties become more relevant instead as the hand approaches the object and motor programs are refined with estimates of the grip forces required to successfully lift the object.
Vision is more than object recognition: In order to interact with the physical world, we estimate object properties such as mass, fragility, or elasticity by sight. The computational basis of this ability is poorly understood. Here, we propose a model based on the statistical appearance of objects, i.e., how they typically move, flow, or fold. We test this idea using a particularly challenging example: estimating the elasticity of bouncing objects. Their complex movements depend on many factors, e.g., elasticity, initial speed, and direction, and thus every object can produce an infinite number of different trajectories. By simulating and analyzing the trajectories of 100k bouncing cubes, we identified and evaluated 23 motion features that could individually or in combination be used to estimate elasticity. Experimentally teasing apart these competing but highly correlated hypotheses, we found that humans represent bouncing objects in terms of several different motion features but rely on just a single one when asked to estimate elasticity. Which feature this is, is determined by the stimulus itself: Humans rely on the duration of motion if the complete trajectory is visible, but on the maximal bounce height if the motion duration is artificially cut short. Our results suggest that observers take into account the computational costs when asked to judge elasticity and thus rely on a robust and efficient heuristic. Our study provides evidence for how such a heuristic can be derived—in an unsupervised manner—from observing the natural variations in many exemplars. Significance Statement How do we perceive the physical properties of objects? Our findings suggest that when tasked with reporting the elasticity of bouncing cubes, observers rely on simple heuristics. Although there are many potential visual cues, surprisingly, humans tend to switch between just a handful of them depending on the characteristics of the stimulus. The heuristics predict not only the broad successes of human elasticity perception but also the striking pattern of errors observers make when we decouple the cues from ground truth. Using a big data approach, we show how the brain could derive such heuristics by observation alone. The findings are likely an example of ‘computational rationality’, in which the brain trades off task demands and relative computational costs.
Primates must accurately estimate the size of objects in their environment to interact with them efficiently. Hong, Yamins, Majaj, and DiCarlo (2016) reported that one could accurately approximate an object's size within an image from the population activity across the macaque inferior temporal (IT) cortex upon brief (100 ms) image presentations. These neural predictions were consistent with human behavioral estimates of object size within the same images—suggesting a linear IT readout model as the leading neural decoding hypothesis for object size estimation in primates. However, perceived and image-based (i.e., retinal) object sizes were highly correlated in the Hong et al. (2016) study. Notably, two objects with identical retinal sizes may be perceived to differ in size when embedded at different locations along a linear perspective (Ponzo illusion). Therefore, this size illusion allows us to perform a stronger test and assess whether the IT-based linear readout model predicts the perceived or retinal size. We created a set of image pairs by placing objects "near" or "far" with respect to a linear perspective background. We performed large-scale neural recordings (2 Utah arrays; n=192 sites) across the macaque IT cortex while the monkey fixated the images for 100 ms. Extending the results of Hong et al. (2016), we observed that approximations of object sizes from the IT responses (~190-205 ms) showed a significant bias ("far"-"near"; Δ=22%; p<0.001) that is qualitatively similar to that measured behaviorally in humans. Interestingly, however, most deep convolutional neural network (DCNN) models that so far best approximate the primate IT responses failed to demonstrate such a bias. Together, our results provide further support for the linear IT readout model of object size perception while exposing a significant explanatory gap in current DCNNs as models of primate vision.
Common everyday materials such as textiles, foodstuffs, soil or skin can have complex, mutable and varied appearances. Under typical viewing conditions, most observers can visually recognize materials effortlessly, and determine many of their properties without touching them. Visual material perception raises many fascinating questions for vision researchers, neuroscientists and philosophers, yet has received little attention compared to the perception of color or shape. Here we discuss some of the challenges that material perception raises and argue that further philosophical thought should be directed to how we see materials.
How humans visually select where to grasp objects is determined by the physical object properties (e.g., size, shape, weight), the degrees of freedom of the arm and hand, as well as the task to be performed. We recently demonstrated that human grasps are near-optimal with respect to a weighted combination of different cost functions that make grasps uncomfortable, unstable, or impossible, e.g., due to unnatural grasp apertures or large torques. Here, we ask whether humans can consciously access these rules. We test if humans can explicitly judge grasp quality derived from rules regarding grasp size, orientation, torque, and visibility. More specifically, we test if grasp quality can be inferred (i) by using visual cues and motor imagery alone, (ii) from watching grasps executed by others, and (iii) through performing grasps, i.e., receiving visual, proprioceptive and haptic feedback. Stimuli were novel objects made of 10 cubes of brass and wood (side length 2.5 cm) in various configurations. On each object, one near-optimal and one sub-optimal grasp were selected based on one cost function (e.g., torque), while the other constraints (grasp size, orientation, and visibility) were kept approximately constant or counterbalanced. Participants were visually cued to the location of the selected grasps on each object and verbally reported which of the two grasps was best. Across three experiments, participants were required to either (i) passively view the static objects and imagine executing the two competing grasps, (ii) passively view videos of other participants grasping the objects, or (iii) actively grasp the objects themselves. Our results show that, for a majority of tested objects, participants could already judge grasp optimality from simply viewing the objects and imagining to grasp them, but were significantly better in the video and grasping session. These findings suggest that humans can determine grasp quality even without performing the grasp—perhaps through motor imagery—and can further refine their understanding of how to correctly grasp an object through sensorimotor feedback but also by passively viewing others grasp objects.
To catch or avoid an object it is crucial to predict its future trajectory. Here, we test how well observers can predict the landing position of a bouncing object and identify the strategies they employ to solve this task. A large database (N=100,000) of short simulations of a cube bouncing in a 3-dimensional room allows us to measure the cube’s typical behavior, while its elasticity, initial orientation, position and velocity are varied. We randomly selected and rendered 240 animations from this database as stimuli. Fourteen observers saw either the first 10, 20, 30, 40, or 50 frames of each animation and had to predict where the cube would eventually come to rest. The number of remaining frames in which the cube was still moving (but that were not presented) varied between 10-107 frames. Observers responded by moving a marker to the predicted position on the floor, indicated their certainty by adjusting its radius, and then saw the remaining frames for feedback. We found that observers make systematic predictions of the cube’s final position and are better than chance for predictions of up to about 60 frames. Unsurprisingly, observers were more accurate the fewer frames (and thus fewer bounces) they had to predict. For their predictions, observers took into account the final (i.e., the last visible) direction along which the cube was moving and often predicted the landing position to be in a similar direction. Thus, the closer the trajectory was to a straight line after disappearance, the more accurate they were. Furthermore, observers predicted longer paths for cubes that moved faster before they disappeared. Presumably, observers use a simplified mental simulation strategy, which—because it is physically inaccurate—accumulates errors over time and therefore produces a larger divergence from the computer-simulated landing position with a growing number of simulation steps.
Visually inferring the elasticity of a bouncing object poses a challenge to the visual system: The observable behavior of the object depends on its elasticity but also on extrinsic factors, such as its initial position and velocity. Estimating elasticity requires disentangling these different contributions to the observed motion. We created 2-second simulations of a cube bouncing in a room and varied the cube's elasticity in 10 steps. The cube's initial position, orientation, and velocity were varied randomly to gain three random samples for each level of elasticity. We systematically limited the visual information by creating three versions of each stimulus: (a) a full rendering of the scene, (b) the cube in a completely black environment, and (c) a rigid version of the cube following the same trajectories but without rotating or deforming (also in a completely black environment). Thirteen observers rated the apparent elasticity of the cubes and the typicality of their motion. Generally, stimuli were judged as less typical if they showed rigid motion without rotations, highly elastic cubes, or unlikely events. Overall, elasticity judgments correlated strongly with the true elasticity but did not show perfect constancy. Yet, importantly, we found similar results for all three stimulus conditions, despite significant differences in their apparent typicality. This suggests that the trajectory alone contains the information required to make elasticity judgments.
We rarely experience difficulty picking up objects, yet of all potential contact points on the surface, only a small proportion yield effective grasps. Here, we present extensive behavioral data alongside a normative model that correctly predicts human precision grasping of unfamiliar 3D objects. We tracked participants' forefinger and thumb as they picked up objects of 10 wood and brass cubes configured to tease apart effects of shape, weight, orientation, and mass distribution. Grasps were highly systematic and consistent across repetitions and participants. We employed these data to construct a model which combines five cost functions related to force closure, torque, natural grasp axis, grasp aperture, and visibility. Even without free parameters, the model predicts individual grasps almost as well as different individuals predict one another's, but fitting weights reveals the relative importance of the different constraints. The model also accurately predicts human grasps on novel 3D-printed objects with more naturalistic geometries and is robust to perturbations in its key parameters. Together, the findings provide a unified account of how we successfully grasp objects of different 3D shape, orientation, mass, and mass distribution.
Humans strongly rely on vision to guide grasping. Visual grasp selection is highly systematic and consistent across repetitions and participants, suggesting that humans employ a common set of constraints when visually selecting grasps. We formalized these constraints as a set of grasp-cost functions related to torque, grasp axis, grasp aperture, and object visibility, which we have shown predict grasping behavior with striking fidelity. Here, we test if humans can explicitly estimate grasp optimality derived from these grasp-cost functions. We additionally ask whether vision alone is sufficient to compute grasp optimality, or whether sensorimotor feedback is required to link vision to action selection. Stimuli were novel objects made of 10 cubes of brass and wood (side length 2.5 cm) in various configurations. On each object, an optimal and a sub-optimal grasp were selected based on one of the cost functions, while cost for the other constraints was kept approximately constant or counterbalanced. Participants were visually cued as to the location of the grasps on each object via colored markers. In a vision-only session, participants were required to judge which of the two grasps they believed to be better, without ever having grasped the object. In a vision-plus-grasp session, participants were required to attempt both grasps on each object, and again indicate which of the two grasps they judged to be better. Participants (N=11) were already able to judge grasp optimality above chance in the vision-only session (65+/−13% correct, p= 0.0035). Additionally, participants were significantly better at judging grasp optimality in the vision-plus-grasp session (77+/−7% correct, p=0.0081). Together, these findings show that humans can consciously access the visuomotor computations underlying grasp selection, and highlight the fundamental role of sensorimotor feedback in linking visual perception to motor control.
When haptically exploring softness, humans use higher peak forces when indenting harder versus softer objects. Here, we investigated the influence of different channels and types of prior knowledge on initial peak forces. Participants explored two stimuli (hard vs. soft) and judged which was softer. In Experiment 1 participants received either semantic (the words "hard" and "soft"), visual (video of indentation), or prior information from recurring presentation (blocks of harder or softer pairs only). In a control condition no prior information was given (randomized presentation). In the recurring condition participants used higher initial forces when exploring harder stimuli. No effects were found in control and semantic conditions. With visual prior information, participants used less force for harder objects. We speculate that these findings reflect differences between implicit knowledge induced by recurring presentation and explicit knowledge induced by visual and semantic information. To test this hypothesis, we investigated whether explicit prior information interferes with implicit information in Experiment 2. Two groups of participants discriminated softness of harder or softer stimuli in two conditions (blocked and randomized). The interference group received additional explicit information during the blocked condition; the implicit-only group did not. Implicit prior information was only used for force adaptation when no additional explicit information was given, whereas explicit interfered with movement adaptation. The integration of prior knowledge only seems possible when implicit prior knowledge is induced-not with explicit knowledge.
Humans exhibit spatial biases when grasping objects. These biases may be due to actors attempting to shorten their reaching movements and therefore minimize energy expenditures. An alternative explanation could be that they arise from actors attempting to minimize the portion of a grasped object occluded from view by the hand. We reanalyze data from a recent study, in which a key condition decouples these two competing hypotheses. The analysis reveals that object visibility, not energy expenditure, most likely accounts for spatial biases observed in human grasping.
Non-rigid objects deform and move in response to external forces. To estimate an object’s softness or elasticity, the visual system has to rapidly disentangle multiple causal contributions. For example, an object deforms strongly either because it is very soft or because the applied force is very large. To investigate how the brain solves this, we simulated and rendered 20 short animations of rigid objects interacting with a non-rigid target. We varied the external force as well as the target’s softness and elasticity and had 15 observers rate the two internal properties. Despite large stimulus variations across external variations, responses were broadly in accordance with the simulated internal properties. However, objects that deformed permanently (i.e. not elastic) were rated as softer. We characterized the visual features of the objects by measuring the deformation, wobbliness, external motion and movement duration of the underlying 3D-meshes. A linear combination of these four features predicts softness perception very well. Next, we simulated over 200.000 animations, massively increasing the variations of internal and external factors. Ten observers rated the softness and elasticity of a small subset of animations. We measured the four features in all 200.000 simulations and fitted a linear regression in order to learn the mappings between visual features and physical material properties. Although the weighted combination of features predicts the physical properties only moderately, the same weights (i.e. without fitting to the perceptual data) predict perceived material properties strikingly well and can account for the perceptual influence of elasticity on softness.