The ability to infer the past from what we perceive in the present is a key capacity of human cognition. Witnessing a broken vase, humans will automatically bring to mind a causal story of what happened. Multiple sources of sensory evidence can support this inference. Seeing the broken pottery tells you something, but hearing the crash tells you even more. In this work, we explore people's inferences about the past from multimodal evidence. We present a physical reasoning paradigm called Plinko. In the prediction task, participants must determine where a ball that is dropped into a box with obstacles will land. A computational model that uses mental simulation in an Intuitive Physics Engine captures participant predictions very well. In the inference task, participants must infer which hole in the box the ball fell from. Across conditions, participants are presented with different combinations of visual and auditory cues, and must combine this information to determine what happened. We develop a sequential sampling model that selectively simulates from promising hypotheses, and demonstrate that this model accurately captures participants' judgments and eye-movements. By coordinating sensory evidence in an underlying causal representation of the physical world, this simulation approach is able to capture complex multimodal inferences that go beyond traditional approaches to multimodal integration.
People regularly make inferences about objects in the world that they cannot see by flexibly integrating information from multiple sources: auditory and visual cues, language, and our prior beliefs and knowledge about the scene. How are we able to so flexibly integrate many sources of information to make sense of the world around us, even if we have no direct knowledge? In this work, we propose a neurosymbolic model that uses neural networks to parse open-ended multimodal inputs and then applies a Bayesian model to integrate different sources of information to evaluate different hypotheses. We evaluate our model with a novel object guessing game called “What's in the Box?” where humans and models watch a video clip of an experimenter shaking boxes and then try to guess the objects inside the boxes. Through a human experiment, we show that our model correlates strongly with human judgments, whereas unimodal ablated models and large multimodal neural model baselines show poor correlation.
Having an internal model of one's attention can be useful for effectively managing limited perceptual and cognitive resources. While previous work has hinted at the existence of an internal model of attention, it is still unknown how rich and flexible this model is, whether it corresponds to one's own attention or to a generic person-invariant schema, and whether it is specified as a list of facts and rules or alternatively as a probabilistic simulation model. To this end, we tested participants' ability to estimate their own behavior in a visual search task with novel displays. In six online experiments (four pre-registered), prospective search time estimates reflected accurate metacognitive knowledge of key findings in the visual search literature, including the set-size effect, higher efficiency of color over conjunction search, and the asymmetric contributions of target and distractor identities to search difficulty. In contrast, estimates were biased to assume serial search, and demonstrated little to no insight into sizeable effects of search asymmetries for basic visual features, and of target-distractor similarity. Together, our findings reveal a complex picture, where internal models of visual search are sensitive to some, but not all, of the factors that make some searches more difficult than others. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Many surface cues support three-dimensional shape perception, but people can sometimes still see shape when these features are missing -- in extreme cases, even when an object is completely occluded, as when covered with a draped cloth. We propose a framework for 3D shape perception that explains perception in both typical and atypical cases as analysis-by-synthesis, or inference in a generative model of image formation: the model integrates intuitive physics to explain how shape can be inferred from deformations it causes to other objects, as in cloth-draping. Behavioral and computational studies comparing this account with several alternatives show that it best matches human observers in both accuracy and response times, and is the only model that correlates significantly with human performance on difficult discriminations. Our results suggest that bottom-up deep neural network models are not fully adequate accounts of human shape perception, and point to how machine vision systems might achieve more human-like robustness.
Effective curiosity-driven learning requires recognizing that the value of evidence for testing hypotheses depends on what other hypotheses are under consideration. Do we intuitively represent the discriminability of hypotheses? Here we show children alternative hypotheses for the contents of a box and then shake the box (or allow children to shake it themselves) so they can hear the sound of the contents. We find that children are able to compare the evidence they hear with imagined evidence they do not hear but might have heard under alternative hypotheses. Children (N = 160; mean: 5 years and 4 months) prefer easier discriminations (Experiments 1-3) and explore longer given harder ones (Experiments 4-7). Across 16 contrasts, children’s exploration time quantitatively tracks the discriminability of heard evidence from an unheard alternative. The results are consistent with the idea that children have an “intuitive psychophysics”: children represent their own perceptual abilities and explore longer when hypotheses are harder to distinguish.
Author(s): Friedman, Yoni; Le, Tuan Anh; Egger, Bernhard; Siegel, Max; Tenenbaum, Josh | Abstract: Humans perceive the world through a rich, object-centric lens. We are able to infer 3D geometry and features of objects from sparse and noisy data. Gestalt rules describe how perceptual stimuli tend to be grouped based on properties like proximity, closure, and continuity. However, it remains an open question how these mechanisms are implemented algorithmically in the brain, and how (or why) they functionally support 3D object perception. Here, we describe a computational model which accounts for the Gestalt principle of Common Fate - grouping stimuli by shared motion statistics. We argue that this mechanism can be explained as bottom-up neural amortized inference in a top-down generative model for object-based scenes. Our generative model places a low-dimensional prior on the motion and shape of objects, while our inference network learns to group feature clusters using inverse renderings of noisily textured objects moving through time, effectively enabling 3D shape perception.
How might we explain these flexible, seemingly effortless judgments? This chapter pre sents an answer centered at the notion of physical object repre sen tations (PORs), a basic system of knowledge that supports perceiving, learning, and reasoning about all the objects in our environment— their shapes, appearances, affordances, substances, and the way they react to forces applied to them. Our goal here is to outline a computational framework for studying the form and content of PORs in the mind and brain. PORs can be considered an interface between perception and cognition, linking what we perceive to how we plan our actions and talk about the world. Despite their fundamental role in perception, many impor tant questions about object repre sen ta tions remain open. What kind of information formats or data structures underlie PORs so as to support the many ways in which humans flexibly and creatively interact with the world? How can properties of objects be inferred from sensory inputs, and how are they represented in neural cir cuits? How can these repre sen ta tions integrate sense data across vision, touch, and audition? After introducing the computational ingredients of POR theory from a reverseengineering perspective, we review recent work that is beginning to answer some of these questions. We focus on three case studies: (1) how PORs can explain human judgments in intuitive physics, across a broad range of physical outcome prediction scenarios; (2) how PORs provide a substrate for physically mediated object shape perception in scenarios where traditional visual cues fail and a natu ral substrate for multimodal (visualhaptic) perception and crossmodal transfer; and (3) how in one domain of highlevel vision— face perception— PORs might be computed by neural cir cuits, and how thinking in terms of PORs suggests a new way to interpret multiple stages of pro cessing in the primate brain.
Humans can successfully interpret images even when they have been distorted by significant image transformations. Such images could aid in differentiating proposed computational architectures for perception because while all proposals predict similar results for typical stimuli (good performance), they differ when confronting atypical stimuli. Here we study two classes of degraded stimuli – Mooney faces and silhouettes of faces – as well as typical faces, in humans and several computational models, with the goal of identifying divergent predictions among the models, evaluating against human judgments, and ultimately informing models of human perception. We find that our top-down inverse rendering model better matches human percepts than either an invariance-based account implemented in a deep neural network, or a neural network trained to perform approximate inverse rendering in a feedforward circuit. 1156 ©2020 The Author(s). This work is licensed under a Creative Commons Attribution 4.0 International License (CC BY).
Humans possess the unique ability of combinatorial generalization in auditory perception: given novel auditory stimuli, humans perform auditory scene analysis and infer causal physical interactions based on prior knowledge. Could we build a computational model that achieves combinatorial generalization? In this paper, we present a case study on box-shaking: having heard only the sound of a single ball moving in a box, we seek to interpret the sound of two or three balls of different materials. To solve this task, we propose a hybrid model with two components: a neural network for perception, and a physical audio engine for simulation. We use the outcome of the network as an initial guess and perform MCMC sampling with the audio engine to improve the result. Combining neural networks with a physical audio engine, our hybrid model achieves combinatorial generalization efficiently and accurately in auditory scene perception.