Over the last few decades, psychologists have developed precise quantitative models of human recall performance in visual working memory (VWM) tasks. However, these models are tailored to a particular class of artificial stimulus displays and simple feature reports from participants (e.g., the color or orientation of a simple object). Our work has two aims. The first is to build models that explain people’s memory errors in continuous report tasks with natural images. Here, we use image generation algorithms to generate continuously varying response alternatives that differ from the stimulus image in natural and complex ways, in order to capture the richness of people’s stored representations. The second aim is to determine whether models that do a good job of explaining memory errors with natural images also explain errors in the more heavily studied domain of artificial displays with simple items. We find that: (i) features taken from state-of-the-art deep encoders predict trial-level difficulty in natural images better than several reasonable baselines; and (ii) the same visual encoders can reproduce set-size effects and response bias curves in the artificial stimulus domains of orientation and color. Moving forward, our approach offers a scalable way to build a more generalized understanding of VWM representations by combining recent advances in both AI and cognitive modeling.
Efficient data compression is essential for capacity-limited systems, such as biological perception and perceptual memory. We hypothesize that the need for efficient compression shapes biological systems in many of the same ways that it shapes engineered systems. If true, then the tools that engineers use to analyze and design systems, namely rate-distortion theory (RDT), can profitably be used to understand human perception and memory. The first portion of this article discusses how three general principles for efficient data compression provide accounts for many important behavioral phenomena and experimental results. We also discuss how these principles are embodied in RDT. The second portion notes that exact RDT methods are computationally feasible only in low-dimensional stimulus spaces. To date, researchers have used deep neural networks to approximately implement RDT in high-dimensional spaces, but these implementations have been limited to tasks in which the sole goal is compression with respect to reconstruction error. Here, we introduce a new deep neural network architecture that approximately implements RDT. An important property of our architecture is that it is trained "end-to-end," operating on raw perceptual input (e.g., pixel values) rather than intermediate levels of abstraction, as is the case with most psychological models. The article's final portion conjectures on how efficient compression can occur in memory over time, thereby providing motivations for multiple memory systems operating at different time scales, and on how efficient compression may explain some attentional phenomena such as RTs in visual search. (PsycInfo Database Record (c) 2020 APA, all rights reserved).
The “resource-rational” approach is ambitious and worthwhile. A shortcoming of the proposed approach is that it fails to constrain what counts as a constraint. As a result, constraints used in different cognitive domains often have nothing in common. We describe an alternative framework that satisfies many of the desiderata of the resource-rational approach, but in a more disciplined manner.
Humans can easily describe, imagine, and, crucially, predict a wide variety of behaviors of liquids-splashing, squirting, gushing, sloshing, soaking, dripping, draining, trickling, pooling, and pouring-despite tremendous variability in their material and dynamical properties. Here we propose and test a computational model of how people perceive and predict these liquid dynamics, based on coarse approximate simulations of fluids as collections of interacting particles. Our model is analogous to a "game engine in the head", drawing on techniques for interactive simulations (as in video games) that optimize for efficiency and natural appearance rather than physical accuracy. In two behavioral experiments, we found that the model accurately captured people's predictions about how liquids flow among complex solid obstacles, and was significantly better than several alternatives based on simple heuristics and deep neural networks. Our model was also able to explain how people's predictions varied as a function of the liquids' properties (e.g., viscosity and stickiness). Together, the model and empirical results extend the recent proposal that human physical scene understanding for the dynamics of rigid, solid objects can be supported by approximate probabilistic simulation, to the more complex and unexplored domain of fluid dynamics.
Human brains are finite, and thus have bounded capacity. An efficient strategy for a capacity-limited agent is to continuously adapt by dynamically reallocating capacity in a task-dependent manner. Here we study this strategy in the context of visual working memory (VWM). People use their VWM stores to remember visual information over seconds or minutes. However, their memory performances are often error-prone, presumably due to VWM capacity limits. We hypothesize that people attempt to be flexible and robust by strategically reallocating their limited VWM capacity based on two factors: (a) the statistical regularities (e.g., stimulus feature means and variances) of the to-be-remembered items, and (b) the requirements of the task that they are attempting to perform. The latter specifies, for example, which types of errors are costly versus irrelevant for task performance. These hypotheses are formalized within a normative computational modeling framework based on rate-distortion theory, an extension of conventional Bayesian approaches that uses information theory to study rate-limited (or capacity-limited) processes. Using images of plants that are naturalistic and precisely controlled, we carried out two sets of experiments. Experiment 1 found that when a stimulus dimension (the widths of plants' leaves) was assigned a distribution, subjects adapted their VWM performances based on this distribution. Experiment 2 found that when one stimulus dimension (e.g., leaf width) was relevant for distinguishing plant categories but another dimension (leaf angle) was irrelevant, subjects' responses in a memory task became relatively more sensitive to the relevant stimulus dimension. Together, these results illustrate the task-dependent robustness of VWM, thereby highlighting the dependence of memory on learning.
Liquids can splash, squirt, gush, slosh, soak, drip, drain, trickle, pool, and be poured–complex behaviors that we can easily distinguish, imagine, describe, and, crucially, predict, despite tremendous diversity among different liquids’ material and dynamical characteristics. This proficiency suggests the brain has a sophisticated cognitive mechanism for reasoning about liquids, yet to date there has been little effort to study this mechanism quantitatively or describe it computationally. Here we find evidence that people’s reasoning about how liquids move is consistent with a computational cognitive model based on approximate probabilistic simulation. In a psychophysical experiment, participants predicted how different liquids would flow around solid obstacles, and their judgments agreed with those of a family of models in which volumes of liquid are represented as collections of interacting particles, within a dynamical fluid simulation. Our model explains people’s accuracy, and their predictions’ sensitivity to liquids of different viscosity. We also explored several models that did not involve simulation, and found they could not account for the experimental data as well. Our results are consistent with previous reports that people’s physical understanding of solid objects is based on simulation, but extends this thesis to the more complex and unexplored domain of reasoning about liquids.
We describe studies of nanoparticle synthesis using oligonucleotides as capping ligands. The oligonucleotides nucleate, grow, and stabilize near-infrared fluorescent, approximately uniform PbS nanocrystals in an aqueous environment. The properties of the resulting particles strongly depend upon the sequences as well as synthesis conditions. Fourier Transform infrared measurements suggest that functional groups on the nucleobases such as carbonyl and amine moieties are responsible for surface passivation, while the phosphate backbone is strained to accommodate nucleobase bonding, preventing irreversible aggregation and thereby stabilizing the colloids. Our theoretical model indicates that oligonucleotide-mediated particle growth relies on the chemical reactivity of the oligonucleotide ligands that saturate dangling bonds of growing clusters, and favorable sequences are those that have the highest surface reactivity with growing particles. The oligonucleotide template approach is facile and versatile, offering a route to produce a range of material compositions for other chalcogenide semiconductor quantum dots and metal oxide nanoparticles.