Working memory is a reconstructive process that requires integrating multiple hierarchical representations of objects. This hierarchical reconstruction allows us to overcome perceptual uncertainty and limited cognitive capacity but yields systematic biases in working memory as individual items are influenced by the ensemble statistics of the scene, or of their particular group. Given the importance of the hierarchical encoding of a display, we aim to characterize what structures people use to encode visual scenes using a nonparametric data-driven approach. In Experiment 1, we examine visuospatial memory for locations by asking participants to recall the locations of objects in a serial reproduction task. We show that people report items in a more compact structure than they initially were and organize them into clustered spatial groups. In Experiment 2, we explicitly introduce discrete color groups, allowing us to test whether the color feature governs the spatial grouping. We find that the spatial structures were color-contingent. By analyzing color groups, we circumvent the grouping uncertainty in Experiment 1 and further reveal that people compress color groups into collinear structures with similar orientations and equidistant spacing. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
The world around us is filled with complex objects, full of color, motion, shape, and texture, and these features seem to be represented separately in the early visual system. Anne Treisman pointed out that binding these separate features together into coherent conscious percepts is a serious challenge, and she argued that selective attention plays a critical role in this process. Treisman also showed that, consistent with this view, outside the focus of attention we suffer from illusory conjunctions: misperceived pairings of features into objects. Here we used Treisman’s logic to study the structure of pre-attentive representations of multipart, multicolor objects, by exploring the patterns of illusory conjunctions that arise outside the focus of attention. We found consistent evidence of some pre-attentive binding of colors to their parts, and weaker evidence of binding multiple colors of the same object. The extent to which such hierarchical binding occurs seems to depend on the geometric structure of multipart objects: Objects whose parts are easier to separate seem to exhibit greater pre-attentive binding. Together, these results suggest that representations outside the focus of attention are not entirely a “shapeless bundles of features,” but preserve some meaningful object structure.
How do people use the structure of items when storing them in visual memory? An iterated learning experiment reveals that people expect visual scenes to be hierarchically organized into self-similar groups, and that this expectation biases visual memory. A series of subsequent experiments on the error structure of recall reveals that the inferred hierarchical structure not only biases what people remember, but also influences how people forget: memory errors can be attributed to different nodes in the hierarchical parse by virtue of the relationships between errors for different objects. Together, these results show not how people use environmental structure to remember displays, but also that the inferred scene hierarchy influences the encoded data structure in memory.
What happens to memories as we forget? They might gradually lose fidelity, lose their associations (and thus be retrieved in response to the incorrect cues), or be completely lost. Typical long-term memory studies assess memory as a binary outcome (correct/incorrect), and cannot distinguish these different kinds of forgetting. Here we assess long-term memory for scalar information, thus allowing us to quantify how different sources of error diminish as we learn, and accumulate as we forget. We trained subjects on visual and verbal continuous quantities (the locations of objects and the distances between major cities, respectively), tested subjects after extended delays, and estimated whether recall errors arose due to imprecise estimates, misassociations, or complete forgetting. Although subjects quickly formed precise memories and retained them for a long time, they were slow to learn correct associations and quick to forget them. These results suggest that long-term recall is especially limited in its ability to form and retain associations.
How much do individuals, compared to the population, know about the distribution of values in the world? Participants reported the prices of consumer goods such as watches and belts and we compared how accurately individuals vs. the overall population knew the mean and dispersion of prices. Although individuals and the population both knew objects’ average prices and relative standard deviations, the population was more sensitive to the absolute standard deviation of prices. In a second experiment, we examined whether individuals’ impoverished distribution knowledge impairs their ability to interpret advertisements. Consistent with people using Bayesian inference, the higher an object’s actual price dispersion, the more participants relied on advertisements; however, this effect is considerably smaller than a simple proportional offset, suggesting again that individuals underestimate dispersion. Thus, despite having a sense of the distribution of real world quantities, individuals tend to know only a fraction of the world distribution.
People seem to compute the ensemble statistics of objects and use this information to support the recall of individual objects in visual working memory. However, there are many different ways that hierarchical structure might be encoded. We examined the format of structured memories by asking subjects to recall the locations of objects arranged in different spatial clustering structures. Consistent with previous investigations of structured visual memory, subjects recalled objects biased toward the center of their clusters. Subjects also recalled locations more accurately when they were arranged in fewer clusters containing more objects, suggesting that subjects used the clustering structure of objects to aid recall. Furthermore, subjects had more difficulty recalling larger relative distances, consistent with subjects encoding the positions of objects relative to clusters and recalling them with magnitude-proportional (Weber) noise. Our results suggest that clustering improved the fidelity of recall by biasing the recall of locations toward cluster centers to compensate for uncertainty and by reducing the magnitude of encoded relative distances.
Although many investigations of visual summary representations ("ensemble statistics") have focused on how people compute the central tendency of stimuli such as average set size (e.g. Ariely, 2001), orientation (e.g. Parks, et al., 2001), or facial emotion (e.g. Haberman & Whitney, 2009), less attention has been given to representations of set heterogeneity. People rapidly extract set variance (Michael, et al., 2013), and the variance of a set affects how ensembles are averged (Corbett et al., 2012; Fouriezos et al., 2008; Im & Halberda, 2013). We investigated the ability to detect changes in the variance of circle sizes across sets, using a staircase algorithm. On each trial subjects (n = 23) were presented first with a pedestal display of circles followed by a test display, and had to judge if the variance of the circle sizes (the logarithm of the circle diameter) of the test display was the same as the pedestal set, or if it had changed (the mean was held constant). In one block of 200 trials the changed test variance increased compared to the pedestal variance, while in the other block of 200 trials the changed test variance decreased compared to the pedestal (block order was counterbalanced). We found that people could detect smaller differences between the pedestal and test variance when the variance had decreased, compared to equivalent changes when the variance increased. Meeting abstract presented at VSS 2015.
Visual working memory stores object features (e.g., locations) according to their statistical structure (Alvarez & Oliva, 2009). When recalling objects, people often use that structure information to compensate for uncertainty about the individual objects (Brady & Alvarez, 2011). Although any stimulus has its own ensemble statistics, people also have expectations from the real world about how objects are organized. Here we try to characterize Gestalt priors about the spatial arrangement of objects in an iterative visual working memory paradigm. We examined visual working memory's priors for locations by asking participants to recall the locations of objects, and then having someone else remember and reproduce those recalled locations. A long sequence of individuals remembering the positions recalled by previous participants yields a Markov chain that will overemphasize the priors that people use to encode object locations (Sanborn & Griffiths, 2008). Across iterations, subjects recalled objects more densely packed (t(9)=8.03, p< .001) and with more similar translational errors (t(9)=9.11, p< .001), suggesting that subjects grouped objects in memory. To determine how subjects grouped objects, we designed a non-parametric clustering algorithm that infers whether objects are parts of clusters or straight lines. The clustering model revealed that subjects increasingly grouped objects as lines, going from using line groupings 2% to 22% of the time. Furthermore, consecutive subjects were more likely to group objects the same way when arranged in lines (t(9)=8.08, p< .001) or very eccentric clusters (t(9)=3.14, Bonferroni corrected p=.024). This suggests that linear arrangements are particularly stable in memory. Our results are consistent with evidence that people use priors from the real world to efficiently encode information in visual working memory (Orhan & Jacobs, 2014). Additionally, the increasing likelihood of people remembering objects as components of lines rather than clusters suggests that these priors aid the perception of higher-level constructs from ensemble statistics. Meeting abstract presented at VSS 2015.
What hierarchical structures do people use to encode visual displays? We examined visual working memory’s priors for locations by asking participants to recall the locations of objects in an iterated learning task. We designed a nonparametric clustering algorithm that infers the clustering structure of objects and encodes individual items within this structure. Over many iterations, participants recalled objects with more similar displacement errors, especially for objects our clustering algorithm grouped together, suggesting that subjects grouped objects in memory. Additionally, participants increasingly remembered objects as lines with similar orientations and lengths, consistent with the Gestalt grouping principles of continuity and similarity. Furthermore, the increasing tendency of participants to remember objects as components of hierarchically organized lines rather than individual objects or clusters suggests that these priors aid the perception of higher-level structures from ensemble statistics.
In an unfamiliar environment, searching for and navigating to a target requires that spatial information be acquired, stored, processed, and retrieved. In a study encompassing all of these processes, participants acted as taxicab drivers who learned to pick up and deliver passengers in a series of small virtual towns. We used data from these experiments to refine and validate MAGELLAN, a cognitive map-based model of spatial learning and wayfinding. MAGELLAN accounts for the shapes of participants' spatial learning curves, which measure their experience-based improvement in navigational efficiency in unfamiliar environments. The model also predicts the ease (or difficulty) with which different environments are learned and, within a given environment, which landmarks will be easy (or difficult) to localize from memory. Using just 2 free parameters, MAGELLAN provides a useful account of how participants' cognitive maps evolve over time with experience, and how participants use the information stored in their cognitive maps to navigate and explore efficiently.
People seem to compute the ensemble statistics of objects and use this information to support the recall of individual objects in visual short-term memory. However, the appropriate grouping of objects into ensembles is not always obvious, and people may need to infer different hierarchical organizations of the objects. These different organizations should determine how ensemble information influences object recall. We tested whether objects' hierarchical structure influences visual short-term memory recall and assessed the encoding scheme people use to represent objects in a hierarchical structure. To address these questions, we asked subjects to recall the locations of objects arranged in different spatial clustering structures. Objects in the same cluster were recalled with similar displacement errors, suggesting that the hierarchical structure induced correlated errors. Furthermore, objects arranged into fewer clusters containing more objects were recalled more accurately. We considered three accounts of this improvement: (a) fewer misassociations of objects to locations, (b) more effective random guessing around cluster centers, and (c) more accurate encoding of object locations. Our analyses suggest that performance improved as objects were more densely clustered because guessing around the cluster centers decreased and object locations were recalled more accurately. One explanation for this pattern is that subjects represented the relative positions of objects using a log encoding. Such a scheme would allow more densely clustered objects to be recalled with greater fidelity. Consequently, we designed a model that represents the locations of objects relative to their clusters and recalls the relative positions with Weber noise on distance. We fit the model to subjects' responses and found for each clustering structure the model was able to accurately predict objects' bias towards clusters and the noise of object locations. Together, these results suggest that denser clustering allows more parsimonious encoding of the object hierarchy, preserving resources for encoding individual objects. Meeting abstract presented at VSS 2014