State-of-the-art multiple object tracking (MOT) models have recently been shown to behave in qualitatively different ways from human observers. They exhibit superhuman performance for large numbers of targets and subhuman performance when targets disappear behind occluders. Here we investigate whether human gaze behavior can help explain differences in human and model behavior. Human subjects watched scenes with objects of various appearances. They tracked a designated subset of the objects, which moved continuously and frequently disappeared behind static black-bar occluders, reporting the designated objects at the end of each trial. We measured eye movements during tracking and tracking accuracy. We found that human gaze behavior is clearly guided by task relevance: designated objects were preferentially fixated. We compared human performance to that of cognitive models inspired by state-of-the-art MOT models with object slots, where each slot represents the model's probabilistic belief about the location and appearance of one object. In our model, incoming observations are unambiguously assigned to slots using the Hungarian algorithm. Locations are tracked probabilistically (given the hard assignment) with one Kalman filter per slot. We equipped the computational models with a fovea, yielding high-precision observations at the center and low-precision observations in the periphery. We found that constraining models to follow the same gaze behavior as humans (imposing the human-measured fixation sequences) best captures human behavioral phenomena. These results demonstrate the importance of gaze behavior, allowing the human visual system to optimally use its limited resources.
A tremendous amount of work on visual attention has helped to characterize *what* we attend to, but has focused less on precisely *how* and *why* attention is allocated to dynamic scenes across time. Nowhere is this contrast more apparent than in multiple object tracking (MOT). Hundreds of papers have explored MOT as a paradigmatic example of selective attention, in part because it so well captures attention as a dynamic process. It is especially ironic, then, that nearly all of this work reduces each MOT trial to a single value (i.e. the number of targets successfully tracked) -- when in reality, each MOT trial presents an experiment unto itself, with constantly shifting attention over time. Here we seek to capture this dynamic ebb and flow of attention at a subsecond resolution, both empirically and computationally. Empirically, observers completed MOT trials during which they also had to detect sporadic momentary probes, as a measure of the moment-by-moment degree of attention being allocated to each object. Computationally, we characterize (for the first time, to our knowledge) an algorithmic architecture of just how and why such dynamic attentional shifts occur. To do so, we introduce a new 'hypothesis driven adaptive computation' model. Whereas previous models employed many MOT-specific assumptions, this new approach generalizes to any task-driven context. It provides a unified account of attention as the dynamic allocation of computing resources, based on task-driven hypotheses about the properties (e.g. location, target status) of each object. Here, this framework was able to explain the observed probe detection performance measured at a subsecond resolution, independent of general spatial factors (such as the proximity of each probe to the MOT targets' centroid). This case study provides a new way to think about attention and how it interfaces with perception in terms of rational resource allocation.
Most work on attention (in terms of both psychophysical experiments and computational modeling) involves selection in static scenes. And even when dynamic displays are used, performance is still typically characterized with only a single variable (such as the number of items correctly tracked in Multiple Object Tracking; MOT). But the allocation of attention in daily life (e.g. during foraging, navigation, or play) involves both objective performance and subjective effort, and both can vary dramatically from moment to moment. Here we attempt to capture this sort of rich temporal ebb and flow of attention in a novel and generalizable adaptive computation architecture. Compute is allocated across both objects (in space) and moments (in time) based on hypothesis-driven inferences that are continuously refined, and during MOT this framework is able to explain both object tracking performance and the subjective sense of trial-by-trial effort.