In virtual reality (VR), we assessed how untrained participants searched for fire sources with the digital twin of a novel augmented-reality (AR) device: a firefighter’s helmet equipped with a heat sensor and an integrated display indicating the heat distribution in its field of view. This was compared to the digital twin of a current state-of-the-art device, a handheld thermal imaging camera. The study had three aims: (i) compare the novel device to the current standard, (ii) demonstrate the usefulness of VR for developing AR devices, (iii) investigate visual search in a complex, realistic task free of visual context. Users detected fire sources faster with the thermal camera than with the helmet display. Responses in target-present trials were faster than in target-absent trials for both devices. Fire localization after detection was numerically faster and more accurate, in particular in the horizontal plane, for the helmet display than for the thermal camera. Search was strongly biased to start on the left-hand side of each room, reminiscent of pseudoneglect in scene viewing. Our study exemplifies how VR can be used to study vision in realistic settings, to foster the development of AR devices, and to obtain results relevant to basic science and applications alike.
Multistable perception occurs in all sensory modalities, and there is ongoing theoretical debate about whether there are overarching mechanisms driving multistability across modalities. Here we study whether multistable percepts are coupled across vision and audition on a moment-by-moment basis. To assess perception simultaneously for both modalities without provoking a dual-task situation, we query auditory perception by direct report, while measuring visual perception indirectly via eye movements. A support-vector-machine (SVM)-based classifier allows us to decode visual perception from the eye-tracking data on a moment-by-moment basis. For each timepoint, we compare visual percept (SVM output) and auditory percept (report) and quantify the co-occurrence of integrated (one-object) or segregated (two-objects) interpretations in the two modalities. Our results show an above-chance coupling of auditory and visual perceptual interpretations. By titrating stimulus parameters toward an approximately symmetric distribution of integrated and segregated percepts for each modality and individual, we minimize the amount of coupling expected by chance. Because of the nature of our task, we can rule out that the coupling stems from postperceptual levels (i.e., decision or response interference). Our results thus indicate moment-by-moment perceptual coupling in the resolution of visual and auditory multistability, lending support to theories that postulate joint mechanisms for multistable perception across the senses.
The decision of whether to cross a road or wait for a car to pass, humans make frequently and effortlessly. Recently, the application of drift-diffusion models (DDMs) on pedestrians’ decision-making has proven useful in modelling crossing behaviour in pedestrian-vehicle interactions. These models consider binary decision-making as an incremental accumulation of noisy evidence over time until one of two choice thresholds (to cross or not) is reached. One open question is whether the assumption of a kinematics-dependent drift-diffusion process, which was made in previous pedestrian crossing DDMs, is justified, with DDM-parameters varying over time according to the developing traffic situation. It is currently unknown whether kinematics-dependent DDMs provide a better model fit than conventional DDMs, which are fitted per condition. Furthermore, previous DDMs have not considered reaction times for the not-crossing option. We address these issues by a novel experimental design combined with modelling. Experimentally, we use a 2-alternative-forced-choice paradigm, where participants view videos of approaching cars from a pedestrian’s perspective and respond whether they want to cross before the car or to wait until the car has passed. Using these data, we perform thorough model comparison between kinematics-dependent and condition-wise fitted DDMs. Our results demonstrate that condition-wise fitted DDMs can show better model fits than kinematics-dependent DDMs as reflected in the mean-squared-errors. The condition-wise fitted models need considerably more parameters, but in some cases still outperform kinematics-dependent DDMs in measures that penalize the parameter number (e.g., Akaike information criterion). Introducing a starting point bias provides support for the novel hypothesis of rapid early evidence build-up from the initial view of the vehicle distance. The drift rates obtained for the condition-wise fitted models align with the assumptions in the kinematics-dependent models, confirming that pedestrians’ decision processes are kinematics-dependent. However, the partial preference for condition-wise fitted models in the model selection suggests that the correct form of kinematics-dependence has not yet been identified for all DDM-parameters, indicating room for improvement of current pedestrian crossing DDMs. Developing more accurate models of human cognitive processes will likely facilitate autonomous vehicles to understand pedestrians’ intentions as well as to show unambiguous human-like behaviour in future traffic interactions with humans.
Which properties of a natural scene affect visual search? We consider the alternative hypotheses that low-level statistics, higher-level statistics, semantics, or layout affect search difficulty in natural scenes. Across three experiments (n n = 20 each), we used four different backgrounds that preserve distinct scene properties: (a) natural scenes (all experiments); (b) 1/f f noise (pink noise, which preserves only low-level statistics and was used in Experiments 1 and 2); (c) textures that preserve low-level and higher-level statistics but not semantics or layout (Experiments 2 and 3); and (d) inverted (upside-down) scenes that preserve statistics and semantics but not layout (Experiment 2). We included "split scenes" that contained different backgrounds left and right of the midline (Experiment 1, natural/noise; Experiment 3, natural/texture). Participants searched for a Gabor patch that occurred at one of six locations (all experiments). Reaction times were faster for targets on noise and slower on inverted images, compared to natural scenes and textures. The N2pc component of the event-related potential, a marker of attentional selection, had a shorter latency and a higher amplitude for targets in noise than for all other backgrounds. The background contralateral to the target had an effect similar to that on the target side: noise led to faster reactions and shorter N2pc latencies than natural scenes, although we observed no difference in N2pc amplitude. There were no interactions between the target side and the non-target side. Together, this shows that-at least when searching simple targets without own semantic content-natural scenes are more effective distractors than noise and that this results from higher-order statistics rather than from semantics or layout.
It is often assumed that rendering an alert signal more salient yields faster responses to this alert. Yet, there might be a trade-off between attracting attention and distracting from task execution. Here we tested this in four behavioral experiments with eye-tracking using an abstract alert-signal paradigm. Participants performed a visual discrimination task (primary task) while occasional alert signals occurred in the visual periphery accompanied by a congruently lateralized tone. Participants had to respond to the alert before proceeding with the primary task. When visual salience (contrast) or auditory salience (tone intensity) of the alert were increased, participants directed their gaze to the alert more quickly. This confirms that more salient alerts attract attention more efficiently. Increasing auditory salience yielded quicker responses for the alert and primary tasks, apparently confirming faster responses altogether. However, increasing visual salience did not yield similar benefits: instead, it increased the time between fixating the alert and responding, as high-salience alerts interfered with alert-task execution. Such task interference by high-salience alert-signals counteracts their more efficient attentional guidance. The design of alert signals must be adapted to a “sweet spot” that optimizes this stimulus-dependent trade-off between maximally rapid attentional orienting and minimal task interference.
Gaze is an important and potent social cue to direct others' attention towards specific locations. However, in many situations, directional symbols, like arrows, fulfill a similar purpose. Motivated by the overarching question how artificial systems can effectively communicate directional information, we conducted two cueing experiments. In both experiments, participants were asked to identify peripheral targets appearing on the screen and respond to them as quickly as possible by a button press. Prior to the appearance of the target, a cue was presented in the center of the screen. In Experiment 1, cues were either faces or arrows that gazed or pointed in one direction, but were non-predictive of the target location. Consistent with earlier studies, we found a reaction time benefit for the side the arrow or the gaze was directed to. Extending beyond earlier research, we found that this effect was indistinguishable between the vertical and the horizontal axis and between faces and arrows. In Experiment 2, we used 100% "counter-predictive" cues; that is, the target always occurred on the side opposite to the direction of gaze or arrow. With cues without inherent directional meaning (color), we controlled for general learning effects. Despite the close quantitative match between non-predictive gaze and non-predictive arrow cues observed in Experiment 1, the reaction-time benefit for counter-predictive arrows over neutral cues is more robust than the corresponding benefit for counter-predictive gaze. This suggests that-if matched for efficacy towards their inherent direction-gaze cues are harder to override or reinterpret than arrows. This difference can be of practical relevance, for example, when designing cues in the context of human-machine interaction.
When humans walk, it is important for them to have some measure of the distance they have traveled. Typically, many cues from different modalities are available, as humans perceive both the environment around them (for example, through vision and haptics) and their own walking. Here, we investigate the contribution of visual cues and nonvisual self-motion cues to distance reproduction when walking on a treadmill through a virtual environment by separately manipulating the speed of a treadmill belt and of the virtual environment. Using mobile eye tracking, we also investigate how our participants sampled the visual information through gaze. We show that, as predicted, both modalities affected how participants (N = 28) reproduced a distance. Participants weighed nonvisual self-motion cues more strongly than visual cues, corresponding also to their respective reliabilities, but with some interindividual variability. Those who looked more toward those parts of the visual scene that contained cues to speed and distance tended also to weigh visual information more strongly, although this correlation was nonsignificant, and participants generally directed their gaze toward visually informative areas of the scene less than expected. As measured by motion capture, participants adjusted their gait patterns to the treadmill speed but not to walked distance. In sum, we show in a naturalistic virtual environment how humans use different sensory modalities when reproducing distances and how the use of these cues differs between participants and depends on information sampling.NEW & NOTEWORTHY Combining virtual reality with treadmill walking, we measured the relative importance of visual cues and nonvisual self-motion cues for distance reproduction. Participants used both cues but put more weight on self-motion; weight on visual cues had a trend to correlate with looking at visually informative areas. Participants overshot distances, especially when self-motion was slow; they adjusted steps to self-motion cues but not to visual cues. Our work thus quantifies the multimodal contributions to distance reproduction.
Screen-based communication increasingly replaces face-to-face interactions. Gaze information is important for nonverbal communication. Therefore, we investigate how humans (“receivers”) estimate gaze direction of others (“senders”) in videoconferencing-like settings, a major example of screen-based communication. In two online experiments, receivers estimated gaze targets – which were known to the experimenters – from images of senders. As in real videoconferencing settings, receivers had no information about the geometry of the senders' setup (camera position, etc.), but had to rely on the information available on screen. In Experiment 1, we found that gaze-target estimates were more closely related to the actual target positions in the horizontal than in the vertical direction, a bias toward the sender's head position, and some advantage of presenting the same sender in succession. For Experiment 2, we created a new database of sender images, in which the senders' head position in the image and gaze-target position were systematically varied. Additionally, images of natural scenes were presented whose content served as potential gaze targets. At large, we replicated the findings of Experiment 1, and found only little effect of image content on the estimates. As gaze is an important part of intuitive human-machine interaction, our findings bear relevance beyond videoconferencing. • Gaze estimation in videoconference settings is investigated in two online experiments. • Differences between horizontal and vertical estimates are found. • Some effect of repeating the same sender (person whose gaze is estimated) is observed. • There is little effect of content that is shared between the sender and receiver (person who estimates the gaze). • A database of over 6000 images of 8 senders looking at different gaze targets from different positions is made available.
Humans can quickly adapt their behavior to changes in the environment. Classical reversal learning tasks mainly measure how well participants can disengage from a previously successful behavior but not how alternative responses are explored. Here, we propose a novel 5-choice reversal learning task with alternating position-reward contingencies to study exploration behavior after a reversal. We compare human exploratory saccade behavior with a prediction obtained from a neuro-computational model of the basal ganglia. A new synaptic plasticity rule for learning the connectivity between the subthalamic nucleus (STN) and external globus pallidus (GPe) results in exploration biases to previously rewarded positions. The model simulations and human data both show that during experimental experience exploration becomes limited to only those positions that have been rewarded in the past. Our study demonstrates how quite complex behavior may result from a simple sub-circuit within the basal ganglia pathways.
Using a map in an unfamiliar environment requires identifying correspondences between elements of the map's allocentric representation and elements in egocentric views. Aligning the map with the environment can be challenging. Virtual reality (VR) allows learning about unfamiliar environments in a sequence of egocentric views that correspond closely to the perspectives and views that are experienced in the actual environment. We compared three methods to prepare for localization and navigation tasks performed by teleoperating a robot in an office building: studying a floor plan of the building and two forms of VR exploration. One group of participants studied a building plan, a second group explored a faithful VR reconstruction of the building from a normal-sized avatar's perspective, and a third group explored the VR from a giant-sized avatar's perspective. All methods contained marked checkpoints. The subsequent tasks were identical for all groups. The self-localization task required indication of the approximate location of the robot in the environment. The navigation task required navigation between checkpoints. Participants took less time to learn with the giant VR perspective and with the floorplan than with the normal VR perspective. Both VR learning methods significantly outperformed the floorplan in the orientation task. Navigation was performed quicker after learning in the giant perspective compared to the normal perspective and the building plan. We conclude that the normal perspective and especially the giant perspective in VR are viable options for preparing for teleoperation in unfamiliar environments when a virtual model of the environment is available.
Objects influence attention allocation; when a location within an object is cued, participants react faster to targets appearing in a different location within this object than on a different object. Despite consistent demonstrations of this object-based effect, there is no agreement regarding its underlying mechanisms. To test the most common hypothesis that attention spreads automatically along the cued object, we utilized a continuous, response-free measurement of attentional allocation that relies on the modulation of the pupillary light response. In Experiments 1 and 2, attentional spreading was not encouraged because the target appeared often (60%) at the cued location and considerably less often at other locations (20% within the same object and 20% on another object). In Experiment 3, spreading was encouraged because the target appeared equally often in one of three possible locations within the cued object (cued-end, middle, uncued-end). In all experiments, we added gray-to-black and gray-to-white luminance gradients to the objects. By cueing the gray ends of the objects, we could track attention. If attention indeed spreads automatically along objects, then pupil size should be greater after the gray-to-dark object is cued because attention spreads towards darker areas of the object than when the gray-to-white object is cued, regardless of the target location probability. However, unequivocal evidence of attention spreading along the object was only found when spreading was encouraged. These findings do not support an automatic spreading of attention. Instead, they suggest that attentional spreading along the object is guided by cue-target contingencies.
In natural scene viewing, participants tend to direct their gaze towards the image center, the "central bias". Unless the head is fixed, gaze shifts to peripheral targets are accomplished by a combination of eye and head movements, but there are substantial individual differences in the propensity to use the head. Here, we address whether inter-individual differences in central bias and head movement propensity are related. In one part of the experiment, participants freely viewed natural scenes of two different sizes under laboratory conditions (limited screen size, sitting with head fixed by chin and forehead rest). In the second part, participants stood in the center of a cylindrical, 240° panoramic screen and were free to move their eyes, head and upper body. In each trial, they directed their gaze from the screen's vertical midline (0°) to a peripheral target that was located at 5 to 70 degrees eccentricity on either side. Two different conditions were used - an exogenous mode (vertical bar appearing at target position) and an endogenous mode (target position verbally instructed at screen center). The central bias was quantified by the median Euclidian distance of fixations to the image center. Across participants, we found a strong correlation in central bias between the two image sizes; this is, central bias scaled with image size. In the peripheral target task, we found a strong correlation between the exogenous and the endogenous mode, indicating that the tasks were a robust measure of head movement propensity. Despite substantial inter-individual variability in both tasks, no significant correlation was found between head movement propensity and central bias. Taken together, our results suggest that central bias in scene viewing on typical screen sizes is predominately determined by visual properties, and individual differences in central bias are not explained by head movement propensity.
Walking is a complex task. To prevent falls and injuries, gait needs to constantly adjust to the environment. This requires information from various sensory systems; in turn, moving through the environment continuously changes available sensory information. Visual information is available from a distance, and therefore most critical when negotiating difficult terrain. To effectively sample visual information, humans adjust their gaze to the terrain or – in laboratory settings – when facing motor perturbations. During activities of daily living, however, only a fraction of sensory and cognitive resources can be devoted to ensuring safe gait. How do humans deal with challenging walking conditions, when they face high cognitive load? Young, healthy participants (N=24) walked on a treadmill through a virtual, but naturalistic environment. Occasionally, their gait was experimentally perturbed, inducing slipping. We varied cognitive load by asking participants in some blocks to count backwards in steps of seven; orthogonally, we varied whether visual cues indicated upcoming perturbations. We replicated earlier findings on how humans adjust their gaze and their gait rapidly and flexibly on various time scales: eye- and head movements responded in a partially compensatory pattern and visual cues mostly affected eye movements. Interestingly, the cognitive task affected mainly head orientation. During the cognitive task, we found no clear signs of a less stable gait nor of a cautious gait mode, but evidence that participants adapted their gait less to the perturbations than without secondary task. In sum, cognitive load affects head orientation and impairs the ability to adjust to gait perturbations.
Gaze is a powerful cue for directing attention. We investigate the interpretation of an abstract figure as gaze modulates its efficacy as an attentional cue. In each trial, two vertical lines on a central disk moved to one side (left or right). Independent of this "feature-cued" side, a target (black disk) subsequently appeared on one side. After 300 trials (phase 1), participants watched a video of a human avatar walking away. For one group, the avatar wore a helmet that visually matched the central disk and looked at black disks to either side. The other group's video was unrelated to the cueing task. After another 300 trials (phase 2), videos were swapped between groups; 300 further trials (phase 3) followed. In all phases, participants responded more quickly for targets appearing on the feature-cued side. There was a significant interaction between group and phase for reaction times: In phase 3, the group who had just watched the avatar with the helmet had a reduced advantage to the feature-cued side. Hence, interpreting the disk as a turning head seen from behind counteracts the cueing by the motion of the disk. This suggests that the mere perceptual interpretation of an abstract stimulus as gaze yields social cueing effects.
The course of pupillary constriction and dilation provides an easy-to-access, inexpensive, and noninvasive readout of brain activity. We propose a new taxon-omy of factors affecting the pupil and link these to associated neural underpin-nings in an ascending hierarchy. In addition to two well-established low-level factors (light level and focal distance), we suggest two further intermediate -level factors, alerting and orienting, and a higher-level factor, executive function-ing. Alerting, orienting, and executive functioning - including their respective un-derlying neural circuitries - overlap with the three principal attentional networks, making pupil size an integrated readout of distinct states of attention. As a now widespread technique, pupillometry is ready to provide meaningful applications and constitutes a viable part of the psychophysiological toolbox.
Sequential auditory scene analysis (ASA) is often studied using sequences of two alternating tones, such as ABAB or ABA_, with "_" denoting a silent gap, and "A" and "B" sine tones differing in frequency (nominally low and high). Many studies implicitly assume that the specific arrangement (ABAB vs ABA_, as well as low-high-low vs high-low-high within ABA_) plays a negligible role, such that decisions about the tone pattern can be governed by other considerations. To explicitly test this assumption, a systematic comparison of different tone patterns for two-tone sequences was performed in three different experiments. Participants were asked to report whether they perceived the sequences as originating from a single sound source (integrated) or from two interleaved sources (segregated). Results indicate that core findings of sequential ASA, such as an effect of frequency separation on the proportion of integrated and segregated percepts, are similar across the different patterns during prolonged listening. However, at sequence onset, the integrated percept was more likely to be reported by the participants in ABA_low-high-low than in ABA_high-low-high sequences. This asymmetry is important for models of sequential ASA, since the formation of percepts at onset is an integral part of understanding how auditory interpretations build up.
When using virtual reality (VR) to study cognitive processes, the user’s sensory input and motor output in the virtual setting should closely resemble the actual physical signals. For extended VR settings, in which users can freely navigate substantial distances, this becomes challenging, as the virtual movement cannot be matched in the real world for space constraints. It has been suggested that performing rotations in the physical world, while keeping translational movements purely virtual, may present a good compromise. However, most such control strategies use head orientation as input to determine the virtual movement direction and/or require additional tracking hardware. Using the head to control the virtual locomotion is problematic when the objective of the VR investigation relates to attention or gaze. For such research topics, it is critical to avoid confounding gaze and head orientation with commands for moving in the VR. Here, we propose a control strategy that combines virtual translation with physical body rotation to control VR locomotion. This leaves the head free from navigational control and allows studying attention, gaze and other cognitive processes in naturalistic settings and large, complex virtual environments. Importantly, we achieve this by standard software and hardware, which is part of typical consumer-grade VR bundles.
Sensory consequences of one's own action are often perceived as less intense, and lead to reduced neural responses, compared to externally generated stimuli. Presumably, such sensory attenuation is due to predictive mechanisms based on the motor command (efference copy). However, sensory attenuation has also been observed outside the context of voluntary action, namely when stimuli are temporally predictable. Here, we aimed at disentangling the effects of motor and temporal predictability-based mechanisms on the attenuation of sensory action consequences. During fMRI data acquisition, participants (N = 25) judged which of two visual stimuli was brighter. In predictable blocks, the stimuli appeared temporally aligned with their button press (active) or aligned with an automatically generated cue (passive). In unpredictable blocks, stimuli were presented with a variable delay after button press/cue, respectively. Eye tracking was performed to investigate pupil-size changes and to ensure proper fixation. Self-generated stimuli were perceived as darker and led to less neural activation in visual areas than their passive counterparts, indicating sensory attenuation for self-generated stimuli independent of temporal predictability. Pupil size was larger during self-generated stimuli, which correlated negatively with the blood oxygenation level dependent (BOLD) response: the larger the pupil, the smaller the BOLD amplitude in visual areas. Our results suggest that sensory attenuation in visual cortex is driven by action-based predictive mechanisms rather than by temporal predictability. This effect may be related to changes in pupil diameter. Altogether, these results emphasize the role of the efference copy in the processing of sensory action consequences.