Auditory scene analysis benefits from spatial separation of the sound sources to be disentangled. Studies characterizing the ability to localize sound sources have revealed that spatial ambiguities can be resolved by slightly moving the head. This implies that the auditory system must take into account how head movements displace the sensors (ears) relative to the sound source, in order to separate the consequences of own movement from actual movement of the sound source. Outside of the typical lab, not only the head can be turned, but also the whole body might move towards or away from the sound source. With modern virtual-reality (VR) technology including near-real-time auditory feedback, it has now become feasible to study the relations between body movement, head movement, and sound localization in a controlled manner. Here we present a VR study in which participants move through a virtual maze to locate a repetitive alarm signal. Participants’ turning decisions provide a behavioral measure of their ability to localize the sound source. In addition, we introduce occasional task-unrelated location deviations in the alarm signal to record an unobtrusive neural measure of localization ability: the Mismatch Negativity (MMN) component extracted from participants’ continuous electroencephalogram (EEG) informs us about participants’ deviance detection at the sensory level. We examine these neural and behavioral correlates of sound localization during simulated locomotion and real head movements. Our results show high behavioral accuracy and the elicitation of a robust MMN component. This indicates faithful extraction of source location although participants move relative to the sound source. Our findings are compatible with a predictive-coding framework, where the effects caused by own movement are taken into account for the interpretation of sensory signals. Yet because the range of physical variation by own movement was narrower than intended, we cannot exclude the possibility that deviance detection rested on simpler sensory mechanisms. We discuss how this constraint can be overcome in future studies to fully exploit our new paradigm for naturalistic, yet controlled studies on sound localization and auditory scene analysis under dynamic listening requirements.
In the vision of a “hybrid society,” the interaction between humans and embodied digital technologies is envisaged to be as smooth and seamless as human–human interaction. In the foreseeable future, however, humans will undoubtedly continue to interact with many machines through designated human–machine interfaces (HMIs). Designing such HMIs for production settings presents a particular challenge, as users may vary in their experience and expertise. Moreover, challenging issues of a specific HMI are often difficult to verbalize and therefore hard to obtain through user report and expert interviews alone. To make interactions through HMIs smooth, it is therefore crucial to evaluate HMIs with state-of-the-art technologies that do not require explicit report. Here, we demonstrate the use of eye tracking for HMI design in an industrial production setting, where smooth human–machine interaction is particularly critical to ensure safe and efficient operation. Using a real-life example, we illustrate how eye tracking allows dissociating users’ difficulties to find a particular interaction item (“search”) from their challenges in realizing that it is indeed the item to be operated (“verification”). We argue that this distinction is crucial for (re-) designing HMIs to optimize usability and that the usefulness of eye tracking extends beyond the specific context to human–machine interaction in general.
The human auditory system rapidly encodes auditory regularities. Evidence comes from the oddball paradigm, in which frequent (standard) sounds are occasionally replaced with a rare (deviant) sound. Deviants relative to standards typically elicit signs of prediction error (e.g., MMN and P3a). It is, however, less clear whether deviants, which also bear predictive information but are encountered less often than standards, might inform auditory prediction. To investigate this, naïve participants listened to sound sequences constructed according to a new, modified version of the oddball paradigm: two kinds of deviants differing in their probability of repetition yield the sound actually following a deviant either conditionally likely or unlikely. As this sound is either the same deviant (repetition) or a standard (no repetition), it is either unlikely or likely with respect to the global stimulus probability at the same time. In an active deviant detection task, we replicated previous behavioural findings, demonstrating that predictive information carried by deviants (conditional probability) is extracted when behaviourally relevant. Our analyses further reveal that respective response time effects increase over the course of the task. However, in a passive listening setting, both MMN and P3a were confined to violations of rules based on global probability, while not being sensitive to conditional probability. Though some sensitivity to conditional probability had been observed in a previous study, these effects were tiny compared to those of global probability. Thus, the auditory system seems to mainly rely on rules that are encountered frequently (standard regularity), at least during passive listening.
In virtual reality (VR), we assessed how untrained participants searched for fire sources with the digital twin of a novel augmented-reality (AR) device: a firefighter’s helmet equipped with a heat sensor and an integrated display indicating the heat distribution in its field of view. This was compared to the digital twin of a current state-of-the-art device, a handheld thermal imaging camera. The study had three aims: (i) compare the novel device to the current standard, (ii) demonstrate the usefulness of VR for developing AR devices, (iii) investigate visual search in a complex, realistic task free of visual context. Users detected fire sources faster with the thermal camera than with the helmet display. Responses in target-present trials were faster than in target-absent trials for both devices. Fire localization after detection was numerically faster and more accurate, in particular in the horizontal plane, for the helmet display than for the thermal camera. Search was strongly biased to start on the left-hand side of each room, reminiscent of pseudoneglect in scene viewing. Our study exemplifies how VR can be used to study vision in realistic settings, to foster the development of AR devices, and to obtain results relevant to basic science and applications alike.
Multistable perception occurs in all sensory modalities, and there is ongoing theoretical debate about whether there are overarching mechanisms driving multistability across modalities. Here we study whether multistable percepts are coupled across vision and audition on a moment-by-moment basis. To assess perception simultaneously for both modalities without provoking a dual-task situation, we query auditory perception by direct report, while measuring visual perception indirectly via eye movements. A support-vector-machine (SVM)-based classifier allows us to decode visual perception from the eye-tracking data on a moment-by-moment basis. For each timepoint, we compare visual percept (SVM output) and auditory percept (report) and quantify the co-occurrence of integrated (one-object) or segregated (two-objects) interpretations in the two modalities. Our results show an above-chance coupling of auditory and visual perceptual interpretations. By titrating stimulus parameters toward an approximately symmetric distribution of integrated and segregated percepts for each modality and individual, we minimize the amount of coupling expected by chance. Because of the nature of our task, we can rule out that the coupling stems from postperceptual levels (i.e., decision or response interference). Our results thus indicate moment-by-moment perceptual coupling in the resolution of visual and auditory multistability, lending support to theories that postulate joint mechanisms for multistable perception across the senses.
The auditory system has an amazing ability to rapidly encode auditory regularities. Evidence comes from the popular oddball paradigm, in which frequent (standard) sounds are occasionally exchanged for rare deviant sounds, which then elicit signs of prediction error based on their unexpectedness (e.g., MMN and P3a). Here, we examine the widely neglected characteristics of deviants being bearers of predictive information themselves; naive participants listened to sound sequences constructed according to a new, modified version of the oddball paradigm including two types of deviants that followed diametrically opposed rules: one deviant sound occurred mostly in pairs (repetition rule), the other deviant sound occurred mostly in isolation (non-repetition rule). Due to this manipulation, the sound following a first deviant (either the same deviant or a standard) was either predictable or unpredictable based on its conditional probability associated with the preceding deviant sound. Our behavioral results from an active deviant detection task replicate previous findings that deviant repetition rules (based on conditional probability) can be extracted when behaviorally relevant. Our electrophysiological findings obtained in a passive listening setting indicate that conditional probability also translates into differential processing at the P3a level. However, MMN was confined to global deviants and was not sensitive to conditional probability. This suggests that higher-level processing concerned with stimulus selection and/or evaluation (reflected in P3a) but not lower-level sensory processing (reflected in MMN) considers rarely encountered rules.
Humans achieve skilled actions by continuously correcting for motor errors or perceptual misjudgments, a process called sensorimotor adaptation. This can occur with the actor both detecting (explicitly) and not detecting the error (implicitly). We investigated how the magnitude of a perturbation and the corresponding error signal each contribute to the detection of a size perturbation during interaction with real-world objects. Participants grasped cuboids of different lengths in a mirror-setup allowing us to present different sizes for seen and felt cuboids, respectively. Visuo-haptic size mismatches (perturbations) were introduced either abruptly or followed a sinusoidal schedule. These schedules dissociated the error signal from the visuo-haptic mismatch: Participants could fully adapt their grip and reduce the error when a perturbation was introduced abruptly and then stayed the same, but not with a constantly changing sinusoidal perturbation. We compared participants’ performance in a two-alternative forced choice (2AFC) task where participants judged these mismatches, and modelled error-correction in grasping movements by looking at changes in maximum grip apertures, measured using motion tracking. We found similar mismatch-detection performance with sinusoidal perturbation schedules and the first trial after an abrupt change, but decreasing performance over further trials for the latter. This is consistent with the idea that reduced error signals following adaptation make it harder to detect perturbations. Error-correction parameters indicated stronger error-correction in abruptly introduced perturbations. However, we saw no correlation between error-correction and overall mismatch-detection performance. This emphasizes the distinct contributions of the perturbation magnitude and the error signal in helping participants detect sensory perturbations.
Which properties of a natural scene affect visual search? We consider the alternative hypotheses that low-level statistics, higher-level statistics, semantics, or layout affect search difficulty in natural scenes. Across three experiments (n n = 20 each), we used four different backgrounds that preserve distinct scene properties: (a) natural scenes (all experiments); (b) 1/f f noise (pink noise, which preserves only low-level statistics and was used in Experiments 1 and 2); (c) textures that preserve low-level and higher-level statistics but not semantics or layout (Experiments 2 and 3); and (d) inverted (upside-down) scenes that preserve statistics and semantics but not layout (Experiment 2). We included "split scenes" that contained different backgrounds left and right of the midline (Experiment 1, natural/noise; Experiment 3, natural/texture). Participants searched for a Gabor patch that occurred at one of six locations (all experiments). Reaction times were faster for targets on noise and slower on inverted images, compared to natural scenes and textures. The N2pc component of the event-related potential, a marker of attentional selection, had a shorter latency and a higher amplitude for targets in noise than for all other backgrounds. The background contralateral to the target had an effect similar to that on the target side: noise led to faster reactions and shorter N2pc latencies than natural scenes, although we observed no difference in N2pc amplitude. There were no interactions between the target side and the non-target side. Together, this shows that-at least when searching simple targets without own semantic content-natural scenes are more effective distractors than noise and that this results from higher-order statistics rather than from semantics or layout.
It is often assumed that rendering an alert signal more salient yields faster responses to this alert. Yet, there might be a trade-off between attracting attention and distracting from task execution. Here we tested this in four behavioral experiments with eye-tracking using an abstract alert-signal paradigm. Participants performed a visual discrimination task (primary task) while occasional alert signals occurred in the visual periphery accompanied by a congruently lateralized tone. Participants had to respond to the alert before proceeding with the primary task. When visual salience (contrast) or auditory salience (tone intensity) of the alert were increased, participants directed their gaze to the alert more quickly. This confirms that more salient alerts attract attention more efficiently. Increasing auditory salience yielded quicker responses for the alert and primary tasks, apparently confirming faster responses altogether. However, increasing visual salience did not yield similar benefits: instead, it increased the time between fixating the alert and responding, as high-salience alerts interfered with alert-task execution. Such task interference by high-salience alert-signals counteracts their more efficient attentional guidance. The design of alert signals must be adapted to a “sweet spot” that optimizes this stimulus-dependent trade-off between maximally rapid attentional orienting and minimal task interference.
When humans walk, it is important for them to have some measure of the distance they have traveled. Typically, many cues from different modalities are available, as humans perceive both the environment around them (for example, through vision and haptics) and their own walking. Here, we investigate the contribution of visual cues and nonvisual self-motion cues to distance reproduction when walking on a treadmill through a virtual environment by separately manipulating the speed of a treadmill belt and of the virtual environment. Using mobile eye tracking, we also investigate how our participants sampled the visual information through gaze. We show that, as predicted, both modalities affected how participants (N = 28) reproduced a distance. Participants weighed nonvisual self-motion cues more strongly than visual cues, corresponding also to their respective reliabilities, but with some interindividual variability. Those who looked more toward those parts of the visual scene that contained cues to speed and distance tended also to weigh visual information more strongly, although this correlation was nonsignificant, and participants generally directed their gaze toward visually informative areas of the scene less than expected. As measured by motion capture, participants adjusted their gait patterns to the treadmill speed but not to walked distance. In sum, we show in a naturalistic virtual environment how humans use different sensory modalities when reproducing distances and how the use of these cues differs between participants and depends on information sampling.NEW & NOTEWORTHY Combining virtual reality with treadmill walking, we measured the relative importance of visual cues and nonvisual self-motion cues for distance reproduction. Participants used both cues but put more weight on self-motion; weight on visual cues had a trend to correlate with looking at visually informative areas. Participants overshot distances, especially when self-motion was slow; they adjusted steps to self-motion cues but not to visual cues. Our work thus quantifies the multimodal contributions to distance reproduction.
Using a map in an unfamiliar environment requires identifying correspondences between elements of the map's allocentric representation and elements in egocentric views. Aligning the map with the environment can be challenging. Virtual reality (VR) allows learning about unfamiliar environments in a sequence of egocentric views that correspond closely to the perspectives and views that are experienced in the actual environment. We compared three methods to prepare for localization and navigation tasks performed by teleoperating a robot in an office building: studying a floor plan of the building and two forms of VR exploration. One group of participants studied a building plan, a second group explored a faithful VR reconstruction of the building from a normal-sized avatar's perspective, and a third group explored the VR from a giant-sized avatar's perspective. All methods contained marked checkpoints. The subsequent tasks were identical for all groups. The self-localization task required indication of the approximate location of the robot in the environment. The navigation task required navigation between checkpoints. Participants took less time to learn with the giant VR perspective and with the floorplan than with the normal VR perspective. Both VR learning methods significantly outperformed the floorplan in the orientation task. Navigation was performed quicker after learning in the giant perspective compared to the normal perspective and the building plan. We conclude that the normal perspective and especially the giant perspective in VR are viable options for preparing for teleoperation in unfamiliar environments when a virtual model of the environment is available.
Actions in the real world have immediate sensory consequences. Mimicking these in digital environments is within reach, but technical constraints usually impose a certain latency (delay) between user actions and system responses. It is important to assess the impact of this latency on the users, ideally with measurement techniques that do not interfere with their digital experience. One such unobtrusive technique is electroencephalography (EEG), which can capture the users’ brain activity associated with motor responses and sensory events by extracting event-related potentials (ERPs) from the continuous EEG recording. Here we exploit the fact that the amplitude of sensory ERP components (specifically, N1 and P2) reflects the degree to which the sensory event was perceived as an expected consequence of an own action (self-generation effect). Participants (N=24) elicit auditory events in a virtual-reality (VR) setting by entering codes on virtual keypads to open doors. In a within-participant design, the delay between user input and sound presentation is manipulated across blocks. Occasionally, the virtual keypad is operated by a simulated robot instead, yielding a control condition with externally generated sounds. Results show that N1 (but not P2) amplitude is reduced for self-generated relative to externally generated sounds, and P2 (but not N1) amplitude is modulated by delay of sound presentation in a graded manner. This dissociation between N1 and P2 effects maps back to basic research on self-generation of sounds. We suggest P2 amplitude as a candidate read-out to assess the quality and immersiveness of digital environments with respect to system latency.
Word stress is demanding for non-native learners of English, partly because speakers from different backgrounds weight perceptual cues to stress like pitch, intensity, and duration differently. Slavic learners of English and particularly those with a fixed stress language background like Czech and Polish have been shown to be less sensitive to stress in their native and non-native languages. In contrast, German English learners are rarely discussed in a word stress context. A comparison of these varieties can reveal differences in the foreign language processing of speakers from two language families. We use electroencephalography (EEG) to explore group differences in word stress cue perception between Slavic and German learners of English. Slavic and German advanced English speakers were examined in passive multi-feature oddball experiments, where they were exposed to the word impact as an unstressed standard and as deviants stressed on the first or second syllable through higher pitch, intensity, or duration. The results revealed a robust Mismatch Negativity (MMN) component of the event-related potential (ERP) in both language groups in response to all conditions, demonstrating sensitivity to stress changes in a non-native language. While both groups showed higher MMN responses to stress changes to the second than the first syllable, this effect was more pronounced for German than for Slavic participants. Such group differences in non-native English word stress perception from the current and previous studies are argued to speak in favor of customizable language technologies and diversified English curricula compensating for non-native perceptual variation.
The human auditory system is believed to represent regularities inherent in auditory information in internal models. Sounds not matching the standard regularity (deviants) elicit prediction error, alerting the system to information not explainable within currently active models. Here, we examine the widely neglected characteristic of deviants bearing predictive information themselves. In a modified version of the oddball paradigm, using higher-order regularities, we set up different expectations regarding the sound following a deviant. Higher-order regularities were defined by the relation of pitch within tone pairs (rather than absolute pitch of individual tones). In a deviant detection task participants listened to oddball sequences including two deviant types following diametrically opposed rules: one occurred mostly in succession (high repetition probability) and the other mostly in isolation (low repetition probability). Participants in Experiment 1 were not informed (naïve), whereas in Experiment 2 they were made aware of the repetition rules. Response times significantly decreased from first to second deviant when repetition probability was high-albeit more in the presence of explicit rule knowledge. There was no evidence of a facilitation effect when repetition probability was low. Significantly more false alarms occurred in response to standards following high compared with low repetition probability deviants, but only in participants aware of the repetition rules. These findings provide evidence that not only deviants violating lower- but also higher-order regularities can inform predictions about auditory events. More generally, they confirm the utility of this new paradigm to gather further insights into the predictive properties of the human brain.
Within recent years, it has become popular to use physiological and expression data to ameliorate inter-cognitive communication between human and machine. One emotion that is highly relevant for this is frustration, which occurs when a user's goal in using a system is failed to be met. This paper presents a latent variable model that estimates frustration in two different driving contexts by continuous subjective frustration rating, facial expressions and frontal alpha asymmetry in the electroencephalogram. We then compare this full model to models with less measurement variables to evaluate which measurements can be left out. Our results show that expression frequency and subjective frustration make important contributions to the model of experienced frustration. This paper presents a proof of concept for using a latent variable model to evaluate collected measures to estimate an experienced emotion. This method can inform researchers which measurements are most informative in different circumstances. Additionally, the method can be used to evaluate how well purely objective measurements (that are the only feasible measurements in most applied settings) perform in comparison to a model including subjective ratings.