Although static localization performance in auditory displays is known to substantially improve as a listener spends more time in the environment, the impact of real-time interactive movement on these tasks is not yet well understood. Accordingly, a training procedure was developed and evaluated to address this question. In a set of experiments, listeners searched for and marked the locations of five virtually spatialized sound sources. The task was performed with and without training. Finally, the listeners performed a second search and mark task to assess the impacts of training. The results indicate that the training procedure maintained or significantly improved localization accuracy. In addition, localization performance did not improve for listeners who did not complete the training procedure.
An important application of cognitive architectures is to provide human performance models that capture psychological mechanisms in a form that can be "programmed" to predict task performance of human-machine system designs. Although many aspects of human performance have been successfully modeled in this approach, accounting for multitalker speech task performance is a novel problem. This article presents a model for performance in a two-talker task that incorporates concepts from psychoacoustics, in particular, masking effects and stream formation.
Speech recognition was measured as a function of the target-to-masker ratio (TMR) with syntactically similar speech maskers. In the first experiment, listeners were instructed to report keywords from the target sentence. Data averaged across listeners showed a plateau in performance below 0 dB TMR when masker and target sentences were from the same talker. In this experiment, some listeners tended to report the target words at all TMRs in accordance with the instructions, while others reported keywords from the louder of the sentences, contrary to the instructions. In the second experiment, stimuli were the same as in the first experiment, but listeners were also instructed to avoid reporting the masker keywords, and a payoff matrix penalizing masker keywords and rewarding target keywords was used. In this experiment, listeners reduced the number of reported masker keywords, and increased the number of reported target keywords overall, and the average data showed a local minimum at 0 dB TMR with same-talker maskers. The best overall performance with a same-talker masker was obtained with a level difference of 9 dB, where listeners achieved near perfect performance when the target was louder, and at least 80% correct performance when the target was the quieter of the two sentences.
An extension of the auditory module in EPIC is introduced to model the two-talker coordinate response measure (CRM) listening task. The construct of an auditory stream is employed as an object in the working memory of EPIC’s cognitive processor. Production rules are developed that execute the two-talker CRM task. Analysis of these rules reveal two sources of possible error in the output of the auditory processor to working memory. Each is explored in turn and the production rules modified to provide a corpus-driven model that accounts for human performance in the listening task.
Engineering the sound quality of an automobile has drawn upon well known perceptual studies of loudness to set design targets for various noise sources, such as wind and power train. Attempts to apply these studies to design target specifications for squeak and rattle noises have met with limited success, primarily because the perceptual response to transient signals differs in several important dimensions from the response to wideband, steady-state noises, where the phase spectrum, for example, contributes little. We present psychophysical results that detail how the perception of a transient’s magnitude spectrum can be affected by its phase spectrum, introduce a computational auditory model, and demonstrate the model’s use in visualizing the transient events.
Virtual auditory environments (VAEs) are created by processing digital sounds such that they convey a 3D location to the listener. This technology has the potential to augment systems in which an operator tracks the positions of targets. Prior work has established that listeners can locate sounds in VAEs, however less is known concerning listener memory for virtual sounds. In this study, three experimental tasks assessed listener recall of sound positions and identities, using free and cued recall, with one or more delays. Overall, accuracy degrades as listeners recall the environment, however when using free recall, listeners exhibited less degradation.
An important application of cognitive architectures is to provide human performance models that capture psychological mechanisms in a form that can be “programmed” to predict task performance of human-machine system designs. While many aspects of human performance have been successfully modeled in this approach, accounting for multi-talker speech task performance is a novel problem. This paper presents a model for performance in a two-talker task that incorporates concepts from the psychoacoustic study of speech perception, in particular, masking effects and stream formation.
Virtual auditory environments (VAEs) can be used to communicate spatial information, with sound sources representing the location of objects. A critical factor in this type of immersive system is the degree to which the participant can interact with the virtual environment. Our prior work has demonstrated that listeners can successfully locate virtual spatialized sounds, delivered over headphones, in a VAE using a mouse and screen to navigate the virtual world. The screen indicates the avatars position on the vertical plane. The present study seeks to determine the effects of plane mapping on listener performance. In the horizontal-plane interface, the listener used a WACOM tablet and pen to navigate the VAE on the horizontal plane. Results suggest that there is no significant performance difference when locating a single sound source. In the multi-source context, it was observed that the time taken to locate the first sound was significantly larger than the time taken to locate the remaining sounds.
This study examined the ability of human listeners to detect the presence and judge the strength of a statistical dependency among the elements comprising sequences of sounds. The statistical dependency was imposed by specifying transition matrices that determined the likelihood of occurrence of the sound elements. Markov chains were constructed from these transition matrices having states that were pure tones/noise bursts that varied along the stimulus dimensions of frequency and/or interaural time difference. Listeners reliably detected the presence of a statistical dependency in sequences of sounds varying along these stimulus dimensions. Furthermore, listeners were able to discriminate the relative strength of the dependency in pairs of successive sound sequences. Random variation along an irrelevant stimulus dimension had small but significant adverse effects on performance. A much greater decrement in performance was found when the sound sequences were concurrent. Likelihood ratios were computed based on the transition matrices to specify Ideal Observer performance for the experimental conditions. Preliminary modeling efforts were made based on degradations of Ideal Observer performance intended to represent human observer limitations. This experimental approach appears to be useful for examining auditory "stream" formation and maintenance over time based on the predictability of the constituent sound elements.
Although the concept of virtual spatial audio has existed for almost twenty-five years, only in the past fifteen years has modern computing technology enabled the real-time processing needed to deliver high-precision spatial audio. Furthermore, the concept of virtually walking through an auditory environment did not exist. The applications of such an interface have numerous potential uses. Spatial audio has the potential to be used in various manners ranging from enhancing sounds delivered in virtual gaming worlds to conveying spatial locations in real-time emergency response systems. To incorporate this technology in real-world systems, various concerns should be addressed. First, to widely incorporate spatial audio into real-world systems, head-related transfer functions (HRTFs) must be inexpensively created for each user. The present study further investigated an HRTF subjective selection procedure previously developed within our research group. Users discriminated auditory cues to subjectively select their preferred HRTF from a publicly available database. Next, the issue of training to find virtual sources was addressed. Listeners participated in a localization training experiment using their selected HRTFs. The training procedure was created from the characterization of successful search strategies in prior auditory search experiments. Search accuracy significantly improved after listeners performed the training procedure. Next, in the investigation of auditory spatial memory, listeners completed three search and recall tasks with differing recall methods. Recall accuracy significantly decreased in tasks that required the storage of sound source configurations in memory. To assess the impacts of practical scenarios, the present work assessed the performance effects of: signal uncertainty, visual augmentation, and different attenuation modeling. Fortunately, source uncertainty did not affect listeners' ability to recall or identify sound sources. The present study also found that the presence of visual reference frames significantly increased recall accuracy. Additionally, the incorporation of drastic attenuation significantly improved environment recall accuracy. Through investigating the aforementioned concerns, the present study made initial footsteps guiding the design of virtual auditory environments that support spatial configuration recall.
Preferences for subjective design qualities, such as shape, are difficult to capture and relate to engineering specifications. The present paper uses Interactive Evolutionary Systems (IES) to locate a human user's most preferred cola bottle shape among a set of parameterised bottle shapes. Several researchers have used IES to identify user preference, but have never independently confirmed that preference. In the present paper, participants used the IGA to select their favourite design from a small design space. The method of paired comparisons was used to characterise preference over the entire design space, showing a 91% agreement with IGA most preferred selections.
Interaction between the listener and their environment in a spatial auditory display plays an important role in creating better situational awareness, resolving front/back and up/down confusions, and improving localization. Prior studies with 6DOF interaction suggest that using either a head tracker or a mouse-driven interface yields similar performance during a navigation and search task in a virtual auditory environment. In this paper, we present a study that compares listener performance in a virtual auditory environment under a static mode condition, and two dynamic conditions (head tracker and mouse) using orientation-only interaction. Results reveal tradeoffs among the conditions and interfaces. While the fastest response time was observed in the static mode, both dynamic conditions resulted in significantly reduced front/back confusions and improved localization accuracy. Training effects and search strategies are discussed.
In the design of spatial auditory displays, listener interactivity can promote greater immersion, better situational awareness, reduced front/back confusion, improved localization, and greater externalization. Interactivity between the listener and their environment has traditionally been achieved using a head tracker interface. However, trackers are expensive, sensitive to calibration, and may not be appropriate for use in all physical environments. Interactivity can be achieved using a number of alternative interfaces. This study compares learning rates and performance in a single-source auditory search tasking for a head-tracker and a mouse/keyboard interface within a single source and multi-source context.