
As broadcasting environments change rapidly to digital, user requirements for next-generation services that surpass the current HDTV service quality become more demanding. The next-generation of broadcasting services will change from HD to UHD and from 5.1 channel audio to more than 10 audio channels, including a height channel for a high quality realistic broadcasting service. In accordance with the estimated trends of future broadcasting services, we propose a 10.2 channel audio format for a Korean UHDTV broadcasting service. It can create almost similar spatial sound images as 22.2 channel audio with half the number of speakers. In this paper, we propose a 10.2 channel audio acquisition system for the creation of UHDTV content, and measurements and preliminary evaluation are carried out to determine whether the performance is acceptable for broadcasting.
As visual display complexity grows, visual cues and alerts may become less salient and therefore less effective. Although the auditory system’s resolution is rather coarse relative to the visual system, there is some evidence for virtual spatialized audio to benefit visual search on a small frontal region, such as a desktop monitor. Two experiments examined if search times could be reduced compared to visual-only search through spatial auditory cues rendered using one of two methods: individualized or generic head-related transferfunctions. Results showed the cue type interacted with display complexity, with larger reductions compared to visual-only search as set size increased. For larger set sizes, individualized cues were significantly better than generic cues overall. Across all set sizes, individualized cues were better than generic cues for cueing eccentric elevations (>±8°). Where performance must be maximized, designers should use individualized virtual audio if at all possible, even in small frontal region within the field of view.
In this paper, we demonstrate a responsive sound installation consisting of computer-linked thermometers and cameras installed in both interior and exterior locations that detect the states of these spaces based on image and temperature data. The system simultaneously produces and modifies sound pictograms in different spaces in order to convey information on motion or temperature changes. We also propose a sound design method that represents sounds made by physical objects using the above-described installation and sound design method.
The following paper introduces a new Layer Based Amplitude Panning algorithm and supporting D4 library of rapid prototyping tools for the 3D time-based data representation using sound. The algorithm is designed to scale and support a broad array of configurations, with particular focus on High Density Loudspeaker Arrays (HDLAs). The supporting rapid prototyping tools are designed to leverage oculocentric strategies to importing, editing, and rendering data, offering an array of innovative approaches to spatial data editing and representation through the use of sound in HDLA scenarios. The ensuing D4 ecosystem aims to address the shortcomings of existing approaches to spatial aural representation of data, offers unique opportunities for furthering research in the spatial data audification and sonification, as well as transportable and scalable spatial media creation and production.
As 3D audio becomes more common place to enhance auditory environments, designers are faced with the challenge of choosing HRTFs for listeners that provide proper audio cues. Subjective selection is a low-cost alternative to expensive HRTF measurement, however little is known concerning whether the preferred HRTFs are similar or if users exhibit random behavior in this task. In addition, PCA (principal component analysis) can be used to decompose HRTFs in representative features, however little is known concerning whether the features have a relevant perceptual basis. 12 listeners completed a subjective selection experiment in which they judged the perceptual quality of 14 HRTFs in terms of elevation, and front-back distinction. PCA was used to decompose the HRTFs and create an HRTF similarity metric. The preferred HRTFs were significantly more similar to each other, the preferred and non-preferred HRTFs were significantly less similar to each other, and in the case of front-back distinction the non-preferred HRTFs were significantly more similar to each other.
This paper presents a brief description of surface electromyography (sEMG), what it can be used for, as well as some of the problems associated with visual displays of sEMG data. Sonifications of sEMG data have shown potential for certain applications in data monitoring and movement training, however there are still challenges related to the design of these sonifications that need to be addressed. Our previous research has shown that different sonification designs resulted in better listener performance for different sEMG evaluation tasks (e.g. identifying muscle activation time vs. muscle exertion level). Based on this finding, we speculated that sonifications may benefit from being designed to be task-specific, and that integrating a task analysis into the sonification design process may help sonification designers identify intuitive and meaningful sonification designs. This paper presents a brief introduction to what a task analysis is, provides an example of how a task analysis can be used to inform sonification design, and outlines future research into a task-analysis-based approach to sonification design.
This paper describes an approach to sonification based on an iPhone app created for multiple users to explore a microtonal scale generated from harmonics using the combination product set method devised by tuning theorist Erv Wilson. The app is intended for performance by a large consort of hand-held mobile phones where phones are played collaboratively in a shared listening space. Audio consisting of handbells and sine tones is synthesised independently on each phone. Sound projection from each phone relies entirely on venue acoustics unaided by mains-powered amplification. It was designed to perform a microtonal composition called Transposed Dekany which takes the form of a chamber concerto in which a consort of players explore the properties of an microtonal scale. The consort subdivides into families of instruments that play in different pitch registers assisted by processes that are enabled and disabled at various stages throughout the performance. The paper outlines Wilson’s method, describes its current implementation and considers hypothetical sonification scenarios for implementation using different data with potential applications in the physical world.
Auditory cues, when coupled with visual objects, have lead to reduced response times in visual search tasks, suggesting that adding auditory information can potentially aid Air Force operators in complex scenarios. These benefits are substantial when the spatial transformations that one has to make are relatively simple i.e., mapping a 3-D auditory space to a 3-D visual scene. The current study focused on listeners’ abilities to map sound surrounding a listener to a 2-D visual space, by measuring performance in localization tasks that required the following responses: 1) Headpointing: turn and face a loudspeaker from where a sound emanated, 2) Tablet: point to an icon representing a loudspeaker displayed in an array on a 2-D GUI or, 3) Hybrid: turn and face the loudspeaker from where a sound emanated and them indicate that location on a 2-D GUI. Results indicated that listeners’ localization errors were small when the response modality was head-pointing, and localization errors roughly doubled when they were asked to make a complex transformation o fauditory-visual space (i.e., while using a hybrid response); surprisingly, the hybrid response technique reduced errors compared to the tablet response conditions. These results have large implications for the design of auditory displays that require listeners to make complex, non-intuitive transformations of auditory-visual space.
The demands of concurrent radio communications in Navy shipboard command centers contribute to the problem of operator information overload and impede personnel optimization goals for new platforms. Motivations for serializing this task and human performance research with virtual, multichannel, rate-accelerated speech in support of this idea are briefly reviewed, and the results of a recent listening study in which participants carried out a Navyrelevant word-spotting task in this context are reported.
Attention redirection trials were carried out using a wearable interface incorporating auditory and visual cues. Visual cues were delivered via the screen on the Recon Jet – a wearable computer resembling a pair of glasses – while auditory cues were delivered over a bone conduction headset. Cueing conditions included the delivery of individual cues, both auditory and visual, and in combination with each other. Results indicate that the use of an auditory cue drastically decreases target acquisition times. This is true especially for targets that fall outside the visual field of view. While auditory cues showed no difference when paired with any of the visual cueing conditions for targets within the field of view of the user, for those outside the field of view a significant improvement in performance was observed. The static visual cue paired with the binaurally spatialised, dynamic auditory cue appeared to provide the best performance in comparison to any other cueing conditions. In the absence of a visual cue, the binaurally spatialised, dynamic auditory cue performed the best.
In this paper we examined the role of informative sound in a simple decision-making game task. A within-subject experiment with 48 participants measured the response time, success rate and number of timeouts of the players in a number of eight-second decision tasks. As time proceeds, the task becomes easier at the risk of players timing out and reducing the overall opportunities they will have to attempt the task. We designed a simple informative sound display that uses a tone that increases in amplitude over the duration of the task. We test player performance in three conditions, no sound (visual-only), constant (non-informative) sound and increasing (informative) sound. We found that the increasing sound display significantly reduced timeouts when compared with the visual only and constant sound versions of the task. This reduction in timeouts did not impair the players’ performance in terms of their success rate nor response time.
While several potential auditory cues responsible for sound source externalization have been identified, less work has gone into providing a simple and robust way of manipulating perceived externalization. The current work describes a simple approach for parametrically modifying individualized head-related transfer function spectra that results in a systematic change in the perceived externalization of a sound source. Methods and results from a subjective evaluation validating the technique are presented, and further discussion relates the current method to previously identified cues for auditory distance perception.
Music psychologists have frequently shown that music affects people’s behaviour. Applying this concept to work-related computing tasks has the potential to lead to improvements in a person’s productivity, efficiency and effectiveness. This paper presents two quantitative experiments exploring whether transcription typing performance is affected when hearing a music accompaniment that includes vocals. The first experiment showed that classifying the typists as either slow or fast ability is important as there were significant interaction effects once this between group factor was included, with the accuracy of fast typists reduced when the music contained vocals. In the second experiment, a Dutch transcription typing task was added to manipulate task difficulty and the volume of playback was included as a between groups independent variable. When typing in Dutch the fast typists’ speed was reduced with louder music. When typing in English the volume of music had little effect on typing speed for either the fast or slow typists. The fast typists achieved lower speeds when the loud volume music contained vocals, but with low volume music the inclusion of vocals in the background music did not have a noticeable affect on typing speed. The presence of vocals in the music reduced the accuracy of the text entry across the whole sample. Overall, these experiments show that the presence of vocals in background music reduces typing performance, but that we might be able to exploit instrumental music to improve performance in tasks involving typing with either low or high volume music.
A pilot study was conducted to explore the potential of sonically-enhanced gestures as controls for future in-vehicle information systems (IVIS). Four concept menu systems were developed using a LEAP Motion and Pure Data: (1) 2x2 with auditory feedback, (2) 2x2 without auditory feedback, (3) 4x4 with auditory feedback, and (4) 4x4 without auditory feedback. Seven participants drove in a simulator while completing simple target-acquisition tasks using each of the four prototype systems. Driving performance and eye glance behavior were collected as well as subjective ratings of workload and system preference. Results from driving performance and eye tracking measures strongly indicate that the 2x2 grids yield better driving safety outcomes than 4x4 grids. Subjective ratings show similar patterns for driver workload and preferences. Auditory feedback led to similar improvements in driving performance and eye glance behavior as well as subjective ratings of workload and preference, compared to visual-only.
Memorable life events are important to form the present selfimage. Looking back on these memories provides an opportunity to ruminate meaning of life and envision future. Integrating the life-log concept and auditory graphs, we have implemented a mobile application, “LifeMusic”, which helps people reflect their memories by listening to their life event sonifcation that is synchronous to these memories. Reflecting the life events through LifeMusic can relieve users of the present and have them journey to the past moments and thus, they can keep balance of emotions in the present life. In the current paper, we describe the implementation and workflow of LifeMusic and briefly discuss focus group results, improvements, and future works.
Here we report early results from an experiment designed to investigate the use of sonification for the learning of a novel perceptual-motor skill. We find that sonification which employs melody is more effective than a strategy which provides only bare timing information. We additionally show that it might be possible to ‘refresh’ learning after performance has waned following training - through passive listening to the sound that would be produced by perfect performance. Implications of these findings are discussed in terms of general motor performance enhancement and sonic feedback design.
People with Autistic Spectrum Disorders (ASD) are known to have difficulty recognizing and expressing emotions, which affects their social integration. Leveraging the recent advances in interactive robot and music therapy approaches, and integrating both, we have designed musical robots that can facilitate social and emotional interactions of children with ASD. Robots communicate with children with ASD while detecting their emotional states and physical activities and then, make real-time sonification based on the interaction data. Given that we envision the use of multiple robots with children, we have adopted a client-server architecture. Each robot and sensing device plays a role as a terminal, while the sonification server processes all the data and generates harmonized sonification. After describing our goals for the use of sonification, we detail the system architecture and on-going research scenarios. We believe that the present paper offers a new perspective on the sonification application for assistive technologies.
The Interaural Time Difference is one of the primary localization cues for 3D sound. However, due to differences in head and ear anthropometry across the population, ITDs related to a sound source at a given location around the head will differ from subject to subject. Furthermore, most individuals do not possess symmetrical traits between the left and right pinnae. This fact may cause an angle-dependent ITD asymmetry between locations mirrored across the left and right hemispheres. This paper describes an exploratory analysis performed on publicly available databases of individually measured HRIRs. The analysis was first performed separately for each dataset in order to explore the impact of different formats and measurement techniques, and then on pooled sets of repositories, in order to obtain statistical information closer to the population values. Asymmetry in ITDs was found to be consistently more prominent in the rear-lateral angles (approximately between 90° and 130° azimuth) across all databases investigated, suggesting the presence of a sensitive region. A significant difference between the peak asymmetry values and the average asymmetry across all angles was found on three out of four examined datasets. These results were further explored by pooling the datasets together, which revealed an asymmetry peak at 110° that also showed significance. Moreover, it was found that within the region of sensitivity the difference between specular ITDs exceeds the just noticeable difference values for perceptual discrimination at all frequency bands. These findings validate the statistical presence of ITD asymmetry in public datasets of individual HRIRs and identify a significant, perceptually-relevant, region of increased asymmetry. Details of these results are of interest for HRIR modeling and personalization techniques, which should consider implementing compensation for asymmetric ITDs when aiming for perceptually accurate binaural displays. This work is part of a larger study aimed at binaural-audio personalization and user-characterization through non-invasive techniques.
Driving is mainly a visual task, leaving other sensory channels open for additional information communication. As the level of automation increases in vehicles, monitoring the state and performance of the driver and vehicle shifts from the secondary to primary task. Auditory channels provide the flexibility to display a wide variety of information to the driver without increasing the workload of driving task. It is important to identify types of auditory displays and sonification strategies that provide integral information necessary for the driving task, and not overload the driver with unnecessary or intrusive data. To this end, we have developed an in-vehicle interactive sonification system using the medium-fidelity simulator and neurophysiological devices. The system is intended to integrate driving performance data and driver affective state data in real-time. The present paper introduces the architecture of our invehicle interactive sonification system and potential sonification strategies for providing feedback to the driver in an intuitive and non-intrusive manner.