Since the creation of the spatially oriented format for acoustics (SOFA, Audio Engineering Society standard AES69), numerous databases of head-related transfer functions (HRTFs) are now available as standardized SOFA files. However, the methodologies for measuring and postprocessing HRTFs vary significantly across laboratories. This leads to objective and perceptual inconsistencies between HRTF databases and makes it challenging to integrate multiple databases into a single repository to facilitate wide-scale research and application. This paper introduces a normalization procedure, applicable to any HRTF data set, aimed at enhancing the consistency across HRTF data sets obtained from different laboratories while preserving the spatial information essential to HRTFs. The proposed approach consists of six processing steps: low-pass filtering, temporal alignment, temporal windowing, diffuse-field equalization, low-frequency extrapolation, and far-field correction. The normalization was evaluated on 17 HRTF data sets of the same dummy head by means of acoustic analyses and auditory simulations and further validated with respect to a database of 54 human subjects. Results show that the proposed normalization improves data set applicability and consistency while maintaining the directional cues within each data set.
The spatially oriented format for acoustics (SOFA), also known as the AES69 standard, is a container file format for spatial acoustic data.SOFA can be used to describe head-related transfer functions (HRTFs), binaural, spatial, and directional room impulse responses (BRIRs, SRIRs, DRIRs), and directivities, among other data.SOFA was introduced in 2015 [1] to ease the exchange of spatial acoustic data.SOFA specifications consider structured data description, data compression, network transfer, and links to complex room geometries or other data in a hierarchical way.Since its introduction, SOFA has been embraced by many institutions and researchers, who have developed SOFA libraries for various programming environments.SOFA's recent revisions AES69-2020 [2] and AES69-2022 (as known as SOFA 2.1) [3] include a continuous representation of Source and Listener directivity composed of spherical harmonic Emitters and Receivers, new conventions describing the directivity of microphones, musical instruments, and loudspeakers, and
Spatially oriented acoustic data can range from a simple set of impulse responses, such as head-related transfer functions, to a large set of multiple-input multiple-output spatial room impulse responses obtained in complex measurements with a microphone array excited by a loudspeaker array at various conditions.The spatially oriented format for acoustics (SOFA), which was standardized by AES Standard 69, provides a format to store and share such data.SOFA takes into account geometric representations of many acoustic scenarios, data compression, network transfer, and a link to complex room geometries and aims at simplifying the development of interfaces for many programming languages.With the recent advancement of SOFA, the format offers new continuous-direction representation of data by means of spherical harmonics and novel conventions representing many measurement scenarios, such as source directivity and multiple-input multiple-output spatial room impulse responses.This article reviews SOFA by first providing an introduction to SOFA and then describing examples that demonstrate the most recent features of the SOFA 2.1 (AES Standard 69-2022).
The late reverberation characteristics of a sound field are often assumed to be perceptually isotropic, meaning that the decay of energy is perceived as equivalent in every direction. In this paper, we employ Ambisonics reproduction methods to reassess how a decaying sound field is analyzed and characterized and our capacity to hear directional characteristics within late reverberation. We propose the use of objective measures to assess the anisotropy characteristics of a decaying sound field. The energy-decay deviation is defined as the difference of the direction-dependent decay from the average decay. A perceptual study demonstrates a positive link between the range of these energy deviations and their audibility. These results suggest that accurate sound reproduction should account for directional properties throughout the decay.
Spatial room impulse responses (SRIRs) measured using spherical microphone arrays are seeing increasingly widespread use in reproducing room reverberation effects on three-dimensional surround sound systems (e.g., higher-order ambisonics) through multi-channel SRIR convolution. However, such measured impulse responses inevitably present a non-negligible noise floor, which may lead to a perceptible "infinite reverberation effect" when convolved with an input sound. Furthermore, individual sensor noise and momentary measurement artefacts may additionally corrupt the resulting impulse response. This paper presents a robust SRIR denoising procedure applicable to impulse responses with diffuse late reverberation tails, which can be modeled by a stochastic process. In such cases, the non-decaying frequency-dependent noise floor may be replaced by a synthesized incoherent tail parameterized by the SRIR's energy decay envelope. It is shown that performing such tail re-synthesis in the spherical harmonic domain, using an independent zero-mean Gaussian noise for each component, preserves both the reverberation tail's frequency-dependent decay as well as its spatial coherence properties. The proposed process is then evaluated through its application to SRIRs measured in real-world conditions, and finally some aspects of performance and consistency verification are discussed.
Directional room impulse responses (DRIR) measured with spherical microphone arrays (SMA) enable the reproduction of room reverberation effects on three-dimensional surround-sound systems (e.g., Higher-Order Ambisonics) through multichannel convolution. However, such measurements inevitably contain a nondecaying noise floor that may produce an audible "infinite reverberation effect" upon convolution. If the late reverberation tail can be considered a diffuse field before reaching the noise floor, the latter may be removed and replaced with an extension of the exponentially-decaying tail synthesized as a zero-mean Gaussian noise. This has previously been shown to preserve the diffuse-field properties of the late reverberation tail when performed in the spherical harmonic domain (SHD). In this paper, we show that in the case of highly anisotropic yet incoherent late fields, the spatial symmetry of the spherical harmonics is not conducive to preserving the energy distribution of the reverberation tail. To remedy this, we propose denoising in an optimized spatial domain obtained by plane-wave decomposition (PWD), and demonstrate that this method equally preserves the incoherence of the late reverberation field.
When a personalized set of head-related transfer functions (HRTFs) is not available, a common solution is identifying a perceptually appropriate substitute from a database. There are various approaches to this selection process whether based on localization cues, subjective evaluations, or anthropomorphic similarities. This study investigates whether HRTF rankings that stem from different selection methods yield comparable results. A perceptual study was carried out using a basic source localization method and a subjective quality judgment method for a common set of eight HRTFs. HRTF rankings were determined according to different metrics from each method for each subject and the respective results were compared. Results indicate a significant and positive mean correlation between certain metrics. The best HRTFs selected according to one method had significant above-average rating scores according to metrics in the second method.
The use of directional room impulse responses (DRIR) measured with spherical microphone arrays (SMA) has become widespread in the reproduction of spatial effects on surround-sound systems through multichannel convolution. However, the measurement of such DRIRs in real-world conditions is inevitably subject to several risk factors, including the presence of a nondecaying noise floor that can produce an infinite reverberation effect when convolved with an input sound. Recent work has focused on model-based techniques for removing this noise floor by replacing it with a re-synthesized prolongation of the measured tail, which has concurrently led to the development of a framework for the spatial analysis of properties. We present here a comprehensive evaluation of the proposed techniques through their application to DRIRs measured in particularly complex spaces, including spatially anisotropic late tails as well as multipleslope decays characteristic of coupled-volume configurations. Following a brief review of the theoretical underpinnings of the tail re-synthesis denoising procedure, the measurement, treatment, analysis, and subsequent denoising of these DRIRs are each detailed and assessed with quantitative metrics. Finally, a discussion of the anisotropic, direction-dependent analysis results obtained is included as the basis for a wider research question on the acoustical considerations behind a stochastic model allowing for spatial variations.
Recent developments in the measurement of directional room impulse responses (DRIR) by spherical microphone arrays (SMA) have led to their extensive use in sound spatialisation. Room reverberation effects can be reproduced in three-dimensional surround sound systems (e.g. Higher-Order Ambisonics) through multi-channel DRIR convolution. However, such measured impulse responses inevitably present a non-negligible noise floor, leading to a perceptible ‘infinite reverberation effect’. Further, individual sensor noise and non-stationary measurement artefacts may additionally corrupt the deconvolved impulse response. This paper presents recent work regarding the implementation of state of the art DRIR analysis and denoising techniques and their application to extensive DRIR databases measured across a highly varied collection of spaces. We first review the basic energy decay relief (EDR) analysis and reverberation tail re-synthesis process, before presenting several novel refinements developed throughout the course of this implementation. Finally, an overview of the results obtained both globally and with respect to particular cases (complex architectural volumes, outdoor spaces, etc.) is included in order to further examine the capabilities of the denoising framework.
Recent technological advances, such as increased CPU/GPU processing speed, along with the miniaturization of devices and sensors, have created new possibilities for integrating immersive technologies in music and performance art. Virtual and Augmented Reality (VR/AR) have become increasingly interesting as mobile device platforms, such as up-to-date smartphones, with necessary CPU resources entered the consumer market. In combination with recent web technologies, any mobile device can simply connect with a browser to a local server to access the latest technology. The web platform also eases the integration of collabora-tive situated media in participatory artwork. In this paper , we present the interactive music improvisation piece 'Border,' premiered in 2018 at the Beyond Festival at the Center for Art and Media Karlsruhe (ZKM). This piece explores the interaction between a performer and the audience using web-based applications-including AR, real-time 3D audio/video streaming, advanced web audio, and gesture-controlled virtual instruments-on smart mobile devices.
Pipe organs are complex timbral synthesisers in an early acousmatic setting, which have always accompanied the evolution of music and technology. The most recent development is digital augmentation: the organ sound is captured, transformed and then played back in real time. The present augmented organ project relies on three main aesthetic principles: microphony, fusion and instrumentality. Microphony means that sounds are captured inside the organ case, close to the pipes. Real-time audio effects are then applied to the internal sounds before they are played back over loudspeakers; the transformed sounds interact with the original sounds of the pipe organ. The fusion principle exploits the blending effect of the acoustic space surrounding the instrument; the room response transforms the sounds of many single-sound sources into a consistent and organ-typical soundscape at the listener’s position. The instrumentality principle restricts electroacoustic processing to organ sounds only, excluding non-organ sound sources or samples. This article proposes a taxonomy of musical effects. It discusses aesthetic questions concerning the perceptual fusion of acoustic and electronic sources. Both extended playing techniques and digital audio can create musical gestures that conjoin the heterogeneous sonic worlds of pipe organs and electronics. This results in a paradoxical listening experience of unity in the diversity: the music is at the same time electroacoustic and instrumental.
Our demonstration presents recent developments of the EVERTims project, an auralization framework for virtual acoustics and real-time room acoustic simulation. The developments presented here concern the complete re-design of the scene graph editor unit, and the C++ implementation of a new spatial renderer based on the JUCE framework. EVERTims now functions as a Blender add-on to support real-time auralization of any 3D room model, both for its creation in Blender and its exploration in the Blender Game Engine. The EVERTims framework is published as open source software.
Spherical microphone arrays (SMAs) and spherical loudspeaker arrays (SLAs) facilitate the study of room acoustics due to the three-dimensional analysis they provide. More recently, systems that combine both arrays, referred to as multiple-input multiple-output (MIMO) systems, have been proposed due to the added spatial diversity they facilitate. The literature provides frameworks for designing SMAs and SLAs separately, including error analysis from which the operating frequency range (OFR) of an array is defined. However, such a framework does not exist for the joint design of a SMA and a SLA that comprise a MIMO system. This paper develops a design framework for MIMO systems based on a model that addresses errors and highlights the importance of a matched design. Expanding on a free-field assumption, errors are incorporated separately for each array and error bounds are defined, facilitating error analysis for the system. The dependency of the error bounds on the SLA and SMA parameters is studied and it is recommended that parameters should be chosen to assure matched OFRs of the arrays in MIMO system design. A design example is provided, demonstrating the superiority of a matched system over an unmatched system in the synthesis of directional room impulse responses.
This paper presents recent developments of the EVERTims project, an auralization framework for virtual acoustics and real-time room acoustic simulation. The EVERTims framework relies on three independent components: a scene graph editor, a room acoustic modeler, and a spatial audio renderer for auralization. The framework was first published and detailed in previous publications. Recent developments presented here concern the complete redesign of the scene graph editor unit, and the C++ implementation of a new spatial renderer based on the JUCE framework. EVERTims now functions as a Blender add-on to support real-time auralization of any 3D room model, both for its creation in Blender and its exploration in the Blender Game Engine. The EVERTims framework is published as open source software: http://evertims.ircam.fr.
This study evaluates several methods for reporting the perceived location of real sound sources. It is well known that the method used for collecting judgments in auditory-localization experiments has a strong influence on the accuracy of a subject's response. Previous works on auditory-localization tasks revealed that egocentric pointing methods (which are based on a body-centered coordinate system) allow for more accurate judgments than verbal reporting or exocentric pointing techniques (which are based on a 2D or 3D reporting device). Three different egocentric methods are compared: the most commonly applied "manual pointing" and "head pointing" methods, and the "proximal pointing" method, which forces the participants to indicate the apparent direction by pointing in the proximal region of the head with a marker held at the fingertips. The two first methods involve a rotation of the body of the participant, whereas the third method only involves movements of the arm(s) and hand(s) with a fixed head. Sound stimuli were presented randomly over 24 loudspeakers that were uniformly distributed on the upper hemisphere around the subject. The merits of the different methods are compared and discussed with regard to localization errors and to practical considerations. Although they show similar trends, each of the different methods affects the pointing accuracy in a specific way. The proximal pointing method, for example, is more accurate for sources located at high elevation angles. However, at rear locations close to the median plane an increased bias appears due to difficulties in performing the motor task to reach these positions. The proximal pointing method shows faster response times, which may be advantageous when planning 3D sound localization experiments.
Spherical microphone and loudspeaker arrays have been widely studied for the acquisition of spatial sound-field information. Recently, a theoretical framework, based on systems that combine both arrays, was presented for the spatial analysis of enclosed sound fields. Such systems are referred to as multiple-input multiple-output (MIMO) systems, and they provide means for an enhanced spatial analysis. However, their performance is limited by errors due to spatial sampling and system model mismatch. The effects of these errors on the system performance were studied recently in theory, without experimental validation. Therefore, the practical usefulness of MIMO systems for roomacoustics analysis has yet to be determined. This paper presents an initial investigation in this direction. MIMO system performance and limitations are first evaluated in a simulation study. The system is then studied experimentally, through the analysis of room impulse responses (RIRs). Experimental validation is achieved in several aspects. First, system properties are studied and compared to previous theoretical results. Then, MIMO processing methods are applied for a spatial analysis of early reflections in the RIR, showing that early room reflections can be identified experimentally. The results of this investigation suggest that MIMO systems can be employed in practice for various applications of room acoustics.
Directional room impulse responses (DRIRs) are typically measured with spherical micro- phone arrays (SMA). When being combined with spherical loudspeaker arrays (SLA), also the directivity of the sound source can be controlled. Such multiple-input multiple-output (MIMO) systems allow for an in-depth analysis of the acoustics of a room. A prototype SMA/SLA system was used to capture 3-D MIMO DRIRs for several measurement points and under different acoustic conditions at the opera hall of the Salzburg Festival. The SLA consists of 64 micro- phones on a rigid sound-hard sphere (25 cm diameter). It captures the 3-D sound field up to the spherical harmonics expansion order N = 7. The SLA consists of 28 speakers (of three different sizes) on a rigid sound-hard sphere (40 cm diameter). It is equipped with an internal tilt motor and mounted on a remote-controlled turntable. The rotated loudspeaker positions chosen for this measurement approximate a Gaussian sampling of order N = 11. To ensure coinciding ranges of operation, the SLA and SMA parameters were matched using an analytical MIMO system model previously published by some of the authors. The measurement data can be analyzed with respect to perceptual and acoustical parameters typically used in architectural acoustics.
This article presents a report on technological and aesthetic practices in the variable-acoustics performance hall, Espace de Projection, at the Institut de Recherche et Coordination Acoustique/Musique. The hall is surrounded by a 350-loudspeaker array for sound-field reproduction using holophonic approaches such as wave-field synthesis and higher-order Ambisonics. First we present the design and implementation of the audio system and discuss the challenges of both hardware and software architectures. This is followed by a discussion of spatial composition techniques, aesthetic approaches, and methodologies for composing computer music for high-density loudspeaker arrays, explored through the paradigmatic examples of pieces produced by two artist-in-research residencies.
The perception of sound by human listeners in a room has been shown to be affected by the spatial attributes of the sound field. These spatial attributes have been studied using microphone and loudspeaker arrays separately. Systems that combine both loudspeaker and microphone arrays, termed multiple-input multiple-output (MIMO) systems, facilitate enhanced spatial analysis compared to systems with a single array, thanks to the simultaneous use of the arrays and the additional spatial diversity. Using MIMO systems, room impulse responses (RIRs) can be presented using matrix notation, which enables a unique study of a sound field’s spatial attributes, employing methods from linear algebra. For example, a matrix’s rank and null space can be studied to reveal spatial information on a room, such as the number of dominant room reflections and their direction of arrival to the microphone array and the direction of radiation from the loudspeaker array. In this contribution, a theory of the spatial analysis of a sound field using a MIMO system comprised of spherical arrays is developed and a simulation study is presented. In the study, tools proposed for processing MIMO RIRs with the aim of revealing valuable information about acoustic reflections paths are evaluated.
The directionality of the radiated sound is very specific to each musical instrument. The underlying radiation mechanisms may, for instance, depend on the structure of the vibrating body (e.g., string and percussion instruments) or on the spatial distribution of the opening holes (e.g., bells and open finger holes for wind instruments). A good knowledge of the radiation pattern of instruments is essential for many applications, such as orchestration, room acoustics, microphone techniques for live sound and recording, and virtual acoustics. In the first part, we will review previous works on sound source radiation measurement and analysis, discuss the underlying acoustic principles, and try to identify common mechanisms of radiation in musical instruments. In the second part, we will illustrate various projects undertaken at IRCAM and dedicated to the measurement and modeling of the directivity of instruments, to the objective and perceptual characterization of room acoustics, and to the real-time synthesis of virtual source radiation for musical performances. For this latter, several approaches are discussed according to the underlying physical formalisms and associated electroacoustic setups (e.g., spherical loudspeaker arrays, wave field synthesis).