
Data replay (DR) methods play a crucial role in building practical and robust continual replay utilization and catastrophic forgetting significantly constrain the accuracy and efficiency of existing DR-based MASR systems, particularly in low-resource settings. To address these challenges, this paper proposes an optimized data and language selection method for DR (DLSDR) and develops an enhanced MASR system based on it named DLSDR-MASR. The approach used introduces two key innovations. First, the authors propose a data selection strategy based on difficulty scores derived from word error rate and metainformation that helps maintain robust recognition performance across both original and target domains. Second, an active language selection mechanism based on linguistic similarity metrics is introduced to enhance cross-lingual transferability and overall system performance. Comprehensive experiments demonstrate that the proposed method effectively mitigates catastrophic forgetting while maintaining competitive performance on sequential learning tasks when integrated with established DR baselines. Compared with the baseline system, DLSDR-MASR reduces the average final word error rate by 3.67% on pretrained languages and 4.28% on new languages.
Audio enhancement aims to improve the perceived quality of audio signals. Initially, to address bandwidth extension, this work proposes AEROMambaP, an efficient variant of the AERO super-resolution architecture in which attention and long short-term memory layers are replaced by the Mamba state-space model, together with a newly developed differentiable perceptual loss derived from the perceptual audio quality measure (PAQM). During training, the architecture requires two to fours times less GPU memory than the baseline; during inference, it achieves a 14 & times; speedup while using one-fifth of the GPU memory. When upsampling both a piano dataset and MUSDB18 from 11.025 kHz to 44.1 kHz, subjective listening tests show that AEROMambaP outperforms AERO by 15% in perceived quality scores. Next, to enhance of audio signals highly compressed by lossy coding, it is further proposed AEROMambaPS & strns;, which applies the same framework but replaces short-time Fourier transform reconstruction losses with the PAQM loss, specifically to enhance MP3-encoded audio at 32 kbps. In listening evaluations, AEROMambaP S & strns; achieves 52% higher quality rating than AEROMambaP when restoring compressed audio. These results demonstrate that PAQM-driven training coupled with lightweight state-space modeling yields high perceptual quality and computational efficiency in both band-limited and compressed audio scenarios.
This study investigates the use of dipole loudspeakers in addition to typical two-way loudspeakers arranged as a classical stereo setup in a room. The dipole loudspeakers are placed on top of the two-way loudspeakers, their null oriented toward the listener to only contribute to the reverberant sound, adding strong early reflections. Loudness perception for playback over the two-way loudspeakers only and playback over both the two-way and the dipole loudspeakers were compared using a two-alternative forced-choice paradigm. Different music and noise stimuli, different relative levels between two-way and dipole loudspeakers, and two listener positions in the room were used. Results show that adding the dipole-reproduced sound can increase the perceived loudness despite keeping the sound pressure level constant. That effect can be linked to changes in the resulting spectrum, energy decay, interaural correlation, and spatial perception when adding dipole loudspeakers. A binaural loudness model based on the energy received at both ears but not considering the interaural correlation was used to predict the measured level offsets for equal loudness. The predictions were accurate for most stimuli but failed for impulsive sounds where the room reverberation was more perceptible.
Implementing higher-order Ambisonics on consumer devices is hindered by their sparse, irregular microphone arrays, which challenge conventional methods with issues like spatial aliasing and ill-conditioning. Based on the established spherical harmonic beamforming framework, this paper proposes two robust encoding approaches: a frequency-domain method with compensation for high-frequency artifacts and a time-domain (TD) method that holistically optimizes broadband finite impulse response filters for enhanced stability. The framework is inherently scalable, allowing on-demand order expansion. Using a measured smartphone array, comprehensive objective and subjective evaluations demonstrate the superiority of the TD method. It excels in signal fidelity, spatial accuracy, and temporal consistency, outperforming baseline and frequency-domain approaches. The TD method also maintains its advantage in adverse conditions, showing remarkable robustness against noise and multisource environments. It provides a practical, high-performance pathway for enabling high-fidelity spatial audio capture on ubiquitous consumer devices.
In the last decade, digital technology and online platforms have revolutionized the music industry. Streaming services and social media have allowed artists to reach global audiences and curate personal, interactive relationships with fans. Traditionally, artist similarity has been assessed through audio analysis, but the industry shift necessitates new methods: social media offer a rich multimodal dataset for analyzing artist similarity, considering follower intersections, communication styles, and content types. This study investigates how Instagram behavior and content correlate with artists’ musical production. An early fusion approach combines visual and textual analysis via vision transformers and Bidirectional Encoder Representations from Transformers models, utilizing a Siamese neural network trained with triplet loss and cosine distance, in order to yield a high-dimensional representation of artist positions in an embedding space. Evaluation through accuracy, precision, and recall confirms that leveraging social media data to assess artist similarity leads to results comparable to those achieved by traditional audio-based models, thus highlighting the potential of social media analysis in understanding commonalities and differences among artists.
The ISO/MPEG international standard on MPEG-I immersive audio was issued in November 2025 by the MPEG Audio group (ISO/IEC JTC 1/SC 29/WG 6). It provides technology for a compressed audio representation and real-time interactive rendering within virtual and augmented reality applications with six degrees of freedom. It enables efficient bitrate management, high-quality storage, and transmission of virtual environments, including audio sources with spatial dimensions and specific radiation traits (such as musical instruments), along with geometric descriptions of acoustically relevant scene details (e.g., walls, doors, sound reflecting, and occluding elements). The audio rendering process incorporates comprehensive modeling of room acoustics and intricate acoustic phenomena, including occlusion, reflection, and diffraction caused by sound obstacles, Doppler effect, and dynamic environment changes triggered by user interactivity. This article presents an overview of background, development, underlying technology, and selected application aspects of the new standard.
This qualitative study examines immersive music-production workflows in professional commercial contexts, predominantly within Dolby Atmos environments. Using an interpretivist design, the study combined 30 semi-structured interviews with nonparticipant observations and analyzed transcripts and field notes through conventional content analysis with triangulation. Findings suggest that practice progresses through three stages: preproduction, production, and postproduction. In preproduction, teams align expectations through experiential listening, plain-language briefs, spatial storyboards, and early validation of stems and metadata. In production, placements of sound sources and room acoustics function as compositional parameters that shape localization and movement, inform recording choices that retain room character and multimicrophone flexibility, and expose limitations of monitoring and rendering tools. In postproduction, creative aims are balanced with the need for consistent translation across playback systems, supported by template-based routing, metadata policies, and explicit bed, object, and low-frequency-effects practices. Across these stages, practiand economic pressures. Reported responses include stakeholder education, format-agnostic workflows, future-proofed session design, and early artist engagement. The study contributes a stage-based model of immersive production practice, a taxonomy of bed and object strategies, and guidance for achieving reliable translation between loudspeaker and headphone playback.
Binaural reproduction of microphone array recordings has become an important technology in the research and consumer sectors. Several commercially available spherical microphone arrays have been introduced over the years along with various methods for binaural rendering of array recordings. Most of these methods have been evaluated individually, typically using only one specific microphone array. However, a comprehensive and systematic perceptual evaluation combining different methods and various microphone arrays is lacking. This study presents the results of a listening experiment comparing the motion-tracked binaural method, various Ambisonic binaural decoders, and the parametric binaural rendering method COMPASS using loudspeaker orchestra recordings with six different microphone arrays from two rooms, the Berliner Philharmonie and a laboratory space resembling a small chamber music venue. The experiment assessed the binaural renderings with respect to overall listening experience and four perceptual attributes from the Spatial Audio Quality Inventory in comparison to a reference recorded with a head and torso simulator. The results provide detailed insights into which rendering method and array combination provides a high overall listening experience while preserving the assessed perceptual attributes externalization, coloration, source position, and presence. Moreover, the results indicate the extent to which the assessed perceptual attributes contribute to overall listening experience.
This paper presents the main ideas behind ctfr, an extensible, user-friendly Python package for efficiently combining time-frequency representations (TFRs) of audio signals into a single representation that captures the best aspects of each, achieving high resolutions in both time and frequency. The authors develop and evaluate algorithmic tweaks and approximation schemes for existing TFR combination methods, with significant performance improvements over baseline implementations. In addition, combined TFRs are employed in training a deep learning system for note transcription from audio performances, showing improved results over traditional TFRs, thus demonstrating the effectiveness of using combination methods in audio processing and music information retrieval pipelines.
The characterization of nonlinear distortion in electronic devices, such as audio amplifiers, is traditionally based on the measurement of spectral components generated by the device and absent from the input signal. However, analysis based on nonlinear system models reveals the existence of additional distortion contributions that coincide with the fundamental frequencies and are therefore neglected by conventional metrics. This work develops an accurate and low-complexity measurement procedure to estimate these "hidden" components based on one of the most general block-oriented nonlinear models, namely the parallel Wiener-Hammerstein model. The procedure enables both the computation of more representative nonlinear distortion metrics, accounting for all distortion contributions, and the analysis of the phenomenon in the time domain. Finally, the study discusses the practical limitations of the method and presents its validation through both numerical simulations and experimental applications on real-world audio amplifiers.
Rectangular horns are essential components in professional audio systems where precise directivity control is important. Although a horn’s directivity is known to be influenced by its finite mounting enclosure, a shape-optimization method to improve directivity control under these realistic conditions has been lacking. To address this, an optimization method based on a hybrid model is developed. The model couples the mode-matching method for the internal sound field of the horn with the boundary element method to accurately compute the modal radiation impedance at the mouth. This provides the mode-matching method with a realistic boundary condition that fully accounts for the enclosure. The optimization method then employs this hybrid model within a gradient-based procedure to design a rectangular horn that maintains constant coverage angles in both horizontal and vertical planes over a wide frequency band. A physical prototype was manufactured, and its directivity measurements show good agreement with simulations, validating the proposed method as a predictive tool for high-performance horn design.
This study investigates the use of dipole loudspeakers in addition to typical two-way loudspeakers arranged as a classical stereo setup in a room. The dipole loudspeakers are placed on top of the two-way loudspeakers, their null oriented toward the listener to only contribute to the reverberant sound, adding strong early reflections. Loudness perception for playback over the two-way loudspeakers only and playback over both the two-way and the dipole loudspeakers were compared using a two-alternative forced-choice paradigm. Different music and noise stimuli, different relative levels between two-way and dipole loudspeakers, and two listener positions in the room were used. Results show that adding the dipole-reproduced sound can increase the perceived loudness despite keeping the sound pressure level constant. That effect can be linked to changes in the resulting spectrum, energy decay, interaural correlation, and spatial perception when adding dipole loudspeakers. A binaural loudness model based on the energy received at both ears but not considering the interaural correlation was used to predict the measured level offsets for equal loudness. The predictions were accurate for most stimuli but failed for impulsive sounds where the room reverberation was more perceptible.
Scattering delay networks (SDNs), a class of artificial reverberators with physically interpretable parameters, provide an efficient synthesis of room acoustics accounting for wall absorption properties. This paper builds upon a previously proposed highly parametrized and real-time implementation of an SDN, exposing octave-band absorption coefficients for each room wall. The individual manipulation of these coefficients can be challenging due to their high dimensionality. A 2D parameter space (2PS) is proposed to facilitate the navigation of the absorption coefficients. The 2PS is obtained using principal component analysis as an initial dimensionality reduction of a dataset of absorption coefficients, followed by a relaxation procedure to create a seamless 2D representation of the coefficients. A twofold evaluation of the proposed 2PS was conducted: (a) the 2PS was compared to the original space and the raw principal component analysis in a reverb matching task for the tuning of SDN coefficients, and (b) a usability test with expert audio professionals supported the potential of the 2PS from a user standpoint.
In extended reality, dynamic rendering updates the auralization as changes occur in the scene. Directivity is a characteristic that may require dynamic rendering for moving sound sources. This study investigates how the reverberant component of a directional virtual loudspeaker, used as a controlled proxy for a human talker, is perceived as it rotates and whether its directivity contributes to the perceived naturalness of the scene's acoustics. An experiment was conducted in a virtual environment in which recorded speech and vocal sounds were reproduced through a rotating virtual loudspeaker that was visually overlaid with a human avatar. Four directivity implementations were evaluated: (1) fully measured directivity, (2) direct-sound directivity only, (3) static reproduction without orientation-dependent directivity, and (4) an incongruent inverted-reverberation condition. Participants rated the perceived naturalness of each scene on a seven-point scale. Results indicated that the most physically realistic rendering was perceived as most natural, while partial or incorrect directivity implementations reduced realism. Analysis of the associated acoustic parameters showed that deviations in interaural level difference and direct-to-reverberant ratio were systematically related to changes in perceived naturalness.
This study addresses potential challenges in evaluating reproduction systems with complex, spatially dynamic audio material because current standardized methods may lead to biased or unreliable results. To investigate this, a listening experiment was conducted comparing two assessment methods-continuous and overall evaluation-applied to two attributes, basic audio quality and surrounding, using spatially dynamic content reproduced on two reproduction systems (a stereo and a 3D surround configuration). To enable comparison with overall evaluations, continuous ratings were summarized using different strategies. Although ratings were mainly influenced by the reproduction system, spatial variation, and program item, the choice of evaluation method also had a significant effect on the scores. This effect varied depending on the attribute and the metric used to summarize the continuous data. In cases where differences were observed, continuous evaluations consistently produced higher scores than overall ratings, regardless of the summary metric. These findings indicate that continuous evaluation can capture perceptual variations over time that are lost in overall ratings, suggesting it can be a useful approach when assessing attributes influenced by spatially dynamic changes.
When employing a finite impulse response filter for phase correction alongside frequency response compensation, various side effects such as preringing and time-domain coloration can negatively impact perceived audio quality. This paper introduces an optimization strategy for the filter's phase response, aiming to balance these effects so they remain below the threshold of hearing. The method divides the phase into frequency bands for stable optimization consistent with psychoacoustic principles. The approach is further extended to address application-specific goals, including impulse-response decay, stereo image width, and interchannel matching. In test convolutions based on multiposition room measurements, the proposed optimization reduced early reflections by about 5 dB, confirming its effectiveness in improving spatial clarity.
In the Interactive Sound Zones for Better Living (ISOBEL) project, a sound zone system was developed to create two personal listening experiences simultaneously in health care settings. The perception and audio quality of experience of sound zones were evaluated through listening tests, where perceptual attributes were rated in different scenarios with variations in reproduction mode, audio bandwidth, and reproduction levels. The tests were conducted in a simulated hospital setting to improve the ecological validity of the findings. All perceptual attributes were found to be significantly influenced by one or more of the selected physical factors. In audio-on-audio scenarios, the use of sound zones significantly decreased the perceived distraction and increased the liking compared with wide-angle reproduction. The influence of reproduction levels and audio bandwidth depended on the use case and audio content. Distraction was found to be the main predictor of the audio quality of experience of sound zones. The results help refine the experimental methodology for future perceptual tests within the project. This study provides a valuable contribution to understanding how to enhance listening experiences simultaneously for multiple users.
Since the creation of the spatially oriented format for acoustics (SOFA, Audio Engineering Society standard AES69), numerous databases of head-related transfer functions (HRTFs) are now available as standardized SOFA files. However, the methodologies for measuring and postprocessing HRTFs vary significantly across laboratories. This leads to objective and perceptual inconsistencies between HRTF databases and makes it challenging to integrate multiple databases into a single repository to facilitate wide-scale research and application. This paper introduces a normalization procedure, applicable to any HRTF data set, aimed at enhancing the consistency across HRTF data sets obtained from different laboratories while preserving the spatial information essential to HRTFs. The proposed approach consists of six processing steps: low-pass filtering, temporal alignment, temporal windowing, diffuse-field equalization, low-frequency extrapolation, and far-field correction. The normalization was evaluated on 17 HRTF data sets of the same dummy head by means of acoustic analyses and auditory simulations and further validated with respect to a database of 54 human subjects. Results show that the proposed normalization improves data set applicability and consistency while maintaining the directional cues within each data set.
Digital equalizers typically aim to emulate analog responses, and the finite digital bandwidthleads to a divergence from the analog response near the Nyquist frequency. This paper developssets of low-pass, band-pass, and high-pass filters that minimize this divergence. The first setis developed using the matched z-transform with numerators designed to control the high-frequency behavior. The second set generalizes the bilinear transform filters to provide a closermatch to the analog counterparts at highQvalues. A symmetric parametric equalizer is alsoderived that generalizes previous designs. This equalizer is then modified to better match theanalog filter response near the Nyquist frequency
The LA-2A compressor is a popular optical dynamic range compressor and is much used in popular music productions, particularly as a vocal compressor. There are many claims about sonic differences between different LA-2A versions, but most of these are anecdotal with no empirical evidence to back them up. This paper explores the objective differences between six hardware LA-2A compressors, three vintage Teletronix units, and three Universal Audio reissues. Firstly, it conducts a series of measurements on six compressors, using frequency response sweeps, total harmonic distortion measurements, and tone burst to explore their attack and release characteristics. The results showed subtle differences, particularly around THD and attack and release. Subsequently, an ABX listen test was conducted with 17 trained listeners to explore the audibility of these differences. The test comprised comparing three songs where the vocal tracks had been compressed using the two most dissimilar LA-2As. The results showed a mixture of significant and nonsignificant results, suggesting that discrimination is context-dependent and difficult to discern in typical music productions.