
This study describes the design and evaluation of a crosstalk cancellation (CTC) system intended to reproduce binaural audio for two listeners simultaneously. A compact loudspeaker array was combined with measured loudspeaker-to-ear transfer functions to generate band limited virtual sources and frequency domain CTC filters. The signal processing framework, including the implementation of the crossover network and the construction of the frequency-dependent CTC matrices, is presented in detail. The resulting filters are examined in both the time and frequency domains to illustrate their behaviour and to show how individual loudspeakers contribute across frequency bands. While the theoretical foundations of the array geometry are established in a companion paper, the present work focuses on the physical implementation of the system and its perceptual evaluation through subjective listening tests in which 25 participants localized virtual sound sources reproduced at two listening positions. After correction for front-back confusions, a mean absolute error of approximately 15° was achieved at both listening positions, with no statistically significant overall difference in accuracy or consistency between the two seats. The results indicate that listeners were generally able to identify the intended source directions, demonstrating that stable spatial cues can be delivered simultaneously at both positions.
Auditory masking occurs when one sound interferes with the detection of another. While knowledge of this sensory phenomenon is growing, much of our understanding remains theoretical or based on terrestrial models. Empirical data become especially important when considering hearing in marine mammals that rely heavily on acoustic cues for foraging, communicating, and avoiding predation. To better understand how anthropogenic noise can influence hearing in otariid, odobenid, and phocid pinnipeds, detection thresholds for tonal sounds were measured for a trained California sea lion, Pacific walrus, bearded seal, and spotted seal within noise of progressively constrained spectral content. The frequency bandwidth of noise that contributed to the masking of a given tone (the "critical bandwidth") was identified at up to six frequencies between 100 and 16 000 Hz. Absolute critical bandwidth values increased with increasing frequency in all four species, although taxon-level trends were apparent. The sea lion and walrus showed similar auditory adaptation to noise, with critical bandwidths ranging from 16%-35% and 21%-43% of center frequency, respectively. Critical bandwidths were narrower in both phocids, ranging from 14%-29% of center frequency across tested frequencies; particularly narrow bandwidths at 500 Hz and below suggest additional auditory specialization for seals at lower frequencies.
A triple-resonant Janus-Helmholtz transducer achieves broadband performance via longitudinal, liquid cavity, and flexural modal coupling. However, uniform in-phase excitation of its four driving modules causes anti-phase superposition of flexural vibration, severely suppressing flexural resonance output and restricting acoustic radiation efficiency. To address this inherent modal cancellation issue, this work proposes a magnetostrictive-piezoelectric hybrid excitation strategy that leverages the intrinsic 90° material vibration phase difference to enable constructive flexural modal superposition while preserving stable longitudinal and liquid cavity resonance responses. Finite element simulations based on the piezoelectric-piezomagnetic analogy verify the enhanced bandwidth and radiation performance of the hybrid configuration over pure piezoelectric and pure magnetostrictive counterparts. A prototype device is fabricated and experimentally characterized, with results showing strong agreement with simulation predictions. The prototype exhibits resonant frequencies at 350, 800, and 1250 Hz, achieving an operational bandwidth exceeding two octaves. Further tests reveal that separate excitation of piezoelectric and magnetostrictive drive modules reduces impedance variation and doubles driving efficiency in comparison to parallel combined excitation, providing reliable guidance for the high-efficiency design of hybrid broadband underwater acoustic transducers.
The small slope approximation is a commonly used model for scattering of waves from rough surfaces. Its limitations are currently unknown for a power-law roughness spectral density, used to model natural terrain and seafloor roughness. This work compares the small slope approximation to the boundary element method, a numerical solution of the Helmholtz boundary integral equations governing acoustic scattering. These comparisons were used to find the validity limits as a function of dimensionless root mean square height, kh, and root mean square slope, s, where k is the wavenumber. The two lowest terms in the small slope series were used, and both the Dirichlet and fluid-fluid boundary conditions were examined. Both kh and s are found to be important parameters for the validity of the small slope approximation. Different spectral exponents have similar maximum valid kh. The maximum s is almost constant as a function of the power law outer scale, but varies as a function of the spectral exponent. Comparisons with other scattering models are discussed, and aspects of scattering from very large roughness are explored.
The Global Ocean Observing System recently designated ocean sound an essential ocean variable to support passive acoustic observation of the ocean. Underwater gliders have the demonstrated ability to fill gaps within usual ocean monitoring networks and are increasingly used for passive acoustic monitoring studies focusing on acoustic detection and identification of soniferous species. Comprehensive understanding of glider-generated noise and flow noise and characterization of their potential impact on acoustic measurements are critical to further development of glider-borne applications to ocean sound measurements. This study investigates flow noise, generated by glider motion through water, and its potential contribution to noise levels measured from underwater gliders, particularly at low frequencies. From glider-borne acoustic recordings collected in the Ross Sea, free from low-frequency shipping noise contribution, this study provides a comprehensive characterization of flow noise at speed through water representative of underwater glider operations (15-63 cm s-1). It quantifies the effects of flow noise on measurements at frequencies ranging from 3 to 100 Hz. Finally, this study demonstrates the possibility to collect unaffected measurements above 20 Hz from an underwater glider and discuss solutions to reduce flow noise or mitigate its effects.
Soundscapes are increasingly understood as multidimensional environments that shape perceptual, emotional, cognitive, and physiological responses. However, most soundscape studies rely on generic designs that may not reflect individual preferences or sensitivities. This study compares personalised soundscapes, constructed by participants to represent specific emotional-cognitive states, with generic soundscapes across perceptual, emotional, cognitive, and physiological domains. Participants created personalised soundscapes representing Relaxation, Pleasant, Energetic, Annoying, and Focus states using a structured design interface. These were evaluated alongside generic counterparts using ISO 12913-based perceptual scales, emotional ratings, a cognitive Stroop task, and physiological measures including electrodermal activity (EDA), photoplethysmography-derived heart rate variability, and electroencephalography. Within-subject non-parametric analyses were applied throughout. Personalised soundscapes were consistently rated as more pleasant, calmer, and more closely aligned with their intended experiential states than generic soundscapes, with parallel effects observed for emotional evaluations. Cognitive results showed a task-dependent pattern: Energetic soundscapes elicited faster Stroop response times and lower Inverse Efficiency Scores than Focus soundscapes, despite Focus being perceived as calmer and more conducive to self-regulation. Physiological responses were largely descriptive; most effects did not survive correction for multiple comparisons. EDA showed selective significant contrasts for Focus soundscapes, while cardiovascular and electroencephalographic measures showed directionally consistent but statistically unreliable trends. Overall, the findings indicate that personalisation enhances perceptual and emotional alignment and yields partially convergent physiological patterns, while highlighting the importance of task context in soundscape evaluation.
This study investigated whether different types of dysarthric speech (with comparable baseline intelligibility) are equally susceptible to background noise. Using intrinsically degraded speech from four individuals with dysarthria, each representing a distinct motor speech disorder subtype, we examined how intelligibility is impacted by masker type (stationary vs fluctuating noise) and signal-to-noise ratio (SNR). Ninety-five listeners completed a listening task with stationary speech-shaped noise or ecologically valid cafeteria noise, across multiple SNRs. Although all degraded signals showed intelligibility declines in noise, the extent and pattern of decline varied, suggesting that not all forms of pathological degradation interact with background noise in the same way. Notably, the typical intelligibility benefit associated with fluctuating maskers (i.e., masking release) was absent for all dysarthric signals. These findings reveal that equivalently intelligible but qualitatively different dysarthric signals do not respond uniformly to environmental challenges. The results have implications for models of speech perception, which must account for complex interactions between intrinsic (speaker-specific) and external (environmental) degradations, as well as for real-world communication and intervention strategies.
Due to the low attenuation in the water, an acoustic wave is currently considered a promising medium for long-distance underwater information transmission. Acoustic vortex (AV) provides an alternative to expand the channel capacity by exploring the orbital angular momentum dimension. In this work, based on the ring-array of sectorial transducers, we propose a scheme for constructing double-ring perfect acoustic vortex (DR-PAV) beams by applying the Fourier transformation of quasi-Bessel acoustic vortex (QB-AV) beams with different radial wave numbers. The radius and ring width of the inner and outer rings of the DR-PAV beam remain nearly constant, and the topological charge (TC) of the inner and outer rings are independently controllable. Subsequently, the inner and outer rings of the DR-PAV beam are further treated as two independent channels for data transmission. At the receiver end, decoding of DR-PAV beams with multiplexed TCs is achieved by using a simplified double-ring receiver array. This multi-channel underwater data transmission scheme is expected to be applied in acoustic communication in the future.
By auditory and visual assessment, we investigate whether architects and acousticians describe their perception of concert hall acoustics in the same language. Four concert halls were represented in controlled in virtual reality environments. A total of 43 participants (23 acousticians, 20 architects-all with experience and interest in concert hall acoustics) participated in the experiment. During the experiment, all the participants were exposed to audio and video excerpts in two different positions across the four concert halls. An individual vocabulary profiling method (IVP) was used to collect the participants' attributes and textual ratings of the virtual spaces. The data were further analyzed via hierarchical multiple factorial analysis. The results reveal significant distinctions in the perceptual vocabulary and evaluation strategies employed by the two groups of professions, quantifying the perceptual dimensions without consensus vocabulary. They reveal a quasi-bidimensional perception space on architects using visual attributes and on acousticians using auditory attributes, but a higher perception space for the inverse modalities. Finally, the correlation between subjective and objective acoustical parameters reveals discrepancies between both expert groups on which objective parameter correlates with subjective balance and timbre, while it shows agreement for other subjective parameters.
This paper presents the underwater radiated noise semi-analytical model (URN-SAM), a SAM kernel-based Python framework for mapping regional URN from maritime traffic using open and standardized datasets. The model combines European Marine Observation and Data Network vessel density layers, General Bathymetric Chart of the Oceans bathymetry, World Ocean Atlas climatologies, class-dependent source levels derived from the JOMOPANS-ECHO formulation, and propagation kernels precomputed from Bellhop simulations to estimate received levels in the 63 and 125 Hz Marine Strategy Framework Directive indicator bands. Applied to the Spanish Levantine-Balearic Marine Demarcation, URN-SAM reproduced coherent large-scale spatial patterns dominated by major shipping corridors, marked seasonal variability linked to monthly traffic inputs, and clear differences between the two frequency bands. Sensitivity tests showed that vertical aggregation is a nontrivial modeling choice: changing the shallowest evaluation depth materially altered annual maps constructed from the maximum across depths and their exceedance statistics. A numerical benchmark against Bellhop showed close agreement for the selected kernel implementation (root mean square error 1.68 dB; Pearson correlation 0.990). Comparisons with hydrophone data and official Spanish Levantine-Balearic Marine Demarcation products indicate that URN-SAM is more robust at regional and offshore scales than in shallow coastal environments. URN-SAM is proposed as a transparent and computationally efficient screening tool for regional diagnosis, hotspot identification, and category-resolved contribution analysis based on aggregated vessel density inputs.
To detect age-related hearing loss (presbycusis), beyond objective audiometric measurements, self-report instruments are commonly employed. However, clear evidence regarding the methodological and measurement quality of self-report instruments for assessing hearing loss in older adults is lacking. This systematic review, for the first time, aims to identify and evaluate self-report instruments for assessing age-related hearing loss and related functional and psychological consequences in everyday life among older adults. Following PRISMA and COSMIN guidelines, we preregistered in PROSPERO (CRD42023449606) and searched PsycInfo, PubMed, Scopus, and Web of Science. Out of 2459 records, we included 97 studies evaluating 20 self-report instruments. These instruments were categorized into three groups based on their main outcome assessed: (a) listening and communication difficulties (five instruments), (b) communication difficulties and related psychosocial consequences (seven instruments), and (c) listening and communication difficulties and other psychosocial consequences (eight instruments). Some promising questionnaires were identified, although the overall psychometric and methodological quality of the evidence was limited. More rigorous validation studies of these instruments are recommended, as their use alongside audiometric tests may facilitate the earlier detection of unaddressed hearing-related difficulties and their functional and psychosocial consequences, while guiding tailored strategies to improve older adults' quality of life.
Surgical mesh is widely used in soft tissue reinforcement treatments as it results in a lower recurrence rate of complications. However, there is a possibility of abnormalities after mesh implantation. These problems can result in chronic discomfort and demand further surgical procedures and postoperative monitoring. Hence, mesh visualization and detection are essential for understanding postoperative outcomes and surgical planning. However, using traditional imaging modalities, these meshes are difficult to visualize after implantation against surrounding tissues. Therefore, it is challenging to locate the implanted mesh accurately. The hypothesis for this work was that ultrasound shear wave elastography (SWE) and Nakagami parametric imaging could be useful methods for localizing and assessing the implanted mesh, as the mesh differs in mechanical stiffness, along with its echogenicity and the statistical distribution of backscattered signals compared to the surrounding tissue. This study demonstrates the potential of Nakagami imaging and SWE for accurately localizing surgical meshes. While both methods allowed for accurate mesh location recognition, Nakagami imaging provided better localization for meshes located at larger depths. According to these findings, this method could especially be helpful for patients with thick abdominal walls. This could enhance the clinical evaluation of hernia mesh implantation and improve diagnostic accuracy.
Classroom noise can substantially interfere with children's access to spoken instruction, yet speech perception is typically assessed using clinic-based measures collected under controlled listening conditions. This study examined (1) how classroom signal-to-noise ratio (SNR) and developmental level influence speech recognition (SR) and listening comprehension (LC), and (2) whether clinic-based SR predicts classroom LC as effectively as classroom-based SR. One hundred eleven normal-hearing students in Grades 2-5 completed SR (WIPI) and LC (TROG-2) tasks in occupied classrooms under multiple noise levels, as well as SR testing in a clinic setting. Generalized linear mixed-effects models were used to estimate acoustic and developmental effects, and receiver operating characteristic analyses compared predictive accuracy. Higher SNR was associated with significant improvements in both SR and LC, and older students consistently outperformed younger students across acoustic conditions. Although SR performance was higher in the clinic than in the classroom, classroom-based SR provided substantially better prediction of classroom LC than clinic-based SR. These findings demonstrate that assessment context shapes observed listening performance and underscore the importance of ecologically valid measures for predicting functional comprehension outcomes in real educational environments.
Passive acoustic surveys of soniferous marine communities can provide useful metrics for characterizing behavior, abundance, and ecological influences; however, they typically offer limited information about the spatial distribution of biological sources. Recent advances in distributed acoustic sensing (DAS) offer the potential to conduct wide-area surveys of bioacoustic activity that would be logistically infeasible with traditional hydrophone measurements. Here, a 2.2 km DAS cable deployed in Narragansett Bay, RI, was used to map the spatial distribution of oyster toadfish chorusing over a 24-h period during their spawning season in June 2025. A localization technique was developed where nearfield beamforming was performed along a rolling subaperture across the entire array, in order to account for the lack of coherence across widely spaced channels. This method allowed for the localization of biological hotspots, which were concentrated along a breakwater at the east side of Narragansett Bay. Long-term, single-hydrophone recordings also reveal seasonal and diel patterns for the same toadfish chorus, demonstrating how temporal data from traditional acoustic sensors can complement spatial DAS measurements.
While underwater acoustical three-dimensional (3D) imaging offers several benefits compared to two-dimensional imaging, the requirement of a uniform planar array (UPA) with a large number of elements for image reconstruction limits its widespread use. The most popular array structure capable of estimating the direction of arrival in 3D space with low complexity is a cross-array, where two linear arrays are placed orthogonally in a symmetric arrangement. Although other orthogonal linear array configurations are feasible, they have not been thoroughly investigated in the context of underwater 3D acoustical imaging. In this work, we studied six different orthogonal array configurations for underwater acoustical 3D imaging. The study proved that an L shaped array placed at the edges of a UPA performs better than any other orthogonal array both in terms of resolution and ability to resolve ambiguity in multiple targets. We also propose a nonlinear beamforming approach using an L-shaped array for real-time 3D image reconstruction. The proposed approach achieves a signal-to-noise ratio of 12.72 dB and a contrast-to-noise ratio of 1.35 while reducing computation time by 94.64% through the use of significantly fewer array elements than conventional delay-and-sum beamforming using a UPA.
The preferences of music ensembles for the architectural and acoustic features of the stages on which they perform have primarily been examined through field studies. Although these studies provide high ecological validity, they are limited in terms of the architectural variations that can be studied and by the confounding effects resulting from the unique features of the halls in which the experiments are conducted. This study took an experimental approach in which six typical architectural features of stage enclosures were first identified through discussions with concert hall design experts. These features were then systematically varied to produce a set of 256 virtual halls. Using latency-free dynamic binaural synthesis in an anechoic environment, chamber music ensembles of up to five musicians were able to perform simultaneously in the same virtual space. The results revealed the influence of architectural features such as stage width and height on perceived Support and Reverberance. Of the many stage acoustic parameters analyzed, only the top-to-horizontal and top-to-side energy ratios were significantly correlated with perceived quality, Transparency, Support, Brightness, and Bass on stage.
Neural binaural synthesis is a promising technique for immersive audio reproduction. However, current methods rely on signal-oriented optimization objectives that treat all time-frequency regions uniformly. This approach conflicts with the selective nature of human auditory perception, as it may overemphasize perceptually masked or sub-threshold components, resulting in audible artifacts in reverberant sound fields. To address this, a perceptual optimization framework inspired by human spatial hearing is introduced. It includes a coherence-gated spatial loss that uses interaural coherence to weight spatial cues according to their reliability. By reducing the contribution of low-coherence regions associated with diffuse reverberation, this mechanism forces the model to prioritize the reconstruction of direct sound and early reflections critical for localization. A multi-scale wavelet loss is also incorporated to measure reconstruction errors across multiple wavelet decomposition levels, addressing the fixed resolution trade-off of the short-time Fourier transform and helping preserve fine-grained temporal transients. The framework was validated on three distinct backbone architectures. Extensive objective evaluations and subjective listening tests under non-head-tracked binaural playback show consistent improvements over state-of-the-art baselines, demonstrating its effectiveness as a model-agnostic training strategy for improving both signal fidelity and spatial accuracy.
Rapid trajectory shifts, complex kinetics, and strong measurement noise are major difficulties that arise in the tracking of underwater passive maneuvering targets. This study suggests an innovative framework of memory-augmented temporal features learning, based on a Long Short-Term Memory (LSTM) model, to overcome these challenges. The LSTM aims to effectively develop and preserve long-term temporal relationships in the formation of the target's state, leading to reliable prediction of location, velocity, and trajectory even during high maneuverability and extensive observed noise. The designed model is configured through state-space physics and investigated under distinct levels of Gaussian measurement distortion. The evaluation of performance is carried out in the mean squared error (MSE) sense to analyze the degree of accuracy. In comparison, the LSTM-based estimation model offers better results than generalized pseudo-Bayesian estimators, including the Interacting Multiple model Extended Kalman Filter (and the Interacting Multiple model Unscented Kalman Filter. The outcomes reveal that the proposed design significantly lowers state estimation deficiencies and shows significant flexibility for various maneuvering behaviors, proving a feasible option for real-time passive tracking in acoustically challenging underwater situations.
The objective of similitude theory is to establish scaling conditions and laws by which the response of a given system can be scaled to infer that of another, distinct system. Exact similitude laws for the structural acoustic response of thin cylindrical shells exist only in the trivial case where length, radius and thickness are scaled equally, because of the coupling between in-plane and transverse deformations. However, under the bending approximation, which applies when bending wavelengths are one order of magnitude smaller than the length and radius of the shell, thickness can be scaled differently to length and radius. This work focuses on similitude laws and conditions for the structural acoustic response of fluid-loaded cylindrical shells, under the bending approximation. The scaling parameters include the shell dimensions, material parameters and the properties of the acoustic domain. A criterion for validity of the bending approximation based on ring frequencies of the reference and scaled in vacuo shells is proposed. In the case of fluid-loaded cylindrical shells, new similitude conditions and laws are derived for exact scaling of the modal radiation impedance. When these conditions are satisfied, satisfactory re-scaling of both the spatially-averaged vibration response and the radiated sound power is obtained.