
Acoustic, rooftop velocity, and window acceleration measurements were obtained during the NASA Artemis II launch at community sites 26–27 km from LC-39B. The instruments were not co-located; observations are site-specific. The acoustic record peaked at 104 dB re 20 μPa (1-s overall sound pressure level); roof-mounted velocity peaked at 2.08 mm/s; and at a residential site, a glass-mounted accelerometer recorded a distinct out-of-plane response near 7.3 Hz, with peak acceleration of 0.36 g. All three showed peak energy in the 2–20 Hz band, identifying a building-component response not captured by conventional community-noise metrics.
Vessel noise is a recognised stressor for marine fauna, yet acoustic pressure from large-scale maritime events in urban coastal settings remains poorly quantified. Using calibrated hydrophone recordings and Automatic Identification System (AIS) data from the DeuteroNoise dataset, we document near-shore noise during the 37th America's Cup (Barcelona, October 2024) against non-event baselines. Broadband levels increased by 5–8 dB re 1 μPa, day–night variability, ∼2.3 dB in the representative summer baseline, collapsed to <1 dB during the event, and 63 and 125 Hz third-octave bands rose by 4 dB during the day and up to 13 dB at night, indicating sustained low-frequency pressure throughout event days.
Beacon acoustic sources can aid long-range underwater navigation since using an inertial navigation system alone becomes inaccurate over long duration. Here, estimates of the relative position of a single hydrophone mounted on a small underwater platform are used only while recording a 2-min-long low-frequency (200-300 Hz) transmission from a beacon source located at depth ∼1100 m and ranges up to ∼148 km in the vicinity of two seamounts. Accounting for the hydrophone motion in the matched-filtering process during the source transmission allows to estimate the bearing of this single beacon source for navigation purposes despite environmental complexity.
Research with younger adults showed that cloned voices are more intelligible than human voices in noise, with a benefit of 13.4%. This study tested whether this benefit extends to 40 middle-aged listeners (45-65 years), as this population may show emerging difficulties with speech-in-noise. Participants recognised sentences by ten human voices and ten voice clones in four noise levels. Cloned voices were 11.8% more intelligible, with benefits enhanced at the two most severe noise levels (15.9% at -6 dB and 17.5% at -3 dB), suggesting cloned speech enhanced perception in middle-aged listeners, potentially by reducing listening effort and compensating for emerging age-related auditory-cognitive decline.
Korean aegyo is a socially recognized childlike speaking style used predominantly in romantic interactions among adults. This study examined vowel space modification in aegyo by analyzing formant frequencies from twelve Seoul Korean speakers who produced identical scripts in aegyo and non-aegyo styles. Results show that aegyo speech features a significant increase in F1 values across vowels and selective fronting of front vowels, leading to some expansion of the vowel space but mainly shifting to higher F1. These findings motivate the hypothesis that adult speakers construct a childlike vocal quality by increasing F1, possibly through a combination of different factors like global vowel lowering and partial fronting.
Excess binaural fusion can be problematic for listeners in noisy environments, leading to fusion of speech from multiple talkers. Previous research found that hearing aid users experience broad binaural fusion, fusing concurrent dichotic vowels with voice pitch differences of up to an octave [Reiss and Molis (2021) J. Assoc. Res. Otolaryngol. 22(4), 443–461]. This study examined whether cochlear implant (CI) users, compared to normal-hearing listeners, also exhibit abnormal fusion of concurrent double vowels in free-field conditions. The results revealed that CI users often fuse and mis-identify vowels even with large pitch differences, suggesting that they face additional challenges in noisy environments due to excessive binaural fusion.
In pulse-echo sonar, the temporal coherence of echoes from the seafloor can provide an acoustic means to observe subtle long- or short-term changes in the seabed. Measurements of scattering from a seabed composed of sand, clay, and gravel in the Gulf of Maine were made over a period of 8 months. A parametric estimation approach, using a piecewise exponential function, is described and applied to the data to more accurately estimate nonstationary decorrelation rates. This approach may aid in quantifying and describing the rate of change in seafloors with mixed composition and in complex ocean environments with multiple forcing mechanisms.
Satellite-derived sea surface remote sensing data enable large-scale subsurface sound speed profile (SSP) inversion; however, considerable inversion errors commonly arise due to the lack of in situ constraints. To address this limitation, this Letter proposes a bidirectional long short-term memory-based framework that integrates sparse observations to improve inversion accuracy. Experiments in the Kuroshio Extension and the Philippine Sea demonstrate that sparse observations yield improved inversion accuracy, with remarkable improvement without sea surface data. Depth-layer sensitivity analysis reveals that optimal sparse sampling depth intervals differ under scenarios with and without sea surface data, offering guidance for rational underwater observation deployment.
In Beijing Mandarin, r-suffixation is a diminutive formation process that adds a rhotic suffix, which typically overwrites syllable-final nasal codas. Previous work by Zhang [Phonology 17, 427-478 (2000)] used nasal airflow data to argue that the categorical preservation of nasality in suffixed /CVŋ + ɻ/ as opposed to /CVn + ɻ/ is driven by salient differences in nasalization in the stem. A nasometry study revealed no significant nasalance difference between CVn and CVŋ stems, contradicting previous aerodynamic evidence. Nevertheless, a significant nasalization contrast remained in r-suffixed forms. These findings suggest that the alternation is reanalyzed as phonological instead of a reflection of phonetic differences in vowel nasalance in the stem.
Monitoring the abundance and distribution of marine mammals is essential for evaluating human and environmental effects. Passive acoustic monitoring offers a non-visual alternative, but estimating animal numbers remains difficult. We introduce Spatial Counting of Animal Numbers (SCAN), a method that counts unique bearings of detected calls to estimate a conservative minimum number of animals present. SCAN is a complementary metric to traditional presence/absence results and calibrated density estimates. The method is demonstrated using sei whale calls recorded by a four-element array on an autonomous glider. SCAN increases the information on marine mammal populations derived from passive acoustic recordings.
This study quantifies the relationship between distributed acoustic sensing (DAS) and traditional hydrophone measurements of ship noise. Analyzing a cargo ship's transit recorded by a DAS array and a nearby hydrophone, a conversion factor of 177.7 ± 3.5 dB (30–50 Hz) was derived, relating DAS strain to hydrophone pressure. A comparison of this conversion factor with a few existing studies highlights the variability in reported values and methodologies, underscoring the need for more extensive analysis and dedicated research to fully characterize DAS calibration.
Using pressure gradient-based methodology, a compact 33-cm, four-element, tetrahedron-shaped hydrophone array is demonstrated to emulate the low-frequency direction of arrival (DOA) finding performance of a vector sensor. Towed from an autonomous surface vehicle in deep water, near the Kelvin and Atlantis II seamount complexes, this array is capable of determining the horizontal DOA (i.e., bearing) of low-frequency broadband (210-310 Hz) transmissions from bottom-moored sources located at a depth of ∼1100 m and ranges up to 176 km away. Both coherent and incoherent DOA finding methods are presented, producing consistent DOA values for signal-to-noise ratios higher than 11 dB for the selected experimental configuration.
Transmission-line models are the simplest physics-based cochlear models that include an explicit representation of the traveling wave, which plays an important role for encoding complex sounds in the auditory periphery. Despite their many attractive features, transmission-line models suffer from well-known shortcomings that prevent them from replicating the experimental data or explicating the physical functioning of the cochlea. The main issue is that they neglect important two-dimensional (2D) hydrodynamic effects, widely known as “pressure focusing,” which enhance the vibration of the cochlear sensory tissue in a wavelength-dependent fashion. Here, we present an efficient method for including 2D pressure-focusing in transmission-line models.
Mandarin apical vowels [ɹˈ] and [ɻˈ] have resisted articulatory classification: are they vocalized prolongations of the preceding sibilant or independent vowel targets? Using ultrasound imaging and generalized additive mixed models with simultaneous confidence intervals, the tongue contour over space and time in two comparisons has been modeled: each vowel against its sibilant onset and the three vocalic realizations against one another. [ɻˈ] continues the /ʂ/ posture while [ɹˈ] diverges from /s/, and the three realizations differ from one another, the two apical vowels from [i] and from each other. The Mandarin “apical vowel” label may subsume more than one articulatory strategy.
While auditory and visual cues jointly support vowel perception, the role of visual information alone in vowel identification remains understudied. This study examines how Taiwan Mandarin speakers identify rounded vowels using visual-only input. Results reveal an asymmetric pattern between /y/ and /u/, with a bias toward /u/ as the default rounded vowel and greater difficulty identifying /y/. Signal detection analyses further indicate a higher response criterion for /y/. This default preference for /u/ appears to reflect phonological distributions in the language or a universal preference for the more peripheral vowel. Varying degrees of visual input did not modulate identification strategies, and visual information provided only limited support for distinguishing between /u/ and /y/.
Successive room impulse response (RIR) measurements exhibit subtle variations, mainly caused by atmospheric fluctuations due to spatiotemporal air-temperature differences. This work extends the image source method by incorporating temperature-induced variability through global deterministic drift and stochastic local volatility. Measurements in a shoebox room show that the proposed model reproduces key behaviors observed in practice, including the structure of differential RIRs, temporal variance of early reflections and coherence of late reverb, which fluctuation-free simulations fail to capture. The method increases realism without increasing the algorithmic complexity, enabling simulations that better reflect the variability encountered by algorithms in real measurement scenarios.
In ocean acoustics, vertical array depth shifts caused by currents lead to missing shallow-water acoustic field data. This study applies a Gaussian process (GP) regression framework with a Normal-Mode-Based Kernel to address the extrapolation problem in transmission loss reconstruction. Based on simulations of the 2024 South China Sea experiment environment and validation with measured data from the SWellEx-96 experiment, the GP approach successfully extrapolates the acoustic field to missing depths. Results demonstrate that, compared to using the transmission loss value at the closest available depth as a reference, the reconstruction error is significantly reduced. Furthermore, by incorporating dominant modal prior information, the Normal-Mode-Based Kernel–GP method maintains reliable performance even for complex sound fields, improving reconstruction stability under array deployment uncertainties.
Psychoacoustic metrics offer the means to analyse sound qualities using sophisticated models of human perception. A substantial drawback is the need for high-resolution acoustic data as input. Quasi-psychoacoustic metrics are proposed to address sound quality analysis when high-resolution acoustic data are unavailable. The proposed metrics aim to approximate sound qualities of loudness, sharpness, and tonal loudness using acoustic data with a much lower time resolution. As such, the metrics are expected to be primarily applicable to sounds featuring relatively slow modulations in intensity. A particular application case is used to validate the metrics: sound quality predictions for unmanned aircraft systems.
In B-mode medical ultrasound, a ringdown artifact is a vertical, linear artifact composed of horizontal bands. In this study, in vitro ringdown artifact models are investigated. Frequency domain spectral analysis of the ringdown artifact shows narrowband peaks, with a difference frequency that matches the spacing of bands in the ringdown artifact. These data are consistent with linear superposition of narrowband waves. Reconstruction of individual focused pulse acquisitions visualizes a triangular ringdown artifact, indicating that the linear shape of the ringdown artifact may itself be an artifact of composite imaging.
This study tested whether targeted formant enhancement improves recognition of natural and N95-muffled IEEE sentences in quiet and four-talker babble. Twenty-one normal-hearing native English listeners identified unmodified, F2-enhanced, and combined F2-F3-enhanced sentences in quiet and at babble. While formant enhancements did not improve recognition in quiet, they significantly facilitated speech recognition in babble, with combined F2-F3 enhancement yielding broader and more consistent benefits than F2 enhancement, especially for N95-muffled speech. These findings support multi-formant enhancement as a promising approach for improving intelligibility of speech produced with and without face masks in noisy environments.