
In fields such as noise control, medical ultrasound, and acoustic communication, the flexible regulation of reflected sound waves has significant application value. In this work, a dual-band acoustic metasurface was designed using a split hollow cuboid with an open-hole plate (OPSHC) structure, which simultaneously achieves the direction control of reflected sound waves in both frequency bands. An OPSHC is a series structural unit, and the two center frequencies are mainly controlled by the diameters of the two openings in the structure and the position of the open-hole plate. Through finite element simulation, the influence of the center frequency of the metasurface and the position of the open-hole plate on the bandwidth of the anomalous reflection was studied. The results show that when the low-frequency center frequency is fixed, the low-frequency bandwidth of the metasurface increases with the increase in the high-frequency center frequency. When the position of the plate is moved, the low-frequency bandwidth increases and the high-frequency bandwidth decreases. This type of metasurface provides a new technical approach for broadband acoustic metasurface applications in noise control and underwater detection systems.
Air-penetrating and noise-canceling constructions are required for numerous noise control issues. High ventilation performance conflicts with effective sound insulation, and vice versa. For this reason, ventilated noise barriers are currently being intensively researched and developed. One of the most popular solutions is the louvered-type barrier, whose acoustic efficiency depends on its geometric parameters as well as the acoustic properties of the louvers. One of the main challenges is optimizing the acoustic impedance of louver surfaces in order to achieve maximum reflection, absorption, or minimum transmission of sound waves. This paper proposes an analytical solution to the diffraction problem of a plane sound wave incident on a periodic array of similar thin screens with arbitrary impedance surfaces. An infinite system of linear equations is derived, and its numerical solution allows us to find the reflection and transmission coefficients. It has been shown that screens with reactive impedance are necessary to achieve maximum sound reflection. On the other hand, dissipative screens are required for minimal sound transmission. Additionally, the absorption properties of the array have been studied. It has been found that there is an optimal impedance value that provides the maximum absorption coefficient.
Modern, electrically operated heat pumps are characterized by a high degree of efficiency and represent an attractive alternative to conventional heating systems. However, the noise emissions from heat pumps installed outside can lead to increasing noise pollution in densely populated residential areas, which represents an obstacle to widespread use. As part of a research project, a heat pump mock-up was built based on an outdoor unit in the Fraunhofer IBP. With this mock-up, investigations have now been carried out with a prototypical silencer-resonator concept. The aim was to reduce the sound power on the outlet side of the heat pump mock-up. To estimate the effect of this silencer-resonator concept for heat pumps, FEM simulations were first carried out using COMSOL Multiphysics (R) with a simplified model. The simulation results validated the silencer-resonator concept for heat pumps and indicated the considerable potential for sound reduction. A measurement was then set up, with which different silencer lengths and absorber thicknesses in the silencer were tested. The measured sound attenuation was higher than the simulated values. The results showed that porous absorbers with sufficient thickness can achieve effective performance in the mid-frequency range. A maximum sound power reduction of 5.7 dB was achieved with the 0.15 m absorber. Additionally, Helmholtz resonators were implemented to attenuate the low-frequency range and tonal peaks. With these resonators sound attenuation was increased to 7.7 dB.
Multichannel audio is a sound field reproduction technology that uses multiple loudspeakers. Object-based audio is a playback method for multichannel audio that enables the construction of sound images at specified positions using coordinates within the playback space. However, the sound image positions must be manually specified by audio content creators, which increases the production workload, especially for works containing many sound images or feature films. We have previously proposed a method to reduce the workload of content creators by constructing sound images based on object positions in visual images. However, a significant challenge remains since depth localization of the sound image is not accurate enough. This paper aims to improve localization accuracy by changing the range of sound image movement along the depth direction. To confirm the localization accuracy of sound images constructed using the proposed method, we conducted a subjective evaluation experiment. The experiment identified the optimal movement range by presenting participants with visual images synchronized with sound images moving across varying spatial scales. Consequently, we were able to identify the range of sound image movement in the depth direction necessary for presenting sound images with high consistency with the visual images.
The Church of St. Francis in Pula, Croatia, is a well-preserved example of Franciscan gothic sacral architecture from the late 13th century. As preaching was highly valued by the Franciscan order as a way of communicating with the faithful, the study is focused on determining whether speech intelligibility in the church would have been adequate for successful communication between priests and their audience. The archaeoacoustic analysis of the church was performed in four stages: (1) in situ acoustic measurements in the present state, (2) development and calibration of the model of the present state based on measurement results, (3) development of the two models of the presumed historical state based on the calibrated model and historical data, and (4) prediction of acoustic conditions in the present and the historical states in terms of reverberation time T30 and of speech intelligibility in terms of speech transmission index STI. The factors considered in the study were (1) acoustics of the church, (2) profile of the audience (friars and the faithful), (3) layout of the audience areas (choir area in the front of the nave for the friars, back area of the nave for the faithful), (4) positions of the speech sources (altar for addressing the friars, pulpit for addressing the faithful), (5) occupancy (unoccupied and fully occupied church), (6) language used in liturgical ceremonies (Latin and native language), and (7) language proficiency of the audience (native speakers, users of a second language). The results show that (1) fair speech intelligibility (STI >= 0.45 for the faithful as native speakers, STI >= 0.50 for friars as non-native speakers of Latin) can be achieved for 50% of the audience in the choir area and for the entire audience in the back area in favourable conditions (fully occupied church, audience addressed from dedicated speaker positions), (2) the position of the pulpit (close to the audience and considerably elevated above it) is more favourable than the position of the altar (remote, barely elevated above the audience), and (3) in unoccupied conditions, fair speech intelligibility can still be achieved in at least 50% of the back audience area with the faithful gathered close to the pulpit, while it is not possible for the front audience area addressed from the altar. The summary conclusion is that the church of St. Francis in its presumed historical layout(s) would fulfil its primary function in a limited capacity. Fair speech intelligibility would likely have been sufficient for the audience to follow liturgical ceremonies conducted in the church, but not without difficulty.
Noise pollution poses a serious threat to human health and well-being, especially in educational environments where concentration and learning are essential. While urban noise has been widely studied, its effects within university settings remain underexplored. This study investigates environmental noise and student perceptions on two campuses of the University of Guadalajara, Mexico-one located in an urban area and the other in a semi-rural setting. Noise levels were measured using the CESVA-SC260 integrating instrument (CESVA Instruments, SLU, Barcelona, Spain), and student perceptions were gathered through a survey. A total of 731 students participated, with 357 from the urban campus and 374 from the semi-rural one. Results showed that noise levels on both campuses frequently exceeded the WHO's recommended limit of 55 dB(A) for educational facilities, with readings between 40.9 and 85.0 dB(A); 89% of measurements surpassed the threshold. Major sources of noise included vehicular traffic, student gatherings, and construction-related machinery. Survey responses indicated that 41% of students perceived noise as a health risk, and 96% reported adverse effects on well-being and identified it as a disruptor of academic tasks. These findings underscore the pressing need for targeted noise management strategies in university environments and call for further research into effective, context-specific interventions that enhances learning conditions.
Speaker change detection (SCD) in long, multi-party meetings is essential for diarization, Automatic speech recognition (ASR), and summarization, and is now often performed in the space of pre-trained speech embeddings. However, unsupervised approaches remain dominant when timely labeled audio is scarce, and their behavior under a unified modeling setup is still not well understood. In this paper, we systematically compare two representative unsupervised approaches on the multi-talker audio meeting corpus: (i) a clustering-based pipeline that segments and clusters embeddings/features and scores boundaries via cluster changes and jump magnitude, and (ii) a multi-scale jump-based detector that measures embedding discontinuities at several window lengths and fuses them via temporal clustering and voting. Using a shared front-end and protocol, we vary the underlying features (ECAPA, WavLM, wav2vec 2.0, MFCC, and log-Mel) and test the model's robustness under additive noise. The results show that embedding choice is crucial and that the two methods offer complementary trade-offs: the pipeline yields low false alarm rates but higher misses, while the multi-scale detector achieves relatively high recall at the cost of many false alarms.
Virtual acoustics enables the creation and simulation of realistic and ecologically valid indoor environments vital for hearing research and audiology. For real-time applications, room acoustics simulation requires simplifications. However, the acoustic level of detail (ALOD) necessary to capture all perceptually relevant effects remains unclear. This study examines the impact of varying ALOD in simulations of three real environments: a living room with a coupled kitchen, a pub, and an underground station. ALOD was varied by generating different numbers of image sources for early reflections, or by excluding geometrical room details specific for each environment. Simulations were perceptually evaluated using headphones in comparison to measured, real binaural room impulse responses, or by using loudspeakers. The perceived overall difference, spatial audio quality differences, plausibility, speech intelligibility, and externalization were assessed. A transient pulse, an electric bass, and a speech token were used as stimuli. The results demonstrate that considerable reductions in acoustic level of detail are perceptually acceptable for communication-oriented scenarios. Speech intelligibility was robust across ALOD levels, whereas broadband transient stimuli revealed increased sensitivity to simplifications. High-ALOD simulations yielded plausibility and externalization ratings comparable to real-room recordings under both headphone and loudspeaker reproduction.
Trentino, a sparsely populated and almost entirely mountainous region in northeastern Italy, has so far received little attention in linguistic studies on soundscapes, which provide an important cultural ecosystem service. This study analyzes the responses of 68 participants-31 from mountain areas and 37 from urban areas-to an open-ended questionnaire adapted from Guastavino, using a mixed-methods approach to investigate: (1) differences in current and ideal soundscape perception between residents of urban and mountain areas in Trentino; (2) how these findings compare with Guastavino's study conducted in a purely urban context; (3) the role of Trentino's multilingual context in shaping the description and understanding of the soundscape. Findings reveal that, in addition to a latent substratum of the dialectal component, differences emerge mainly in the description of ideal soundscapes. Urban participants evaluate human sounds more negatively and use metonymic expressions for mechanical noises. Mountain participants align their ideal soundscape more closely with their lived experience, often identifying the sound source rather than the sound itself. Tranquility and silence are central values across both groups for the ideal soundscape and for the current one, cognitively linked to natural environments, which therefore remains a cultural legacy to be preserved.
Low-frequency hydrophones are used to detect underwater low-frequency acoustic signals and are widely applied in marine science, resource exploration, environmental monitoring, and military operations. Their primary advantage lies in the fact that low-frequency acoustic waves experience less attenuation in water, enabling long-distance detection. This characteristic makes them indispensable for long-range and wide-area sensing. In this study, a piston-structured hydrophone using a stack of lead zirconate titanate (PZT) piezoelectric ceramic sheets is designed. Finite element simulation analysis is used to derive the output voltage variation in the piezoelectric ceramic stack as a function of its thickness and end-face diameter. The piston-structured hydrophone is then designed accordingly. Results show that the piston structure, combined with the longitudinal stacking of PZT piezoelectric ceramic sheets, enhances the sensitivity of the piezoelectric hydrophone. The prepared hydrophone has a directivity of 360 degrees in the operating frequency range of 1 Hz to 1 kHz, as well as a flat frequency response and high sensitivity of -161 dB. These research results indicate that the proposed sonar design provides valuable reference for the development of low-frequency sonar with higher sensitivity, which is of great significance to the development of marine science.
Enclosed learning spaces, e.g., classrooms, are used in most schools. Open learning spaces, which enable teaching more than one group of students at a time, have become increasingly popular. A recent survey showed that acoustic satisfaction was lower among teachers working in open learning spaces. Our purpose was to compare the acoustic conditions of these learning space types. We investigated the room acoustic quality of 73 learning spaces in 20 schools. Ten schools involved only enclosed and ten both open and enclosed learning spaces. Measurements concerned speech transmission index, STI, background noise level, LAeq, and reverberation time, T. Variation in results in both learning space types was rather large. In enclosed learning spaces, STI varied within 0.64-0.83, LAeq within 25-47 dB, and T within 0.34-0.82 s. The corresponding variations in open learning spaces were 0.47-0.91, 29-44 dB, and 0.44-0.72 s. The differences between enclosed and open learning spaces were surprisingly small. Due to the different intended uses of these space types, Finnish target values are tighter for open than for enclosed learning spaces. These target values were fulfilled in 56% of enclosed and 9% of open learning spaces. The more frequent violation of target values in open learning spaces was due to the STI being too large at longer distances. Our study provides suggestive evidence that the room acoustic conditions are worse in open than enclosed learning spaces. Further research is needed to prove whether room acoustic conditions could explain worse acoustic satisfaction in teachers.
The use of ultrasonic guided waves (UGWs) is an efficient damage monitoring technique. Due to their characteristics of a wide monitoring range and low power consumption, UGWs have been widely applied in various structural health monitoring fields. In practice, the transducers and coupling agents used for UGW excitation and reception are prone to failure due to service environmental factors, resulting in abnormal UGW signals. To ensure reliable damage monitoring, this paper proposed an abnormal UGW signal identification method based on the UGW reconstruction errors. First, a multi-scale progressive reconstruction network (MPRN) is proposed to accurately reconstruct normal UGW signals. Leveraging the inherent differences between normal and anomalous UGW signal characteristics, the reconstruction errors increase significantly when abnormal UGW signals are input into the MPRN, which has been trained exclusively on normal data. This discrepancy in reconstruction errors enables the identification of abnormal signals. The experimental results show that sensor failure causes frequency shifts in the received UGW signals. When reconstructing normal UGW signals, the proposed MPRN achieves high fidelity, with an average NRMSE as low as 0.0036 and an average PSNR as high as 40.04 dB. In contrast, when reconstructing abnormal UGW signals, the average NRMSE is no lower than 0.62, and the average PSNR is no higher than 16.67 dB. The proposed reconstruction-error-based abnormal UGW signal identification method achieves a maximum accuracy of 93.43%.
This study investigates the influence of microgravity on the fundamental frequency (F0) of astronauts' speech. A speech corpus was compiled, including recordings in microgravity and on Earth, matched by speaker and content. The signal processing methodology included filtering with consideration of human auditory perception, segmentation of speech fragments, F0 estimation using digital signal processing techniques, and visualization through fundamental frequency dynamics plots. Results revealed a consistent increase in F0 for most astronauts under microgravity, with maximum values of 450 Hz for female speakers and 245 Hz for male speakers. Elevated F0 levels were observed for approximately 86% of the total duration of speech fragments recorded in microgravity, compared with 14% on Earth. These findings confirm that microgravity affects the speech apparatus and acoustic characteristics of voice. Practical implications include adapting voice-controlled systems and automatic speech recognition for space environments, monitoring crew condition, and studying speech physiology under extreme conditions.
The high peak-to-average power ratio (PAPR) in classical high-speed digital data transmission systems with orthogonal frequency division multiplexing (OFDM) limits energy efficiency and communication range. This paper proposes a method for randomizing OFDM signals via frequency coding using synthesized pseudorandom sequences with improved autocorrelation properties, obtained through machine learning, to minimize PAPR in complex, non-stationary hydroacoustic channels for communicating with underwater robotic systems. A neural network architecture was developed and trained to generate codes of up to 150 elements long based on an analysis of patterns in previously found best short sequences. The obtained class of OFDM signals does not require regular and accurate estimation of channel parameters while remaining resistant to various types of impulse noise, Doppler shifts, and significant multipath interference typical of the underwater environment. The attained spectral efficiency values (up to 0.5 bits/s/Hz) are relatively high for existing hydroacoustic communication systems. It has been shown that the peak power of such multi-frequency information transmission systems can be effectively reduced by an average of 5-10 dB, which allows for an increase in the communication range compared to classical OFDM methods in non-stationary hydrological conditions at acceptable bit error rates (from 10-2 to 10-3 and less). The effectiveness of the proposed methods of randomization with synthesized codes and frequency coding for OFDM signals was confirmed by field experiments at sea on the shelf, over distances of up to 4.2 km, with sea waves of up to 2-3 Beaufort units and mutual movement of the transmitter and receiver.
This study investigates how listeners perceive consonance and dissonance in dyads composed of simple (sine) tones, focusing on the effects of frequency ratio ($R$) and mean frequency ($F$). Seventy adult participants - categorized by musical training, gender, and age group - rated randomly ordered dyads using binary preference responses (``like'' or ``dislike''). Dyads represented standard Western intervals but were constructed with sine tones rather than musical notes, preserving interval ratios while varying absolute pitch. Statistical analyses reveal a consistent decrease in preference with increasing mean frequency, regardless of interval class or participant group. Octaves, fifths, fourths, and sixths showed a nearly linear decline in preference with increasing $F$. Major seconds were among the least preferred. Musicians rated octaves and certain consonant intervals more positively than non-musicians, while gender and age groups exhibited different sensitivity to high frequencies. The findings suggest that both interval structure and pitch range shape the perception of consonance in simple-tone dyads, with possible psychoacoustic explanations involving frequency sensitivity and auditory fatigue at higher frequencies.
Sound reproduction is the electro-mechanical re-creation of sound waves using analogue and digital audio equipment. Although identical reproduction of a sound is implied to be acoustically identical, numerous fixed and variable conditions are affecting the acoustic result. To arrive at a better understanding of the causes and the extent of deviations in sound reproduction, differences in the amplitude, phase and frequency of a sound signal at various stages in the process of reproduction were measured and compared under a set of controlled conditions, one of them being the presence of a human subject in the acoustic environment. Deviations in acoustic reproduction were found to be significantly smaller than ± 0.1 dB amplitude and ± 1 degree phase shift when comparing trials recorded on the same day. Deviations significantly increased greater than 5 times the amplitude and 6 times the phase shift when comparing trials recorded on different days. Deviations further increased significantly with greater than 59 times the amplitude and 32 times the phase shift with a human subject present in the acoustic environment. For the first time, it was shown that the human body does not always absorb, but can also amplify sound energy. The degree of either absorption or amplification per frequency shows consistent variance in response to the subject and changes in the stimulus, indicating a non-linear relationship between the observed deviations and the presence of the human subject. The findings of the present study may serve as a reference for acoustic standards and corrective methods improving the accuracy and predictability of sound reproduction and its applications in measurement, diagnostics and therapeutic methods.
Impulsive noise poses a significant challenge to broadband feedforward active noise control (ANC) systems, particularly in sensitive environments such as infant incubators. This paper presents an adaptive impulsive noise cancellation approach based on the Kalman filter, designed to improve noise attenuation performance under nonstationary and impulsive interference. The proposed framework integrates impulsive noise detection with a Kalman filter-based suppression scheme. Simulation studies are conducted to evaluate the performance of the combined system in comparison to traditional ANC methods, such as Filtered-x Least Mean Square (FxLMS) and Filtered-x Normalized LMS (FxNLMS). Results demonstrate that the Kalman filter can effectively reduce the influence of impulsive disturbances without degrading overall broadband noise cancellation. A case study involving an infant incubator illustrates the practical effectiveness and robustness of the proposed technique in a real-world healthcare application. The findings support the integration of Kalman filter-based adaptive control in future ANC designs targeting impulsive noise environments.
Using efficient voice alarms to ensure safe evacuation is important during emergencies, especially for the elderly. Factors that have important influence on speech perceptions have been investigated for several years. However, relatively few studies have specifically explored the key factors influencing perceptions of voice alarms in emergency situations. This study investigated the combined effects of speech rate (SR), signal-to-noise ratio (SNR), and reverberation time (RT) on older people's perception of voice alarms. Thirty older adults were invited to evaluate speech intelligibility, listening difficulty, and perceived urgency after hearing 48 different voice alarm conditions. For comparison, 25 young adults were also recruited in the same experiment. The results for older adults showed that: (1) When SR increased, speech intelligibility significantly decreased, and listening difficulty significantly increased. Perceived urgency reached its maximum at the normal speech rate for older adults, in contrast to young adults, for whom urgency was greatest at the fast speech rate. (2) With the rising SNR, speech intelligibility and perceived urgency significantly increased, and listening difficulty significantly decreased. In contrast, with the rising RT, speech intelligibility and perceived urgency significantly decreased, while listening difficulty significantly increased. (3) RT exerted a relatively stronger independent influence on speech intelligibility and listening difficulty among older adults compared to young adults, which tended not to be substantially moderated by SR or SNR. The interactive effect of SR and RT on perceived urgency was significant for older people, but not significant for young people. These findings provide referential strategies for designing efficient voice alarms for the elderly.
This study investigated how test room acoustic conditions relate to listening comprehension performance in a high-stakes English as a foreign language (EFL) assessment context. Using score data (n = 2532) from five TOEFL ITP test sessions conducted between 2021 and 2025 at a private university in Chiba, Japan, we compared performance across three lecture halls with documented differences in reverberation time (RT) and Speech Transmission Index (STI). Each listening score was linked to an approximated seat-based STI value, while grammar/reading scores were used to account for baseline proficiency. Linear mixed-effects modeling analyses indicated that examinees in the least favorable acoustic environment (RT0.5–2kHz 1.51 s, STI 0.60) obtained lower listening scores than those in rooms with shorter RT (0.93 s, 0.79 s) and higher STI (0.69, 0.67), respectively. Subgroup analyses revealed a significant effect at the CEFR-J B1.1 level, though the room and B1.1 effects showed modest estimated marginal mean differences (EMMDiff) roughly corresponding to 2–3 points on the total scale. Seat-based STI analyses also showed significant EMMDiff, with approximately 3–7 total score point differences observed between categories F (0.52–0.55) and ≥D (≥0.60). While the dataset was limited to one institution and the sample distribution limited generalizability of the findings, the study offers empirical findings that can inform future research and discussions on equitable listening assessment practices.
Elastic closed-cell porous material is widely applied as a class of light sound insulation product. However, it is difficult to accurately predict its soundproof property due to the occurrence of the closed cells. Therefore, a combined theoretical model of Biot’s theory and acoustic field equations has been developed to predict the sound transmission loss (STL) in the mass control region. Five NBR-PVC closed-cell composites with different parameters were selected to verify the prediction model. Their STL measurement values were compared with the data calculated separately by the theoretical model and the Mass Law, whether under normal incidence or under random incidence. The results show that the Mass Law overestimates the sound insulation values of closed-cell porous material. STL prediction values from the theoretical model have more acceptable agreements to the measurement data than those from the Mass Law. The average deviation rates of the theoretical model are less than 4% under the normal incidence condition and are about 2.9% under the random incidence condition.