
The article presents a room acoustics study conducted for a multi-purpose hall, which is currently in the design phase. The hall is part of the Cultural Center building, which will be constructed in Zagórz, a small town located in the southern part of Poland in the Podkarpacie region. The architectural design was developed by Studio Święciński Architekci, based in Krosno. The results of the acoustic prediction for the hall were obtained using the open-source software I-SIMPA. To perform the acoustic study, the multi-purpose hall was geometrically modeled as a 3D shape in acoustical scale; sound absorption and scattering coefficients for the internal surfaces were then assigned along with the sound source and receivers in the audience area. Several acoustic parameters (T_Sabine, T_Eyring, T30, EDT, C80, D50, G, STI) were calculated in pre and post-treatment conditions, to simulate the acoustic efficiency of the conference hall in accordance with its future use.
The ultrasound waves are recognized as an effective green technology with outstanding potentials in various applications including the dehydration processes of food materials. However, the application of this technology is still confined at the laboratory levels due to limited power capacity of available ultrasound systems in the large-scale airborne applications. The extensive plate radiators are essential parts that can effectively increase the amplitude of oscillations generated by piezo-electric transducers and induce high-intensity acoustic pressure field in the surrounding airflow. Using an analytical model, the geometric parameters in a simple flat circular radiator were evaluated. However, this configuration propagated counter-phase radiations with diminished overall ultrasound effect. In this study, a stepped profile with specific geometrical dimensions is proposed to modify the response of excited extensive plate. The modal analysis showed that for 20,259 Hz, just close to the system’s resonance frequency, the plate generated directed and focused vibrations. Furthermore, the design provided a more convenient distribution of natural frequencies to eliminate the risk of harmful interaction of mode shapes.
This work mainly focuses on a deepwater soundscape study in the Nansen Basin, Central Arctic Ocean, by using Ambient Noise Measurement System (ANMS) data, which was deployed as part of the Nansen Environmental and Remote Sensing Center (NERSC-4) mooring from September 2019 to July 2020. The main objective of the work is to study the sources of underwater noise of the Central Arctic Ocean (CAO) along with their characteristic patterns. Also in this work, the seasonal variation of the sound pressure level of CAO is studied. The seasonal variation of the sound pressure level of the CAO is maximum in summer, and the minimum seasonal variation was observed during spring. The difference in seasonal variability in 1/3 octave band frequencies between summer and spring was 4.1 dB (re 1 μPa²/Hz). Shipping and marine species are the most dominant noise sources of the Central Arctic Ocean. Bowhead whales and bearded seals are the most common marine species observed in this location. The bearded seal vocalisation is more intense between 300 and 700 Hz. Bowhead whales produce two different types of simple calls i.e., sound patterns : one is S-moans, and the other is V-moans. The S-moan patterns are mainlydominated by low frequencies i.e., below 600 Hz. The V-moans patterns are dominated in the frequency range of 300 to 1230 Hz. The killer whales are also observed in the study location, which represents a shifting Arctic environment. The soundscape generated by killer whales is observed between 160 Hz and 2 kHz.
In recent years, virtual reality (VR) has been an attractive alternative to traditional methods of conducting hearing tests, offering both a natural way to indicate sound direction and complete digital data recording. The aim of this study was to analyze how the use of VR goggles and motion controllers affects the accuracy of localizing acoustic stimuli and the subjective comfort of participants, compared to a WEB-based application operated with a mouse. The study was performed for 16 azimuthal samples (0° elevation) and 8 elevation samples at azimuth 90°. Mean absolute angular errors (MAAE) were analyzed, and the paired-samples t-tests were performed. The results show that using the VR interface can reduce horizontal MAAE by 7° compared to the WEB-based interface. In the elevation tasks, no statistically significant differences were observed. Respondents rated VR the highest in terms of intuitiveness, comfort, speed, and perceived precision. These findings confirm the potential of VR to be used in localization tests, especially in the horizontal plane, and suggest directions for further research on HRTF personalization and interface optimization for elevation tasks.
The relationships between human voice parameters and body dimensions have been previously described, but the connections between voice and face geometry remain poorly researched. This study aims to determine the relationships between face dimensions and acoustic parameters in both sexes and examines 111 adult participants (30 males). Each participant undergoes voice recording, which includes five sustained vowels, along with anthropometric measurements of the neck, head, and face regions. Comparisons between voice parameters and the head, face, and neck regions are conducted employing Pearson's correlation coefficients (r) and a multiple linear regression model. The results reveal significant relationships between head, neck, face dimensions and acoustic parameters in both sexes. Males with higher noses, greater head circumferences, and wider faces tend to have lower formants and more stable voices. Females with larger head circumferences had lower formant values, and those with greater neck circumferences tend to have more stable voices. Also, females with increased nose height have a lower fourth formant (F4). Moreover, females with wider faces, noses, and jaws tend to have less rough voices (lower jitter) and longer maximum phonation time (MPT). These findings may be useful for scientists and law enforcement authorities in creating algorithms that build face models based on voice signals.
A conventional cone loudspeaker has a limited capacity for creating the impression of spatiality, while a distributed mode loudspeaker (DML) has an inherent ability to evoke it. DMLs have their specific drawbacks, but some of these can be compensated for. A key question arises - is it a cone loudspeaker or a compensated DML that is preferred by listeners? A listening experiment with carefully controlled conditions was carried out to answer this question; 30 subjects participated. The participants evaluated three stereo systems: one based on a DML speaker (with its power response equalized) and two conventional two-way active systems. Two perceptual attributes were evaluated: overall preference and spatial impression. A graded pairwise comparison was used as an experimental paradigm; the results were analyzed according to the law of comparative judgment. The findings indicated that, even though the DMLs achieved slightly lower ratings than the conventional systems on average, the perceptual differences were very small. This was confirmed by the hypothesis testing that was performed on the raw results of the pairwise comparisons. tening experiment.
Automated fish welfare monitoring in intensive aquaculture is hindered by environmental noise, individual variability, and data scarcity. These challenges have not been fully resolved by existing deep learning approaches. Traditional computer-vision methods are constrained by underwater turbidity, whereas Passive Acoustic Monitoring (PAM) offers a promising non-invasive alternative for assessing aquatic environments. This study proposes Fisher-SPD, a lightweight, geometry-aware framework for classifying the acoustic behaviour of cage-farmed Larimichthys crocea. A Fisher-score mechanism adaptively selects discriminative frequency bands, effectively filtering complex broadband noise commonly found in commercial sea cages. Acoustic segments are modelled as Symmetric Positive Definite (SPD) covariance matrices and mapped onto a linear tangent space via the Log-Euclidean Metric, preserving intrinsic statistical structure even under data-scarce conditions. Furthermore, a Physics-Consistency Masking mechanism applies source-level physical priors as hard inference-time constraints to robustly suppress false positives originating from background interference. Under a strict Leave-One-Subject-Out Cross-Validation (LOSO-CV) protocol, the framework achieved a mean zero-shot accuracy of 89.22% ± 2.59%, significantly outperforming ResNet-18 and other deep learning baselines, while maintaining a low inference latency of only 3.60 ms on edge devices. Through few-shot domain calibration, the system achieved 92.86% accuracy in a dual-fish overlapping-source environment. Ultimately, this framework provides a robust, data-efficient solution for real-time stress detection and welfare monitoring in modern intensive aquaculture.
This study investigates the system-level acoustic response of tramway track renewal under real urban operating conditions, with emphasis on the confounding role of operating speed in before–after comparisons. Pass-by noise measurements were conducted on a tram route in Poznan before (2022) and after (2025) renewal, when a conventional ballasted track was replaced by a ballastless system. During the same period, mean pass-by speed increased from 25.9 to 58.1 km/h. Energetic, psychoacoustic, and spectral indicators were analysed using direct comparisons, quartile-based speed sensitivities, and an integrated year–speed–tram type regression model. The direct comparison showed an increase in LAeq of +3.37 dB, while SEL remained nearly unchanged (Δ ≈ −0.06 dB), indicating that higher instantaneous levels were largely offset by shorter pass-by durations. Spectral analysis revealed lower measured levels of approximately 4–11 dB in the 125–200 Hz range and localised increases of up to 7–11 dB around 1.6–2 kHz. Quartile-based per-km/h derivatives indicated lower 2025 local sensitivities for LAeq, SEL, and loudness N in the pooled dataset. However, speed-doubling sensitivities showed that this apparent flattening was partly scale-dependent and not uniform across quantiles. The integrated model provided supportive evidence for changed metric–speed relationships, particularly for LAeq, although common-reference slope estimates were partly extrapolative because of limited speed overlap. The results demonstrate the usefulness of speed-sensitive analysis for interpreting before–after tramway noise measurements under changing operating regimes and should be regarded as system-level field-condition responses rather than isolated causal effects of track renewal.
Accurate sediment classification is crucial for advancing marine research, environmental monitoring, and sustainable seabed use. However, acquiring large amounts of labeled data in such settings is often challenging, expensive, and time-consuming. To address this limitation, a semi-supervised learning framework has been proposed that leverages convolutional neural networks for sediment classification using both labeled and unlabeled data. The approach utilizes pseudo-labeling, where confident predictions on unlabeled samples are iteratively incorporated into training to enhance model generalization. The method is applied and evaluated on a dataset that includes multi-modal inputs such as a digital elevation model and multibeam sonar backscatter data. Experimental results indicate that semi-supervised learning with convolutional neural networks can achieve high classification accuracy in scenarios characterized by limited labeled data and a large volume of unlabeled data. This approach highlights the potential of deep learning combined with semi-supervised strategies for efficient underwater environment classification.
Air conditioning is an important contributor to in-vehicle noise. In order to achieve low noise optimization, investigations were performed in a typical air conditioning unit. Computational Fluid Dynamics and Computational Aeroacoustics were combined to simulate the noise characteristics. Steady and transient simulations were conducted to identify noise sources, revealing that the dipole noise is mainly distributed on the fan blades and the volute tongue, and the fan contributes more significantly to the far-field noise. Based on Response Surface Methodology (RSM), three key structural parameters of the fan-volute tongue clearance, volute tongue radius, and blade outlet angle-were selected as design variables, while the far-field A-weighted sound pressure level (SPL) was set as the response index. The final optimization parameter is: volute tongue clearance of 13.36 mm, volute tongue radius of 11.83 mm, and blade outlet angle of 110°. The far-field sound pressure level (SPL) is reduced from 61.8 dB to 55.5 dB, with a relative reduction of 10.16%. This research validates the effectiveness and reliability of the proposed method, providing an efficient and systematic technical reference for the noise reduction design of automotive air conditioning systems.
Recent deep-learning based speech enhancement algorithms have many applications in the areas of noise reduction, de-reverberation, bandwidth extension, echo cancellation, to name a few. Packet loss is also one of the main causes of voice quality degradation in VoIP calls. Currently, generative adversarial networks (GANs) have shown a strong ability in image generation, and many of those models also work well in speech tasks. In this work, we propose a light-weight model based on GAN to handle the task of audio packet loss concealment. Specifically, we use a U-shaped network operating in the time-frequency domain as a generator, which is trained by a Mel-GAN discriminator with multi-loss. In addition, to enhance the model’s performance under unfavorable channels, we introduce noise and bandwidth loss in the training data. The experiments show that our method outperforms the baseline in both objective and subjective metrics under an ideal channel with no other distortions, and it still largely maintains its performance in the presence of noise and bandwidth loss.
The structure of the acousto-optic delay line and possible areas of its application are discussed. The necessity of determining the cutoff frequency of the acousto-optic delay line in all areas of its application is substantiated. The relationship between the cutoff frequency of an electric circuit and its time constant is discussed. The obtained result is extrapolated to the acousto-optic delay line, in which the cutoff frequency is formed by a completely different mechanism. The mechanism of cutoff frequency formation in the acousto-optic delay line is discussed. It is shown that the cutoff frequency in the acousto-optic delay line is formed due to the finite velocity of interaction of the acoustic wave with the light beam in the photoelastic medium. Two methods for measuring the cutoff frequency of the acousto-optic delay line are discussed: the method of system-parametric measurement and the method of cursor measurement. The system-parametric method for measuring the cutoff frequency of the acousto-optic delay line is implemented based on the known values of the laser beam diameter and the velocity of propagation of the acoustic wave in the photoelastic medium. The cursor method for measuring the cutoff frequency of the acousto-optic delay line is implemented based on the parameters of the oscillogram of its response to the input action in the form of a rectangular pulse. Theoretical and experimental aspects of the application of these methods are discussed. Corresponding numerical examples and experimental results are given.
These studies focus on acoustical parameters of steel flat-oval ducts as a function of their roughness. The four types of steel ducts were measured: raw steel, galvanised steel, painted steel, and aluminium as the reference one. The roughness of the duct was measured, and roughness parameters were specified. The sound power level was obtained on the specially constructed stand test with an outlet to the reverberation room. Insertion losses to evaluate the acoustic attenuation performance of the studied steel ducts were obtained. In the present study, an aluminium duct, which is very smooth with minimal airflow friction, was treated as a low-noise object ('silencer'). These studies have shown that for each of the tested steel ducts, the self-noise is higher than for the aluminium duct. The largest differences in this self-noise were observed at a velocity of 12 m/s for the galvanised duct and the raw steel duct compared to the aluminium duct. Insertion losses in straight ducts are consistent with literature and are very low for flat-oval steel ducts. Aluminium duct performs better acoustically than the other ducts studied at lower velocities; however, as airflow velocity increases, the differences in acoustic performance between the materials become less pronounced. This suggests that aerodynamic effects dominate over material surface treatments at higher velocities.
To investigate the principal components of acoustic emission (AE) signals and the damage modes of polypropylene fiber (PPF)-reinforced recycled concrete, ten groups of specimens with coarse aggregate (CA) replacement rates of 0% and 25 % and with different particle sizes, are designed and fabricated. Uniaxial compression AE tests are conducted to obtain AE parameters during the fracture process of PPF-reinforced recycled concrete. In this study, the Pearson correlation coefficient is employed to investigate the correlations among AE parameters. Then, principal component analysis (PCA) is performed on the AE signals to conduct dimensionality reduction of the multi-dimensional data. On this basis, the optimal number of clusters for the principal components of AE signals is determined based on the silhouette coefficient. Finally, the k-means clustering algorithm is introduced to perform cluster analysis on the principal components of AE signals of PPF-reinforced recycled concrete. The clustering results are compared with each other to explore the characteristics of each cluster and to identify the corresponding damage mode for each cluster. The discriminability of AE parameters with respect to damage modes is also investigated. The research findings can provide a reference for predicting the fracture mechanism of PPF-reinforced recycled concrete.
Acoustic scattering scale models often fail to meet the acoustic similarity design requirements due to limitations in fabrication technology, testing facilities, and safe transportation, which restrict the accurate extrapolation of acoustic scattering characteristics between scaled models and full-scale ships. To overcome this challenge, the present study applies highlight model theory to perform acoustic similarity analysis and to correct the local target strength of simple objects in the model based on overall acoustic scattering correction. A novel method for correcting the target strength of non-proportional scaled models is proposed. The method is validated using various model geometries, including ellipsoids, finite-length cylinders, truncated elliptical cones, and complex structures. Additionally, the plate element method is employed for target strength correction and scaling conversion analysis for non-proportional scaled models. The study highlights the variation in target strength due to changes in geometric dimensions and demonstrates the effectiveness of the proposed correction method. The results indicate that the proposed correction approach allows for more accurate extrapolation of target strength from non-proportional scaled models to full-scale prototypes, thereby better satisfying the requirements of practical engineering applications.
Reverberation constitutes a primary interference for active sonar signals, particularly the intense reverberation originating from the reflection of the incident signal. Sharing the same generation mechanism as the target echo, it severely hampers the extraction and analysis of the target signal. To enhance signal processing capabilities under strong reverberation, this paper proposes a sparse dictionary construction method based on multi-order Fractional Fourier Transform (FRFT) domain feature fusion. This method exploits the distinctive characteristics exhibited by target echoes and strong reverberation signals across different fractional transform domains to discriminate between them. It constructs sparse sub-dictionaries using these distinct fractional orders, trains the weight of each sub-dictionary via an adaptive gradient optimization strategy to achieve sparse representation of the signal, suppresses the strong interference in the sparse domain, and reconstructs the target signal through a reconstruction process, thereby achieving the goal of extracting the target signal while suppressing strong interference. Results from processing lake trial data demonstrate that the proposed method can effectively extract target echo signals amidst strong reverberation, with the signal-to-reverberation ratio improvement consistently no less than 2.1 dB and reaching up to 15.6 dB. This method provides an effective approach for the processing and analysis of weak underwater signals.
High-quality speech communication is often compromised by background noise, reducing intelligibility and perceived quality. We investigate data-efficient few-shot transfer of a Speech Enhancement Generative Adversarial Network (SEGAN) to a new noise domain. Starting from a generator pretrained on VoiceBank–DEMAND, we adapt the model to MiniLibriMix using only 300 paired noisy–clean examples. To prevent overfitting and catastrophic forgetting, we introduce SAFE (Stable Adversarial Few-shot Enhancement), a three-fold stabilisation strategy with (i) exponential-moving-average (EMA) weight averaging, (ii) L2-SP weight anchoring to the source-domain parameters, and (iii) a teacher–student consistency loss. SAFE maintains VoiceBank performance (PESQ ≈ 1.84; STOI ≈ 90 %) and, after an optional perceptual fine-tuning stage (MR-STFT + adversarial), yields substantial target-domain gains on MiniLibriMix (PESQ 1.11 → 1.26, STOI 71.5 % → 81.5 %) with only a minor source-domain trade-off in STOI. Ablation experiments demonstrate that EMA provides the strongest stabilising effect, while L2‑SP and consistency regularisation offer complementary benefits. These results suggest that stable few‑shot adaptation can make lightweight time‑domain speech enhancers practical for rapid deployment in novel acoustic environments.
In the paper the theoretical modeling of ultrasonic testing of railway rails with high scanning speed is considered. The model for the calculation of the ultrasonic field generated by the ultrasonic transducers and the pulse echo amplitude received after wave reflection at the defect is developed. The model is based on well-established principles of elastodynamic theory: the Rayleigh-Sommerfeld integral, the Auld reciprocity relation, and the Kirchhoff approximation. It forms the basis for design of computer program to simulate ultrasonic inspections of railway rails with automated mobile systems. The major innovation introduced in the model is taking into account the high scanning speed of the ultrasonic probes over the rail head and the limited repetition rate of the ultrasonic system. The mentioned aspects of the high-speed rail testing require the revision of one of the basic paradigms of the current ultrasonic models, which assume that the scanning speed of the ultrasonic probe is negligible in comparison to the speed of ultrasonic waves propagating in the tested material. Actually, when scanning rails at a speed of 120 km/h, the ultrasonic probe can change its position up to 5 mm between transmitting and receiving ultrasonic pulses reflected from defects located in the rail foot. Such a shift in probe position is not negligible and should be considered in calculations. As a consequence, the ultrasonic system's slow repetition rate and fast scanning speed can make it less likely that certain rail flaws will be found. To quantitatively examine the severity of these phenomena, the new ultrasonic model and related simulation software was developed.
Indian Sign Language (ISL) is vital for communication among India's hearing-impaired community. However, the lack of standardised datasets and reliable identification frameworks has hampered the use of ISL in modern assistive technology. This paper presents a deep learning-based solution to robust ISL alphabet identification, with an emphasis on both accuracy and practical use. A curated static ISL alphabet collection was created by combining authoritative visual references from the official Indian Sign Language website and the Ramakrishna Mission Vivekananda Educational and Research Institute (RKMVERI). Multiple deep learning models were trained and assessed, including CNN, ResNet-50, DenseNet-121, VGG16, MobileNetV2, and EfficientNet-B0, with a new hybrid CNN-ResNet architecture outperforming the others. 98% classification accuracy is achieved by the suggested approach, outperforming individual baseline models. Furthermore, the framework is expanded to support real-time applications, combining webcam-based capture with immediate conversion of recognized signs to textual and synthesized vocal output. Comprehensive performance evaluation, including confusion matrix analysis and ROC curves, demonstrates the solution's durability and practical applicability. This research enhances accessibility, promotes inclusive education, and prepares the path for scalable sign language translation systems in real-world human-machine interaction scenarios by enabling accurate and real-time ISL recognition with voice feedback.
ISO 12913 standards provide a unified framework for describing and assessing soundscapes, yet the absence of a Polish translation has so far limited their practical use. This paper presents the first application of a validated Polish version of the ISO 12913-2 perceptual attributes, enabling full cross-language comparability of results. Whereas Polish research has traditionally focused on noise annoyance and broad judgements of acoustic comfort or discomfort, we outline the complete ISO-compliant assessment procedure, which combines: a soundwalk, questionnaires and audio-visual recording. The study was conducted at eight diverse urban locations in Poznań, Poland. Participants rated the soundscapes using eight attributes: przyjemne, tętniące życiem, bogate w wydarzenia, chaotyczne, dokuczliwe, monotonne, ubogie w wydarzenia, spokojne. Each rating set is mapped to a point in the two-dimensional pleasantness-eventfulness space defined in ISO 12913-3, facilitating visual comparison of locations and the identification of design needs. Results reveal pronounced perceptual differences between spatial typologies and demonstrate that the standardized approach provides richer, multidimensional information about the acoustic environment than conventional noise indicators. The proposed methodology establishes a reference framework for Polish soundscape studies and can support the creation of more people-friendly urban acoustic environments.