Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay speech detection from multi-channel recordings. Derived from classical beamforming over discrete azimuth and elevation grids, acoustic maps encode directional energy distributions that reflect physical differences between human speech radiation and loudspeaker-based replay. A lightweight convolutional neural network is designed to operate on this representation, achieving competitive performance on the ReMASC dataset with approximately 6k trainable parameters. Experimental results show that acoustic maps provide a compact and physically interpretable feature space for replay attack detection across different devices and acoustic environments.
Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions. In this work, we decompose simulated RIRs into four variants (full, direct-only, no-late, and no-early) using the mixing time estimated from the echo density function as the boundary between early reflections and late reverberation. We define four calibration scenarios, from fully calibrated (synchronised capture, known source level) to fully uncalibrated (arbitrary onset, unknown level), and evaluate all combinations on a matched dataset. Results show that without time calibration, mean absolute error (MAE) increases to 1.29 m and the model extracts reverberation-based cues, with early reflections emerging as the most informative component. Further analysis against DRR, C_50, and T_60 confirms that estimation accuracy improves with stronger early energy and degrades in highly reverberant environments. When time calibration is available, the model achieves a MAE of 0.14 m by extracting the propagation delay alone, regardless of the RIR content.
In the context of digital transformation in the occupational health and safety sector, this research focuses on the efficacy of virtual reality training tools in ensuring workplace safety during maintenance activities. A comprehensive methodology, including focus groups and workshops, is utilized to define a study protocol and evaluate the proposed virtual reality solution. The study suggests that immersive virtual reality training, which integrates the Skill-Rule-Knowledge framework, has the potential to enhance workers’ ability to manage unforeseen situations. It emphasizes the importance of practical, scenario-based training and outlines a detailed evaluation process. The conclusion highlights the need for ongoing validation and future steps to extend the application to supervisors, fostering improved health and safety management in the workplace.
In recent years, the development of interconnected devices has expanded in many fields, from infotainment to education and industrial applications. This trend has been accelerated by the increased number of sensors and accessibility to powerful hardware and software. One area that significantly benefits from these advancements is Teleoperated Driving (TD). In this scenario, a controller drives safely a vehicle from remote leveraging sensors data generated onboard the vehicle, and exchanged via Vehicle-to-Everything (V2X) communications. In this work, we tackle the problem of detecting the presence of cars and pedestrians from point cloud data to enable safe TD operations. More specifically, we exploit the SELMA dataset, a multimodal, open-source, synthetic dataset for autonomous driving, that we expanded by including the ground-truth bounding boxes of 3D objects to support object detection. We analyze the performance of state-of-the-art compression algorithms and object detectors under several metrics, including compression efficiency, (de)compression and inference time, and detection accuracy. Moreover, we measure the impact of compression and detection on the V2X network in terms of data rate and latency with respect to 3GPP requirements for TD applications.
Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection robustness, existing datasets and methods predominantly rely on single-channel recordings. In this work, we introduce an acoustic simulation framework designed to simulate multi-channel replay speech configurations using publicly available resources. Our setup models both genuine and spoofed speech across varied environments, including realistic microphone and loudspeaker impulse responses, room acoustics, and noise conditions. The framework employs measured loudspeaker directionalities during the replay attack to improve the realism of the simulation. We define two spoofing settings, which simulate whether a reverberant or an anechoic speech is used in the replay scenario, and evaluate the impact of omnidirectional and diffuse noise on detection performance. Using the state-of-the-art M-ALRAD model for replay speech detection, we demonstrate that synthetic data can support the generalization capabilities of the detector across unseen enclosures.
In this work, we investigate the generalization of a multi-channel learning-based replay speech detector, which employs adaptive beamforming and detection, across different microphone arrays. In general, deep neural network-based microphone array processing techniques generalize poorly to unseen array types, i.e., showing a significant training-test mismatch of performance. We employ the ReMASC dataset to analyze performance degradation due to inter- and intra-device mismatches, assessing both single- and multi-channel configurations. Furthermore, we explore fine-tuning to mitigate the performance loss when transitioning to unseen microphone arrays. Our findings reveal that array mismatches significantly decrease detection accuracy, with intra-device generalization being more robust than inter-device. However, fine-tuning with as little as ten minutes of target data can effectively recover performance, providing insights for practical deployment of replay detection systems in heterogeneous automatic speaker verification environments.
Replay attacks belong to the class of severe threats against voice-controlled systems, exploiting the easy accessibility of speech signals by recorded and replayed speech to grant unauthorized access to sensitive data. In this work, we propose a multi-channel neural network architecture called M-ALRAD for the detection of replay attacks based on spatial audio features. This approach integrates a learnable adaptive beamformer with a convolutional recurrent neural network, allowing for joint optimization of spatial filtering and classification. Experiments have been carried out on the ReMASC dataset, which is a state-of-the-art multi-channel replay speech detection dataset encompassing four microphones with diverse array configurations and four environments. Results on the ReMASC dataset show the superiority of the approach compared to the state-of-the-art and yield substantial improvements for challenging acoustic environments. In addition, we demonstrate that our approach is able to better generalize to unseen environments with respect to prior studies.
In this paper, a semi-supervised approach for the classification of audio signals under domain shift of ICME 2024 Grand Challenge is presented. In more detail, a low-complexity attention-based convolutional neural network is introduced for the identification of the scene. Specifically, it exploits the log-Mel spectrogram and the Waveg-ram learning-based time-frequency representation. Experimental results on a portion of the challenge development dataset show outstanding performance. The proposed approach achieved a macro-accuracy performance of 63.1%, outperforming the baseline by 3.1%. Code, model, and pre-trained weights are available at https://github.com/michaelneri/ICME2024RM3Team.
In this letter, a novel deep neural network, designed to enhance the efficiency and effectiveness of unsupervised sound anomaly detection, is presented. The proposed model exploits an attention module and separable convolutions to identify salient time-frequency patterns in audio data to discriminate between normal and anomalous sounds with reduced computational complexity. The approach is validated through extensive experiments using the Task 2 dataset of the DCASE 2020 challenge. Results demonstrate superior performance in terms of anomaly detection accuracy while having fewer parameters than state-of-the-art methods.
Training in the Occupational Safety and Health (OSH) sector is crucial for minimizing workplace hazards and ensuring employee well-being. Virtual Reality (VR) emerges as a training tool that can enhance learning outcomes and simulate hazardous scenarios safely. However, several aspects must be taken into consideration when implementing VR-based training solutions. The paper investigates how to effectively design, develop, integrate, and validate a VR OSH training tool. To this aim, a comprehensive guideline of 9 key elements articulated in 29 items is proposed. Every element and item is retrieved from analyzing the existing literature on the topic using a systematic approach. The result is a comprehensive guide to consider all these aspects from the outset of design, in a cohesive, complete, and tailored manner. This formalization is intended to facilitate the advancement of research and implementation of these solutions, which to date have been largely confined to prototypes or lack real practical application.
In recent years, video data has been extensively used for surveillance purposes. Anyway, if a fight can be recognized by everyone, abnormal sounds could pass unnoticed. Moreover, if a dangerous event is not in our line of sight, the only cue that can be exploited is the sound produced by the threat. With the adoption of Artificial Intelligence-based techniques, it is possible to detect these anomalies by inspecting the videos acquired by cameras located inside the bus. However, this scenario is complex for several reasons since video analysis is computationally expensive, requiring costly hardware equipment for processing and storage. Moreover, videos suffer from occlusions and luminance variations, making the system not suitable in all situations. To this aim, the objective of my Ph.D. is to propose a data-driven framework that can detect if an audio recording is anomalous and, if this is the case, to identify which and where dangerous events are occurring. The architecture I propose, denoted as Coarse-to-Fine, is composed of two elements. The first is responsible for modeling the normal background of a target environment in an unsupervised fashion. If an anomalous audio is detected, a second element focuses on what, when, and where the anomaly occurs.
Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where distances are discretized into distinct categories, enabling smoother model training and achieving higher accuracy but imposing restrictions on the precision of the obtained sound source position. Towards this direction, in this paper we propose a novel approach for continuous distance estimation from audio signals using a convolutional recurrent neural network with an attention module. The attention mechanism enables the model to focus on relevant temporal and spectral features, enhancing its ability to capture fine-grained distance-related information. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using audio recordings in controlled environments with three levels of realism (synthetic room impulse response, measured response with convolved speech, and real recordings) on four datasets (our synthetic dataset, QMULTIMIT, VoiceHome-2, and STARSS23). Experimental results show that the model achieves an absolute error of 0.11 meters in a noiseless synthetic scenario. Moreover, the results showed an absolute error of about 1.30 meters in the hybrid scenario. The algorithm's performance in the real scenario, where unpredictable environmental factors and noise are prevalent, yields an absolute error of approximately 0.50 meters. For reproducible research purposes we make model, code, and synthetic datasets available at https://github.com/michaelneri/audio-distance-estimation
This paper examines the enhancement of occupational safety and health (OSH) training in manufacturing through the integration of Safety-II principles within a Virtual Reality (VR) training framework, applying the Analysis, Design, Development, Implementation, and Evaluation (ADDIE) model. Amidst the digital transformation in manufacturing, innovative training methods such as VR have become instrumental in improving operational safety and efficiency. However, the incorporation of Safety-II, a philosophy emphasizing the complexity of organizational systems and the need for resilience, has not been systematically applied to VR training designs. Through a series of focus group sessions, this study presents a new method for creating VR training materials. Designed to foster systemic thinking, resilience, proactivity, learning, flexibility, and leverage human performance variability, these courses are in line with Safety-II's transition from merely preventing negative outcomes to understanding and increasing positive capacities. The results lead to a newintegration of the ADDIE model within a Safety-II framework for VR-based OSH training, enhanced by Skill-Rules-Knowledge (SRK) specific guiding principles.
The research conducted within the audio signal processing field is increasingly focusing on environmental sound classification. This paper presents a low-complexity Fully Convolutional Network composed of two parallel branches. These branches are responsible for extracting features from the Cadence Frequency Diagram representation and the Chebychev moments, respectively. By utilizing both domains of machine-and deep-learning, the proposed pipeline takes advantage of the unique characteristics of each. The key strength of this architecture lies in its reduced number of layers and parameters, as well as its ability to efficiently compute the Cadence Frequency Diagram and Chebychev moments. The effectiveness of the proposed pipeline is demonstrated through various tests conducted on two audio datasets, namely UrbanSound8K and ESC-50.
We introduce the novel task of continuous-valued speaker distance estimation which focuses on estimating non-discrete distances between a sound source and microphone, based on audio captured by the microphone. A novel learning-based approach for estimating speaker distance in reverberant environments from a single omnidirectional microphone is proposed. Using common acoustic features, such as the magnitude and phase of the audio spectrogram, with a convolutional recurrent neural network results in errors on the order of centimeters in noiseless audios. Experiments are carried out by means of an image-source room simulator with convolved speeches from a public dataset. An ablation study is performed to demonstrate the effectiveness of the proposed feature set. Finally, a study of the impact of real background noise, extracted from the WHAM! dataset at different signal-to-noise ratios highlights the discrepancy between noisy and noiseless scenarios, underlining the difficulty of the problem.
Artificial Intelligence techniques are being applied in the quality assessment of immersive multimedia content, such as virtual and augmented reality scenarios. The immersive nature of these applications poses a unique challenge to traditional quality assessment methods. In fact, estimating user acceptance of immersive technologies is complex due to multiple aspects, such as usability, enjoyment, and cyber sickness. Artificial Intelligence-based approaches offer a promising solution to this problem, enabling objective evaluations of immersive multimedia such as spatial audios, point clouds, and light field images. This work presents an overview of different artificial intelligence techniques that have been used for quality assessments of immersive multimedia content, including machine learning algorithms, deep learning, and computer vision. The advantages of these techniques and some examples of practical application are provided. Future works are presented, underlining the possible outcomes of a Ph.D. study in this field.
Accurate detection and classification of objects in 3D point clouds is a central problem in several applications such as autonomous navigation and augmented/virtual reality scenarios. In this paper we present a deep learning strategy for 3D object detection for railway applications based on the VoxelNet model. Due to the lack of publicly available annotated data, we created a virtual railway environment for generating a synthetic annotated railway point cloud dataset. This approach allows to model shapes and locations of target landmarks such as traffic lights and railtracks. The achieved results show that our network learns an effective representation of railway landmarks using only raw LiDAR point clouds, leading to encouraging results and possible future implementations in this research field. We also made the annotated dataset available to the research community at https://gitlab.com/michael.neri/sard-synthetic-annotated-railway-dataset.
This paper proposes a machine learning-based architecture for audio signals classification based on a joint exploitation of the Chebychev moments and the Mel-Frequency Cepstrum Coefficients. The procedure starts with the computation of the Mel-spectrogram of the recorded audio signals; then, Chebychev moments are obtained projecting the Cadence Frequency Diagram derived from the Mel-spectrogram into the base of Chebychev moments. These moments are then concatenated with the Mel-Frequency Cepstrum Coefficients to form the final feature vector. By doing so, the architecture exploits the peculiarities of the discrete Chebychev moments such as their symmetry characteristics. The effectiveness of the procedure is assessed on two challenging datasets, UrbanSound8K and ESC-50.
The objective of a sound event detector is to recognize anomalies in an audio clip and return their onset and offset. However, detecting sound events in noisy environments is a challenging task. This is due to the fact that in a real audio signal several sound sources co-exist. Moreover, the characteristics of polyphonic audios are different from isolated recordings. It is also necessary to consider the presence of noise (e.g. thermal and environmental). In this contribution, we present a sound anomaly detection system based on a fully convolutional network which exploits image spatial filtering and an Atrous Spatial Pyramid Pooling module. To cope with the lack of datasets specifically designed for sound event detection, a dataset for the specific application of noisy bus environments has been designed. The dataset has been obtained by mixing background audio files, recorded in a real environment, with anomalous events extracted from monophonic collections of labelled audios. The performances of the proposed system have been evaluated through segment-based metrics such as error rate, recall, and F1-Score. Moreover, robustness and precision have been evaluated through four different tests. The analysis of the results shows that the proposed sound event detector outperforms both state-of-the-art methods and general purpose deep learning-solutions.
Satellite-based positioning has been selected as one of the key game changers for the evolution of the European Rail Traffic Management System, introducing strict accuracy, integrity and continuity requirements. However, GNSSs are vulnerable to several degradations that impair the fulfillment of the performance demands. For this reason, we propose the integration of GNSS and on-board cameras for train positioning. More specifically, we present a positioning framework based on local landmarks which requires semantic segmentation, that is a pixel-wise classification of the input images. In this work, we describe the positioning framework and evaluate the semantic segmentation approach on a publicly available dataset. The achieved results outperform state-of-the-art approaches and, although the semantic segmentation issue in the railway environment has already been addressed, to the best of our knowledge this is the first approach which exploits these techniques for train positioning purposes.