This paper focuses on distributed node-specific signal estimation in topology-unconstrained wireless acoustic sensor networks (WASNs) where sensor nodes only transmit fused versions of their local sensor signals. For this task, the topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) algorithm (TI-DANSE) has previously been proposed. It converges towards the centralized signal estimation solution in non-fully connected and time-varying network topologies. However, the applicability of TI-DANSE in real-world scenarios is limited due to its slow convergence. The latter results from the fact that, in TI-DANSE, nodes only have access to the in-network sum of all fused signals in the WASN. We address this low convergence speed issue by introducing an improved TI-DANSE algorithm, referred to as TI-DANSE$<^>+$. The TI-DANSE$<^>+$ algorithm outperforms TI-DANSE in terms of convergence speed by letting the updating node use each partial in-network sum of fused signals (coming from its neighbors) separately, when updating its estimation parameters. In this way, the number of available degrees of freedom in the optimization problem at the updating node is increased, leading to faster convergence. This separate use of incoming partial in-network sums is further exploited by combining TI-DANSE$<^>+$ with a tree-pruning strategy that maximizes the number of neighbors at the updating node. In fully connected WASNs, it is observed that TI-DANSE$<^>+$ converges as fast as the original DANSE algorithm (the latter only defined for fully connected WASNs) while using peer-to-peer data transmission instead of broadcasting and thus saving communication bandwidth. If link failures occur, the convergence of TI-DANSE$<^>+$ towards the centralized solution is preserved without any change in its formulation. Altogether, the proposed TI-DANSE$<^>+$ algorithm can be viewed as an all-round alternative to DANSE and TI-DANSE which (i) merges the advantages of both, (ii) reconciliates their differences into a single formulation, and (iii) shows advantages of its own in terms of communication bandwidth usage. The convergence properties and signal estimation performance of TI-DANSE$<^>+$ are demonstrated through speech enhancement experiments in simulated topology-unconstrained WASNs.
This paper reviews the current state and emerging trends in synthetic speech detection. It outlines the main data-driven approaches, discusses the advantages and drawbacks of focusing future research solely on neural encoding detection, and offers recommendations for promising research directions. Unlike works that introduce new detection methods or datasets, this paper aims to guide future state-of-the-art research in the field and to highlight the risk of overcommitting to approaches that may not stand the test of time.
Although artificial heads based on the human anatomy are typically used to measure binaural cues in various scenarios, they usually do not provide sufficient acoustic insulation that becomes important, e.g., when measuring at low frequencies or with hearing protectors. For these purposes, an acoustic test fixture (ATF) with high acoustic insulation should be used. However, such ATFs do not reflect the human anatomy, which may affect measurements of binaural cues and other spatially separated sounds. This study investigated differences in interaural time and level differences measured with an artificial head with human anatomy and the GRAS ATF 45CA. The measurements took place inside an anechoic chamber with a circular array of 48 loudspeakers. Moreover, a mountable head consisting of two 3D-printed shells for the ATF based on a KEMAR head was designed to approximate plausible head anatomy while preserving the advantages of an ATF. Binaural cues were also measured with the ATF and attached head shells. The results indicate that differences in binaural cues between the artificial head and ATF largely depend on incident angle and frequency. The differences could be reduced with the additional head shells which make them a reasonable add-on for the ATF when measuring binaural cues.
This paper synthesizes insights from a scoping workshop funded by the Volkswagen Foundation that focused on reading competencies in the digital age. The purpose of the workshop was to discuss the current state of research, identify actionable next steps and outline considerations for future work in the context of learning to read in a digital world. The conference brought together German researchers with a wide range of backgrounds and expertise in reading research including computer science, psychology, language processing and educational science. Presentations and discussions during the conference centred broadly on potentials and challenges of learning to read and reading support in the context of digitalization—focusing on issues related to artificial intelligence and app-based learning. Results of the workshop and subsequent writing sessions include a synthesis of key considerations and core questions for the field of research on digital approaches to support reading. Findings highlight the need for the field at large to identify and develop digital tools for diagnosing readers’ abilities and providing them with the support they need, and to do so with increased interdisciplinary collaboration, open science practices and communication among scholars.