ImportanceThe assessment of opioid withdrawal in the neonate, or neonatal opioid withdrawal syndrome (NOWS), is problematic because current assessment methods are based on subjective observer ratings. Crying is a distinctive component of NOWS assessment tools and can be measured objectively using acoustic analysis.ObjectiveTo evaluate the feasibility of using newborn cry acoustics (acoustics referring to the physical properties of sound) as an objective biobehavioral marker of NOWS.Design, Setting, and ParticipantsThis prospective controlled cohort study assessed whether acoustic analysis of neonate cries could predict which infants would receive pharmacological treatment for NOWS. A total of 177 full-term neonates exposed and not exposed to opioids were recruited from Women & Infants Hospital of Rhode Island between August 8, 2016, and March 18, 2020. Cry recordings were processed for 118 neonates, and 65 neonates were included in the final analyses. Neonates exposed to opioids were monitored for signs of NOWS using the Finnegan Neonatal Abstinence Scoring Tool administered every 3 hours as part of a 5-day observation period during which audio was recorded continuously to capture crying. Crying of healthy neonates was recorded before hospital discharge during routine handling (eg, diaper changes).ExposuresThe primary exposure was prenatal opioid exposure as determined by maternal receipt of medication-assisted treatment with methadone or buprenorphine.Main Outcomes and MeasuresNeonates were stratified by prenatal opioid exposure and receipt of pharmacological treatment for NOWS before discharge from the hospital. In total, 775 hours of audio were collected and trimmed into 2.5 hours of usable cries, then acoustically analyzed (using 2 separate acoustic analyzers). Cross-validated supervised machine learning methods (combining the Boruta algorithm and a random forest classifier) were used to identify relevant acoustic parameters and predict pharmacological treatment for NOWS.ResultsFinal analyses included 65 neonates (mean [SD] gestational age at birth, 36.6 [1.1] weeks; 36 [55.4%] female; 50 [76.9%] White) with usable cry recordings. Of those, 19 neonates received pharmacological treatment for NOWS, 7 neonates were exposed to opioids but did not receive pharmacological treatment for NOWS, and 39 healthy neonates were not exposed to opioids. The mean of the predictions of random forest classifiers predicted receipt of pharmacological treatment for NOWS with high diagnostic accuracy (area under the curve, 0.90 [95% CI, 0.83-0.98]; accuracy, 0.85 [95% CI, 0.74-0.92]; sensitivity, 0.89 [95% CI, 0.67-0.99]; specificity, 0.83 [95% CI, 0.69-0.92]).Conclusions and RelevanceIn this study, newborn acoustic cry analysis had potential as an objective measure of opioid withdrawal. These findings suggest that acoustic cry analysis using machine learning could improve the assessment, diagnosis, and management of NOWS and facilitate standardized care for these infants.
The Steered Response Power using the Phase Transform weight (SRP-PHAT) has been shown to be robust in noisy and reverberant conditions. Also, volume contraction has been applied effectively to trap the global maximum for densely-hilly 3-D spaces like the SRP. However, previous methods have suffered from the presence of peaks representing multiple talkers in close proximity as is likely in a conversational cocktail-party setting. We present a volume contraction algorithm called Multi-Stage Rejection Sampling (MSRS) for detection of multiple peaks in the SRP-PHAT space. Our method not only circumvents sorting - a computationally expensive step in volume contraction algorithms - but also automatically divides a search volume into sub-volumes for robust detection of multiple peaks. We discuss some modifications to the standard SRP-PHAT functional and present results using all real-room data for baseline white-noise, an eight-speaker teleconferencing setup and a fully unconstrained cocktail-party situation containing about 21 persons in the room.
Technology improvements, hardware, software and algorithmic, have made the use of a largeaperture microphone array cost effective. In this paper we present real, measured results for our wired, 128microphone array that surrounds a focal area (room) of about 7Mx5M. While it was necessary to evaluate the performance of the array offline using the array’s recording feature, we ensured that all the algorithms still run in real time on our host 6-core Intel I7 PC system. The processing stages include a multiple-source locator, one or more beamformers, an isolation algorithm for the beamformed speech outputs, Kalman-Filter tracking of sources, and a source “labeler” based on pitch and spectrum. All real-room data are used in the evaluations ranging from white noise from stationary speakers to a totally unconstrained group of 21 moving talkers in a conversational cocktail-party situation. Evaluation is presented in terms of location-determination performance. Thus we present our multiple-talker location-determination system in detail, which contains algorithms from previously described work improved for performance and to run in real time. However, our focus is experimental and results show that a large array system can perform very usefully even in a real, unrestricted cocktail-party situation. Results indicate our system is very effective for single talkers in a reverberant room, similarly effective for multiple talkers in a stationary teleconferencing environment, beneficial for an eight human talker multi-conversation and surprisingly applicable to a truly unconstrained cocktail-party situation that had about 21 persons in the room.
Our past research has focused on wired microphone arrays having a large number of elements and wide aperture. The wiring for these large arrays is tedious and only a tree structure with intermediate multiplexing of signals works well. A wireless array alleviates this immense wiring issue, albeit other difficulties are introduced including interferences in the wireless transmissions, synchronization, the need for automatic self-calibration, and module power, size and mounting. This paper introduces our first successful wireless array system, HMA-III. It discusses the desired attributes for such an array and the tradeoffs made to assemble a first working system. Results for the automatic self-calibration are presented.
Purpose: In this article, the authors describe and validate the performance of a modern acoustic analyzer specifically designed for infant cry analysis.Method: Utilizing known algorithms, the authors developed a method to extract acoustic parameters describing infant cries from standard digital audio files. They used a frame rate of 25 ms with a frame advance of 12.5 ms. Cepstral-based acoustic analysis proceeded in 2 phases, computing frame-level data and then organizing and summarizing this information within cry utterances. Using signal detection methods, the authors evaluated the accuracy of the automated system to determine voicing and to detect fundamental frequency (F-0) as compared to voiced segments and pitch periods manually coded from spectrogram displays.Results: The system detected F-0 with 88% to 95% accuracy, depending on tolerances set at 10 to 20 Hz. Receiver operating characteristic analyses demonstrated very high accuracy at detecting voicing characteristics in the cry samples.Conclusions: This article describes an automated infant cry analyzer with high accuracy to detect important acoustic features of cry. A unique and important aspect of this work is the rigorous testing of the system's accuracy as compared to ground-truth manual coding. The resulting system has implications for basic and applied research on infant cry development.
Large-aperture microphone arrays can be used to capture and enhance speech from individual talkers in noisy, multi-talker, and reverberant environments. However, they must be calibrated, often more than once, to obtain accurate 3-dimensional coordinates for all microphones. Direct-measurement techniques, such as using a measuring tape or a laser-based tool are cumbersome and time-consuming. Some previous methods that used acoustic signals for array calibration required bulky hardware and/or fixed, known source locations. Others, which allowed more flexible source placement, often have issues with real data, have reported results for 2D only, or work only for small arrays. This paper describes a complete and robust method for automatic calibration using acoustic signals which is simple, repeatable, accurate, and has been shown to work for a real system. The method requires only a single transducer (speaker) with a microphone attached above its center. The unit is freely moved around the focal volume of the microphone array generating a single long recording from all the microphones. After that, the system is completely automatic. We describe the free source method (FrSM), validate its effectiveness and present accuracy results against measured ground truth. The performance of FrSM is compared to that from several other methods for a real 128-microphone array.
Large-aperture microphone arrays can be used to capture and enhance speech from individual talkers in noisy, multi-talker, and reverberant environments. An important factor for their use is array calibration. Due to the size and complexity of these arrays, direct-measuring techniques such as using a measuring tape or a laser-based tool are too cumbersome and time consuming. Previous methods that used acoustic signals for array calibration worked well for smaller arrays, but they had the disadvantage of establishing a coordinate system based on the sound sources whose positions had to be measured precisely. This not only limited the number of source locations that could be used for calibration but also made inevitable re-calibrations tedious. This paper describes a new and robust method for automatic calibration using acoustic signals that not only establishes a fixed coordinate system on the microphones but also significantly simplifies the process in general. This method requires only a single speaker with a microphone attached above its center. The unit is freely moved around the focal volume of the microphone array for a single long recording from all the microphones. We describe the method, validate its effectiveness and compare the results against measured ground truth and some previous methods for a real 128-microphone array.
Cross-correlation-based methods have been used extensively in the task of locating multiple, simultaneous sound sources in adverse environments. In this paper, we present a low-cost prealignment enhancement to fix the temporal drawback of the cross-correlation functionals by aligning all the microphone signals to ensure they correlate to the same temporal event. We further introduce a new functional, the steered-response power of the minimum-variance distortionless response using the phase transform (MVDR-PHAT), for multiple-source detection and localization. Experimental results using real data of a 10-talker recording in an adverse room show the improvements of the proposed functional and the prealignment enhancement over traditional techniques in detecting and locating simultaneously active talkers.
Two great contemporaries in New York City, both driven and stubborn, played dramatic roles in the growth of the electronics industry. They were giants in stature, although one was dedicated to science and the other to business. They were friends for over 20 years and grew inimical as older men. Sarnoff received more honors in his day, but others could not maintain the edifice he built during his leadership. Armstrong's inventions have endured and have set some basis for most of modern day analog and digital broadcasting.
A method is presented for using mean fundamental frequency to measure talker similarity in real time from conversational speech in a noisy, reverberant room. This talker-similarity function is designed with the ultimate goal of real-time talker labeling in mind. A large-aperture array of wallmounted microphones is used, and talkers are allowed to enter the room without providing prior enrollment data. Because noise conditions vary widely at different positions and orientations in the room, it is desirable to use features that are unaffected by environmental conditions so that training can be performed on clean speech. Fundamental frequency is a natural feature choice because it is more robust to noise than other spectral parameters. The extraction of identity information from very short segments of speech was achieved by taking into account the within-talker and betweentalker variability of mean fundamental frequency. Models trained with clean speech were verified to be accurate when tested on speech recorded from freely-walking human talkers with distant microphones in a noisy room. The experiments show that mean fundamental frequency alone, while rarely sufficient for confirming that two talkers are the same, is very effective for rejecting an identity match.
As an avid reader, crazy-gadget-loving engineer, and history enthusiast, I bought a US$1 paperback when I was a graduate student in 1969 titled Man of High Fidelity, the 1969 paperback version of the 1956 biography of Edwin Howard Armstrong by Lawrence Lessing. The story told within was so fascinating that I needed to learn more, so I borrowed a copy (it had been given to an outstanding Radio Corporation of America (RCA) researcher by Sarnoff himself) of the commissioned autobiography of David Sarnoff written by his cousin, Eugene Lyons, published in 1966. It was clear that these two giants of their time had lives that intermingled over more than 40 years.
Beamforming techniques are applied to microphone arrays with the aim of separating sources and improving intelligibility, by means of spatial filtering. The non-stationary nature of speech implies the use of adaptive beamformers and several solutions have been implemented. Furthermore, interfering signals coming from the same direction as the target signal, cannot be filtered by the beamformer. The method presented in this paper is an alternative to adaptive beamforming, combining a simple delay-and-sum beamformer with a time-frequency masking method based on phase information. The beamformer is steered to the desired source and a function related to the phase differences between the steered signals at the microphones is evaluated to reject any interference that passed through the beamformer. Thus, the algorithm does not need to constantly adapt the filter coefficients and takes advantage of both beamforming properties and time-frequency separation techniques. The separation performance of the method has been evaluated in a noisy and reverberant environment using different arrays, talkers and scenarios. Real data are used to show the performance of the real-time algorithm when isolating one of the sources in the mixture.
The problem addressed is generally in the family of speaker verification, but its conditions and requirement s are very different from what has been published in this area. We want to label approximately five to ten moving talkers in the reverberant and noisy focal area of a large-aperture microphone array, in real time, from short segments of conversational audio. Current real-time systems use only spatial informat ion, which is inadequate when talkers move while silent. Given th e dynamic noise conditions in this kind of environment, it is very difficult to collect sufficient training data for a conventional algorithm. The proposed algorithm is easily implementablein real time and is trained offline using only noise-free data. This is possible because a set of robust features is used that represent the differences between a pair of speech segments . Also, the output of the algorithm is a probability, rather th an a measure requiring threshold tuning, which makes on-the-fly decisions straightforward. The algorithm is introduced and its properties are demonstrated. While quite simple, the algor ithm has been shown to outperform some commonly-used talkerdistance metrics for real data taken from a noisy-environment, array system. This microphone-array dataset is available t o other researchers at www.lems.brown.edu/array/data/movingta lkers/. EDICS Categories: SPE-SPKR, AUD-LMAP
Extracting a high-quality speech signal of a single source from a multiple-source input in an adverse environment has always been a challenge for microphone-array processing. Three major approaches have been proposed to tackle this problem: blind-source separation (BSS), beamforming (BF), and computational auditory scene analysis (CASA). Combinations of the CASA and BF, BSS and BF also have been introduced. In this paper, we propose a new algorithm which utilizes the null-steering beamformer minimum-variance distortionless response (MVDR) using the proven-robust phase transform (MVDR-PHAT) and the CASA framework that closely mimics human hearing perception. Experimental results using real data recorded in a room with high background and reverberation noise indicated the improved performance of the proposed algorithm compared to those of traditional beamforming algorithms and an SRP-PHAT-based source-separation algorithm recently described at ICASSP 2010.
Two new methods for locating multiple sound sources using a single segment of data from a large-aperture microphone array are presented. Both methods employ the proven-robust steered response power using the phase transform (SRP-PHAT) as a functional. To cluster the data points into highly probable regions containing global peaks, the first method fits a Gaussian mixture model (GMM), whereas the second one sequentially finds the points with highest SRP-PHAT values that most likely represent different clusters. Then the low-cost global optimization method, stochastic region contraction (SRC), is applied to each cluster to find the global peaks. We test the two methods using real data from five simultaneous talkers in a room with high noise and reverberation. Results are presented and discussed.
An important application for microphone arrays is to extract high-quality output from a single wideband source in multi-source and adverse environments. Methods based on blind-source separation and beamforming have been proposed in the literature. In this paper, we propose a new algorithm for isolating sources which may be considered an alternate approach to adaptive beamforming. The proposed method assigns every time-frequency point to one of the sources in the environment or as background noise using an SRP-PHAT-based discriminator. It then creates the desired source's output by spectrally subtracting the inappropriate time-frequency points. After giving the details of the procedure, we show results from real data from our large-aperture microphone array. Though, clearly a relatively small subset, listening results to date for two or three talkers have yielded high-quality isolations. For two talkers, in particular, results are presented which indicate that the method correctly identifies and removes about 90% of those time-frequency points where the interfering signal is dominant, without affecting the desired signal significantly.
Knowing the orientation of a talker in the focal area of a large-aperture microphone array enables the development of better beamforming algorithms (to obtain higher-quality speech output), improves source-location/tracking algorithms, and allows better selection and control of cameras in a video conference situation. Measurements in an anechoic room (e.g., Chu and Warnock, 2002) have quantified the average frequency-dependent magnitude (source radiation pattern) of the human speech source showing a front-to-back difference in magnitude that increases with frequency by about 8 dB/decade reaching about 18 dB at 8000 Hz. These amplitude differences, while severely masked by both coherent and noncoherent noise in a real environment, are the most extractable phenomena from a talker's orientation when compared to other phenomena such as phase differences due to the source or effects due to diffraction at the mouth. In this paper, we propose a robust, source-radiation-pattern-based method for extraction of the azimuth angle of a single talker for whom an accurate point-source location estimate is known. The method requires no a priori training and has been tested in more than 100 situations with real human talkers having various locations and orientations in a room equipped with a large aperture microphone array. We compare these results against earlier published algorithms and find that the method proposed herein is the most robust and is sufficient to be considered for a real time system.
Computational cost has been an issue for the proven robust source localization algorithm, steered response power (SRP) using the phase transform (SRP-PHAT). Some proposed computation reduction algorithms degrade under high noise and reverberant conditions. Some require at least 10% the cost of a full SRP-PHAT gridsearch. In ICASSP 2007, we introduced a robust, low-cost global optimization technique, stochastic region contraction (SRC). In this paper, we present another algorithm, stochastic particle filtering (SPF), which uses SRC's initialization and is a kind of Importance Sampling technique. In this paper, the SRP is computed using a modification to the conventional PHAT, namely Ã-PHAT. Extensive experiments using real data and simulated data are shown. The results indicate that, while maintaining the desirable accuracy of the full search, this method reduces the cost to about half the cost of SRC (0.03% the cost of full search), thus making SRP-PHAT more practical for real-time applications.
John Adcock合作论文数FXPAL6