The directivity of the human voice has been studied since the early twentieth century using different measurement systems with progressively higher spatial resolution. Artificial heads and mouths have been used because of their ability to repeat a given sound production, hence allowing sequential measurements of directivity with a reduced number of microphones and still achieving high spatial resolution. Unlike most artificial heads, whose external geometry is abstracted, this study uses a custom 3D-printed head with detailed geometry and three different mouth openings, all based on 3D scans from magnetic resonance imaging data. The impulse response measurements were performed using a 3D robotic arm, resulting in directivity data with 5 degrees resolution in both azimuth and elevation. The measured directivity patterns are consistent with previous research on energy distribution in space over angles (azimuth and elevation) and frequency, with a higher spatial resolution and for different mouth shapes. The resulting data set is made available in several standardized file formats to facilitate accurate voice directivity simulations in virtual acoustic environments.
The human voice is directional by nature, and its directivity has been the subject of extensive study. Accurate data on voice directivity are fundamental in any work associating natural human speech and its interaction with architectural spaces. In this research, we present and analyze results obtained using a novel high-resolution measurement setup. The system features an array of 180 MEMS microphones arranged on a horizontal circle with a 3-m diameter, enabling measurement of sound sources with a 2° resolution in the azimuthal plane. Test subjects were positioned at the center of the array, with the location of their mouths precisely calibrated using a multi-camera and laser alignment system. Data were collected from 24 talkers (5 female, 19 male) articulating fluent English speech. The analysis, conducted across different frequency bands, explores variations between sexes and compares findings with previous studies. We discuss the relevance of the resolution provided by the system and examine its performance within the context of existing literature. Potential directions for future work are also outlined.
The first published attempt to measure human voice directivity in 1929 involved a single microphone and a rotating chair.Since then, a number of experiments were conducted, increasing the level of detail of the sound field representation according to the available equipment.Most experiments assumed repeatability of voice production by the human talker over an iterative process, which can hardly be guaranteed even for trained subjects.Hence using a rotating device to retrieve the 3D directivity of a human talker is questioned.All measurement positions should preferably be recorded simultaneously to ensure an optimal consistency of the obtained directivity patterns.The present experimental setup is composed of 256 MEMS microphones located on a spherical structure of 1.80 m radius.The characteristics (sensitivity, frequency range, and dynamics) of those compact devices are presented and allow for measurements of voice directivity from 50 Hz up to 20 kHz.A first step of characterization is conducted with a reproducible sound source composed of 12 Aurasound loudspeakers producing controlled directivity.Then, human talkers are recorded in the system.Far-field directivity functions are estimated using a spherical wave propagation model.The obtained results are consistent with the previous literature and provide an extended angular accuracy.
The far-field directivity function (FFDF) of a sound source characterises the angular dependence of the acoustic fields far away from this source. This function can be estimated by performing a spherical wave expansion (SWE) of the sound field from pressure measurements distributed around the source. The truncation order of the SWE is a critical parameter that depends on the number of measurement points and on the spatial sampling scheme used for the measurements. Unfortunately, for irregular sampling schemes, no theoretical result allows to determine a truncation order. In this work, a surrounding cuboid microphone array composed of 256 MEMS microphones is deployed. A cross-validation framework is used in order to determine an optimal truncation order for the SWE of the field radiated by a test source. The quality of the reconstructed field is assessed on the frequency range 100-5000 Hz. Finally, the estimated SWEs and far-field directivities are compared to an analytical model of the test source.
Human vocal folds are highly deformable non-linear oscillators. During phonation, they stretch up to 50% under the complex action of laryngeal muscles. Exploring the fluid/structure/acoustic interactions on a human-scale replica to study the role of the laryngeal muscles remains a challenge. For that purpose, we designed a novel in vitro testbed to control vocal-folds pre-phonatory deformation. The testbed was used to study the vibration and the sound production of vocal-fold replicas made of (i) silicone elastomers commonly used in voice research and (ii) a gelatin-based hydrogel we recently optimized to approximate the mechanics of vocal folds during finite strains under tension, compression and shear loadings. The geometrical and mechanical parameters measured during the experiments emphasized the effect of the vocal-fold material and pre-stretch on the vibration patterns and sounds. In particular, increasing the material stiffness increases glottal flow resistance, subglottal pressure required to sustain oscillations and vibratory fundamental frequency. In addition, although the hydrogel vocal folds only oscillate at low frequencies (close to 60 Hz), the subglottal pressure they require for that purpose is realistic (within the range 0.5–2 kPa), as well as their glottal opening and contact during a vibration cycle. The results also evidence the effect of adhesion forces on vibration and sound production.
Several places, for example theaters, auditoria or even churches, present an interesting acoustic feature: a non-linear sound decay (Fig.1). This phenomenon is provided under specific conditions by various architectural volumes, which are acoustically linked to each other. Acousticians have been interested in understanding the relationship of acoustic fields in these volumes, and their interactions. Knowledge on these phenomena has led architects and acousticians to design concert halls based on the coupled volume principle in last decades 1 , with more or less success. The concert halls built in Lucerne, Switzerland 2 and more recently in Suzhou, China 3 and soon in Paris, France 4 are relevant examples. Therefore precise analysis of sound energy decay in such places is necessary. Moreover standardization of an analysis method still does not clearly exist for non-linear sound energy decays 5 . Furthermore a fine knowledge of the sound energy decay is necessary to estimate the influence of variations of reverberation on the perception of a listener in such places 6 . While analytical models 7-10 , that describe non-linear decays for step or impulse response, provide smooth decay curves, the ones from measurements or numerical simulations are generally more jagged than the latter, and hence more difficult to analyze. An accurate analysis method has to be robust in order to point out the relevant characteristics, despite the fluctuations of decay curves. This paper will first present sound decay models and different analysis methods developed in recent years, based on two different principles. Then a new method for estimating non-linear decay characteristics is proposed. Finally some problematic issues will be raised concerning the use of Schroeder backward integration 11 for non-linear decays.
. It currently allows to study the vibromechanical behaviour of extensible isotropic vocal folds in fluid-structure interaction. This article describes the different components of the testbed and presents the first results obtained on homogeneous and isotropic folds. The larynx replica consists of a deformable silicone-rubber envelope into which expandable vocal fold replicas can be inserted for testing. The material and structural properties of the folds are adjustable. Three silicone-rubber with various stiffnesses were tested, as well as a cross-linked hydrogel. The self-oscillation of the pleats is controlled by the upstream air flow. The tested folds self-oscillated over a wide range of flow rates and for various degrees of stretching.
A perceptual threshold related to spatial resolution of the human voice directivity was determined through a listening test of similarity (MUSHRA). Directivity data of an artificial talking head measured at high spatial resolution (spherical harmonics order 35) was the input of a room acoustics simulation software (RAVEN) to build sound stimuli in various room acoustic conditions and source–receiver arrangements, with different voices. Results showed that, at spherical harmonics order 8 and above, the voice signal was not anymore perceived as significantly different from the greatest resolution. An analytical model was proposed and showed good agreement with the listening test results.
: An in-vitro testbed was designed to study the vibromechanical behaviour of biomimetic and extensible vocal folds. The paper describes the several steps of its conception. It consists of a deformable laryngeal envelope in which stretchable vocal-fold replicas of adjustable material and structural properties can be inserted for testing. The folds are able to oscillate for a wide range of aerodynamic conditions and material elongations.
Performers, whether instrumentalists or singers, tend to develop rather individual adaptation patterns across room acoustical conditions. While instrumentalists mostly adapt in terms of tempo, singers tend to emphasize variations of loudness and timbral colour in conjunction with related room-acoustical parameters, namely Stage Support and Bass Ratio. Adaptation to room acoustics has also been investigated on talkers in terms of vocal effort, explicitly mentioned by participants and correlated to voice intensity. Vocal effort while talking was found to be larger in rooms that lack early reflections. Such adaptation should be reflected on glottal-source parameters. Electroglottographic measurements enable to non-invasively assess several features of glottal behaviour, such as contact quotient (duration of glottal contact over a cycle) or its counterpart, open quotient. In speech, contact quotient was found to be correlated to vocal effort. In singing, it was found to be strongly related to vocal intensity in laryngeal mechanism M1, and to fundamental frequency in mechanism M2. The present study aims at investigating the singing adaptation process in rooms with various acoustics, by assessing voice production by means of acoustical and electroglottographical in-situ measurements.
Previous research has showed that singers and instrumentalist musicians tend to adapt their sound production to the acoustics of the venue in which they perform. Studying such an adaptation process is not trivial since many parameters are to be taken into account. For example, various performances can dier due to aspects of the halls other than acoustic, the interaction with the audience, or the physical and psychological state of the performer. A specific methodology has been used to study how singers vary their sound production when they perform in dierent acoustical spaces within a short period of time. Virtual concert halls were proposed to four singers who could hear themselves in reverberant sound fields by means of a real-time convolution process and dynamic binaural synthesis. The signals captured by a microphone in near-field and by an electroglottography unit were analysed and yielded several parameters related to voice production, namely vocal intensity, fundamental frequency, and glottal contact quotient. These voice-quality parameters were compared to room acoustical parameters from the eight dierent rooms proposed to the singers. Results showed that the singers reacted in specific manners by varying their voice quality and being sensitive to dierent room acoustical parameters.
To investigate the influence of room acoustics on singing, four lyrical singers (soprano, mezzo-soprano, tenor, baritone) performed four musical pieces in eight different venues (from dry studio to reverberant church). In addition to vocal intensity measured by a near-field microphone, glottal behavior (vibratory fundamental frequency and contact quotient) was assessed by electroglottography. Statistical linear mixed models showed that the variance in vocal performance was partly explained by room acoustics. Complementary to previous results on voice musical features influenced by timbre and level of the room's response, voice production parameters were mostly influenced by spatial aspects of the room's response.
Jean Schoentgen合作论文数National Fund for Scientific Research, Belgium1