
In this study, five singers performed under expressive and neutral conditions in front of a listener, while biometric signals including electroencephalography (EEG), galvanic skin response (GSR), and electrocardiography (ECG) were recorded from both performers and audience, alongside high-quality audio recordings of the singers, and emotion annotations. Audio features were extracted and analyzed. Physiological signals were processed to derive markers of relaxation, heart rate variability, and arousal. Four machine learning models were trained to classify expressive versus neutral performances, achieving up to 82% cross-validated accuracy, with pitch, loudness, and spectral features emerging as the most informative predictors. Multimodal analysis revealed additional discriminative patterns, as performers exhibited higher EEG-based relaxation during expressive performances, while audience heart rates reflected heightened arousal. Furthermore, emotional alignment between performers and listeners was observed predominantly in the expressive condition. Audio and behavioral features obtained from the multimodal analysis were consistent with the broader dataset collected within the witheFlow project, which included 16 instrumentalists, allowing for direct comparison with the current study. These findings indicate that vocal expressivity can be reliably quantified using combined audio and physiological analyses, providing a foundation for future studies on emotional communication in music performance.
Voice performance can be characterized with respect to dynamics and tonal range by measurement of a voice range profile (VRP). Usually two series, a s oft and a loud s cale of the sustained phoneme /a:/ from the lowest to the highest possible pitch are recorded and plotted in a diagram from which the extreme values are extracted (lowest, highest, softest, loudest sound). We propose a more detailed analysis of the voice ranges by extraction of seven markers from the diagram, allowing an assessment of the development of the
Oboists form vowels in their mouths while playing to improve the response of certain notes or to change the timbre. We developed an impedance measurement device that allows the measurement of the intraoral impedance while playing the oboe. Several students were measured while playing different pitches and articulating the vowels /e:/, /i:/ and /o:/. For some notes, the signal spectra show a strong influence on the strength of the partial tones under different articulation conditions, while some notes are not affected by this.
This paper presents a finite element model of vocal fold dynamics with muscle activation-dependent parameters. The geometry is parametrized analogously to lumped-parameter formulations, capturing how intrinsic laryngeal muscle activations affect vocal fold length and layer depths. Elastic properties are modeled as polynomial functions of muscle activation, and the polynomial coefficients are identified from data of modal frequencies as functions of activation. The approach provides a simple yet extensible framework that bridges the gap between lumped and highly detailed finite element models, making it well suited for characterizing vocal fold dynamics when activation-dependent data are available.
Enrico Caruso made almost 500 between 1902 and 1920. He recorded recordings many of the arias and songs from his repertoire several times. This enables us to compare recordings from different periods of his career and to create a vocal profile that covers almost his entire active artistic career. From a voice doctors' perspective, this is interesting in two respects. On the one hand, from the perspective of vocal physiology, we can understand was Caruso's technique, which considered novel by contemporary critics. On the other hand, from phoniatric point of view, Caruso's voice can be assessed to determine whether his numerousS health problems, including two operations on his vocal cords, are audible as vocal damage, as contemporary critics who heard him live described such acoustically perceptible limitations. Since Caruso's original recordings have several technical limitations, the analysis was carried out by an expert rating with 21 special focus on vocal technique and the parameters of hoarseness and dysodia.
Infant cry analysis is an emerging field in biomedical engineering for assessing neonatal health. It offers a non-invasive, simple, and low-cost method to detect physical and physiological conditions. This work explores its potential as reliable tool for early pathology detection, supporting timely medical intervention. A deep learning approach based on a convolutional neural network (CNN) was developed using spectrograms and Mel Frequency Cepstral Coefficients (MFCC) as input features. Two strategies were evaluated: direct use of crying signals from the database and preprocessing by removing silent segments before MFCC extraction. Unlike binary schemes, this study addressed a multiclass classification problem, categorizing crying signals into asphyxia, deafness, hyperbilirubinemia, hypothyroidism, and healthy infants. The CNN achieved high classification accuracy per class: 0.94 for asphyxia, 1.00 for deafness, 0.92 for hyperbilirubinemia, 0.95 for hypothyroidism, and 1.00 for healthy infants. Model performance was assessed with accuracy, sensitivity, and specificity across multiple classes. These results show that CNNs can achieve strong performance without extensive preprocessing.
Objective: The aim of this study was to apply motor speech diadochokinetic assessment in voice evaluation of patients with Parkinson's disease (PD). Methods: Six patients with PD were evaluated with a voice assessment protocol, including the self-assessment questionnaire, the perceptual GRB Scale, the INFVo rating scale for substitution voices, the acoustic analysis of jitter, shimmer, glottal-to-noise excitation ratio (GNE), and the motor speech diado-chokinetic parameters (DDK, DDK standard deviation, DDK jitter, and mean syllable length (MSL) in syllables with anterior, middle, and posterior articulation). The UPDRS scale for PD severity was measured as well. Results: t he m edian v alues o f D DK, D DK S D, a nd DDK Jitter were 3.99 S/s (IQR: 2.60-4.44 S/s), 2.38 S/s (IQR: 2.12-3.39 S/s), and 2.56% (IQR: 2.15-2.72 %), respectively. In the considered PD patients, DDK was significantly higher than the normality range, p=0.0277. DDK standard deviation directly correlated with the same perceptive parameter (rho=0.88, p=0.0198), as well as with the tremor (rho= 0.971, p=0.0012), and the intelligibility scores (rho= 0.899, p=0.0149). Conclusions: despite the limited sample size, these preliminary results suggest a promising role of motor speech diadochokinetic parameters in the clinical assessment of patients with PD. However, further studies on larger cohorts are necessary.
Vibrato, a periodic variation in frequency and intensity, is a commonly found property of musical performance, adding a sense of flexibility and richness. Prior studies comparing vibrato in different singing styles employed mainly laboratory recordings, focusing on two principle features, rate and extent. The current study examined vibrato characteristics of classical, jazz and pop singing, employing commercial recordings of prominent artists. This provided a highly ecological description of artistic style. The presence of vibrato was determined perceptually by a panel of five professional musicians. Time-frequency analysis was then performed on pitch contours of sustained notes. A range of thresholds was applied to the time-frequency matrix to improve SNR. Several measures were derived from the thresholded matrix: vibrato extent, rate, consistency and onset time. Vibrato rate and extent were found to be higher in opera singing, compared to jazz and pop singing styles. Vibrato onset was found to be delayed mainly in jazz and pop singing. These results reinforce the importance of vibrato as a stylistically defining parameter and may serve as a tool in future research and vocal pedagogy, as well as in automatic recognition of singing styles.
The aim of this study was to propose an updated panel of acoustic parameters to measure the effectiveness of Botulin toxin laryngeal injections for adductor spasmodic dysphonia (AdSD), including the motor speech diadochokinetic assessment. Study Design. This is a case report study Methods: Two cases diagnosed with adductor spasmodic dysphonia were evaluated at baseline, and one and three months after Botulinum toxin injection in the thyroarytenoid muscle. The voice protocol assessment included the self-assessment questionnaire, the traditional perceptual GRB Scale and the INFVo rating scale for substitution voices, the acoustic analysis of jitter%, shimmer%, glottalto-noise excitation ratio, and the motor speech diadochokinetic parameters. Results: In both clinical cases, the motor speech diadochokinetic parameters showed an improvement one month after the treatment compared to the baseline, and a subsequent decrease 3 months after treatment, in agreement with the trends showed by the self-assessment questionnaire. Conversely, the traditional perceptual and acoustic parameters did not demonstrate to be consistent in both cases with the self-assessment results. Conclusions: The results suggested that motor speech diadochokinetics parameters could be relevant to assess the outcome of Botulinum toxin treatment for AdSD
Let "voice range profile" (a.k.a, "phonetogram") be the term for a graph of the maximum phonatory range of a voice on the f(o)Chi SPL plane, i.e., a closed contour. Let "voice map" be the term for a map of a scalar metric over some relevant range, not necessarily to the extremes, on that same plane, i.e., a 2D scalar field. For imaging several metrics, one voice map can have several "layers", all derived from the same recording. This paradigm for collection and collation of voice data is useful, because it accounts for how the chosen metrics vary systematically with f(o) and SPL. Both f(o) and SPL are influential and typically nonlinear covariates of other voice metrics. Not accounting for them can obscure the effects of an intervention. Here we summarize some central concepts, rationales and techniques related to voice mapping.
Pathological speech, caused by dysphonia or produced via electro-larynx devices, often suffers from poor intelligibility and unnatural prosody. In this paper, we investigate the potential of four state-of-the-art voice conversion models: FreeVC, QuickVC, LLVC, and XVC for restoring healthy-sounding speech. All models are fine-tuned on Austrian-German datasets and evaluated using objective and subjective metrics. Results show substantial gains in intelligibility, naturalness, and perceived vocal health. QuickVC, FreeVC, and XVC perform similarly and achieve the highest preference scores, exceeding unprocessed pathological speech by up to 200%. These findings highlight the potential to improve communication for individuals with voice disorders and motivate further development of efficient, high-quality conversion systems.
In this work I report on the findings of a survey of vibrato parameters in three different large datasets of old recordings of classical western music. Each dataset is designed with different criteria and different purposes. The trend of decreasing rate and increasing rate with recording time is confirmed also in new data. Moreover, a small but significant difference of -0.3 +/- 0.07 Hz in vibrato rate between male and female singers is found, male vibrating slower, and overall decrease of vibrato rate correlated with age in five famous tenors.
This study investigated the impact of overtone flute training on female voice production. Ten vocally untrained women completed a five-week intervention with the shepherd's overtone flute, which requires airflow control without phonation. The aim was to test whether improved breath management influences vocal outcomes influence vocal outcomes. Voice recordings were collected before and after training. Tasks included habitual and stage reading, gradual call, glissando, and full singing voice range profiles (VRPs) at three dynamic levels. Long-term average spectrum (LTAS) was calculated from voiced segments. Comparisons were made using nonparametric tests. Results showed significant increases in maximum phonation time, mean and maximum SPL in stage reading, and maximum f(0) in gradual call. Glissando and singing tasks revealed wider dynamic ranges, caused by higher SPL maxima and lower minima. LTAS analysis confirmed spectral strengthening above 2 kHz, most evident in stage speech, glissando, and medium-dynamic singing. These changes align with earlier findings linking high-frequency energy to epilaryngeal tube adjustments. In conclusion, overtone flute training, despite being voiceless, produced measurable vocal improvements. The intervention enhanced breath control, and reinforced resonance in spectral regions important for projection. These outcomes suggest that overtone flute exercises could be valuable in voice education and rehabilitation.
This study explored potential behavioral indicators of phonation in two individuals with Autism Spectrum Disorder (ASD), focusing on desynchronization signatures derived from EEG-band analysis. Longitudinal data revealed significant differences in the synchronization index (SI) during attentive vocalization, with normotypical controls exhibiting marked desynchronization, most notably among males. While the majority of phonations fell within central quartiles, select outliers demonstrated heightened desynchronization. Correlational findings linked increased desynchronization to greater pitch variability, along with higher kurtosis in males and greater skewness in females, indicating distinctive phonation qualities across cases. The contrast between the male's imitative speech and the female's spontaneous vocalizations underscores the difficulty of standardizing protocols, reinforcing the importance of tailored, empathetic approaches to vocal data collection in ASD research.
In recent years, the combination of physiological and psychological domains has gained increasing interest for capturing emotional and stress-related changes, especially through quantitative analysis of voice and speech. This study explores the relationship between personality traits and stress-induced vocal alterations within an Ecological Momentary Assessment (EMA) framework. Thirty-six male undergraduate students completed baseline assessments of personality traits (PID-5) and provided voice recordings before and after an academic exam. Acoustic analyses focused on sustained vowels and a standardized sentence, extracting acoustic features and mel-frequency cepstral coefficients. The results revealed significant vocal changes, including increased fundamental frequency, alterations in articulation and prosody, along with distinct acoustic patterns associated with higher scores in specific PID-5 dimensions. These preliminary findings suggest that integrating psychological and physiological data can enhance the understanding and monitoring of emotional states and stress responses, paving the way for future development of acoustic quantitative analysis within EMA paradigms as a tool to support identification of potential voice-based biomarkers for stress and psychopathology vulnerability
This a rticle p resents t he d esign o f a n experimental protocol created to identify phonetic biomarkers of Alzheimer's disease (AD) in Russian speakers. Due to the uniqueness of the prosodic system of the Russian language and its lack of study in the context of Alzheimer's disease, the proposed protocol aims to fill this gap. The study is based on a cross-sectional design and includes recording of three types of speech activity: spontaneous conversation, reading a semantically ambiguous poem by Eduard Uspensky, and an ironic p oem by Samuil Marshak, which allows for a multifaceted analysis of various levels of speech production. The recording is carried out at the participants' homes using professional equipment to ensure environmental validity and high sound quality. The data analysis plan involves the automated extraction of a wide range of acoustic parameters (tempo, pauses, frequency, and phonation characteristics, etc.) followed by statistical processing and the use of machine learning methods to build classification models. It is expected that the implementation of this protocol will make it possible to identify reliable and reproducible speech biomarkers, which will form the basis for the creation of objective disease screening and monitoring tools in the Russian-speaking population.
This study questions the notion of creaking in singing. A variety of distorted sounds have been recorded by two singers who are also singing teachers from the international Sing&Scream team. Based on audio and electroglottographic recordings together with endoscopic views, it is shown that creaking is a vocal effect that can be added to sung sounds, independently of the laryngeal mechanism in use. While glottal pulse train and acoustical signal are greatly modified by adding a creaking effect to a sustained sound, the supraglottic adjustments evidenced on endoscopic videos are subtle. Lateral supraglottic constriction seems to be a control mechanism to introduce creaking and increase its degree of irregularity. Pharyngeal constriction has also been evidenced in female singing.
In this longitudinal study aspects of vocal performance in music teachers (n=42) are investigated at the beginning and at the end of their university studies, as well as at least ten years after graduation. We used voice range profiles (VRP) investigating both the singing and speaking voice to characterize the development. Whereas VRP parameters show a clear increase in vocal abilities during training, the third measurement during working life reveals a stagnation or even regression in the vocal performance represented in the VRP.
This paper argues that the perception of human personality through voice is deeply connected to the principles of music perception. By analyzing speech melody through the lens of harmony, melodic contour, and acoustic expression, we can uncover the specific vocal cues that signal personality traits such as neuroticism and extraversion.
The acoustic and kinematic analyses of speech represent non-invasive and cost-ffective tools to support clinicians in the assessment of bulbar dysfunction, which is particularly challenging in amyotrophic Lateral sclerosis (ALS). In this study, we applied multimodal speech analysis to identify suitable acoustic kinematic biomarkers and potentially able to detect bulbar impairment, particularly the presence of dysphagia in its early stages. Our results revealed clear distinctions between dysphagic and non-dysphagic individuals with ALS during connected speech, with the third formant and speech timing metrics being the most significant features (p < 0.001). Our findings suggest that a reduced F3 and increased duration of pauses during connected speech may serve important biomarkers for detecting the onset of dysphagia in ALS.