Standard endoscopy of vocal folds is in general limited to two-dimensional imaging. Laser-based 3D imaging offers not only absolute measurements but also the possibility of assessing all three spatial directions. However, due to human inter-individuality, a fixed grid configuration (with fixed edge length and spot size) does not necessarily provide the best coverage and resolution. We present a liquid lens optical design for a diffractive spot array generator with dynamic adjustment capabilities for both array size and spot size. The tunable nature of the liquid lenses enables precise control over the spot array generated by a diffractive optical element (DOE). The first liquid lens controls the spot divergence in the observation plane, while the second liquid lens adjusts the zoom factor. The optical configuration provides a dynamic range of 1.8 with respect to array size, significantly enhancing adaptability in imaging across various applications.
Objectives This study investigates the use of sustained phonations recorded during high-speed videoendoscopy (HSV) for machine learning-based assessment of hoarseness severity (H). The performance of this approach is compared with conventional recordings obtained during voice therapy to evaluate key differences and limitations of HSV-derived acoustic recordings. Methods A database of 617 voice recordings with a duration of 250 ms was gathered during HSV examination (HS). Two databases comprising 809 vowels recorded during voice therapy were used for comparison, examining recording durations of 1 second (VT-1) and 250 ms (VT-2). A total of 490 features were extracted, including perturbation and noise characteristics, spectral and cepstral coefficients, as well as features based on modulation spectrum, nonlinear dynamic analysis, entropy, and empirical mode decomposition. Model development focused on selecting a minimal-optimal feature subset and suitable classification algorithms. Recordings were classified into two groups of hoarseness based on auditory-perceptual ratings by experts, yielding a continuous hoarseness score yˆ. Model performance was evaluated based on classification accuracy, correlation between predicted scores yˆ∈[0,1] and subjective ratings H∈{0,1,2,3}, and correlation between the relative change in quantitative and subjective ratings. Results Logistic regression combined with five acoustic features achieved a classification accuracy of 0.863 (VT-1), 0.847 (VT-2), and 0.742 (HS) on the test sets. A correlation of 0.797 (VT-1), 0.763 (VT-2), and 0.637 (HS) was obtained between yˆ and H, respectively. For 21 test subjects with two recordings, the model yielded a correlation of 0.592 (VT-1), 0.486 (VT-2), and 0.088 (HS) between ∆yˆ and ∆H. Conclusion While acoustic signals recorded during HSV show potential for quantitative hoarseness assessment, they are less reliable than voice therapy recordings due to practical challenges associated with oral laryngeal examination. Addressing these limitations, for example, through the use of flexible nasal endoscopy, could improve the quality of HSV-derived acoustic recordings and voice assessments.
Objective: This study investigates relationships between the oscillation behavior of the medial and superior vocal fold (VF) surfaces during sustained phonation in a human cadaver hemilarynx. Methods: An experimental test stand synchronously captured the medial and superior VF surfaces of a human ex vivo hemilarynx during sustained phonation using two high-speed camera setups in 24 experimental settings. The 3D coordinates of the medial VF surface were reconstructed by triangulation of sewn-in marker points, while laser-based reconstruction was used for the superior VF surface. Correlation analysis and linear regression were used to quantify the connections of the mean and maximal vertical and lateral VF displacements and the VF velocities. Additionally, stepwise linear regression was used to analyze the impact of the measurement variables mean flow rate, adduction and elongation. Results: Strong linear relationships between all of the tested corresponding parameter pairs of the superior and medial VF surfaces were found (p<.001). Mean and maximum vertical displacements of the medial surface were both approximately 50% of the superior surface. The mean lateral displacements for the medial surface were 12% below the superior surface but 12% higher for the maximum values. The mean and maximum VF velocities were 32% and 36% lower for the medial surface. Conclusion: The suggested multi-modal test stand allows efficient, comprehensive analysis of human hemilarynges and provides promising information about the interaction of the different VF areas and opens up the systematic analysis of multiple hemilarynges. Significance: In future, our results could integrate into ENT diagnostics using 3D laryngoscopy where the hidden medial VF surface dynamics may be predicted from the observable superior surface.
Simulating and representing the phonation process is a computationally intense problem. In this study, we address this issue using implicit neural representations to determine the possibilities of saving computational load by representing computational fluid dynamics simulations through continuous functions represented in a deep neural network. Our work demonstrates the feasibility of using implicit neural representations of a laryngeal aerodynamic simulation containing about 180 × 106 data points within a single neural network. Additionally, we show that with only 20% of the simulated data, we can restore the original resolution with implicit neural representations, showing only nuanced differences compared to the original simulation. We are also confident that with the proposed approach, we can further lower the representation in space and time in future work.
For clinical assessment of the voice production process, evaluation of the acoustic voice signal as well as the laryngeal dynamics (i.e., the vocal fold oscillations), producing the source signal in the larynx, is highly important. Applying artificial intelligence (AI) methods in a clinical environment has become more and more popular during the last decade. Such AI methods are applied in a variety of clinical domains to enhance data visualisation, separate physiological from pathological data, predict risk factors for diseases and recurrences for cancer or quantify the severity of pathologies and disorders.
Voice production has been an area of interest in science since ancient times, and although advancing research has improved our understanding of the anatomy and function of the larynx, there is still little general consensus on these two topics. This review aims to outline the main developments in this field and highlight the areas where further research is needed. The most important hypotheses are presented and discussed highlighting the four main lines of research in the anatomy of the human larynx and their most important findings: (1) the arrangement of the muscle fibers of the thyroarytenoid muscle is not parallel to the vocal folds in the internal part (vocalis muscle), leading to altered properties during contraction; (2) the histological structure of the human vocal cords differs from other striated muscles; (3) there is a specialized type of heavy myosin chains in the larynx; and (4) the neuromuscular system of the larynx has specific structures that form the basis of an intrinsic laryngeal nervous system. These approaches are discussed in the context of current physiological models of vocal fold vibration, and new avenues of investigation are proposed.
Background: During the Covid-19 pandemic, singing activities were restricted due to several super-spreading events which have been observed during rehearsals and vocal performances. However, it has not been clarified how the aerosol dispersion, which has been assumed to be the leading transmission factor, could be reduced by masks which are specially designed for singers. Material and Methods: 12 professional singers (10 of the Bavarian Radio-Chorus and two freelancers, 7 females and 5 males) were asked to sing the melody of the ode of joy of Beethovens 9. symphony Freude schoener Goetterfunken, Tochter aus Elysium in D-major without masks and afterwards with five different singers masks, all distinctive in their material and proportions. Every task was conducted after inhaling the basic liquid from an e-cigarette. The aerosol dispersion was recorded by three high-definition video cameras during and after the task. The cloud was segmented and the dispersion was analyzed for all three spatial dimensions. Further, the subjects were asked to rate the practicability of wearing the tested masks during singing activities using a questionnaire. Results: Concerning the median distances of dispersion, all masks were able to decrease the impulse dispersion of the aerosols to the front. In contrast, the dispersion to the sides and to the top was increased. The evaluation revealed that most of the subjects would reject performing a concert with any of the masks. Conclusion: Although, the results exhibit that the tested masks could be able to reduce the radius of aerosol expulsion for virus-laden aerosol particles, there are more improvements necessary to enable the practical implementations for professional singing.
Vocal fold (VF) vibrations are the primary source of human phonation. High-speed video (HSV) endoscopy enables the computation of descriptive VF parameters for assessment of physiological properties of laryngeal dynamics, i.e., the vibration of the VFs. However, underlying biomechanical factors responsible for physiological and disordered VF vibrations cannot be accessed. In contrast, physically based numerical VF models reveal insights into the organ’s oscillations, which remain inaccessible through endoscopy. To estimate biomechanical properties, previous research has fitted subglottal pressure-driven mass–spring–damper systems, as inverse problem to the HSV-recorded VF trajectories, by global optimization of the numerical model. A neural network trained on the numerical model may be used as a substitute for computationally expensive optimization, yielding a fast evaluating surrogate of the biomechanical inverse problem. This paper proposes a convolutional recurrent neural network (CRNN)-based architecture trained on regression of a physiological-based biomechanical six-mass model (6 MM). To compare with previous research, the underlying biomechanical factor “subglottal pressure” prediction was tested against 288 HSV ex vivo porcine recordings. The contributions of this work are two-fold: first, the presented CRNN with the 6 MM handles multiple trajectories along the VFs, which allows for investigations on local changes in VF characteristics. Second, the network was trained to reproduce further important biomechanical model parameters like VF mass and stiffness on synthetic data. Unlike in a previous work, the network in this study is therefore an entire surrogate of the inverse problem, which allowed for explicit computation of the fitted model using our approach. The presented approach achieves a best-case mean absolute error (MAE) of 133 Pa (13.9%) in subglottal pressure prediction with 76.6% correlation on experimental data and a re-estimated fundamental frequency MAE of 15.9 Hz (9.9%). In-detail training analysis revealed subglottal pressure as the most learnable parameter. With the physiological-based model design and advances in fast parameter prediction, this work is a next step in biomechanical VF model fitting and the estimation of laryngeal kinematics.
Objective:The objective of this study is to evaluate three-dimensional vertical motion of the superior surface of the vocal folds in vivo in (a) typically developing children as a function of vocal frequency variations and (b) a child with vocal nodules. Methods:A custom developed laser endoscope coupled with high-speed videoendoscopy was used to obtain 3D parameters from 2 healthy children, one child with vocal nodules, and 23 vocally healthy adults (females = 11, males = 12). Parameters of amplitude (mm), maximum opening/closing velocity (mm/s), and mean opening/closing velocity (mm/s) were computed for the lateral and vertical vibratory motion along the anterior, middle, and posterior sections of the vocal folds were computed. Results:We provide for the first time, absolute measurements of vertical amplitude and maximum/ mean velocity during the opening and closing phases, in vivo in children. Overall, the vertical motion was larger in vocally normal children compared with the lateral motion, especially along the visible posterior section of the vocal folds and during low pitch phonation. The opening phase dynamics were consistently large along the posterior section in the child with vocal nodules. Conclusions:The study findings establish the feasibility of capturing 3D motion in a clinical setting and provide proof of concept for the application of the proposed 3D laser in the pediatric population. Future large sample size studies are needed to establish the diagnostic potential of examining the closing phase vertical motion to evaluate vibratory development in children with normal voice and investigating the opening phase vertical motion in children with nodules. Level of Evidence:N/A.
Objective To systematically evaluate the evidence for the reliability, sensitivity and specificity of existing measures of vowel-initial voice onset. Methods A literature search was conducted across electronic databases for published studies (MEDLINE, EMBASE, Scopus, Web of Science, CINAHL, PubMed Central, IEEE Xplore) and grey literature (ProQuest for unpublished dissertations) measuring vowel onset. Eligibility criteria included research of any study design type or context focused on measuring human voice onset on an initial vowel. Two independent reviewers were involved at each stage of title and abstract screening, data extraction and analysis. Data extracted included measures used, their reliability, sensitivity and specificity. Risk of bias and certainty of evidence was assessed using GRADE as the data of interest was extracted. Results The search retrieved 6,983 records. Titles and abstracts were screened against the inclusion criteria by two independent reviewers, with a third reviewer responsible for conflict resolution. Thirty-five papers were included in the review, which identified five categories of voice onset measurement: auditory perceptual, acoustic, aerodynamic, physiological and visual imaging. Reliability was explored in 14 papers with varied reliability ratings, while sensitivity was rarely assessed, and no assessment of specificity was conducted across any of the included records. Certainty of evidence ranged from very low to moderate with high variability in methodology and voice onset measures used. Conclusions A range of vowel-initial voice onset measurements have been applied throughout the literature, however, there is a lack of evidence regarding their sensitivity, specificity and reliability in the detection and discrimination of voice onset types. Heterogeneity in study populations and methods used preclude conclusions on the most valid measures. There is a clear need for standardisation of research methodology, and for future studies to examine the practicality of these measures in research and clinical settings.
Objectives There has been the assumption that whispering may impact vocal function, leading to the widespread recommendation against its practice after phonosurgery. However, the extent to which whispering affects vocal function and vocal fold oscillation patterns remains unclear.Methods 10 vocally healthy subjects (5 male, 5 female) were instructed to forcefully whisper a standardized text for 10 min at a sound level of 70 dB(A), measured at a microphone distance of 30 cm to the mouth. Prior to and following the whisper loading, the dysphonia severity index was assessed. Simultaneously, recordings of high speed videolaryngoscopy (HSV), electroglottography, and audio signals during sustained phonation on the vowel /i/ (250 Hz for females and 125 Hz for males) were analyzed after segmentation of the HSV material.Results The pre-post analysis revealed only minor changes after the intervention. These changes included a rise in minimum intensity, an increase in the glottal area waveform-derived open quotient, and the glottal gap index. However, no statistically significant changes were observed in the harmonic-to-noise-ratio, the glottal- to-noise-excitation-ratio, and the electroglottographic open quotient.Conclusion Overall, the study suggests that there are only small effects on vocal function in consequence of a forced whisper loading.
The primary acoustic signal of the voice is generated by the complex oscillation of the vocal folds (VFs), whereby physicians can barely examine the medial VF surface due to its anatomical inaccessibility. In this study, we investigated possibilities to infer medial surface dynamics by analyzing correlations in the oscillatory behavior of the superior and medial VF surfaces of four human hemilarynges, each in 24 different combinations of flow rate, VF adduction, and elongation. The two surfaces were recorded synchronously during sustained phonation using two high-speed camera setups and were subsequently 3D-reconstructed. The 3D surface parameters of mean and maximum velocities and displacements and general phonation parameters were calculated. The VF oscillations were also analyzed using empirical eigenfunctions (EEFs) and mucosal wave propagation, calculated from medial surface trajectories. Strong linear correlations were found between the 3D parameters of the superior and medial VF surfaces, ranging from 0.8 to 0.95. The linear regressions showed similar values for the maximum velocities at all hemilarynges (0.69–0.9), indicating the most promising parameter for predicting the medial surface. Since excessive VF velocities are suspected to cause phono-trauma and VF polyps, this parameter could provide added value to laryngeal diagnostics in the future.
OBJECTIVE:To examine changes in lateral and vertical vibratory motion along the anterior, middle, and posterior sections of the vocal folds, as a function of vocal frequency variations. METHODS:Absolute measurements of vocal fold surface dynamics from high-speed videoendoscopy with custom laser endoscope were made on 23 vocally healthy adults during sustained /i:/ production at 10%, 20%, and 80% of pitch range. The 3D parameters of amplitude (mm), maximum velocity opening/closing (mm/s), and mean velocity opening/closing (mm/s) were computed for the lateral and vertical vibratory motion along the anterior, middle, and posterior sections of the vocal folds. Linear mixed model analysis was conducted to evaluate the differences in (a) vocal frequency levels (high vs. normal vs. low pitch), (b) axis level (vertical vs. lateral), (c) position level (anterior vs. middle vs. posterior), and (d) gender differences (male vs. female). RESULTS:Overall, the superior surface vertical motion of the vocal fold is greater compared with the lateral motion, especially in males. Along the superior surface, the mean and maximum closing velocities are greater posteriorly for low pitch. The location (anterior, middle, and posterior) along the superior surface is relevant only for vocal fold closing rather than opening, as the dynamics are different along the various locations. CONCLUSIONS:The study highlights the significance of assessing the vertical motion of the superior surface of the vocal fold to understand the complex dynamics of voice production. LEVEL OF EVIDENCE:NA Laryngoscope, 134:3267-3276, 2024.
The onset of phonation can be used as a biomarker for quantitative voice assessment and identification of phonation type, e.g., breathy, modal, and creaky. The onset of phonation can be calculated considering different modalities: acoustic, electroglottography (EGG), airflow, and glottal area waveforms. Considering the potential of voice onset in clinical applications, we have been developing the Voice Onset Analysis Tool (VOAT), a computer-aided program written in Python for the semi-automatic detection of vowel phonation onset. VOAT can detect the vowel onset on recordings with multiple phonations. Furthermore, the user can manually select the segments from the recording where the onset is present or apply a voice activity detection for automatic segmentation. VOAT provides filtering options to facilitate the detection and allows the saving of results in an Excel file for further processing. Our software also allows batch processing to detect vowel onset on several files automatically.
Human vocalization is a complex process that is still only partially understood. Previous studies have suggested the possibility of a localized neuromuscular network of the larynx. Here we investigate this structure in human dissection specimens using multiple immunofluorescence and transmission electron microscopy (TEM). In the area of the pars interna of the thyroarytenoid muscle, muscle fibers are present that are clearly differentiated from skeletal or cardiac muscle cells and show an intermediate ultrastructure. In addition, intramuscular neurons are present that are detectable by both electron and fluorescence microscopy and may have a sensory function in a local neuronal network. Also, several types of sensory and motor synapses are detectable and distributed throughout the pars interna of the thyroarytenoid muscle, with multisynaptic muscle fibers being a common feature. These findings suggest the existence of a previously unrecognized type of muscle fiber coupled to an intramuscular neuronal network, the presence of which could explain functional peculiarities at the laryngeal level.
AbstractVoice production of humans and most mammals is governed by the MyoElastic-AeroDynamic (MEAD) principle, where an air stream is modulated by self-sustained vocal fold oscillation to generate audible air pressure fluctuations. An alternative mechanism is found in ultrasonic vocalizations of rodents, which are established by an aeroacoustic (AA) phenomenon without vibration of laryngeal tissue. Previously, some authors argued that high-pitched human vocalization is also produced by the AA principle. Here, we investigate the so-called “whistle register” voice production in nine professional female operatic sopranos singing a scale from C6 (≈ 1047 Hz) to G6 (≈ 1568 Hz). Super-high-speed videolaryngoscopy revealed vocal fold collision in all participants, with closed quotients from 30 to 73%. Computational modeling showed that the biomechanical requirements to produce such high-pitched voice would be an increased contraction of the cricothyroid muscle, vocal fold strain of about 50%, and high subglottal pressure. Our data suggest that high-pitched operatic soprano singing uses the MEAD mechanism. Consequently, the commonly used term “whistle register” does not reflect the physical principle of a whistle with regard to voice generation in high pitched classical singing.
Purpose: Individual prediction of treatment response is crucial for personalized treatment in multimodal approaches against head-and-neck squamous cell carcinoma (HNSCC). So far, no reliable predictive parameters for treatment schemes containing immunotherapy have been identified. This study aims to predict treatment response to induction chemo-immunotherapy based on the peripheral blood immune status in patients with locally advanced HNSCC. Methods: The peripheral blood immune phenotype was assessed in whole blood samples in patients treated in the phase II CheckRad-CD8 trial as part of the pre-planned translational research program. Blood samples were analyzed by multicolor flow cytometry before (T1) and after (T2) induction chemo-immunotherapy with cisplatin/docetaxel/durvalumab/tremelimumab. Machine Learning techniques were used to predict pathological complete response (pCR) after induction therapy. Results: The tested classifier methods (LDA, SVM, LR, RF, DT, and XGBoost) allowed a distinct prediction of pCR. Highest accuracy was achieved with a low number of features represented as principal components. Immune parameters obtained from the absolute difference (lT2-T1l) allowed the best prediction of pCR. In general, less than 30 parameters and at most 10 principal components were needed for highly accurate predictions. Across several datasets, cells of the innate immune system such as polymorphonuclear cells, monocytes, and plasmacytoid dendritic cells are most prominent. Conclusions: Our analyses imply that alterations of the innate immune cell distribution in the peripheral blood following induction chemo-immuno-therapy is highly predictive for pCR in HNSCC.
Electrophysiological studies of the larynx expose the mechanisms by which voice production is controlled. Previous studies have revealed certain phenomena during laryngeal oscillations that suggest a complex control mechanism. Starting from the principle of agonist-antagonist muscular pairing, the aim of this study was to gain a deeper insight into the function of the cricothyroid (CT) and thyroarytenoid (TA) muscles, both central to voice production. Electromyographic recordings were used to determine the response of the two muscles to different stimulation situations in an ex vivo animal model of the denervated larynx of pigs (n=26). Using a set of different experiments, it was shown that when one muscle (CT or TA muscle) was electrically stimulated, a response was observed in the other muscle, which in the otherwise-denervated larynx, was caused only by the applied stimulation and exhibited the characteristics of compound action potentials. This response was reproducible in all larynxes examined and was present bidirectionally. No response was registered in the absence of stimulation. The results show the existence of coactivation of the CT and TA muscles in the absence of external innervation hinting at the presence of a localized neuronal network of the larynx that has not been described previously. Further morphological investigation is needed to determine the presence of this internal laryngeal neuronal network.
Introduction. Group singing has been associated with higher transmission risks via exhaled and spread aerosols in the CoVID19 pandemic. For this reason, many musical activities, such as rehearsals and lessons, but also voice therapy sessions, have been restricted in many countries. Consequently, transmission risks and pathways have been studied, such as aerosol amounts generated by exhalation tasks, convectional flows in rooms, or the impulse dispersion of different kinds of phonation. The use of water resistance exercises such as those utilizing LAX VOX (R), are common in voice lessons and as vocal warm-ups. With this context, this study investigates the impulse dispersion characteristics of aerosols during a voiced water resistance exercise in comparison to normal singing. Methods. Twelve professional singers (six male, six female) were asked to phonate a stable pitch through a silicone tube into a bottle filled with water, holding the end of the tube 5 cm below the surface. Before performing the tasks, the singers inhaled the vapor consisting of 0.5 L base liquid from an e-cigarette. The exhaled gas cloud coming out of the bottle was recorded in all three spatial directions and the dispersion was measured as a function of time. Results. At the end of the phonation task, the median distance to the front was 0.55 m and the median of the lateral expansion of the cloud was 0.89 m, the maximum to the front reached 0.88 m, and the maximum of lateral expansion 1.05 m. For the upwards direction of the clouds a median of 1.00 m and a maximum of 1.34 m from the mouth were measured. Three seconds after the end of the task, the medians were declining. Conclusion. The exhaled aerosol cloud can expand despite the obstacle of the water when using LAX VOX (R) during phonation.
Introduction. Due to increased aerosol generation during singing, choir rehearsals were widely prohibited in the course of the CoVID-19 pandemic. Most studies on aerosol generation and dispersion focus on professional singers. However, it has not been clari fied if these data are also representative for amateur singers. Methods. Nine non-professional singers (four male, five female) were asked to perform five tasks; speaking (T+), singing a text softly (MT-) and loudly (MT+), singing on the vowel [(sic)] (M+) and singing with a N95 mask (MT+N95). Before performing the tasks, the singers were asked to inhale 0.5 L vapor produced by an e-cigarette consisting of the basic liquid. The spread of the exhaled vapor was recorded in all three dimensions by high-de fini-tion cameras and the impulse dispersion was detected as a function of time. Results. Regarding the median dispersion to the front, all tasks showed comparable distances from 0.69 m to 0.82 m at the end of the tasks. However, the maximum aerosol dispersion showed a larger variety among different subjects or tasks, respectively. Especially in the M+ task a maximum distance of 1.96 m to the front was reached by a single subject. Although singing with a N95 mask resulted in a slightly increased median dispersion to the front, the maximum dispersion was decreased from 1.47 m (MT+) to 1.04 m (MT+N95). Conclusion. The maximum dispersion distance to the front of 1.96 m at the end of the M+ task and 1.47 m at the end of the MT+ task showed higher values in comparison to professional singers. Differences in phonation, articulation and mouth opening could lead to greater impulse dispersion. Singing in loud phonation with a N95 mask reduced the maximum impulse dispersion to the front to 1.04 m. Taking all results into consideration, a slightly larger safety distance should be necessary for non-professional singers.
Tino Haderlein合作论文数Department of Computer Science, Friedrich-Alexander-Universität13