Purpose: This study investigated the intelligibility and processing characteristics of Mandarin sentences synthesized based on a large-scale language model (LLM), compared to natural speech, specifically for deaf adults with cochlear implants (CIs). Method: Fifty participants, 25 Mandarin-speaking deaf adults with CIs and 25 Mandarin-speaking adults with normal hearing (NH), were enrolled in the study. Sentence recognition rates and reaction times were measured using sentences from the Mandarin speech perception database presented as six different speech types. These included original natural speech by a single female speaker at slow and normal speaking rates as well as synthetic speech generated by the LLM-based text-to-speech (TTS) model in two different voices (male and female) at two speaking rates (slow and normal). Linear mixed-effect models analyzed how sentence recognition accuracy and reaction time varied with listener group and speech type. Results: Participants with CIs demonstrated significantly lower sentence recognition rates and longer reaction times than NH participants for both natural and synthetic speech. Participants with CIs exhibited greater individual variability than NH listeners in terms of sentence recognition performance and reaction times. Changes in speaking rate significantly impacted the recognition rates of participants with CIs. Crucially, when controlling for speaking rate differences, there was no significant difference in sentence recognition rates or reaction times between natural and LLM-based synthetic speech for participants with CIs. Conclusions: The results suggest that LLM-based synthetic speech could achieve intelligibility levels indistinguishable from natural speech for CI recipients under comparable speaking rate conditions, similar to its performance for NH listeners. This study offers valuable insights for the future application of advanced TTS models in speech perception assessment and auditory rehabilitation for patients with CIs, including the potential development of personalized training materials. Supplemental Material: https://doi.org/10.23641/asha.32736135
Characterizing the vocal tract spectral envelope in high-pitched vowels is challenging due to sparse harmonics and source-filter coupling. This study proposes an analytical framework using a shared glottal driving proxy to guide source-tract decoupling. Speech is decomposed into multi-band Hilbert envelopes; first-order differentiation and non-linear rectification are then applied to isolate transient energy increments during glottal closure. Non-negative matrix factorization captures the cross-band temporal dynamics of these increments, yielding a macroscopic temporal proxy of glottal excitation. This proxy acts as a physical constraint within an auto-regressive with exogenous input model to estimate the underlying vocal tract transfer function. To validate the estimated envelopes, formant estimation error serves as an objective metric on the OPENGLOT database. Particularly under high-pitched conditions where the fundamental frequency (F0) exceeds 300 Hz, the overall mean first formant (F1) estimation error across all evaluated vowels and parameter combinations is 11.43%, yielding an approximate 30% relative reduction from the traditional linear predictive coding baseline (16.51%). The results confirm that this approach provides a complementary pathway for vocal tract envelope modeling.
The hypopharynx is crucial in defining the unique acoustic properties contributing to speaker identity. However, the acoustic influence of gender-specific morphological variations in the hypopharynx remains insufficiently explored. Previous studies have largely focused on the acoustic effects of the hypopharynx in male subjects but few for females. This study investigated gender-related acoustic differences in the hypopharynx using 3D vocal tract models reconstructed from the MRI datasets of three male and four female Chinese subjects. Finite-difference time-domain simulations were conducted to examine the acoustic contributions of the laryngeal cavity and bilateral piriform fossae through cavity occlusion experiments. By comparing vocal tract transfer functions and sound pressure distributions before and after occlusion, the study observed notable differences between the male and female subjects. The female subjects exhibited higher frequency of laryngeal cavity resonance and piriform fossae anti-resonance than the male subjects. Hypopharyngeal cavity occlusion in the female subjects showed broader frequency variability and greater amplitude fluctuations. The acoustic interaction between the laryngeal cavity and piriform fossae was found to be stronger in the female subjects. These differences may be attributable to the comparatively shorter vocal tract length and smaller hypopharyngeal dimensions observed in the female subjects.
Mandarin seven-sound test was introduced to bridge a frequency gap from 6000 Hz to 8000 Hz identified in the standard Mandarin version of the Ling six-sound test. However, until the present research, there are few relevant studies to confirm the feasibility of this test or to investigate its potential for use across various pronunciation communities. In this study, eight speakers were invited to record the audio material for Mandarin seven-sound test. Subsequently, 31 adult participants took the Mandarin seven-sound test using these recordings. The results showed that individuals with hearing loss performed significantly poorer than those with normal hearing in terms of both reaction time and accuracy rate. Specifically, within the hearing-loss group, the accuracy rates for the /s/ and /(sic)/ sounds were notably lower compared to other phonemes. Additionally, this study found that the demographic characteristics of the speakers, such as gender and age, did not significantly influence the accuracy rate for either group. These results indicate that Mandarin seven-sound test not only effectively addresses the frequency gap of the Ling six-sound test but also maintains a high level of consistency across different speakers. Moreover, the online testing system has proven to be strong user-friendliness and provides significant value for clinical speech audiometry, speech therapy, and the development of hearing aid technology.
Acoustic characteristics of speech exhibit variability across individuals, while preserving shared phonetic information to listeners. In this paper, the general time-frequency pattern of individual speaker characteristics is discussed based on our previous research. The main target here is set at speaker-specific acoustic effects of the vocal tract in both higher and lower frequency ranges. To address the under-explored phenomena, two experiments were conducted. Firstly, simulations based on the transmission line model are used to explore how resonances in higher frequencies vary with different hypopharyngeal-cavity shapes. Secondly, speech signals emitted from the mouth and nostrils are recorded separately to observe potential factors for spectral irregularity in lower frequencies. From our findings, a time-frequency model of individual speaker characteristics is proposed that provides insights into how individuality is manifested in speech spectral patterns.
The study prepared Al2O3–MgO based castables bonded by hydratable alumina (HA) instead of calcium aluminate cement (CAC) for the working lining of Si-killed stainless steel ladles. The microstructure, phase composition, mechanical properties, and slag resistance of castables were investigated by SEM, XRD, and thermodynamic software FactSage®. The results indicated that the HA bonded castables showed superior hot flexural strength, thermal shock resistance and slag resistance than the CAC bonded castables, due to the optimized pore characteristics, less liquid content, and higher liquid viscosity of the castable matrix and the formation of a continuous insulating layer.
The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding must be manually selected during inference. Due to the fact that only the matched text and speech are used in the training process, using unmatched text and speech for inference would cause the model to synthesize speech with low content quality. In this study, we propose to mitigate these two problems by using multiple reference audios and style embedding constraints rather than using only the target audio. Multiple reference audios are automatically selected using the sentence similarity determined by Bidirectional Encoder Representations from Transformers (BERT). In addition, we use ''target'' style embedding from a Pre-trained encoder as a constraint by considering the mutual information between the predicted and ''target'' style embedding. The experimental results show that the proposed model can improve the speech naturalness and content quality with multiple reference audios and can also outperform the baseline model in ABX preference tests of style similarity.
相比铝酸盐水泥(CAC),水合氧化铝(HA)具有更好的高温性能,可用作铝熔铸用内衬材料的结合剂.笔者综述了HA在铝熔铸内衬材料养护过程和热处理过程中水化产物的组成和结构变化规律,讨论了热处理过程中HA水化产物演变与内衬材料宏观性能的联系,可为高性能浇注料的发展提供有益参考.
Vocal-tract area function is a one-dimensional representation of the vocal tract, in which speech signals are interpreted according to their place vs. area patterns. Recent work on deriving vocal-tract area functions from volumetric vocal-tract data is successful for the main tract part, whereas the region near the tract ends lacks accuracy due to the use of a planar grid system on the wedge-shaped tract opening. This study employs a spe-cial treatment on the anterior tract part using curved grid planes with a gradual evolution of convexity, which is applied to cross-sectioning the anterior tract regions including the post-incisor cavity, inter-dental channel, and lip tube. With the method, volumetric MRI data for vowels /a/ and /i/ were processed to describe the articulatory configuration in those regions. The results revealed that the anterior tract regions are observed as identifiable tract segments with a natural-shaped final opening. Thus, our proposed area function scheme promises nearly complete descriptions of articulatory configuration together with a smooth interface for sound radiation with minor modifications.
以电熔镁砂、电熔尖晶石和金属铝为原料,酚醛树脂为结合剂,制成MgO-MA和含有5%(w)金属Al的MgO-MA-Al试样.在110℃干燥24 h后,一部分试样在密闭匣钵中于1600℃保温5 h热处理,分析其物相组成和显微结构;另一部分试样于氮气气氛1600℃保温5 h进行抗钢渣试验.结果表明:1)在密闭匣钵于1600℃保温5 h热处理过程中,含有5%(w)金属Al的MgO-MA-Al试样因强还原性单质Al的存在促进了MgO向Mg(g)转化,从而促进MgO致密层的形成.2)在氮气气氛1600℃保温5 h进行抗渣试验的过程中,同样会发生试样外层致密化的过程.这使试样的抗渣渗透性和抗渣侵蚀性得到提高.
以板状刚玉、α-Al2 O3微粉和铝酸钙水泥为原料制备刚玉质浇注料试样.将试样于25℃分别在不同的湿度养护3 d,再在110℃烘干24 h后分别于800、1100、1450和1620℃保温3 h进行热处理,研究了养护湿度(40%、70%和100%)对其性能、物相组成、显微结构的影响.结果表明:随着养护湿度的降低,浇注料的脱模强度、烘干强度及中高温热处理后的常温抗折强度增大,而显气孔率及热震后的抗折强度保持率逐渐降低.另外,随着养护湿度降低,铝酸钙水泥的水化程度更高,1620℃热处理后浇注料的骨料与基质结合更加紧密.
State-of-the-art neural text-to-speech (TTS) networks are trained with a large amount of speech data, which significantly improves the quality of synthetic speech compared with traditional approaches. However, the prosody and controllability of the generated speech is still insufficient, especially in tonal languages. Moreover, the generated prosody is solely defined by the input text, which does not allow for different styles for the same sentence or words. In this study, we extended Tacotron2 with a pitch prediction task to capture discrete pitch-related representations. Specifically, the learned pitch-related suprasegmental information is fed simultaneously with traditional character features into the decoder to generate final Mel spectrogram. Experiments show that the proposed method can improve the quality of the generated speech (mean opinion score of 4.37 vs. 4.22). Moreover, we demonstrated that we can easily achieve word-level pitch control during generation by changing local pitch-related representations before passing them to the decoder network.
Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent context. However, 2D lip easily suffers from various face orientations. To this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent contexts. A new regional feedback module (RFM) is proposed to obtain attentions in different lip regions. Besides, prior knowledge of lip motion is investigated to complement RFM, where landmark-level and frame-level features are merged to form a better feature representation. Moreover, we present two methods, i.e., coordinate transformation and face posture correction to pre-process the LSD-AV dataset, which contains 68 speakers and 146 sentences per speaker. The evaluation results on this dataset demonstrate that our proposed 3LMNet is superior to the baseline models, i.e., LSTM, VGG-16 and ResNet-34, and outperforms the state-of-the-art using 2D lip image as well as the 3D face. The code of this work is released at https://github.com/wutong18/Three-Dimensional-Lip- Motion-Network-for-Text-Independent-Speaker-Recognition.
End-to-end speech synthesis demonstrates remarkable performance in monolingual speech, whereas code-switching (CS) speech synthesis remains a challenge owing to the sparsity of data and diverse syntactic structures across languages. Previous studies show that large mixed-lingual corpora are essential for effective learning text/language representations and target speaker information. In this study, we propose a method using three independent encoders (text, language, and speaker), which requires only a small amount of mixed-lingual data to realize the CS speech synthesis of Mandarin and English. Additionally, to distinguish between Mandarin and English, we investigate two text-representation methods: (1) the implicit method, which uses Pinyin and the CMU 1 1 http://www.speech.cs.cmu.edu/cgi-bin/cmudict dictionary to represent both languages; and (2) the explicit method, which uses language markers i.e., masks, to differentiate the languages. Through our proposed method, we can improve synthesized speech in terms of quality and speaker similarity using a small amount of mixed-lingual data. In addition, the experimental results demonstrate that the proposed method achieves performance improvement of 0.06 in terms of the mean opinion score and absolute improvement of 0.64% in terms of the character error rate compared to the baseline method.
The quality of multispeaker text-to-speech (TTS) is composed of speech naturalness and speaker similarity. The current multispeaker TTS based on speaker embeddings extracted by speaker verification (SV) or speaker recognition (SR) models has made significant progress in speaker similarity of synthesized speech. SV/SR tasks build the speaker space based on the differences between speakers in the training set and thus extract speaker embeddings that can improve speaker similarity; however, they deteriorate the naturalness of synthetic speech since such embeddings lost speech dynamics to some extent. Unlike SV/SR-based systems, the automatic speech recognition (ASR) encoder outputs contain relatively complete speech information, such as speaker information, timbre, and prosody. Therefore, we propose an ASR-based synthesis framework to extract speech embeddings using an ASR encoder to improve multispeaker TTS quality, especially for speech naturalness. To enable the ASR system to learn the speaker characteristics better, we explicitly feed the speaker-id to the training label. The experimental results show that the speech embeddings extracted by the proposed method have good speaker characteristics and beneficial acoustic information for speech naturalness. The proposed method significantly improves the naturalness and similarity of multispeaker TTS.
There have been reports that strength of hydratable alumina (HA)-bonded castables without silica fume drops significantly at 600 degrees C and decreases substantially again at 1000 degrees C. But the strength variation of the HA-bonded castables during the intermediate temperature range has not been investigated and elaborated from the perspective of phase evolution and microstructural change in the castables. In this work, the relationship between the change in the strength of castables and the microstructural characteristics of the HA-bonded castables was investigated. The phase and microstructure evolution of HA-bonded castables between 110 degrees C and 1250 degrees C were investigated by X-ray diffraction (XRD), scanning electron microscopy (SEM), and thermogravimetric analysis (TG). It has been found that strength drops of the HA-bonded castables during heating process do not mainly happen at the temperature at which HA hydrates decompose, but at the temperature at which the structure of dehydrated HA hydrates disintegrates.
The combination of the recently proposed LPCNet vocoder and a seq-to-seq acoustic model, i.e., Tacotron, has successfully achieved lightweight speech synthesis systems. However, the quality of synthesized speech is often unstable because the precision of the pitch parameters predicted by acoustic models is insufficient, especially for some tonal languages like Chinese and Japanese. In this paper, we propose an end-to-end speech synthesis system, TacoLPCNet, by conditioning LPCNet on Mel spectrogram predictions. First, we extend LPCNet for the Mel spectrogram instead of using explicit pitch information and pitch-related network. Furthermore, we optimize the system by model pruning, multi-frame inference, and increasing frame length, to enable it to meet the conditions required for real-time applications. The objective and subjective evaluation results for various languages show that the proposed system is more stable for tonal languages within the proposed optimization strategies. The experimental results also verify that our model improves synthesis runtime by 3.12 times than that of the baseline on a standard CPU while maintaining naturalness.
Purpose The primary purpose of this study was to explore the audiovisual speech perception strategies.80.23.47 adopted by normal-hearing and deaf people in processing familiar and unfamiliar languages. Our primary hypothesis was that they would adopt different perception strategies due to different sensory experiences at an early age, limitations of the physical device, and the developmental gap of language, and others. Method Thirty normal-hearing adults and 33 prelingually deaf adults participated in the study. They were asked to perform judgment and listening tasks while watching videos of a Uygur–Mandarin bilingual speaker in a familiar language (Standard Chinese) or an unfamiliar language (Modern Uygur) while their eye movements were recorded by eye-tracking technology. Results Task had a slight influence on the distribution of selective attention, whereas subject and language had significant influences. To be specific, the normal-hearing and the d10eaf participants mainly gazed at the speaker's eyes and mouth, respectively, in the experiment; moreover, while the normal-hearing participants had to stare longer at the speaker's mouth when they confronted with the unfamiliar language Modern Uygur, the deaf participant did not change their attention allocation pattern when perceiving the two languages. Conclusions Normal-hearing and deaf adults adopt different audiovisual speech perception strategies: Normal-hearing adults mainly look at the eyes, and deaf adults mainly look at the mouth. Additionally, language and task can also modulate the speech perception strategy.
End-to-end text-to-speech (TTS) can synthesize monolingual speech with high naturalness and intelligibility. Recently, the end-to-end model has also been used in code-switching (CS) TTS and performs well on naturalness, intelligibility and speaker consistency. However, existing systems rely on skillful bilingual speakers to build a CS mix-lingual data set with a high Language-Mix-Ratio (LMR), while simply mixing monolingual data sets results in accent problems. To reduce the cost of recording and maintain the speaker consistency, in this paper, we investigate an effective method to use a low LMR imbalanced mix-lingual data set. Experiments show that it is possible to construct a CS TTS system with a low LMR imbalanced mix-lingual data set with diverse input text presentations, meanwhile produce acceptable synthetic CS speech with more than 4.0 Mean Opinion Score (MOS). We also find that the result will be improved if the mix-lingual data set is augmented with monolingual English data.
To advance the study of lip-reading recognition in accordance with Chinese pronunciation norms, we carefully investigated Mandarin tone recognition based on visual information, in contrast to that of the previous character-based Chinese lip reading technique. In this paper, we mainly studied the vowel tonal transformation in Chinese pronunciation and designed a lightweight skipping convolution network framework (SCNet). And, the experimental results showed that the SCNet was sensitive to the more detailed description of the pitch change than that of the traditional model and achieved a better tone recognition effect and outstanding antiinterference performance. In addition, we conducted a more detailed study on the assistance of the deep texture information in lip-reading recognition. We found that the deep texture information has a significant effect on tone recognition, and the possibility of multimodal lip reading in Chinese tone recognition was confirmed. Similarly, we verified the role of the SCNet syllable tone recognition and found that the vowel and syllable tone recognition accuracy of our model was as high as 97.3%, which also showed the robustness of our proposed method for Chinese tone recognition and it can be widely used for tone recognition.