A systematic approach is proposed to synthesizing personalized spontaneous speech using a small-sized unsegmented speech corpus of the target speaker. First, an automatic segmentation algorithm is employed to segment and label a collected semispontaneous speech corpus of the target speaker. Then, a pretrained average voice model is adapted to the voice model of the target speaker by using the segmented data. A postfilter based on modulation spectrum is adopted to further improve the speaker similarity of the synthesized speech as well as alleviate the over-smoothing problem of the synthesized speech. For generating spontaneous speech, a smoothing method applied at the prosodic word level is proposed to improve speech fluency. For objective evaluation on spontaneous speech segmentation, the segmentation accuracy of the proposed method is superior to that of Viterbi-based forced alignment. The results of subjective listening test also show that the proposed method can improve the spontaneity and speaker similarity of the synthesized speech compared to the maximum likelihood linear regression based speaker adaptation method.
In control vector-based expressive speech synthesis, the emotion/style control vector defined in the categorical (CAT) emotion space is uneasy to be precisely defined by the user to synthesize the speech with the desired emotion/style. This paper applies the arousal-valence (AV) space to the multiple regression hidden semi-Markov model (MRHSMM)-based synthesis framework for expressive speech synthesis. In this study, the user can designate a specific emotion by defining the AV values in the AV space. The multidimensional scaling (MDS) method is adopted to project the AV emotion space and the categorical (CAT) emotion space onto their corresponding orthogonal coordinate systems. A transformation approach is thus proposed to transform the AV values to the emotion control vector in CAT emotion space for MRHSMM-based expressive speech synthesis. In the synthesis phase given the input text and desired emotion, with the transformed emotion control vector, the speech with the desired emotion is generated from the MRHSMMs. Experimental result shows the proposed method is helpful for the user to easily and precisely determine the desired emotion for expressive speech synthesis.
This study proposes a hybrid approach to natural-sounding speech synthesis based on candidate expansion, unit selection, and prosody adjustment using a small corpus. The proposed method is more specific to tonal language, in particular Mandarin. In conventional speech synthesis studies, the quality of synthesized speech depends heavily on the size of the speech corpus. However, it is highly time-consuming and labor-intensive to prepare a large labeled corpus. In this work, candidate expansion is proposed to retrieve potential candidates that are unlikely to be retrieved using only linguistic features. The optimal unit sequence is then obtained from the expanded candidates by using the proposed unit selection mechanism at the phoneme and prosodic word levels. Finally, a prosodic word-level prosody adjustment is proposed to improve the continuity and smoothness of the prosody of the synthesized speech. To evaluate the proposed method, the Tsing-Hua corpus of speech synthesis was adopted. The results of an objective evaluation demonstrate the effectiveness of candidate expansion and the improvement of the continuity and smoothness of the prosody of the synthesized speech. The results of a subjective evaluation also show the proposed system could synthesize the speech with improved quality and naturalness, in particular for a small-sized or resource-limited corpus.
In conventional HMM-based speech synthesis, the algorithm for generating a high-quality reading style (neutral) speech has been well investigated. However, the human-like expressive speech synthesis is still rather far from practicability, which is caused by many factors. One of the influential factors is that the speech variability caused by speaker's arousal is rarely emphasized in speech synthesis. Accordingly, this paper proposed a novel speech synthesis method considering the speech variability. Two major advantages are highlighted by considering the speech variability. The first advantage is that the proposed method is capable of generating the time-variant human-like and expressive speech. The second one is to increase the diversity of expressive speech and to improve the drawback of traditional speech synthesis system with the monotonous characteristics of speech. The experimental result shows that the proposed method can improve the diversity capability of synthetic speech and successfully achieve the more expressive speech compare to conventional HTS one.
In the home care scenarios, the feedback voices are often presented in reading style (neutral); however, users prefer known voices, such as those of family members, friends or celebrities. For this purpose, the speaker adaptation or voice conversion is usually adopted. However, these techniques demand a well-tagged speaker-dependent or parallel corpus, and building such a corpus is a time-consuming task. Moreover, users don't have the ability to deal with such a database by themselves. Therefore, this study presents a voice-customizable text-to-speech system which allows the system voice can be changed easily by users. In this system, parallel corpus generation, voice conversion based on decision tree and linear multivariate regression (LMR), and HMM-based speech synthesis are integrated. The advantage of this system is that the target speaker needs to record only a small amount of speech data according to a pre-designed balanced text corpus. Thus, the voice conversion models can be automatically estimated and stored. In subjective and objective evaluations, the operation convenience or friendliness of the proposed system is better than the baseline system.
This paper presents a human-robot interactive system for senior companion based on cloud computing infrastructure. The proposed senior companion robot system (SCRS) is designed based on cloud computing network. In the server side, two cloud services are proposed, 1) the web-based user remote management service (WURMS) for remote robot control; 2) the robotic multimodal interactive computation services (RMICS) for providing the human-robot operation interfaces including the speech/sound recognition, speaker identification, face identification, sound source estimation and text to speech (TTS). In the robot client side, the behavior model is designed to use WURMS and RMICS services. In the experiments, two robots called “Robert” and “Davinci” are designed to evaluate the SCRS's capability. With using only low-cost and low-power CPUs (Intel Atom N450), both of the two robots can still work wirelessly for real-time human-robot interaction. Finally, we design five senior companion scenarios, and the experimental average MOS (Mean Opinion Score) is 4.16.
As the growth of economy and technology has become increasingly rapid, mental care is getting more important today. However, recent movements, such green technologies, place more emphasis on environmental issues but less on mental care. Therefore, this paper presents a newly emerging technology called orange computing for mental care applications. Orange computing refers to health, happiness, and physiopsychological care computing, which focuses on designing algorithms and systems for enhancing body and mind balance. The representative color for orange computing comes from a harmonic fusion of passion, love, and happiness. In order to demonstrate the main idea of orange computing, in this work we also discuss affective signal processing, one of the important topics in orange computing. Besides, a case study on the emotion speech recognition system for emotion care services is presented at the end of the paper. Experimental results show that the proposed system can enhance the accuracy rate by 4.44% on average. In contrast with the baselines, the proposed system can identify emotion speeches more precisely, subsequently demonstrating the feasibility and effectiveness of the system.
In recent years, emotion-aware human-machine interactions have become an important issue. Most of the traditional researches focused on the use of different features and classification methods to improve emotion recognition rates. However, they still cannot recognize detailed and various emotions. Accordingly, in this paper, an emotion recognition system, which combines the acoustic and textual features from speech, is proposed to detect seven emotional states: Joy, sadness, anger, fear, surprise, worry and disgust, respectively. The AdaBoost approach is also used to learn and classify each emotional state. The experimental result shows that the emotion recognition accuracy of the proposed system is better than that of traditional approaches.
This paper brings together speech recognition, emotion inference and virtual agents to implement a system for student interaction in an educational environment. By analyzing the capture speech, we can perceive an indication of the emotion status of the target student. Using the inference results, an agent can choose the suitable dialogue to interact with the student. Our experiments indicate that there is a lot of potential for such a system to be applied in other situations as well.