谷歌浏览器插件
订阅小程序
在清言上使用

Text Enhancement for Paragraph Processing in End-to-End Code-switching TTS

2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)(2021)

引用 0|浏览13
暂无评分
摘要
Current end-to-end code-switching Text-to-Speech (US) can already generate high quality two languages speech in the same utterance with single speaker bilingual corpora. When the speakers of the bilingual corpora are different, the naturalness and consistency of the code-switching TTS will be poor. The cross-lingual embedding layers structure we proposed makes similar syllables in different languages relevant, thus improving the naturalness and consistency of generated speech. In the end-to-end code-switching TTS, there exists problem of prosody instability when synthesizing paragraph text. The text enhancement method we proposed makes the input contain prosodic information and sentencelevel context information, thus improving the prosody stability of paragraph text. Experimental results demonstrate the effectiveness of the proposed methods in the naturalness, consistency, and prosody stability. In addition to Mandarin dand English, we also apply these methods to Shanghaiese and Cantonese corpora, proving that the methods we proposed can be extended to other languages to build end-to-end codeswitching US system.
更多
查看译文
关键词
code-switching,end-to-end speech synthesis,text enhancement,prosodic boundary,cross-lingual
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要