2025 13th International Conference on Orange Technology (ICOT)(2025)
Department of Electrical Engineering
被引用0|浏览0
摘要
We propose a modular pipeline that converts Chinese text into Taiwanese Hokkien speech and synthesizes a photorealistic talking-face video. The system couples a Taiwanese-aware TTS front end (hybrid word segmentation, iTaigi Tâi-lô romanization with dynamic tone annotation, tone-aware Tacotron2) with a Seed-VC–style timbre converter and a two-stage speech-driven talking-face synthesizer using cross-modal fusion, temporal modeling, and diffusion rendering. On an internal set, the speech reaches MOS 4.602 and CER 2.01%. On talking-face benchmarks, we obtain FID 29.18, CPBD 0.498, and LSE-D 7.23, surpassing representative baselines in realism, sharpness, and lip-sync. The design is lightweight and modular, runs on a single RTX 3080, and is transferable to other low-resource tonal languages.
更多
查看译文
关键词
Text-to-Speech,Voice Conversion,Talking Face Generation