We present Diff-KATKG, a novel diffusion-based framework for high-fidelity talking head generation, jointly driven by facial keypoints and action units (AUs). To enable fine-grained motion control under sparse driving conditions, we design a cross-attention-based fusion module that fuses keypoint and AU features into a unified embedding, which serves as the conditioning input to the noise prediction network of the diffusion model. This joint representation effectively captures both pose and expression dynamics, enabling expressive and controllable video synthesis. To further enhance temporal coherence, we introduce a cross-frame feature aggregation strategy that leverages spatiotemporal dependencies from previously generated frames to guide the denoising process. This facilitates smoother transitions and more natural motion across frames. Benefiting from the progressive denoising mechanism of the diffusion model, our approach achieves detailed and stable frame reconstruction, significantly improving perceptual realism and temporal consistency.
更多
查看译文
Chat Paper
正在生成论文摘要
关键词
Talking head generation,Joint keypoint and AUs guidance,Diffusion model,Cross-attention-based fusion,Temporal consistency