Diffusion models can effectively generate intricate images by gradually refining noise into detailed visual data. Diffusion inversion maps a real-world image to a feature space to achieve image attribute manipulation. Existing diffusion inversion techniques cannot effectively integrate semantic and visual information, resulting in insufficient control precision. This paper introduces the Cross-Modal Contrastive Inversion (CMCI) method to enhance single-image attribute manipulation for pre-trained diffusion models. CMCI learns cross-modal conditioning embedding through contrastive regularization, combining visual and textual information to guide denoising. Our method helps diffusion models generate images semantically matched with textual descriptions while preserving the integrity of the source image. CMCI can be applied to single-video attribute manipulation tasks by capturing both static features and dynamic temporal changes through our contrastive regularization. Comprehensive experiments are performed, and the results suggest that CMCI outperforms competing inversion methods in terms of both inversion accuracy and manipulation capabilities.
更多
查看译文
关键词
Modeling,Videos,Learning (artificial intelligence),Noise reduction,Text to image,Educational institutions,Visualization,Diffusion models,Training,Computers,diffusion inversion,attribute manipulation