Image appearance transfer plays a significant role in interior design by allowing designers to efficiently explore different design styles according to clients' preferences. Given as input an example image and a target scene, the goal is to efficiently produce scene images that exhibit the desired appearance while maintaining the harmonization of the entire scene. For interior designs composed of multiple objects, the key to object-aware appearance transfer is to prevent scrappy appearance features from being distributed all over the generated images. In this paper, we utilize a pre-trained vision�language model (VLM) and a text-to-image generative model to solve the object-aware appearance transfer task. Specifically, we propose a VLM-assisted align-and-complement strategy using scene graph representation to determine object appearance with comprehensive considerations in the target scene. In addition, when conducting multi-object appearance transfer, we propose a multiple contrastive loss to learn object-aware appearance features from single examples and to manipulate the compositional conditions to precisely control the transfer. We have constructed an evaluation dataset and performed comparative experiments to demonstrate the effectiveness of our method of object-aware appearance transfer for interior designs. Both qualitative and quantitative evaluations demonstrate that our method successfully transfers object appearances from an example image to target scenes without feature leakage, achieving superior visual effects to competing solutions.