This paper introduces a novel Contrastive-Aligned Diffusion model (CA Diff), a cross-modal latent-space alignment and diffusion-based generative framework for multi-objective shape design. The framework jointly embeds cross-modal constraints, the structural geometry and corresponding performance, into a unified latent space through a contrastive learning module. Based on this aligned representation, a two-stage diffusion model can produce shapes that satisfy multi-objective performance constraints while maintaining diversity. To validate this new framework, a case of airfoil generation is studied. The results demonstrate that, under multi-objective aerodynamic constraints, CA Diff significantly reduces the generation error compared with non-aligned conditional diffusion models, achieving a drag-coefficient MAE of 10-4. Meanwhile, it possesses the ability to generate diverse shapes.