Generating novel and functional molecules is an essential task in drug discovery, particularly in addressing the critical challenges of antibiotic resistance and the scarcity of effective treatments for major diseases such as cancer. Three-dimensional (3D) structure design can directly reflect a molecule’s biological function, while its complexity and topology irregularity make de novo 3D molecule generation highly difficult under valid geometric constraints. Conditional or controllable design is a promising solution for this challenging task in a more accurate and quick expectation manner. In this study, we propose TDmol-a Text-guided De novo 3D molecule generation approach based on a multimodal diffusion model. TDmol designs a new two-stage modality alignment contrastive deep learning pipeline to extract the cross-modality shared knowledge and unique features from different modalities, including molecular textual descriptions and molecular 2D/3D structures, enabling conditional and controllable molecule generation. To the best of our knowledge, TDmol is the first to realize text-3D modality alignment for conditional molecule generation. Experimental results demonstrate that TDmol achieves a significant performance enhancement for generating valid and reliable 3D molecular structures. This work highlights the potential of multimodal foundation models in digital medicine, offering a scalable and generalizable framework for 3D molecule generation that can be adapted across diverse clinical settings. A user-friendly webserver of TDmol has been deployed for academic use at http://www.csbio.sjtu.edu.cn/bioinf/TDMol.