Generative Exploration of Fragment-Based Chemical Space Via Large Language Models Enables the Discovery of Potent Leads for Targets Lacking Bioactive Ligands | AMiner
Generative Exploration of Fragment-Based Chemical Space Via Large Language Models Enables the Discovery of Potent Leads for Targets Lacking Bioactive Ligands
Lead discovery is a critical step in drug development, yet current computational methods often face a trade-off between exploring novel chemical space and leveraging target-specific knowledge priors. Score-oriented algorithms can support broad exploration but lack medicinal-chemistry guidance, whereas knowledge-driven design benefits from target-specific knowledge while often staying close to existing chemotypes, limiting their utility in the pivotal real-world scenario of designing leads for targets that lack bioactive ligands. Here, we present a lead-discovery framework, AutoLeadDesign, that couples large language model (LLMs) reasoning with fragment-based chemical exploration without requiring target-specific ligand templates. Through a mutually reinforcing feedback loop between the LLM and chemical fragments, AutoLeadDesign enables target-aware and interpretable exploration of chemically diverse scaffolds while preserving critical binding interactions. Computational analysis reveals that AutoLeadDesign mitigates the tendency towards local optimization often observed in LLM-based molecular design. Analysis of the design trajectories indicates that the design process shares similar mechanisms to medicinal-chemistry strategies reported in expert-led studies (specifically fragment linking, merging, and growing), supporting the interpretability of the binding modes. Furthermore, we synthesized and experimentally validated leads targeting two therapeutically relevant targets: SARS-CoV-2 PLpro and the oncogenic KRAS G12D mutant. One designed inhibitor, PLP011, showed potent antiviral activity by suppressing viral replication and restoring the viability of infected cells. Another candidate, KP032, achieved a half-maximal inhibitory concentration (IC 50 ) of 50.24 nM against the challenging target KRAS G12D. Finally, we released the fragment libraries targeting the therapeutic proteins investigated in this study. These libraries are intended to serve as fragment-based priors, providing medicinal chemists and future machine learning frameworks with essential structural insights for targeted drug design. By combining LLM-based biochemical reasoning with fragment-grounded exploration, AutoLeadDesign provides a practical strategy for more adaptive and interpretable lead discovery.