Proteins in complex with small-molecule ligands represent the core of structure-based drug discovery. However, three-dimensional representations are absent from most deep-learning-based generative models. Here, we present a graph-based generative modeling technology that encodes explicit 3D protein-ligand contacts within a relational graph architecture and evaluate its behavior using the dopamine D2 receptor (DD2R) as a model system. The models combine a conditional variational autoencoder that allows for activity-specific molecule generation with putative contact generation that provides predictions of molecular interactions within the target-binding pocket. We show that molecules generated with our 3D procedure are more compatible with the DD2R-binding pocket than those produced by a comparable ligand-based 2D generative method, as measured by docking scores, expected stereochemistry, and recoverability in commercial chemical databases. Predicted protein-ligand contacts were found to be among the highest-ranked docking poses with a high recovery rate. Overall, this work shows how the structural context of a protein target can enhance the generation of small molecules within a realistic binding environment.
Biomedical foundation models, trained on diverse sources of small molecule data, hold great potential for accelerating drug discovery. However, their complex nature often presents a barrier for researchers seeking scientific insights and drug candidate generation. SPARK addresses this challenge by providing a user-friendly, web-based interface that empowers researchers to leverage these powerful models in their scientific workflows. Through SPARK, users can specify target proteins and desired molecule properties, adjust pre-trained models for tailored inferences, generate lists of potential drug candidates, analyze and compare molecules through interactive visualizations, and filter candidates based on key metrics (e.g., toxicity). By seamlessly integrating human knowledge and biomedical AI models' capabilities through an interactive web-based system, SPARK can improve the efficiency of collaboration between human experts and AI, thereby accelerating drug candidate discovery and ultimately leading to breakthroughs in finding cures for various diseases.
Cytotoxic-T-lymphocyte (CTL) mediated control of HIV-1 is enhanced by targeting highly networked epitopes in complex with human-leukocyte-antigen-class-I (HLA-I). However, the extent to which the presenting HLA allele contributes to this process is unknown. Here we examine the CTL response to QW9, a highly networked epitope presented by the disease-protective HLA-B57 and disease-neutral HLA-B53. Despite robust targeting of QW9 in persons expressing either allele, T cell receptor (TCR) cross-recognition of the naturally occurring variant QW9_S3T is consistently reduced when presented by HLA-B53 but not by HLA-B57. Crystal structures show substantial conformational changes from QW9-HLA to QW9_S3T-HLA by both alleles. The TCR-QW9-B53 ternary complex structure manifests how the QW9-B53 can elicit effective CTLs and suggests sterically hindered cross-recognition by QW9_S3T-B53. We observe populations of cross-reactive TCRs for B57, but not B53 and also find greater peptide-HLA stability for B57 in comparison to B53. These data demonstrate differential impacts of HLAs on TCR cross-recognition and antigen presentation of a naturally arising variant, with important implications for vaccine design.
Immunologic recognition of peptide antigens bound to class I major histocompatibility complex (MHC) molecules is essential to both novel immunotherapeutic development and human health at large. Current methods for predicting antigen peptide immunogenicity rely primarily on simple sequence representations, which allow for some understanding of immunogenic features but provide inadequate consideration of the full scale of molecular mechanisms tied to peptide recognition. We here characterize contributions that unsupervised and supervised artificial intelligence (AI) methods can make toward understanding and predicting MHC(HLA-A2)-peptide complex immunogenicity when applied to large ensembles of molecular dynamics simulations. We first show that an unsupervised AI method allows us to identify subtle features that drive immunogenicity differences between a cancer neoantigen and its wild-type peptide counterpart. Next, we demonstrate that a supervised AI method for class I MHC(HLA-A2)-peptide complex classification significantly outperforms a sequence model on small datasets corrected for trivial sequence correlations. Furthermore, we show that both unsupervised and supervised approaches reveal determinants of immunogenicity based on time-dependent molecular fluctuations and anchor position dynamics outside the MHC binding groove. We discuss implications of these structural and dynamic immunogenicity correlates for the induction of T cell responses and therapeutic T cell receptor design.
Recent work showed that active site rather than full-protein-sequence information improves predictive performance in kinase-ligand binding affinity prediction. To refine the notion of an "active site", we here propose and compare multiple definitions. We report significant evidence that our novel definition is superior to previous definitions and better models of ATP-noncompetitive inhibitors. Moreover, we leverage the discontiguity of the active site sequence to motivate novel protein-sequence augmentation strategies and find that combining them further improves performance.
The highly infectious SARS-CoV-2 variant B.1.351 that first emerged in South Africa with triple mutations (N501Y, K417N, and E484K) is globally worrisome. It is known that N501Y and E484K can enhance binding between the coronavirus receptor domain (RBD) and human ACE2. However, the K417N mutation appears to be unfavorable as it removes one interfacial salt bridge. Here, we show that despite the decrease in binding affinity (1.48 kcal/mol) between RBD and ACE2, the K417N mutation abolishes a buried interfacial salt bridge between the RBD and neutralizing antibody CB6. This substantially reduces their binding energy by 9.59 kcal/mol, thus facilitating the process by which the variant efficiently eludes CB6 (including many other antibodies). Our theoretical predictions agree with existing experimental findings. Harnessing the revealed molecular mechanisms makes it possible to redesign therapeutic antibodies, thus making them more efficacious.
Recent advances in deep learning have enabled the development of large-scale multimodal models for virtual screening and de novo molecular design. The human kinome with its abundant sequence and inhibitor data presents an attractive opportunity to develop proteochemometric models that exploit the size and internal diversity of this family of targets. Here we challenge a standard practice in sequence-based affinity prediction models: instead of leveraging the full primary structure of proteins, each target is represented by a sequence of 29 residues defining the ATP binding site. In kinase-ligand binding affinity prediction, our results show that the reduced active site sequence representation is not only computationally more efficient but consistently yields significantly higher performance than the full primary structure. This trend persists across different models, datasets, performance metrics and holds true when predicting affinity for both unseen ligands and kinases. Our interpretability analysis further demonstrates that, even without supervision, the full sequence model can learn to focus on the active site residues to a higher extent. We then investigate a de novo molecular design task and find that the active site provides benefits in the computational efficiency, but otherwise, both kinase representations yield similar optimized affinities (for both SMILES and SELFIES-based molecular generators). Our work challenges the assumption that full primary structure is indispensable for modelling human kinases. We hope that these results will inspire additional investigation into hybrid mechanistic-DL modeling approaches to support the identification and optimization of kinase inhibitors’ candidates.
Proteins in complex with small molecule ligands represent the core of structure-based drug discovery. However, three-dimensional representations are absent from most deep-learning-based generative models. We here present a graph-based generative modeling technology that encodes explicit 3D protein-ligand contacts within a relational graph architecture. The models combine a conditional variational autoencoder that allows for activity-specific molecule generation with putative contact generation that provides predictions of molecular interactions within the target binding pocket. We show that molecules generated with our 3D procedure are more compatible with the binding pocket of the dopamine D2 receptor than those produced by a comparable ligand-based 2D generative method, as measured by docking scores, expected stereochemistry, and recoverability in commercial chemical databases. Predicted protein-ligand contacts were found among highest-ranked docking poses with a high recovery rate. This work shows how the structural context of a protein target can be used to enhance molecule generation.
Recently, severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) variants (B.1.1.7 and B.1351) have emerged harbouring mutations that make them highly contagious. The N501Y mutation within the receptor‐binding domain (RBD) of the spike protein of these SARS‐CoV‐2 variants may enhance binding to the human angiotensin‐converting enzyme 2 (hACE2). However, no molecular explanation for such an enhanced affinity has so far been provided. Here, using all‐atom molecular dynamics simulations, we show that Y501 in the mutated RBD can be well‐coordinated by Y41 and K353 in hACE2 through hydrophobic interactions, which may increase the overall binding affinity of the RBD for hACE2 by approximately 0.81 kcal·mol−1. The binding dynamics revealed in our study may provide a working model to facilitate the design of more effective antibodies.
Coronavirus disease 2019 (COVID-19) is an ongoing global pandemic caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), with very limited treatments so far. Demonstrated with good druggability, two major proteases of SARS-CoV-2, namely main protease (Mpro) and papain-like protease (PLpro) that are essential for viral maturation, have become the targets for many newly designed inhibitors. Unlike Mpro that has been heavily investigated, PLpro is not well-studied so far. Here, we carried out the in silico high-throughput screening of all FDA-approved drugs via the flexible docking simulation for potential inhibitors of PLpro and explored the molecular mechanism of binding between a known inhibitor rac5c and PLpro. Our results, from molecular dynamics simulation, show that the chances of drug repurposing for PLpro might be low. On the other hand, our long (about 450 ns) MD simulation confirms that rac5c can be bound stably inside the substrate-binding site of PLpro and unveils the molecular mechanism of binding for the rac5c-PLpro complex. The latter may help perform further structural optimization and design potent leads for inhibiting PLpro.
Coronavirus disease 2019 (COVID-19) has been an ongoing global pandemic for over a year. Recently, an emergent SARS-CoV-2 variant (B.1.1.7) with an unusually large number of mutations had become highly contagious and wide-spreading in United Kingdom. From genome analysis, the N501Y mutation within the receptor binding domain (RBD) of the SARS-CoV-2’s spike protein might have enhanced the viral protein’s binding with the human angiotensin converting enzyme 2 (hACE2). The latter is the prelude for the virus’ entry into host cells. So far, the molecular mechanism of this enhanced binding is still elusive, which prevents us from assessing its effects on existing therapeutic antibodies. Using all atom molecular dynamics simulations, we demonstrated that Y501 in mutated RBD can be well coordinated by Y41 and K353 in hACE2 through hydrophobic interactions, increasing the overall binding affinity between RBD and hACE2 by about 0.81 kcal/mol. We further explored how the N501Y mutation might affect the binding between a neutralizing antibody (CB6) and RBD. We expect that our work can help researchers design proper measures responding to this urgent virus mutation, such as adding a modified/new neutralizing antibody specifically targeting at this variant in the therapeutic antibody cocktail.
The highly infectious SARS-CoV-2 variant B.1.617 with double mutations E484Q and L452R in the receptor binding domain (RBD) of SARS-CoV-2’s spike protein is worrisome. Demonstrated in crystal structures, the residues 452 and 484 in RBD are not in direct contact with interfacial residues in the angiotensin converting enzyme 2 (ACE2). This suggests that albeit there are some possibly nonlocal effects, the E484Q and L452R mutations might not significantly affect RBD’s binding with ACE2, which is an important step for viral entry into host cells. Thus, without the known molecular mechanism, these two successful mutations (from the point of view of SARS-CoV-2) can be hypothesized to evade human antibodies. Using in silico all-atom molecular dynamics (MD) simulation as well as deep learning (DL) approaches, here we show that these two mutations significantly reduce the binding affinity between RBD and the antibody LY-CoV555 (also named as Bamlanivimab) that was proven to be efficacious for neutralizing the wide-type SARS-CoV-2. With the revealed molecular mechanism on how L452R and E484K evade LY-CoV555, we expect that more specific therapeutic antibodies can be accordingly designed and/or a precision mixing of antibodies can be achieved in a cocktail treatment for patients infected with the variant B.1.617.
Since the beginning of the COVID-19 pandemic, scientists across the globe are racing to find a cure for the highly contagious infectious disease caused by the SARS-CoV-2 virus. Despite many promising ongoing progress, there are currently no FDA approved drug to treat infected patients. Recently, the crowdsourcing of drug discovery for inhibiting the main protease (Mpro) of SARS-CoV-2 have yielded a plenty of drug fragments resolved inside the active site of Mpro via the crystallography method. Following the principle of fragment-based drug design (FBDD), we are motivated to design a potent drug candidate (named B19) by merging three fragments JFM, U0P, and HWH. Through extensive all-atom molecular dynamics simulation and molecular docking, we found that B19 among all designed ones is most stable inside the Mpro's active site and the binding free energy of B19 is comparable to or even a little better than that of a native protein ligand processed by Mpro. Our promising results suggest that B19 and its derivatives can potentially be efficacious drug candidates for COVID-19.
The newly emerging Kappa, Delta, and Lambda SARS-CoV-2 variants are worrisome, characterized with the double mutations E484Q/L452R, T478K/L452R, and F490S/L452Q, respectively, in their receptor binding domains (RBDs) of the spike proteins. As revealed in crystal structures, most of these residues (e.g., 452 and 484 in RBDs) are not in direct contact with interfacial residues in the angiotensin-converting enzyme 2 (ACE2). This suggests that albeit there are some possibly nonlocal effects, these mutations might not significantly affect RBD’s binding with ACE2, which is an important step for viral entry into host cells. Thus, without knowing the molecular mechanism, these successful mutations (from the point of view of SARS-CoV-2) may be hypothesized to evade human antibodies. Using all-atom molecular dynamics (MD) simulation, here, we show that the E484Q/L452R mutations significantly reduce the binding affinity between the RBD of the Kappa variant and the antibody LY-CoV555 (also named as Bamlanivimab), which was efficacious for neutralizing the wild-type SARS-CoV-2. To verify simulation results, we further carried out experiments with both pseudovirions- and live virus-based neutralization assays and demonstrated that LY-CoV555 completely lost neutralizing activity against the L452R/E484Q mutant. Similarly, we show that mutations in the Delta and Lambda variants can also destabilize the RBD’s binding with LY-CoV555. With the revealed molecular mechanism on how these variants evade LY-CoV555, we expect that more specific therapeutic antibodies can be accordingly designed and/or a precise mixing of antibodies can be achieved as a cocktail treatment for patients infected with these variants.
Coronavirus disease 2019 (COVID-19) is an ongoing global pandemic and there are currently no FDA approved medicines for treatment or prevention. Inspired by promising outcomes for convalescent plasma treatment, developing antibody drugs (biologics) to block SARS-CoV-2 infection has been the focus of drug discovery, along with tremendous efforts in repurposing small-molecule drugs. In the last several months, experimentally, many human neutralizing monoclonal antibodies (mAbs) were successfully extracted from plasma of recovered COVID-19 patients. Currently, several mAbs targeting the SARS-CoV-2's spike protein (Spro) are in clinical trials. With known atomic structures of mAb-Spro complex, it becomes possible to in silico investigate the molecular mechanism of mAb's binding with Spro and design more potent mAbs through protein mutagenesis studies, complementary to existing experimental efforts. Leveraging superb computing power nowadays, we propose a fully automated in silico protocol for quickly identifying possible mutations in a mAb (e.g.~CB6) to enhance its binding affinity with Spro for the design of more efficacious therapeutic mAbs.
Recently, phosphorene, a novel two-dimensional nanomaterial with a puckered surface morphology, was shown to exhibit cytotoxicity, but its underlying molecular mechanisms remain unknown. Herein, using large scale molecular dynamics simulations, we show that phosphorene nanosheets can penetrate into and extract large amounts of phospholipids from the cell membranes due to the strong dispersion interaction between phosphorene and lipid molecules, which would reduce cell viability. The extracted phospholipid molecules are aligned along the wrinkle direction of the phosphorene nanosheet because of its unique puckered structure. Our results also reveal that small phosphorene nanosheets penetrate into the cell membrane in a specific direction which is determined by the size and surface topography of phosphorene and the thickness of the membrane. These findings might shed light on understanding phosphorene's cytotoxicity and would be helpful for the future potential biomedical applications of phosphorene, such as biosensors and antibacterial agents.
We present a simple, modular graph-based convolutional neural network that takes structural information from protein-ligand complexes as input to generate models for activity and binding mode prediction. Complex structures are generated by a standard docking procedure and fed into a dual-graph architecture that includes separate subnetworks for the ligand bonded topology and the ligand-protein contact map. Recent work has indicated that data set bias drives many past promising results derived from combining deep learning and docking. Our dual-graph network allows contributions from ligand identity that give rise to such biases to be distinguished from effects of protein-ligand interactions on classification. We show that our neural network is capable of learning from protein structural information when, as in the case of binding mode prediction, an unbiased data set is constructed. We next develop a deep learning model for binding mode prediction that uses docking ranking as input in combination with docking structures. This strategy mirrors past consensus models and outperforms a baseline docking program (AutoDock Vina) in a variety of tests, including on cross-docking data sets that mimic real-world docking use cases. Furthermore, the magnitudes of network predictions serve as reliable measures of model confidence.
We applied the flexible docking method to rank-order all FDA-approved drugs as inhibitors for the papain-like protease (PLpro) of SRAS-CoV-2. We also evaluated these results using molecular dynamics (MD) simulations. From MD simulations, we unveiled the molecular mechanism for a known inhibitor rac5c's binding with PLpro.
The 20S proteasome, the catalytic core particle of 26S proteasome, degrades a wide range of intracellular proteins, which is essential for many cellular processes. Herein, we have found that the 20S proteasome activity is either up- or down-regulated by introducing gold nanoclusters (AuNCs) coated with nine peptide tails in two different forms, AuNC(-) and AuNC(+), each encoding five consecutive negatively or positively charged amino acids. Molecular dynamics simulations reveal that AuNC(-) and AuNC(+) bind to different surfaces of the 20S proteasome, and respectively facilitate or hinder the opening of the central gate of 20S proteasome for substrate access to the internal active site for protein degradation. Furthermore, the addition of AuNC(-) induces protective effects in a cell model of Parkinson's disease, by up-regulating the proteasome activity under the condition of reduced ATP production, and enhancing the degradation of overexpressed ee-synuclein, thereby attenuating the loss of cell viability. Our findings suggest the potential application of gold nanoclusters for treating neurodegenerative diseases. (C) 2020 Elsevier Ltd. All rights reserved.
X-ray-responsive nanocarriers for anticancer drug delivery have shown great promise for enhancing the efficacy of chemoradiotherapy. A critical challenge remains for development of such radiation-controlled drug delivery systems (DDSs), which is to minimize the required X-ray dose for triggering the cargo release. Herein, we design and fabricate an effective DDS based on diselenide block copolymers (as nanocarrier), which can be triggered to release their cargo with a reduced radiation dose of 2 Gy due to their sensitivity to both X-ray and the high level of reactive oxygen species (ROS) in the microenvironment of cancer cells. The underlying molecular mechanism is further illustrated by proton nuclear magnetic resonance (1H NMR) experiments and density functional theory (DFT) calculations. In vivo experiments on tumor-bearing mice validated that the loaded drugs are effectively delivered to the tumor site and exert remarkable antitumor effects (minimum tumor volume/weight) along with X-ray. Furthermore, the diselenide nanocarriers exhibit no noticeable cytotoxicity. These findings provide new insights for the de novo design of radiation-controlled DDSs for cancer chemoradiotherapy.