Within the last few decades, increases in computational resources have contributed enormously to the progress of science and engineering (S & E). To continue making rapid advancements, the S & E community must be able to access computing resources. One way to provide such resources is through High-Performance Computing (HPC) centers. Many academic research institutions offer their own HPC Centers but struggle to make the computing resources easily accessible and user-friendly. Here we present SHABU, a RESTful Web API framework that enables S & E communities to access resources from Boston University's Shared Computing Center (SCC). The SHABU requirements are derived from the use cases described in this work.
An increasing number of medically important proteins are challenging drug targets because their binding sites are too shallow or too polar, are cryptic and thus not detectable without a bound ligand or located in a protein-protein interface. While such proteins may not bind druglike small molecules with sufficiently high affinity, they are frequently druggable using novel therapeutic modalities. The need for such modalities can be determined by experimental or computational fragment based methods. Computational mapping by mixed solvent molecular dynamics simulations or the FTMap server can be used to determine binding hot spots. The strength and location of the hot spots provide very useful information for selecting potentially successful approaches to drug discovery.
Despite the growing number of G protein-coupled receptor (GPCR) structures, only 39 structures have been cocrystallized with allosteric inhibitors. These structures have been studied by protein mapping using the FTMap server, which determines the clustering of small organic probe molecules distributed on the protein surface. The method has found druggable sites overlapping with the cocrystallized allosteric ligands in 21 GPCR structures. Mapping of Alphafold2 generated models of these proteins confirms that the same sites can be identified without the presence of bound ligands. We then mapped the 394 GPCR X-ray structures available at the time of the analysis (September 2020). Results show that for each of the 21 structures with bound ligands there exist many other GPCRs that have a strong binding hot spot at the same location, suggesting potential allosteric sites in a large variety of GPCRs. These sites cluster at nine distinct locations, and each can be found in many different proteins. However, ligands binding at the same location generally show little or no similarity, and the amino acid residues interacting with these ligands also differ. Results confirm the possibility of specifically targeting these sites across GPCRs for allosteric modulation and help to identify the most likely binding sites among the limited number of potential locations. The FTMap server is available free of charge for academic and governmental use at https://ftmap.bu.edu/.
Botulinum neurotoxins (BoNTs) are extremely toxic and have been deemed a Tier 1 potential bioterrorism agent. The most potent and persistent of the BoNTs is the "A" serotype, with strategies to counter its etiology focused on designing small-molecule inhibitors of its light chain (LC), a zinc-dependent metalloprotease. The successful structure-based drug design of inhibitors has been confounded as the LC is highly flexible with significant morphological changes occurring upon inhibitor binding. To achieve greater success, previous and new cocrystal structures were evaluated from the standpoint of inhibitor enantioselectivity and their effect on active-site morphology. Based upon these structural insights, we designed inhibitors that were predicted to take advantage of π-π stacking interactions present in a cryptic hydrophobic subpocket. Structure-activity relationships were defined, and X-ray crystal structures and docking models were examined to rationalize the observed potency differences between inhibitors.
We have used crystal structures and molecular modeling to evaluate inhibitor binding modes and design a series of compounds to take advantage of a new, cryptic, hydrophobic sub-pocket. This is a classical SBDD approach to improving enzyme/inhibitor interactions.
Fragment-based drug design has introduced a bottom-up process for drug development, with improved sampling of chemical space and increased effectiveness in early drug discovery. Here, we combine the use of pharmacophores, the most general concept of representing drug-target interactions with the theory of protein hotspots, to develop a design protocol for fragment libraries. The SpotXplorer approach compiles small fragment libraries that maximize the coverage of experimentally confirmed binding pharmacophores at the most preferred hotspots. The efficiency of this approach is demonstrated with a pilot library of 96 fragment-sized compounds (SpotXplorer0) that is validated on popular target classes and emerging drug targets. Biochemical screening against a set of GPCRs and proteases retrieves compounds containing an average of 70% of known pharmacophores for these targets. More importantly, SpotXplorer0 screening identifies confirmed hits against recently established challenging targets such as the histone methyltransferase SETD2, the main protease (3CLPro) and the NSP3 macrodomain of SARS-CoV-2.
Many proteins in their unbound structures have cryptic sites that are not appropriately sized for drug binding. We consider here 32 proteins from the recently published CryptoSite set with validated cryptic sites, and study whether the sites remain cryptic in all available X-ray structures of the proteins solved without any ligand bound near the sites. It was shown that only few of these proteins have binding pockets that never form without ligand binding. Sites that are cryptic in some structures but spontaneously form in others are also rare. In most proteins the forming of pockets is affected by mutations or ligand binding at locations far from the cryptic site. To further explore these mechanisms, we applied adiabatic biased molecular dynamics simulations to guide the proteins from their ligand-free structures to ligand-bound conformations, and studied the distribution of druggability scores of the pockets located at the cryptic sites.
Binding hot spots are regions of proteins that, due to their potentially high contribution to the binding free energy, have high propensity to bind small molecules. We present benchmark sets for testing computational methods for the identification of binding hot spots with emphasis on fragment-based ligand discovery. Each protein structure in the set binds a fragment, which is extended into larger ligands in other structures without substantial change in its binding mode. Structures of the same proteins without any bound ligand are also collected to form an unbound benchmark. We also discuss a set developed by Astex Pharmaceuticals for the validation of hot and warm spots for fragment binding. The set is based on the assumption that a fragment that occurs in diverse ligands in the same subpocket identifies a binding hot spot. Since this set includes only ligand-bound proteins, we added a set with unbound structures. All four sets were tested using FTMap, a computational analogue of fragment screening experiments to form a baseline for testing other prediction methods, and differences among the sets are discussed.
Development of small molecule inhibitors of protein-protein interactions (PPIs) is hampered by our poor understanding of the druggability of PPI target sites. Here, we describe the combined application of alanine-scanning mutagenesis, fragment screening, and FTMap computational hot spot mapping to evaluate the energetics and druggability of the highly charged PPI interface between Kelch-like ECH-associated protein 1 (KEAP1) and nuclear factor erythroid 2 like 2 (Nrf2), an important drug target. FTMap identifies four binding energy hot spots at the active site. Only two of these are exploited by Nrf2, which alanine scanning of both proteins shows to bind primarily through E79 and E82 interacting with KEAP1 residues S363, R380, R415, R483, and S508. We identify fragment hits and obtain X-ray complex structures for three fragments via crystal soaking using a new crystal form of KEAP1. Combining these results provides a comprehensive and quantitative picture of the origins of binding energy at the interface. Our findings additionally reveal non-native interactions that might be exploited in the design of uncharged synthetic ligands to occupy the same site on KEAP1 that has evolved to bind the highly charged DEETGE binding loop of Nrf2. These include pi-stacking with KEAP1 Y525 and interactions at an FTMap-identified hot spot deep in the binding site. Finally, we discuss how the complementary information provided by alanine-scanning mutagenesis, fragment screening, and computational hot spot mapping can be integrated to more comprehensively evaluate PPI druggability.
Allosteric modulation of G protein-coupled receptors represent a promising mechanism of pharmacological intervention. Dramatic developments witnessed in the structural biology of membrane proteins continue to reveal that the binding sites of allosteric modulators are widely distributed, including along protein surfaces. Here we restrict consideration to intrahelical and intracellular sites together with allosteric conformational locks, and show that the protein mapping tools FTMap and FTSite identify 83% and 88% of such experimentally confirmed allosteric sites within the three strongest sites found. The methods were also able to find partially hidden allosteric sites that were not fully formed in X-ray structures crystallized in the absence of allosteric ligands. These results confirm that the intrahelical sites capable of binding drug like allosteric modulators are among the strongest ligand recognition sites in a large fraction of GPCRs and suggest that both FTMap and FTSite are useful tools for identifying allosteric sites and to aid in the design of such compounds in a range of GPCR targets.
•Many proteins have cryptic binding sites that are not detectable in ligand-free structures.•Cryptic sites can provide druggable targets and have recently received growing attention.•Detection methods based on molecular dynamics are becoming computationally feasible.•Only a few of the transient pockets obtained by molecular dynamics are predicted to be druggable.•Prediction methods can be improved by fragment docking and machine learning.
Molecular dynamics (MD) simulations of proteins reveal the existence of many transient surface pockets; however, the factors determining what small subset of these represent druggable or functionally relevant ligand binding sites, called "cryptic sites," are not understood. Here, we examine multiple X-ray structures for a set of proteins with validated cryptic sites, using the computational hot spot identification tool FTMap. The results show that cryptic sites in ligand-free structures generally have a strong binding energy hot spot very close by. As expected, regions around cryptic sites exhibit above-average flexibility, and close to 50% of the proteins studied here have unbound structures that could accommodate the ligand without clashes. Nevertheless, the strong hot spot neighboring each cryptic site is almost always exploited by the bound ligand, suggesting that binding may frequently involve an induced fit component. We additionally evaluated the structural basis for cryptic site formation, by comparing unbound to bound structures. Cryptic sites are most frequently occluded in the unbound structure by intrusion of loops (22.5%), side chains (19.4%), or in some cases entire helices (5.4%), but motions that create sites that are too open can also eliminate pockets (19.4%). The flexibility of cryptic sites frequently leads to missing side chains or loops (12%) that are particularly evident in low resolution crystal structures. An interesting observation is that cryptic sites formed solely by the movement of side chains, or of backbone segments with fewer than five residues, result only in low affinity binding sites with limited use for drug discovery.
Many proteins in their unbound structures lack surface pockets appropriately sized for drug binding. Hence, a variety of experimental and computational tools have been developed for the identification of cryptic sites that are not evident in the unbound protein but form upon ligand binding, and can provide tractable drug target sites. The goal of this review is to discuss the definition, detection, and druggability of such sites, and their potential value for drug discovery. Novel methods based on molecular dynamics simulations are particularly promising and yield a large number of transient pockets, but it has been shown that only a minority of such sites are generally capable of binding ligands with substantial affinity. Based on recent studies, current methodology can be improved by combining molecular dynamics with fragment docking and machine learning approaches.
To test the ability of molecular simulations to accurately predict the solution-state conformational properties of peptidomimetics, we examined a test set of 18 cyclic RGD peptides selected from the literature, including the anticancer drug candidate cilengitide, whose favorable binding affinity to integrin has been ascribed to its pre-organization in solution. For each design, we performed all-atom replica-exchange molecular dynamics simulations over several microseconds and compared the results to extensive published NMR data. We find excellent agreement with experimental NOE distance restraints, suggesting that molecular simulation can be a useful tool for the computational design of pre-organized solution-state structure. Moreover, our analysis of conformational populations estimates that, despite the potential for increased flexibility due to backbone amide isomerizaton, N-methylation provides about 0.5 kcal/mol of reduced conformational entropy to cyclic RGD peptides. The combination of pre-organization and binding-site compatibility explains the strong binding affinity of cilengitide to integrin.