Macrocyclic peptides are increasingly used to bridge the gap between small molecules and biologics. Their biological activity and pharmacological properties are governed by conformation, making a detailed understanding of the conformational landscape essential for rational design. Although nuclear magnetic resonance (NMR) is the primary method for determining structures in solution, it remains challenging to apply, especially to N-methylated peptides, because of the scarcity of informative observables such as NOEs, J-couplings, and RDCs. To address this challenge, we present a new method for conformational analysis, which we call CSCAN (Chemical Shift Conformational ANalysis), based solely on comparing experimental 1H and 13C NMR chemical shifts with theoretical values computed via an efficient workflow combining extensive conformational searches, deep-learning-based geometry refinement, and DFT calculations. Using two pharmaceutically relevant macrocyclic peptides─Cyclosporin A and Aureobasidin A─as examples, we demonstrate that this approach can reliably identify conformations consistent with both X-ray crystallographic and NMR-derived reference structures. Furthermore, it addresses limitations of conventional conformational analyses by detecting differences between solution conformations and X-ray structures, as well as improving the assignment of solution conformations determined from NOE and RDC data.
Systematic optimization of large macrocyclic peptide ligands is a serious challenge. Here, we describe an approach for lead-optimization using the PD-1/PD-L1 system as a retrospective example of moving from initial lead compound to clinical candidate. We show how conformational restraints can be derived by exploiting NMR data to identify low-energy solution ensembles of a lead compound. Such restraints can be used to focus conformational search for analogs in order to accurately predict bound ligand poses through molecular docking and thereby estimate ligand strain and protein-ligand intermolecular binding energy. We also describe an analogous ligand-based approach that employs molecular similarity optimization to predict bound poses. Both approaches are shown to be effective for prioritizing lead-compound analogs. Surprisingly, relatively small ligand modifications, which may have minimal effects on predicted bound pose or intermolecular interactions, often lead to large changes in estimated strain that have dominating effects on overall binding energy estimates. Effective macrocyclic conformational search is crucial, whether in the context of NMR-based restraints, X-ray ligand refinement, partial torsional restraint for docking/ligand-similarity calculations or agnostic search for nominal global minima. Lead optimization for peptidic macrocycles can be made more productive using a multi-disciplinary approach that combines biophysical data with practical and efficient computational methods.
Scaffold replacement as part of an optimization process that requires maintenance of potency, desirable biodistribution, metabolic stability, and considerations of synthesis at very large scale is a complex challenge. Here, we consider a set of over 1000 time-stamped compounds, beginning with a macrocyclic natural-product lead and ending with a broad-spectrum crop anti-fungal. We demonstrate the application of the QuanSA 3D-QSAR method employing an active learning procedure that combines two types of molecular selection. The first identifies compounds predicted to be most active of those most likely to be well-covered by the model. The second identifies compounds predicted to be most informative based on exhibiting low predicted activity but showing high 3D similarity to a highly active nearest-neighbor training molecule. Beginning with just 100 compounds, using a deterministic and automatic procedure, five rounds of 20-compound selection and model refinement identifies the binding metabolic form of florylpicoxamid. We show how iterative refinement broadens the domain of applicability of the successive models while also enhancing predictive accuracy. We also demonstrate how a simple method requiring very sparse data can be used to generate relevant ideas for synthetic candidates.
So-called "cross-docking" is the prediction of the bound configuration of small-molecule ligands that differ from the cognate ligand of a protein co-crystal structure. This is a much more challenging problem than re-docking the cognate ligand, particularly when the new ligand is structurally dissimilar from prior known ones. We have updated the previously introduced PINC ("PINC Is Not Cognate") benchmark which introduced the idea of temporal segregation to measure cross-docking performance. The temporal set encompasses 846 future ligands for ten targets based on information from the earliest 25% of X-ray co-crystal structures known for each target. Here, we extend the benchmark to include thirteen targets where the bound poses of 128 macrocyclic ligands are to be predicted based on knowledge from structures of bound non-macrocyclic ligands. Performance was roughly equivalent for both the temporally-split non-macrocyclic ligand set and the macrocycle prediction set. Using standard and fully automatic protocols for the Surflex-Dock and ForceGen methods, across the combined 974 non-macrocyclic and macrocyclic ligands, the top-scoring pose family was correct 68% of the time, with the top-two pose families achieving a 79% success rate. Correct poses among all those predicted were identified 92% of the time. These success rates far exceeded those observed for the alternative methods AutoDock Vina and Gnina on both sets.
The diffusion learning method, DiffDock, for docking small-molecule ligands into protein binding sites was recently introduced. Results included comparisons to more conventional docking approaches, with DiffDock showing superior performance. Here, we employ a fully automatic workflow using the Surflex-Dock methods to generate a fair baseline for conventional docking approaches. Results were generated for the common and expected situation where a binding site location is known and also for the condition of an unknown binding site. For the known binding site condition, Surflex-Dock success rates at 2.0 Angstroms RMSD far exceeded those for DiffDock (Top-1/Top-5 success rates, respectively, were 68/81 similar success rates (67/73 condition, and results for AutoDock Vina and Gnina followed this pattern. For the unknown binding site condition, using an automated method to identify multiple binding pockets, Surflex-Dock success rates again exceeded those of DiffDock, but by a somewhat lesser margin. DiffDock made use of roughly 17,000 co-crystal structures for learning (98 structures) for a training set in order to predict on 363 test cases (2 PDBBind 2020) from 2019 forward. DiffDock's performance was inextricably linked with the presence of near-neighbor cases of close to identical protein-ligand complexes in the training set for over half of the test set cases. DiffDock exhibited a 40 percentage point difference on near-neighbor cases (two-thirds of all test cases) compared with cases with no near-neighbor training case. DiffDock has apparently encoded a type of table-lookup during its learning process, rendering meaningful applications beyond its reach. Further, it does not perform even close to competitively with a competently run modern docking workflow.
Rapamycin, a well-known macrocyclic natural product withmyriadbiological activities, has been the subject of intense study sinceits first isolation and characterization over five decades ago. Rapamycinhas been found to adopt a single conformation in the solid state (bothwhen protein bound and uncomplexed) and exists as a mixture of twoconformations in solution. Early work established that the major conformerin solution is the trans amide isomer but left the minor conformermostly uncharacterized. Since that time, it has been widely acceptedthat the minor conformer of rapamycin is the cis amide, based solelyon analogy to FK-506, another potent immunosuppressive compound withsome shared key structural elements. To address this long-standingand unresolved question, the solution structure of the minor conformerof rapamycin was investigated using a combination of NMR techniquesand computational methods and determined to be a trans amide specieswith rotation about the ester linkage.
The internal conformational strain incurred by ligands upon binding a target site has a critical impact on binding affinity, and expectations about the magnitude of ligand strain guide conformational search protocols. Estimates for bound ligand strain begin with modeled ligand atomic coordinates from X-ray co-crystal structures. By deriving low-energy conformational ensembles to fit X-ray diffraction data, calculated strain energies are substantially reduced compared with prior approaches. We show that the distribution of expected global strain energy values is dependent on molecular size in a superlinear manner. The distribution of strain energy follows a rectified normal distribution whose mean and variance are related to conformational complexity. The modeled strain distribution closely matches calculated strain values from experimental data comprising over 3000 protein-ligand complexes. The distributional model has direct implications for conformational search protocols as well as for directions in molecular design.
Aureobasidin A (abA) is a natural depsipeptide that inhibits inositol phosphorylceramide (IPC) synthases with significant broad-spectrum antifungal activity. abA is known to have two distinct conformations in solution corresponding to trans- and cis-proline (Pro) amide bond rotamers. While the trans-Pro conformation has been studied extensively, cis-Pro conformers have remained elusive. Conformational properties of cyclic peptides are known to strongly affect both potency and cell permeability, making a comprehensive characterization of abA conformation highly desirable. Here, we report a high-resolution 3D structure of the cis-Pro conformer of aureobasidin A elucidated for the first time using a recently developed NMR-driven computational approach. This approach utilizes ForceGen's advanced conformational sampling of cyclic peptides augmented by sparse distance and torsion angle constraints derived from NMR data. The obtained 3D conformational structure of cis-Pro abA has been validated using anisotropic residual dipolar coupling measurements. Support for the biological relevance of both the cis-Pro and trans-Pro abA configurations was obtained through molecular similarity experiments, which showed a significant 3D similarity between NMR-restrained abA conformational ensembles and another IPC synthase inhibitor, pleofungin A. Such ligand-based comparisons can further our understanding of the important steric and electrostatic characteristics of abA and can be utilized in the design of future therapeutics.
We present results on the extent to which physics-based simulation (exemplified by FEP+) and focused machine learning (exemplified by QuanSA) are complementary for ligand affinity prediction. For both methods, predictions of activity for LFA-1 inhibitors from a medicinal chemistry lead optimization project were accurate within the applicable domain of each approach. A hybrid model that combined predictions by both approaches by simple averaging performed better than either method, with respect to both ranking and absolute pKi values. Two publicly available FEP+ benchmarks, covering 16 diverse biological targets, were used to test the generality of the synergy. By identifying training data specifically focused on relevant ligands, accurate QuanSA models were derived using ligand activity data known at the time of the original series publications. Results across the 16 benchmark targets demonstrated significant improvements both for ranking and for absolute pKi values using hybrid predictions that combined the FEP+ and QuanSA predicted affinity values. The results argue for a combined approach for affinity prediction that makes use of physics-driven methods as well as those driven by machine learning, each applied carefully on appropriate compounds, with hybrid prediction strategies being employed where possible.
Macrocyclic peptides are an important modality in drug discovery, but molecular design is limited due to the complexity of their conformational landscape. To better understand conformational propensities, global strain energies were estimated for 156 protein-macrocyclic peptide cocrystal structures. Unexpectedly large strain energies were observed when the bound-state conformations were modeled with positional restraints. Instead, low-energy conformer ensembles were generated using xGen that fit experimental X-ray electron density maps and gave reasonable strain energy estimates. The ensembles featured significant conformational adjustments while still fitting the electron density as well or better than the original coordinates. Strain estimates suggest the interaction energy in protein-ligand complexes can offset a greater amount of strain for macrocyclic peptides than for small molecules and non-peptidic macrocycles. Across all molecular classes, the approximate upper bound on global strain energies had the same relationship with molecular size, and bound-state ensembles from xGen yielded favorable binding energy estimates.
Using the DUD-E+ benchmark, we explore the impact of using a single protein pocket or ligand for virtual screening compared with using ensembles of alternative pockets, ligands, and sets thereof. For both structure-based and ligand-based approaches, the precise characterization of the binding site in question had a significant impact on screening performance. Using the single original DUD-E protein, Surflex-Dock yielded mean ROC area of 0.81 ± 0.11. Using the cognate ligand instead, with the eSim method for screening, yielded 0.77 ± 0.14. Moving to ensembles of five protein pocket variants increased docking performance to 0.84 ± 0.09. Results for the analogous ligand-based approach (using the five crystallographically aligned cognate ligands) was 0.83 ± 0.11. Using the same ligands, but making use of an automatically generated mutual alignment, yielded mean AUC nearly as good as from single-structure docking: 0.80 ± 0.12. Detailed results and statistical analyses show that structure- and ligand-based methods are complementary and can be fruitfully combined to enhance screening efficiency. A hybrid approach combining ensemble docking with eSim-based screening produced the best and most consistent performance (mean ROC area of 0.89 ± 0.08 and 1% early enrichment of 46-fold). Based on results from both the docking and ligand-similarity approaches, it is clearly unwise to make use of a single arbitrarily chosen protein structure for docking or single ligand query for similarity-based screening.
We report a new method for X-ray density ligand fitting and refinement that is suitable for a wide variety of small-molecule ligands, including macrocycles. The approach (called “xGen”) augments a force field energy calculation with an electron density fitting restraint that yields an energy reward during the restrained conformational search. The resulting conformer pools balance goodness-of-fit with ligand strain. Real-space refinement from pre-existing ligand coordinates of 150 macrocycles resulted in occupancy-weighted conformational ensembles that exhibited low strain energy. The xGen ensembles improved upon electron density fit compared with the PDB reference coordinates without making use of atom-specific B-factors. Similarly, on nonmacrocycles, de novo fitting produced occupancy-weighted ensembles of many conformers that were generally better-quality density fits than the deposited primary/alternate conformational pairs. The results suggest ubiquitous low-energy ligand conformational ensembles in X-ray diffraction data and provide an alternative to using B-factors as model parameters.
We introduce a new method for rapid computation of 3D molecular similarity that combines electrostatic field comparison with comparison of molecular surface-shape and directional hydrogen-bonding preferences (called “eSim”). Rather than employing heuristic “colors” or user-defined molecular feature types to represent conformation-dependent molecular electrostatics, eSim calculates the similarity of the electrostatic fields of two molecules (in addition to shape and hydrogen-bonding). We present detailed virtual screening performance data on the standard 102 target DUD-E set. In its moderately fast screening mode, eSim running on a single computing core is capable of processing over 60 molecules per second. In this mode, eSim performed significantly better than all alternate methods for which full DUD-E data were available (mean ROC area of 0.74, p $$< 10^{-9}$$, by paired t-test, compared with the best performing alternate method). In addition, for 92 targets of the DUD-E set where multiple ligand-bound crystal structures were available, screening performance was assessed using alternate ligands or sets thereof (in their bound poses) as similarity targets. Using the joint alignment of five ligands for each protein target, mean ROC area exceeded 0.82 for the 92 targets. Design-focused application of ligand similarity methods depends on accurate predictions of geometric molecular relationships. We comprehensively assessed pose prediction accuracy by curating nearly 400,000 bound ligand pose pairs across the DUD-E targets. Overall, beginning from agnostic initial poses, we observed an 80% success rate for RMSD $$\le 2.0$$ Å among the top 20 predicted eSim poses. These examples were split roughly 50/50 into cases with high direct atomic overlap (where a shared scaffold exists between a pair) and low direct atomic overlap (where where a ligand pair has dissimilar scaffolds but largely occupies the same space). Within the high direct atomic overlap subset, the pose prediction success rate was 93%. For the more challenging subset (where dissimilar scaffolds are to be aligned), the success rate was 70%. The eSim approach enables both large-scale screening and rational design of ligands and is rooted in physically meaningful, non-heuristic, molecular comparisons.
ForceGen is a template-free, non-stochastic approach for 2D to 3D structure generation and conformational elaboration for small molecules, including both non-macrocycles and macrocycles. For conformational search of non-macrocycles, ForceGen is both faster and more accurate than the best of all tested methods on a very large, independently curated benchmark of 2859 PDB ligands. In this study, the primary results are on macrocycles, including results for 431 unique examples from four separate benchmarks. These include complex peptide and peptide-like cases that can form networks of internal hydrogen bonds. By making use of new physical movements (“flips” of near-linear sub-cycles and explicit formation of hydrogen bonds), ForceGen exhibited statistically significantly better performance for overall RMS deviation from experimental coordinates than all other approaches. The algorithmic approach offers natural parallelization across multiple computing-cores. On a modest multi-core workstation, for all but the most complex macrocycles, median wall-clock times were generally under a minute in fast search mode and under 2 min using thorough search. On the most complex cases (roughly cyclic decapeptides and larger) explicit exploration of likely hydrogen bonding networks yielded marked improvements, but with calculation times increasing to several minutes and in some cases to roughly an hour for fast search. In complex cases, utilization of NMR data to constrain conformational search produces accurate conformational ensembles representative of solution state macrocycle behavior. On macrocycles of typical complexity (up to 21 rotatable macrocyclic and exocyclic bonds), design-focused macrocycle optimization can be practically supported by computational chemistry at interactive time-scales, with conformational ensemble accuracy equaling what is seen with non-macrocyclic ligands. For more complex macrocycles, inclusion of sparse biophysical data is a helpful adjunct to computation.
We introduce the QuanSA method for inducing physically meaningful field-based models of ligand binding pockets based on structure-activity data alone. The method is closely related to the QMOD approach, substituting a learned scoring field for a pocket constructed of molecular fragments. The problem of mutual ligand alignment is addressed in a general way, and optimal model parameters and ligand poses are identified through multiple-instance machine learning. We provide algorithmic details along with performance results on sixteen structure-activity data sets covering many pharmaceutically relevant targets. In particular, we show how models initially induced from small data sets can extrapolatively identify potent new ligands with novel underlying scaffolds with very high specificity. Further, we show that combining predictions from QuanSA models with those from physics-based simulation approaches is synergistic. QuanSA predictions yield binding affinities, explicit estimates of ligand strain, associated ligand pose families, and estimates of structural novelty and confidence. The method is applicable for fine-grained lead optimization as well as potent new lead identification.
We introduce the ForceGen method for 3D structure generation and conformer elaboration of drug-like small molecules. ForceGen is novel, avoiding use of distance geometry, molecular templates, or simulation-oriented stochastic sampling. The method is primarily driven by the molecular force field, implemented using an extension of MMFF94s and a partial charge estimator based on electronegativity-equalization. The force field is coupled to algorithms for direct sampling of realistic physical movements made by small molecules. Results are presented on a standard benchmark from the Cambridge Crystallographic Database of 480 drug-like small molecules, including full structure generation from SMILES strings. Reproduction of protein-bound crystallographic ligand poses is demonstrated on four carefully curated data sets: the ConfGen Set (667 ligands), the PINC cross-docking benchmark (1062 ligands), a large set of macrocyclic ligands (182 total with typical ring sizes of 12–23 atoms), and a commonly used benchmark for evaluating macrocycle conformer generation (30 ligands total). Results compare favorably to alternative methods, and performance on macrocyclic compounds approaches that observed on non-macrocycles while yielding a roughly 100-fold speed improvement over alternative MD-based methods with comparable performance.