Our machine learning methodology, partial order optimum likelihood (POOL) is used to predict biochemically active amino acids in the three-dimensional structures of proteins. Computed electrostatic and chemical properties of individual amino acids serve as input features. Our most recent applications of POOL are described. From predicted local sites of biochemical activity, the biochemical functions of structural genomics proteins of unknown function are predicted by local structure matching of predicted spatial arrays of active amino acids with those of proteins of known function.
Enzymes catalyze reactions under mild conditions that might otherwise require extreme conditions such as high temperature or strong acid or base. To achieve this, amino acid side chains that are weak Brønsted acids or bases in the free amino acid become strong acids and bases in the active sites of many enzymes. Here we report on a set of 30 enzymes that represent all six major EC classes and a variety of different folds for which experimental studies of the mechanistic roles of the catalytic residues have been reported in the literature. Using macromolecular electrostatics techniques, it is shown that the catalytic aspartate and glutamate residues are strongly coupled to at least one other aspartate or glutamate residue, and often to multiple other carboxylate residues, with intrinsic pKa differences less than 1 pH unit. Sometimes these catalytic acidic residues are also coupled to a histidine residue, such that the intrinsic pKa of the acidic residue is higher than that of the histidine. Most catalytic lysine residues studied here are strongly coupled to tyrosine or cysteine residues, wherein the intrinsic pKa of the anion-forming residue is higher than that of the lysine. Some catalytic lysines are also coupled to other lysines with intrinsic pKa differences within 1 pH unit. Some general principles about the design of enzyme active sites are presented. These principles have significant implications for the engineering of novel enzymes. Supported by NSF CHE-1905214 (MJO).
The COVID-19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small-molecule drugs that are widely available, including in low- and middle-income countries, is an ongoing challenge. In this work, we report the results of a community effort, the “Billion molecules against Covid-19 challenge”, to identify small-molecule inhibitors against SARS-CoV-2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 potentially active molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (Nsp12 domain), and (alpha) spike protein S. Overall, 27 potential inhibitors were experimentally confirmed by binding-, cleavage-, and/or viral suppression assays and are presented here. All results are freely available and can be taken further downstream without IP restrictions. Overall, we show the effectiveness of computational techniques, community efforts, and communication across research fields (i.e., protein expression and crystallography, in silico modeling, synthesis and biological assays) to accelerate the early phases of drug discovery.
The COVID‐19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small‐molecule drugs that are widely available, including in low‐ and middle‐income countries, is an ongoing challenge. In this work, we report the results of an open science community effort, the “Billion molecules against COVID‐19 challenge”, to identify small‐molecule inhibitors against SARS‐CoV‐2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for biological activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (only the Nsp12 domain), and (alpha) spike protein S. Overall, 27 compounds with weak inhibition/binding were experimentally identified by binding‐, cleavage‐, and/or viral suppression assays and are presented here. Open science approaches such as the one presented here contribute to the knowledge base of future drug discovery efforts in finding better SARS‐CoV‐2 treatments.
The computed electrostatic and proton transfer properties are studied for 20 enzymes that represent all six major enzyme commission classes and a variety of different folds. The properties of aspartate, glutamate, and lysine residues that have been previously experimentally determined to be catalytically active are reported. The catalytic aspartate and glutamate residues studied here are strongly coupled to at least one other aspartate or glutamate residue and often to multiple other carboxylate residues with intrinsic pKa differences less than 1 pH unit. Sometimes these catalytic acidic residues are also coupled to a histidine residue, such that the intrinsic pKa of the acidic residue is higher than that of the histidine. All catalytic lysine residues studied here are strongly coupled to tyrosine or cysteine residues, wherein the intrinsic pKa of the anion-forming residue is higher than that of the lysine. Some catalytic lysines are also coupled to other lysines with intrinsic pKa differences within 1 pH unit. Some evidence of the possible types of interactions that facilitate nucleophilicity is discussed. The interactions reported here provide important clues about how side chain functional groups that are weak Brønsted acids or bases for the free amino acid in solution can achieve catalytic potency and become strong acids, bases or nucleophiles in the enzymatic environment.
Metabotropic glutamate receptor 2 (mGluR2) is a therapeutic target for several neuropsychiatric disorders. An mGluR2 function in etiology could be unveiled by positron emission tomography (PET). In this regard, 5-(2-fluoro-4-[11C]methoxyphenyl)-2,2-dimethyl-3,4-dihydro-2H-pyrano[2,3-b]pyridine-7-carboxamide ([11C]13, [11C]mG2N001), a potent negative allosteric modulator (NAM), was developed to support this endeavor. [11C]13 was synthesized via the O-[11C]methylation of phenol 24 with a high molar activity of 212 ± 76 GBq/μmol (n = 5) and excellent radiochemical purity (>99%). PET imaging of [11C]13 in rats demonstrated its superior brain heterogeneity and reduced accumulation with pretreatment of mGluR2 NAMs, VU6001966 (9) and MNI-137 (26), the extent of which revealed a time-dependent drug effect of the blocking agents. In a nonhuman primate, [11C]13 selectively accumulated in mGluR2-rich regions and resulted in high-contrast brain images. Therefore, [11C]13 is a potential candidate for translational PET imaging of the mGluR2 function.
An array of triazolopyridines based on JNJ-46356479 (6) were synthesized as potential positron emission tomography radiotracers for metabotropic glutamate receptor 2 (mGluR2). The selected candidates 8-10 featured enhanced positive allosteric modulator (PAM) activity (20-fold max.) and mGluR2 agonist activity (25-fold max.) compared to compound 6 in the cAMP GloSensor assays. Radiolabeling of compounds 8 and 9 (mG2P026) was achieved via Cu-mediated radiofluorination with satisfactory radiochemical yield, >5% (non-decay-corrected); high molar activity, >180 GBq/μmol; and excellent radiochemical purity, >98%. Preliminary characterization of [18F]8 and [18F]9 in rats confirmed their excellent brain permeability and binding kinetics. Further evaluation of [18F]9 in a non-human primate confirmed its superior brain heterogeneity in mapping mGluR2 and higher affinity than [18F]6. Pretreatment with different classes of PAMs in rats and a primate led to similarly enhanced brain uptake of [18F]9. As a selective ligand, [18F]9 has the potential to be developed for translational studies.
Three protein targets from SARS-CoV-2, the viral pathogen that causes COVID-19, are studied: the main protease, the 2'-O-RNA methyltransferase, and the nucleocapsid (N) protein. For the main protease, the nucleophilicity of the catalytic cysteine C145 is enabled by coupling to three histidine residues, H163 and H164 and catalytic dyad partner H41. These electrostatic couplings enable significant population of the deprotonated state of C145. For the RNA methyltransferase, the catalytic lysine K6968 that serves as a Brønsted base has significant population of its deprotonated state via strong coupling with K6844 and Y6845. For the main protease, Partial Order Optimum Likelihood (POOL) predicts two clusters of biochemically active residues; one includes the catalytic H41 and C145 and neighboring residues. The other surrounds a second pocket adjacent to the catalytic site and includes S1 residues F140, L141, H163, E166, and H172 and also S2 residue D187. This secondary recognition site could serve as an alternative target for the design of molecular probes. From in silico screening of library compounds, ligands with predicted affinity for the secondary site are reported. For the NSP16-NSP10 complex that comprises the RNA methyltransferase, three different sites are predicted. One is the catalytic core at the conserved K-D-K-E motif that includes catalytic residues D6928, K6968, and E7001 plus K6844. The second site surrounds the catalytic core and consists of Y6845, C6849, I6866, H6867, F6868, V6894, D6895, D6897, I6926, S6927, Y6930, and K6935. The third is located at the heterodimer interface. Ligands predicted to have high affinity for the first or second sites are reported. Three sites are also predicted for the nucleocapsid protein. This work uncovers key interactions that contribute to the function of the three viral proteins and also suggests alternative sites for ligand design.
Most of the 15,000+ reported Structural Genomics (SG) protein structures in the PDB are of unknown or uncertain biochemical function. Here we present a powerful approach based on computed chemical properties of the individual residues in a protein structure. Local arrays of predicted active residues for sets of proteins of known function are matched with those of SG proteins. For instance, a superfamily consists of proteins with similar structure but multiple different kinds of biochemical function, including different types of reactivity as well as different substrate specificities. Each functional type has a different set of active amino acids. Typically active residues include those in the first layer that make direct contact with the substrate molecule(s) and also some residues in the second and third layers that play supporting roles in the catalytic process. Graph Representation of Active Sites for Prediction of Function (GRASP-Func) establishes local arrays of predicted residues that are common to proteins of the same function. Predicted local arrays for each SG protein are aligned against those of the known members of each functional family. Cases of predicted misannotation, where our prediction differs from the originally assigned function, are especially interesting. Experimental testing is performed by direct biochemical assays to test our predictions. For instance, we show that our predictions are confirmed that RV0760c from Mycobacterium tuberculosis has ketosteroid isomerase activity, whereas NP_103587.1 from Mesorhizobium loti does not. Some cases involving local site matches across different fold types for glycoside hydrolases are also presented. Supported by NSF grant CHE-1905214.
The COVID‐19 pandemic, caused by the Severe Acute Respiratory Coronavirus 2 (SARS‐CoV‐2) virus, first started in the Wuhan region of Hubei, China, and has quickly spread to 191 countries and territories, infecting more than 86.4 million people, and resulting in 1.87 global deaths as of January 6th. With SARS‐CoV‐2’s genomic sequence and protein structures deciphered and updated rapidly, clinical treatments and vaccine developments have proceeded simultaneously as researchers attempt to learn more about the infecting mechanisms of this virus. Among these attempts, computational drug screening for SARS‐CoV‐2 has potential for: (1) narrowing down billions of chemical compounds into a list of possible high‐affinity ligands for SARS‐CoV‐2 protein targets, (2) providing information about the activities of SARS‐CoV‐2 proteins, (3) offering possible treatments, and (4) assisting in scientific knowledge to fight against future coronavirus infections. In this work, computational ligand screening for SARS‐CoV‐2 is a combination of site prediction using machine learning technology Partial Order Optimum Likelihood (POOL) and molecular docking. Among the techniques deployed, the machine learning technology POOL was developed by us and assists in the drug screening process for SARS‐CoV‐2 by predicting targeted protein sites, including those that are not the obvious catalytic sites, such as exosites, allosteric sites, and other interaction sites. Results will be presented for the SARS‐CoV‐2 main protease, non‐structural protein 1 (Nsp1), non‐structural protein 9 (Nsp9), and non‐structural protein 15 (Nsp15). Compounds are taken from a variety of libraries, including the ZINC and Enamine databases. Protein structures are downloaded from the Protein Data Bank (www.rcsb.org). Molecular dynamic structure simulations are used to generate structures for ensemble docking.
The virus SARS-CoV-2, the cause of the current COVID-19 pandemic, is not well understood. It is critical to understand how the viral proteins function and how their function may be modulated. Inhibitors that target these enzymes serve as potential therapeutic interventions against COVID-19. This work uses artificial intelligence methods developed by us to find sites that other methods may not find and therefore, identify potential exosites, allosteric sites, or other sites of interaction in the structures of viral proteins to serve as new targets for the development of antiviral agents. Large datasets of natural and synthetic compounds are computationally searched for molecules that fit into these alternative sites, and any compounds that fit will be targeted for experimental testing for their ability to inhibit the functions of these viral enzymes. This project uses the unique Partial Order Optimum Likelihood (POOL) machine learning method developed by us to predict multiple types of binding sites in SARS-CoV-2 proteins, including catalytic sites, allosteric sites, and other interaction sites. Molecular dynamics simulations are used to generate conformations for ensemble docking. Compounds from large molecular libraries are computationally docked into the predicted sites to identify potentially strong binding ligands. We have identified approximately 10000 potential ligands for more than 50 SARS-CoV-2 proteins to date. Candidate ligands to selected SARS-CoV-2 proteins are experimentally tested in vitro for binding affinity and the effect of the best-predicted inhibitors on catalytic activities determined by direct biochemical assays. Compound libraries for the study include selected compounds from the ZINC and Enamine databases; Chemical Abstract Service database compounds and COVID-specific libraries from Enamine and Life Chemicals.
In the current COVID-19 pandemic, it is critical to understand, as swiftly as possible, how the viral proteins function and how their function might be modulated. The machine learning method Partial Order Optimum Likelihood (POOL) is used to predict binding sites in protein structures from SARS-CoV-2, the virus that causes COVID-19. Using the 3D structure of each protein as input, POOL uses computed electrostatic and chemical properties to predict the amino acids that are biochemically active, including residues in catalytic sites, allosteric sites, and other secondary sites. Docking studies are then performed to predict ligands that bind to each of these predicted sites. For instance, for the x-ray crystal structures of the main protease, POOL predicts two sites: the known catalytic site containing the catalytic dyad His41 and Cys145 and a second nearby site on an adjacent face of the protein surface. The x-ray crystal structure of the SARS-CoV-2 2'-O-ribose RNA methyltransferase (NSP16) protein has been reported in complex with its activating partner NSP10 and with two bound ligands, S-adenosylmethionine (SAM) and β-D-fructopyranose (BDF). POOL predicts three binding sites, including the catalytic SAM-binding site, the BDF binding site on the opposite side, and a third site adjacent to the catalytic / SAM-binding site. Predicted binding ligands (including selected compounds from the ZINC and Enamine databases, Chemical Abstract Service database compounds, and COVID-specific libraries from Enamine and Life Chemicals) are reported for several SARS-CoV-2 proteins. Kinetics assays to test for catalytic activity of the main protease and of 2'-O-ribose RNA methyltransferase in the presence of predicted binding ligands with high scores are underway. Theoretical and experimental methods are aimed at identifying molecules having inhibitory effects on the function of viral proteins. Supported by NSF CHE-2030180.
Three benzimidazole derivatives (13-15) have been synthetized as potential PET imaging ligands for mGluR2 in the brain. Of these compounds, 13 exhibits potent binding affinity (IC50 = 7.6 ± 0.9 nM), PAM activity (EC50 = 51.2 nM), and excellent selectivity against other mGluR subtypes (> 100-fold). [11C]13 was synthesized via O-[11C]methylation of its phenol precursor 25 with [11C]methyl iodide. The achieved radiochemical yield was 20 ± 2% (n = 10, decay-corrected) based on [C]CO2 with radiochemical purity > 98% and molar activity 98±30 GBq/μmol EOS. Ex vivo biodistribution studies revealed reversible accumulation of [11C]13 and hepatobiliary and urinary excretions. PET imaging studies in rats demonstrated that [11C]13 accumulated in the mGluR2-rich brain regions. Pre-administration of mGluR2-selective PAM, 17 reduced the brain uptake of [11C]13, indicating a selective binding. However, pre-administration of 13 significantly enhanced [11C]13 uptake in the brain. Therefore, [11C]13 is both a potential PET imaging ligand for mGluR2 and a drug candidate for the treatment of CNS disorders. Corresponding Authors: A-L.B.: phone, 617-726-3709; abrownell@mgh.harvard.edu, ZZ.: phone, 617-643-4887; zzhang@nmr.mgh.harvard.edu. #Author Contributions G.Y. and X.Q. contributed equally. The manuscript was written through contributions of all authors. All authors have given approval to the final version of the manuscript. ASSOCIATED CONTENT Supporting Information The detailed in vitro assays and molecular modeling work are described in the Supporting Information. The authors declare no conflict interest. HHS Public Access Author manuscript J Med Chem. Author manuscript; available in PMC 2021 November 29. Published in final edited form as: J Med Chem. 2020 October 22; 63(20): 12060–12072. doi:10.1021/acs.jmedchem.0c01394. A uhor M anscript
Three benzimidazole derivatives (13-15) have been synthetized as potential positron emission tomography (PET) imaging ligands for mGluR2 in the brain. Of these compounds, 13 exhibits potent binding affinity (IC50 = 7.6 ± 0.9 nM), positive allosteric modulator (PAM) activity (EC50 = 51.2 nM), and excellent selectivity against other mGluR subtypes (>100-fold). [11C]13 was synthesized via O-[11C]methylation of its phenol precursor 25 with [11C]methyl iodide. The achieved radiochemical yield was 20 ± 2% (n = 10, decay-corrected) based on [11C]CO2 with a radiochemical purity of >98% and molar activity of 98 ± 30 GBq/μmol EOS. Ex vivo biodistribution studies revealed reversible accumulation of [11C]13 and hepatobiliary and urinary excretions. PET imaging studies in rats demonstrated that [11C]13 accumulated in the mGluR2-rich brain regions. Pre-administration of mGluR2-selective PAM, 17 reduced the brain uptake of [11C]13, indicating a selective binding. Therefore, [11C]13 is a potential PET imaging ligand for mGluR2 in different central nervous system-related conditions.
This work addresses the need for chemical tools that can selectively form cross-links. Contemporary thiol-selective cross-linkers, for example, modify all accessible thiols, but only form cross-links between a subset. The resulting terminal "dead-end" modifications of lone thiols are toxic, confound cross-linking-based studies of macromolecular structure, and are an undesired, and currently unavoidable, byproduct in polymer synthesis. Using the thiol pair of Cu/Zn-superoxide dismutase (SOD1), we demonstrated that cyclic disulfides, including the drug/nutritional supplement lipoic acid, efficiently cross-linked thiol pairs but avoided dead-end modifications. Thiolate-directed nucleophilic attack upon the cyclic disulfide resulted in thiol-disulfide exchange and ring cleavage. The resulting disulfide-tethered terminal thiolate moiety either directed the reverse reaction, releasing the cyclic disulfide, or participated in oxidative disulfide (cross-link) formation. We hypothesized, and confirmed with density functional theory (DFT) calculations, that mono- S-oxo derivatives of cyclic disulfides formed a terminal sulfenic acid upon ring cleavage that obviated the previously rate-limiting step, thiol oxidation, and accelerated the new rate-determining step, ring cleavage. Our calculations suggest that the origin of accelerated ring cleavage is improved frontier molecular orbital overlap in the thiolate-disulfide interchange transition. Five- to seven-membered cyclic thiosulfinates were synthesized and efficiently cross-linked up to 104-fold faster than their cyclic disulfide precursors; functioned in the presence of biological concentrations of glutathione; and acted as cell-permeable, potent, tolerable, intracellular cross-linkers. This new class of thiol cross-linkers exhibited click-like attributes including, high yields driven by the enthalpies of disulfide and water formation, orthogonality with common functional groups, water-compatibility, and ring strain-dependence.