The neural network-based program AlphaFold2 (AF2) provides high accuracy structure prediction for a large fraction of globular proteins. An important question is whether these models are accurate enough for reliably docking small ligands. Several recent papers and the results of CASP15 reveal that local conformational errors reduce the success rates of direct ligand docking. Here, we focus on the ability of the models to conserve the location of binding hot spots, regions on the protein surface that significantly contribute to the binding free energy of the protein-ligand interaction. Clusters of hot spots predict the location and even the druggability of binding sites, and hence are important for computational drug discovery. The hot spots are determined by protein mapping that is based on the distribution of small fragment-sized probes on the protein surface and is less sensitive to local conformation than docking. Mapping models taken from the AlphaFold Protein Structure Database show that identifying binding sites is more reliable than docking, but the success rates are still 5% to 10% lower than based on mapping X-ray structures. The drop in accuracy is particularly large for models of multidomain proteins. However, both the model binding sites and the mapping results can be substantially improved by generating AF2 models for the ligand binding domains of interest rather than the entire proteins and even more if using forced sampling with multiple initial seeds. The mapping of such models tends to reach the accuracy of results obtained by mapping the X-ray structures.
Acetylcholinesterase (AChE) is an important enzyme that hydrolyzes the neurotransmitter acetylcholine. Inhibiting AChE represents one of the therapeutic strategies for treating the symptoms of Alzheimer's disease (AD). To date, only several AD drugs have been approved by the U.S. Food and Drug Administration, with donepezil being the most commonly used. In this study, we synthesized 46 donepezil-based analogs via a onestep microwave-assisted synthetic route and evaluated them in vitro using an AChE inhibition assay in order to better understand structural requirements important for the inhibition. Our structure-activity relationship studies revealed that various halogens, electron-withdrawing, and electron-donating groups placed in the orthoand meta-positions on the phenyl moiety of donepezil are generally well-tolerated, yielding potent AChE inhibitors in the low nanomolar range similar to donepezil, while polysubstitutions led to moderate or significant decreases in inhibition potency.
In computational biology, accurate prediction of phosphopeptide-protein complex structures is essential for understanding cellular functions and advancing drug discovery and personalized medicine. While AlphaFold has significantly improved protein structure prediction, it faces accuracy challenges in predicting structures of complexes involving phosphopeptides possibly due to structural variations introduced by phosphorylation in the peptide component. Our study addresses this limitation by refining AlphaFold to improve its accuracy in modeling these complex structures. We employed weighted metrics for a comprehensive evaluation across various protein families. The enhanced model notably outperforms the original AlphaFold, showing a substantial increase in the weighted average local distance difference test (lDDT) scores for peptides: from 52.74 to 76.51 in the Top 1 model and from 56.32 to 77.91 in the Top 5 model. These advancements not only deepen our understanding of the role of phosphorylation in cellular signaling but also have extensive implications for biological research and the development of innovative therapies.### Competing Interest StatementThe authors have declared no competing interest.
The precise prediction of Major Histocompatibility Complex (MHC)-peptide complex structures is pivotal for understanding cellular immune responses and advancing vaccine design. In this study, we enhanced AlphaFold's capabilities by fine-tuning it with a specialized dataset comprised by exclusively high-resolution MHC-peptide crystal structures. This tailored approach aimed to address the generalist nature of AlphaFold's original training, which, while broad-ranging, lacked the granularity necessary for the high-precision demands of MHC-peptide interaction prediction. A comparative analysis was conducted against the homology-modeling-based method Pandora [13], as well as the AlphaFold multimer model [8]. Our results demonstrate that our fine-tuned model outperforms both in terms of RMSD (median value is 0.65 Å) but also provides enhanced predicted lDDT scores, offering a more reliable assessment of the predicted structures. These advances have substantial implications for computational immunology, potentially accelerating the development of novel therapeutics and vaccines by providing a more precise computational lens through which to view MHC-peptide interactions.
The goal of this paper is predicting the conformational distributions of ligand binding sites using the AlphaFold2 (AF2) protein structure prediction program with stochastic subsampling of the multiple sequence alignment (MSA). We explored the opening of cryptic ligand binding sites in 16 proteins, where the closed and open conformations define the expected extreme points of the conformational variation. Due to the many structures of these proteins in the Protein Data Bank (PDB), we were able to study whether the distribution of X-ray structures affects the distribution of AF2 models. We have found that AF2 generates both a cluster of open and a cluster of closed models for proteins that have comparable numbers of open and closed structures in the PDB and not too many other conformations. This was observed even with default MSA parameters, thus without further subsampling. In contrast, with the exception of a single protein, AF2 did not yield multiple clusters of conformations for proteins that had imbalanced numbers of open and closed structures in the PDB, or had substantial numbers of other structures. Subsampling improved the results only for a single protein, but very shallow MSA led to incorrect structures. The ability of generating both open and closed conformations for six out of the 16 proteins agrees with the success rates of similar studies reported in the literature. However, we showed that this partial success is due to AF2 “remembering” the conformational distributions in the PDB and that the approach fails to predict rarely seen conformations.
One often observes small but measurable differences in the diffraction data measured from different crystals of a single protein. These differences might reflect structural differences in the protein and may reveal the natural dynamism of the molecule in solution. Partitioning these mixed-state data into single-state clusters is a critical step that could extract information about the dynamic behavior of proteins from hundreds or thousands of single-crystal data sets. Mixed-state data can be obtained deliberately (through intentional perturbation) or inadvertently (while attempting to measure highly redundant single-crystal data). To the extent that different states adopt different molecular structures, one expects to observe differences in the crystals; each of the polystates will create a polymorph of the crystals. After mixed-state diffraction data have been measured, deliberately or inadvertently, the challenge is to sort the data into clusters that may represent relevant biological polystates. Here, this problem is addressed using a simple multi-factor clustering approach that classifies each data set using independent observables, thereby assigning each data set to the correct location in conformational space. This procedure is illustrated using two independent observables, unit-cell parameters and intensities, to cluster mixed-state data from chymotrypsinogen (ChTg) crystals. It is observed that the data populate an arc of the reaction trajectory as ChTg is converted into chymotrypsin.
Carbapenem antibiotics are the drugs of choice for treatment of deadly infections caused by Gram-negative bacteria. However, their efficacy is severely compromised by the wide spread of carbapenem-hydrolyzing class D β-lactamases (CHDLs).
An important question is how well the models submitted to CASP retain the properties of target structures. We investigate several properties related to binding. First we explore the binding of small molecules as probes, and count the number of interactions between each residue and such probes, resulting in a binding fingerprint. The similarity between two fingerprints, one for the X‐ray structure and the other for a model, is determined by calculating their correlation coefficient. The fingerprint similarity weakly correlates with global measures of accuracy, and GDT_TS higher than 80 is a necessary but not sufficient condition for the conservation of surface binding properties. The advantage of this approach is that it can be carried out without information on potential ligands and their binding sites. The latter information was available for a few targets, and we explored whether the CASP14 models can be used to predict binding sites and to dock small ligands. Finally, we tested the ability of models to reproduce protein–protein interactions by docking both the X‐ray structures and the models to their interaction partners in complexes. The analysis showed that in CASP14 the quality of individual domain models is approaching that offered by X‐ray crystallography, and hence such models can be successfully used for the identification of binding and regulatory sites, as well as for assembling obligatory protein–protein complexes. Success of ligand docking, however, often depends on fine details of the binding interface, and thus may require accounting for conformational changes by simulation methods.
Objective: To investigate the osteoblastogenic activity of the ethyl acetate (EtOAc) extract of Smilax glabra Roxb roots and its major active compound astilbin. Methods: Astilbin was isolated from EtOAc extract using silica gel chromatography combined with fraction crystallization. Chemical structure of astilbin was determined by analysis of the spectroscopic data in comparison with the literature. MTT method was used to detect the toxicity. Alkaline phosphatase (ALP) activity was determined by the spectrophotometric method at 405 nm using p-nitrophenyl phosphate as a substrate. Calcium deposition was stained with alizarin red-S, distained with cetylpyridium chloride, and quantified at 562 nm. In silico model for astilbin-ALP interaction was analyzed using AutoDock 4.2.6. The changes in expression of osteoblast differentiation related genes were determined using quantitative real-time PCR. Results: Both the EtOAc extract and astilbin had no toxicity toward osteoblast MC3T3-E1 cells at 5.0, 10, 25, and 50 μg/mL. At 25 μg/ mL, they enhanced ALP activity and mineralization of osteoblasts up to 30% and 55% for the EtOAc extract and 22% and 41% for astilbin, respectively. Molecular docking analysis of astilbin-ALP interaction revealed Arg167, Asp320, His324, and His437 were key residues participating in hydrophobic interaction; meanwhile, His434 and Thr436 residues were involved in hydrogen bond formation in the active site of human tissue-nonspecific ALP. Moreover, the expression level of genes opn, col1, osx, and runx2 were up-regulated in astilbin treated samples with the fold changes as 2.2; 3.7; 4.1; 2.3, respectively at 10 μg/mL (P<0.05). Conclusions: The EtOAc extract and its major compound astilbin exhibit osteoblastogenic activity by up-regulating important markers for bone cell differentiation. It could be a new and promising osteogenic agent with dual actions for therapeutic applications.
Commercial carbapenem antibiotics are being used to treat multidrug resistant (MDR) and extensively drug resistant (XDR) tuberculosis. Like other β-lactams, carbapenems are irreversible inhibitors of serine d,d-transpeptidases involved in peptidoglycan biosynthesis. In addition to d,d-transpeptidases, mycobacteria also utilize nonhomologous cysteine l,d-transpeptidases (Ldts) to cross-link the stem peptides of peptidoglycan, and carbapenems form long-lived acyl-enzymes with Ldts. Commercial carbapenems are C2 modifications of a common scaffold. This study describes the synthesis of a series of atypical, C5α modifications of the carbapenem scaffold, microbiological evaluation against Mycobacterium tuberculosis (Mtb) and the nontuberculous mycobacterial species, Mycobacterium abscessus (Mab), as well as acylation of an important mycobacterial target Ldt, LdtMt2. In vitro evaluation of these C5α-modified carbapenems revealed compounds with standalone (i.e., in the absence of a β-lactamase inhibitor) minimum inhibitory concentrations (MICs) superior to meropenem-clavulanate for Mtb, and meropenem-avibactam for Mab. Time-kill kinetics assays showed better killing (2–4 log decrease) of Mtb and Mab with lower concentrations of compound 10a as compared to meropenem. Although susceptibility of clinical isolates to meropenem varied by nearly 100-fold, 10a maintained excellent activity against all Mtb and Mab strains. High resolution mass spectrometry revealed that 10a acylates LdtMt2 at a rate comparable to meropenem, but subsequently undergoes an unprecedented carbapenem fragmentation, leading to an acyl-enzyme with mass of Δm = +86 Da. Rationale for the divergence of the nonhydrolytic fragmentation of the LdtMt2 acyl-enzymes is proposed. The observed activity illustrates the potential of novel atypical carbapenems as prospective candidates for treatment of Mtb and Mab infections.
We study the models submitted to round 12 of the Critical Assessment of protein Structure Prediction (CASP) experiment to assess how well the binding properties are conserved when the X-ray structures of the target proteins are replaced by their models. To explore small molecule binding we generate distributions of molecular probes – which are fragment-sized organic molecules of varying size, shape, and polarity – around the protein, and count the number of interactions between each residue and the probes, resulting in a vector of interactions we call a binding fingerprint. The similarity between two fingerprints, one for the X-ray structure and the other for a model of the protein, is determined by calculating the correlation coefficient between the two vectors. The resulting correlation coefficients are shown to correlate with global measures of accuracy established in CASP, and the relationship yields an accuracy threshold that has to be reached for meaningful binding surface conservation. The clusters formed by the probe molecules reliably predict binding hot spots and ligand binding sites in both X-ray structures and reasonably accurate models of the target, but ensembles of models may be needed for assessing the availability of proper binding pockets. We explored ligand docking to the few targets that had bound ligands in the X-ray structure. More targets were available to assess the ability of the models to reproduce protein–protein interactions by docking both the X-ray structures and models to their interaction partners in complexes. It was shown that this application is more difficult than finding small ligand binding sites, and the success rates heavily depend on the local structure in the potential interface. In particular, predicted conformations of flexible loops are frequently incorrect in otherwise highly accurate models, and may prevent predicting correct protein–protein interactions.
In macromolecular crystallography, higher flux, smaller beams, and faster detectors open the door to experiments with very large numbers of very small samples that can reveal polymorphs and dynamics but require re-engineering of approaches to the clustering of images both at synchrotrons and XFELs (X-ray free electron lasers). The need for the management of orders of magnitude more images and limitations of file systems favor a transition from simple one-file-per-image systems such as CBF to image container systems such as HDF5. This further increases the load on computers and networks and requires a re-examination of the presentation of metadata. In this paper, we discuss three important components of this problem-improved approaches to the clustering of images to better support experiments on polymorphs and dynamics, recent and upcoming changes in metadata for Eiger images, and software to rapidly validate images in the revised Eiger format.
Although α-diazo-β-ketoesters are synthetically versatile intermediates, methodology for introducing this functionality into complex molecules is still limited, most frequently involving a carboxylic acid precursor, which is then activated and transformed into a β-ketoester, with the diazo group being subsequently added with a diazo transfer reagent. While introducing this highly functional moiety in a convergent one step process would be ideal, such an objective is limited by the relatively few studies which address functionalization of the α-diazo-β-ketoester at the γ-position. In the present investigation, we evaluate strategies, both new and established, for functionalizing α-diazo-β-ketoesters, particularly with regard to generating compounds prospectively useful in the synthesis of C1-substituted carbapenems. We report the first δ-aldehydo-α-diazo-β-ketoester as well as a method for its oxidation to the corresponding methyl ester, and the formation of a new substituted pyrazole under basic conditions.