Given proteins' fundamental importance in human health and catalysis, the relationships between protein sequence, structure, dynamics, and function have become a topic of great interest. One way to extract information from proteins is to compute the local energetic frustration of their native state. Traditionally, energetic frustration calculations require protein structures as a starting point. However, using a single protein structure to evaluate the energetic frustration for a given amino acid sequence does not always fully represent the protein's structural ensemble. Therefore, we have developed a sequence-based method to evaluate energetic frustration in proteins using direct coupling analysis and statistical potentials. Our approach exhibits significant agreement with established structure-based frustration methods in terms of their mutual agreement with crystallographic B-factor. Moreover, our sequence-based method shows elevated precision in classifying high B-factor residues, suggesting that it has some robustness to unstructured regions of proteins.
Determining protein structures at an atomic level remains a significant challenge in structural biology. We introduce RecCrysFormer, a hybrid model that exploits the strengths of transformers with the aim of integrating experimental and ML approaches to protein structure determination from crystallographic data. RecCrysFormer leverages Patterson maps and incorporates known standardized partial structures of amino acid residues to directly predict electron density maps, which are essential for constructing detailed atomic models through crystallographic refinement processes. RecCrysFormer benefits from a "recycling" training regimen that iteratively incorporates results from crystallographic refinements and previous training runs as additional inputs in the formof template maps. Using a preliminary dataset of synthetic peptide fragments based on Protein Data Bank, RecCrysFormer achieves good accuracy in structural predictions and shows robustness against variations in crystal parameters, such as unit cell dimensions and angles.
The photoreaction and commensurate structural changes of a chromophore within biological photoreceptors elicit conformational transitions of the protein promoting the switch between deactivated and activated states. We investigated how this coupling is achieved in a bacterial phytochrome variant, Agp2-PAiRFP2. Contrary to classical protein crystallography, which only allows probing (cryo-trapped) stable states, we have used time-resolved serial femtosecond x-ray crystallography (tr-SFX) and pump-probe techniques with various illumination and delay times with respect to photoexcitation of the parent Pfr state. Thus, structural data for seven time frames were sorted into groups of molecular events along the reaction coordinate. They range from chromophore isomerization to the formation of Meta-F, the intermediate that precedes the functional relevant secondary structure transition of the tongue. Structural data for the early events were used to calculate the photoisomerization pathway to complement the experimental data. Late events allow identifying the molecular switch that is linked to the intramolecular proton transfer as a prerequisite for the following structural transitions.
About 100 years ago, the field of structural biology was born, led by James B. Sumner who recognized that enzymes were molecules with specific functions. In its contemporary form structural biology is used to interpret and understand molecular and cellular function, to design drugs, and to advance biotechnology in general.
Protein structure determination has long been one of the primary challenges of structural biology, to which deep machine learning (ML)-based approaches have increasingly been applied. However, these ML models generally do not directly incorporate the experimental measurements, such as X-ray crystallographic diffraction data. To this end, we explore an approach that more tightly couples these traditional crystallographic and recent ML-based methods by training a hybrid 3D vision transformer and convolutional network on inputs from both domains. We make use of two distinct input constructs: Patterson maps, which are directly obtainable from crystallographic data, and `partial structure' template maps derived from predicted structures deposited in the AlphaFold Protein Structure Database with subsequently omitted residues. With these, we predict electron-density maps that are then post-processed into atomic models through standard crystallographic refinement processes. Introducing an initial data set of small protein fragments taken from Protein Data Bank entries and placing them in hypothetical crystal settings, we demonstrate that our method is effective at both improving the phases of the crystallographic structure factors and completing the regions missing from partial structure templates, as well as improving the agreement of the electron-density maps with the ground-truth atomic structures.
Structural and functional studies of the carminomycin 4-O-methyltransferase DnrK are described, with an emphasis on interrogating the acceptor substrate scope of DnrK. Specifically, the evaluation of 100 structurally and functionally diverse natural products and natural product mimetics revealed an array of pharmacophores as productive DnrK substrates. Representative newly identified DnrK substrates from this study included anthracyclines, angucyclines, anthraquinone-fused enediynes, flavonoids, pyranonaphthoquinones, and polyketides. The ligand-bound structure of DnrK bound to a non-native fluorescent hydroxycoumarin acceptor, 4-methylumbelliferone, along with corresponding DnrK kinetic parameters for 4-methylumbelliferone and native acceptor carminomycin are also reported for the first time. The demonstrated unique permissivity of DnrK highlights the potential for DnrK as a new tool in future biocatalytic and/or strain engineering applications. In addition, the comparative bioactivity assessment (cancer cell line cytotoxicity, 4E-BP1 phosphorylation, and axolotl embryo tail regeneration) of a select set of DnrK substrates/products highlights the ability of anthracycline 4-O-methylation to dictate diverse functional outcomes.
Determining the atomic-level structure of a protein has been a decades-long challenge. However, recent advances in transformers and related neural network architectures have enabled researchers to significantly improve solutions to this problem. These methods use large datasets of sequence information and corresponding known protein template structures, if available. Yet, such methods only focus on sequence information. Other available prior knowledge could also be utilized, such as constructs derived from x-ray crystallography experiments and the known structures of the most common conformations of amino acid residues, which we refer to as partial structures. To the best of our knowledge, we propose the first transformer-based model that directly utilizes experimental protein crystallographic data and partial structure information to calculate electron density maps of proteins. In particular, we use Patterson maps, which can be directly obtained from x-ray crystallography experimental data, thus bypassing the well-known crystallographic phase problem. We demonstrate that our method, CrysFormer, achieves precise predictions on two synthetic datasets of peptide fragments in crystalline forms, one with two residues per unit cell and the other with fifteen. These predictions can then be used to generate accurate atomic models using established crystallographic refinement programs.
The most abundant natural collagens form heterotrimeric triple helices. Synthetic mimics of collagen heterotrimers have been found to fold slowly, even compared to the already slow rates of homotrimeric helices. These prolonged folding rates are not understood. Here we compare the stabilities, specificities and folding rates of three heterotrimeric collagen mimics designed through a computationally assisted approach. The crystal structure of one ABC-type heterotrimer verified a well-controlled composition and register and elucidated the geometry of pairwise cation-π and axial and lateral salt bridges in the assembly. This collagen heterotrimer folds much faster (hours versus days) than comparable, well-designed systems. Circular dichroism and NMR data suggest the folding is frustrated by unproductive, competing heterotrimer species and these species must unwind before refolding into the thermodynamically favoured assembly. The heterotrimeric collagen folding rate is inhibited by the introduction of preformed competing triple-helical assemblies, which suggests that slow heterotrimer folding kinetics are dominated by the frustration of the energy landscape caused by competing triple helices.
Adenylate kinase is a ubiquitous enzyme in living systems and undergoes dramatic conformational changes during its catalytic cycle. For these reasons, it is widely studied by genetic, biochemical, and biophysical methods, both experimental and theoretical. We have determined the basic crystal structures of three differently liganded states of adenylate kinase from Methanotorrus igneus, a hyperthermophilic organism whose adenylate kinase is a homotrimeric oligomer. The multiple copies of each protomer in the asymmetric unit of the crystal provide a unique opportunity to study the variation in the structure and were further analyzed using advanced crystallographic refinement methods and analysis tools to reveal conformational heterogeneity and, thus, implied dynamic behaviors in the catalytic cycle.
Determining the structure of a protein has been a decades-long open question. A protein's three-dimensional structure often poses nontrivial computation costs, when classical simulation algorithms are utilized. Advances in the transformer neural network architecture -- such as AlphaFold2 -- achieve significant improvements for this problem, by learning from a large dataset of sequence information and corresponding protein structures. Yet, such methods only focus on sequence information; other available prior knowledge, such as protein crystallography and partial structure of amino acids, could be potentially utilized. To the best of our knowledge, we propose the first transformer-based model that directly utilizes protein crystallography and partial structure information to predict the electron density maps of proteins. Via two new datasets of peptide fragments (2-residue and 15-residue) , we demonstrate our method, dubbed \texttt{CrysFormer}, can achieve accurate predictions, based on a much smaller dataset size and with reduced computation costs.
Views Icon Views Article contents Figures & tables Video Audio Supplementary Data Peer Review Share Icon Share Twitter Facebook Reddit LinkedIn Tools Icon Tools Reprints and Permissions Cite Icon Cite Search Site Citation George N. Phillips; And now some updates for SDY readers from the Editor….. Struct Dyn 1 May 2023; 10 (3): 030401. https://doi.org/10.1063/4.0000198 Download citation file: Ris (Zotero) Reference Manager EasyBib Bookends Mendeley Papers EndNote RefWorks BibTex toolbar search Search Dropdown Menu toolbar search search input Search input auto suggest filter your search All ContentAmerican Crystallographic Association IncStructural Dynamics Search Advanced Search |Citation Search
We present an in-depth analysis of selected CASP15 targets, focusing on their biological and functional significance. The authors of the structures identify and discuss key protein features and evaluate how effectively these aspects were captured in the submitted predictions. While the overall ability to predict three-dimensional protein structures continues to impress, reproducing uncommon features not previously observed in experimental structures is still a challenge. Furthermore, instances with conformational flexibility and large multimeric complexes highlight the need for novel scoring strategies to better emphasize biologically relevant structural regions. Looking ahead, closer integration of computational and experimental techniques will play a key role in determining the next challenges to be unraveled in the field of structural molecular biology.
The enediynes are structurally characterized by a 1,5-diyne-3-ene motif within a 9- or 10-membered enediyne core. The anthraquinone-fused enediynes (AFEs) are a subclass of 10-membered enediynes that contain an anthraquinone moiety fused to the enediyne core as exemplified by dynemicins and tiancimycins. A conserved iterative type I polyketide synthase (PKSE) is known to initiate the biosynthesis of all enediyne cores, and evidence has recently been reported to suggest that the anthraquinone moiety also originates from the PKSE product. However, the identity of the PKSE product that is converted to the enediyne core or anthraquinone moiety has not been established. Here, we report the utilization of recombinant E. coli coexpressing various combinations of genes that encode a PKSE and a thioesterase (TE) from either 9- or 10-membered enediyne biosynthetic gene clusters to chemically complement Δ PKSE mutant strains of the producers of dynemicins and tiancimycins. Additionally, 13 C-labeling experiments were performed to track the fate of the PKSE/TE product in the Δ PKSE mutants. These studies reveal that 1,3,5,7,9,11,13-pentadecaheptaene is the nascent, discrete product of the PKSE/TE that is converted to the enediyne core. Furthermore, a second molecule of 1,3,5,7,9,11,13-pentadecaheptaene is demonstrated to serve as the precursor of the anthraquinone moiety. The results establish a unified biosynthetic paradigm for AFEs, solidify an unprecedented biosynthetic logic for aromatic polyketides, and have implications for the biosynthesis of not only AFEs but all enediynes.
Natural products are a valuable source of pharmaceuticals, providing a majority of the small-molecule drugs in use today. However, their production through organic synthesis or in heterologous hosts can be difficult and time-consuming. Therefore, to allow for easier screening and production of natural products, we demonstrated the use of a cell-free protein synthesis system to partially assemble natural products in vitro using S-Adenosyl Methionine (SAM)-dependent methyltransferase enzyme reactions. The tea caffeine synthase, TCS1, was utilized to synthesize caffeine within a cell-free protein synthesis system. Cell-free systems also provide the benefit of allowing the use of substrates that would normally be toxic in a cellular environment to synthesize novel products. However, TCS1 is unable to utilize a compound like S-adenosyl ethionine as a cofactor to create ethylated caffeine analogs. The automation and reduced metabolic engineering requirements of cell-free protein synthesis systems, in combination with other synthesis methods, may enable the more efficient generation of new compounds. Graphical Abstract.
The general de novo solution of the crystallographic phase problem is difficult and only possible under certain conditions. This paper develops an initial pathway to a deep learning neural network approach for the phase problem in protein crystallography, based on a synthetic dataset of small fragments derived from a large well curated subset of solved structures in the Protein Data Bank (PDB). In particular, electron-density estimates of simple artificial systems are produced directly from corresponding Patterson maps using a convolutional neural network architecture as a proof of concept.
For decades, researchers have been determined to elucidate essential enzymatic functions on the atomic lengths scale by tracing atomic positions in real time. Our work builds on new possibilities unleashed by mix-and-inject serial crystallography (MISC) 1-5 at X-ray free electron laser facilities. In this approach, enzymatic reactions are triggered by mixing substrate or ligand solutions with enzyme microcrystals 6 . Here, we report in atomic detail and with millisecond time-resolution how the Mycobacterium tuberculosis enzyme BlaC is inhibited by sulbactam (SUB). Our results reveal ligand binding heterogeneity, ligand gating 7-9 , cooperativity, induced fit 10,11 and conformational selection 11-13 all from the same set of MISC data, detailing how SUB approaches the catalytic clefts and binds to the enzyme non-covalently before reacting to a trans- enamine. This was made possible in part by the application of the singular value decomposition 14 to the MISC data using a newly developed program that remains functional even if unit cell parameters change during the reaction.
The Diels-Alder cycloaddition is one of the most powerful approaches in organic synthesis and is often used in the synthesis of important pharmaceuticals. Yet, strictly controlling the stereoselectivity of the Diels-Alder reactions is challenging, and great efforts are needed to construct complex molecules with desired chirality via organocatalysis or transition-metal strategies. Nature has evolved different types of enzymes to exquisitely control cyclization stereochemistry; however, most of the reported Diels-Alderases have been shown to only facilitate the energetically favourable diastereoselective cycloadditions. Here we report the discovery and characterization of CtdP, a member of a new class of bifunctional oxidoreductase/Diels-Alderase, which was previously annotated as an NmrA-like transcriptional regulator. We demonstrate that CtdP catalyses the inherently disfavoured cycloaddition to form the bicyclo[2.2.2]diazaoctane scaffold with a strict α-anti-selectivity. Guided by computational studies, we reveal a NADP+/NADPH-dependent redox mechanism for the CtdP-catalysed inverse electron demand Diels-Alder cycloaddition, which serves as the first example of a bifunctional Diels-Alderase that utilizes this mechanism.
Marformycins are anti-infective natural products isolated from a deep sea sediment-derived Streptomyces drozdowiczii strain.These cyclodepsipetides contain O-methyl-D-Tyr.Liu et al., (Org.Lett.2015, 17, 1509-1512) identified a SAM-dependent O-methyltransferase, MfnG, in the marformycins biosynthetic gene cluster and found it capable of methylating the phenoic oxygen of both D-Tyr and L-Tyr in vitro.To better understand this enzyme's structural recognition and function, we have determined the MfnG structure using X-ray crystallography.Despite adding S-adenosyl-L-methionine (SAM/AdoMet) to the protein during crystallization, we found the spent product, S-Adenosyl-L-homocysteine (SAH/AdoHcy), bound.Since the SAH is unreactive, we were able to soak in L-Tyrosine to obtain a structure with the methyl doner product (SAH) and a methyl acceptor substrate (L-Tyr).We found MfnG could crystalize from a number of different screening conditions and that these crystals had different unit cell parameters.To date, we have phased 5 forms (2 forms in P212121forms in P21 and a P1 form), which contain one to four dimers (2-8 protomers) per asymmetric unit.Here we compare the packing arrangement in these different crystal packing forms.
Dynemicin is an enediyne natural product from Micromonospora chersina ATCC53710. Access to the biosynthetic gene cluster of dynemicin has enabled the in vitro study of gene products within the cluster to decipher their roles in assembling this unique molecule. This paper reports the crystal structure of DynF, the gene product of one of the genes within the biosynthetic gene cluster of dynemicin. DynF is revealed to be a dimeric eight-stranded β-barrel structure with palmitic acid bound within a cavity. The presence of palmitic acid suggests that DynF may be involved in binding the precursor polyene heptaene, which is central to the synthesis of the ten-membered ring of the enediyne core.