We report the discovery of a new class of local minima that has severely limited the accuracy of macromolecular models. Termed density misfit barrier traps, these minima explain much of the poor fit between macromolecular models and experimental data relative to that of smaller molecules: not just high R factors, but distorted chemical geometry. We postulated that proteins exist as an ensemble of conformations that each have good geometry, but refinement algorithms have been unable to converge to them due to a tangling phenomenon arising from these traps. To demonstrate, a synthetic ground truth data set was generated, consisting of a 2-member ensemble with excellent geometry. A series of starting models, each trapped in increasingly difficult local minima, were prepared, a unified validation score defined, and an open Challenge issued. This Challenge inspired algorithms for escaping such traps, and new programs have been released that are expected to substantially improve the accuracy of macromolecular ensemble models.
Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, key challenges persist in modeling side-chain packing, condition-dependent conformational changes and biomolecular interactions, largely because of limited high-quality training data. At the same time, emerging experimental techniques such as cryo-electron microscopy (cryo-EM), cryo-electron tomography (cryo-ET) and high-throughput crystallography are generating vast amounts of structural information but converting these data into mechanistically interpretable atomic models often remains difficult. Here we show that integrating experimental measurements directly into protein structure prediction can overcome these limitations. We introduce ROCKET, an augmentation of AlphaFold2 that refines predicted structures using cryo-EM, cryo-ET and X-ray crystallography data. By optimizing structures in the space of coevolutionary embeddings rather than Cartesian coordinates, ROCKET captures biologically meaningful structural variation that is inaccessible to AlphaFold2 alone and to existing automated modeling approaches, especially when the signal-to-noise ratio is low. ROCKET enables scalable, automated model building without retraining and provides a general framework for integrating experimental observables with biomolecular machine learning.
Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, challenges persist in capturing sidechain packing, condition-dependent conformational dynamics, and biomolecular interactions, primarily due to scarcity of high-quality training data. Emerging techniques, including cryo-electron tomography (cryo-ET) and high-throughput crystallography, promise vast new sources of structural data, but translating experimental observations into mechanistically interpretable atomic models remains a key bottleneck. Here, we address these challenges by improving the efficiency of structural analysis through combining experimental measurements with a landmark protein structure prediction method - AlphaFold2. We present an augmentation of AlphaFold2, ROCKET, that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and demonstrate that this approach captures biologically important structural variation that AlphaFold2 does not. By performing structure optimization in the space of coevolutionary embeddings, rather than Cartesian coordinates, ROCKET automates difficult modeling tasks, such as flips of functional loops and domain rearrangements, that are beyond the scope of current state-of-the-art methods and, in some instances, even manual human modeling. The ability to efficiently sample these barrier-crossing rearrangements unlocks a new horizon for scalable and automated model building. Crucially, ROCKET does not require retraining of AlphaFold2 and is readily adaptable to multimers, ligand-cofolding, and other data modalities. Conversely, our differentiable crystallographic and cryo-EM target functions are capable of augmenting other structure prediction methods. ROCKET thus provides an extensible framework for the integration of experimental observables with biomolecular machine learning.
Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, challenges persist in capturing sidechain packing, condition-dependent conformational dynamics, and biomolecular interactions, primarily due to scarcity of high-quality training data. Emerging techniques, including cryo-electron tomography (cryo-ET) and high-throughput crystallography, promise vast new sources of structural data, but translating experimental observations into mechanistically interpretable atomic models remains a key bottleneck. Here, we address these challenges by improving the efficiency of structural analysis through combining experimental measurements with a landmark protein structure prediction method – AlphaFold2. We present an augmentation of AlphaFold2, ROCKET, that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and demonstrate that this approach captures biologically important structural variation that AlphaFold2 does not. By performing structure optimization in the space of coevolutionary embeddings, rather than Cartesian coordinates, ROCKET automates difficult modeling tasks, such as flips of functional loops and domain rearrangements at low resolution. ROCKET does not require retraining of AlphaFold2 and is readily adaptable to other data modalities. This new type of structure refinement that optimizes latent representations in evolutionary space could unlock possibilities for high-throughput ligand screening, assemblies solved at low resolution, and conformational landscapes.
Giant tortoises exhibit exceptional longevity, often exceeding the human lifespan. To understand the genomic and epigenomic basis of their longevity, we analyzed the DNA sequence and methylome of Jonathan, an Aldabra giant tortoise (Aldabrachelys gigantea), estimated to be 192 years old. Relative to other giant tortoises (Aldabrachelys gigantea and Chelonoidis abingdonii), we found Jonathan has gene variants in pathways associated with aging, including DNA repair and telomere regulation. Consistent with his advanced age, Jonathan has significant age-related changes in DNA methylation and methylation entropy, compared with a 5-year-old Aldabra individual. Notably, we found that low entropy regions in Jonathan's methylome were enriched for genes involved in the electron transport chain. This suggests that high-fidelity transcription of these genes may be crucial for extreme longevity. With this data, we propose a model for aging, that links efficient mitochondrial energy production with nuclear maintenance of low methylation entropy. ### Competing Interest Statement The Regents of the University of California are the sole owner of patents and patent applications directed at epigenetic biomarkers for which Steve Horvath is a named inventor; SH is a founder and paid consultant of the non-profit Epigenetic Clock Development Foundation that licenses these patents. SH is a Principal Investigator at the Altos Labs, Cambridge Institute of Science. The other authors declare no competing interests.
The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.
Advances in machine learning have enabled sufficiently accurate predictions of protein structure to be used in macromolecular structure determination with crystallography and cryo-electron microscopy data. The Phenix software suite has AlphaFold predictions integrated into an automated pipeline that can start with an amino acid sequence and data, and automatically perform model-building and refinement to return a protein model fitted into the data. Due to the steep technical requirements of running AlphaFold efficiently, we have implemented a Phenix-AlphaFold webservice that enables all Phenix users to run AlphaFold predictions remotely from the Phenix GUI starting with the official 1.21 release. This webservice will be improved based on how it is used by the research community and the future research directions for Phenix.
Fluorescent proteins (FPs) are versatile biomarkers that facilitate effective detection and tracking of macromolecules of interest in real time. Engineered FPs such as superfolder green fluorescent protein (sfGFP) and superfolder Cherry (sfCherry) have exceptional refolding capability capable of delivering fluorescent readout in harsh environments where most proteins lose their native functions. Our recent work on the development of a split FP from a species of strawberry anemone, Corynactis californica, delivered pairs of fragments with up to threefold faster complementation than split GFP. We present the biophysical, biochemical, and structural characteristics of five full-length variants derived from these split C. californica GFP (ccGFP). These ccGFP variants are more tolerant under chemical denaturation with up to 8 kcal/mol lower unfolding free energy than that of the sfGFP. It is likely that some of these ccGFP variants could be suitable as biomarkers under more adverse environments where sfGFP fails to survive. A structural analysis suggests explanations of the variations in stabilities among the ccGFP variants.
Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.
Experimental structure determination can be accelerated with AI-based structure prediction methods such as AlphaFold. Here we present an automatic procedure requiring only sequence information and crystallographic data that uses AlphaFold predictions to produce an electron density map and a structural model. Iterating through cycles of structure prediction is a key element of our procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. We applied this procedure to X-ray data for 215 structures released by the Protein Data Bank in a recent 6-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2Å. Predictions from our iterative template-guided prediction procedure were more accurate than those obtained without templates. We suggest a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization.
Fast, reliable docking of models into cryo-EM maps requires understanding of the errors in the maps and the models. Likelihood-based approaches to errors have proven to be powerful and adaptable in experimental structural biology, finding applications in both crystallography and cryo-EM. Indeed, previous crystallographic work on the errors in structural models is directly applicable to likelihood targets in cryo-EM. Likelihood targets in Fourier space are derived here to characterize, based on the comparison of half-maps, the direction- and resolution-dependent variation in the strength of both signal and noise in the data. Because the signal depends on local features, the signal and noise are analysed in local regions of the cryo-EM reconstruction. The likelihood analysis extends to prediction of the signal that will be achieved in any docking calculation for a model of specified quality and completeness. A related calculation generalizes a previous measure of the information gained by making the cryo-EM reconstruction.
Atomic model refinement at low resolution is often a challenging task. This is mostly because the experimental data are not sufficiently detailed to be described by atomic models. To make refinement practical and ensure that a refined atomic model is geometrically meaningful, additional information needs to be used such as restraints on Ramachandran plot distributions or residue side-chain rotameric states. However, using Ramachandran plots or rotameric states as refinement targets diminishes the validating power of these tools. Therefore, finding additional model-validation criteria that are not used or are difficult to use as refinement goals is desirable. Hydrogen bonds are one of the important noncovalent interactions that shape and maintain protein structure. These interactions can be characterized by a specific geometry of hydrogen donor and acceptor atoms. Systematic analysis of these geometries performed for quality-filtered high-resolution models of proteins from the Protein Data Bank shows that they have a distinct and a conserved distribution. Here, it is demonstrated how this information can be used for atomic model validation.
BpeB and BpeF are multidrug efflux transporters from Burkholderia pseudomallei that enable multidrug resistance. Here, we report the crystal structures of BpeB and BpeF at 2.94 Å and 3.0 Å resolution, respectively. BpeB was found as an asymmetric trimer, consistent with the widely-accepted functional rotation mechanism for this type of transporter. One of the monomers has a distinct structure that we interpret as an intermediate along this functional cycle. Additionally, a detergent molecule bound in a previously undescribed binding site provides insights into substrate translocation through the pathway. BpeF shares structural similarities with the crystal structure of OqxB from Klebsiella pneumoniae , where both are symmetric trimers composed of three “binding”-state monomers. The structures of BpeB and BpeF further our understanding of the functional mechanisms of transporters belonging to the HAE1-RND superfamily.
Machine-learning prediction algorithms such as AlphaFold and RoseTTAFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including new experimental information such as a density map, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt on the basis of experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We show that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for interpretation of crystallographic and electron cryo-microscopy maps.
: AI-based methods such as AlphaFold have raised the possibility of using predicted models in place of experimentally-determined structures. Here we assess the accuracy of AlphaFold predictions by comparing them to density maps obtained from automated redeterminations of recent crystal structures and to the corresponding deposited models. Some 25 AlphaFold predictions match experimental maps closely, but most differ on a global scale through distortion and domain orientation and on a local scale in backbone and side-chain conformation. Such differences occur even in parts of AlphaFold models that were predicted with high confidence. Generally, the dissimilarities exceed those between high-resolution pairs of structures containing the same components but determined in different space groups. 30 Therefore, while AlphaFold predictions are useful hypotheses about protein structures, experimental information remains essential for creating an accurate model. One-Sentence Summary: AlphaFold predictions can be very accurate but should be treated as hypotheses as even high-confidence parts can be inconsistent with experimental data.
The ability to create an AlphaFold model for any sequence in a few minutes changes every protein crystal structure determination into a molecular replacement problem and every protein cryo-EM structure determination into a docking problem.Making this even more transformative is the ability to iteratively improve AlphaFold modeling by docking an AlphaFold model into density, rebuilding it, and using the rebuilt model as a template for further AlphaFold model generation (Terwilliger et al., 2022).These features of AlphaFold will make structure determination by crystallography and cryo-EM easier and more powerful than ever before, but do not fundamentally change the importance of the experiment.Anyone can carry out these steps easily using Phenix and free cloud-based Google Colab notebooks.
The ability of Mycobacterium tuberculosis (Mtb) to persist in its host may enable an evolutionary advantage for drug resistant variants to emerge. A potential strategy to prevent persistence and gain drug efficacy is to directly target the activity of enzymes that are crucial for persistence. We present a method for expedited discovery and structure-based design of lead compounds by targeting the hypoxia-associated enzyme L-alanine dehydrogenase (AlaDH). Biochemical and structural analyses of AlaDH confirmed binding of nucleoside derivatives and showed a site adjacent to the nucleoside binding pocket that can confer specificity to putative inhibitors. Using a combination of dye-ligand affinity chromatography, enzyme kinetics and protein crystallographic studies, we show the development and validation of drug prototypes. Crystal structures of AlaDH-inhibitor complexes with variations at the N6 position of the adenyl-moiety of the inhibitor provide insight into the molecular basis for the specificity of these compounds. We describe a drug-designing pipeline that aims to block Mtb to proliferate upon re-oxygenation by specifically blocking NAD accessibility to AlaDH. The collective approach to drug discovery was further evaluated through in silico analyses providing additional insight into an efficient drug development strategy that can be further assessed with the incorporation of in vivo studies.
Machine learning prediction algorithms such as AlphaFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including experimental information, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt based on experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We find that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for crystallographic and electron cryo-microscopy map interpretation.AlphaFold modeling can be improved synergistically by including information from experimental density maps.
Regulation of bacteriophage gene expression involves repressor proteins that bind and downregulate early lytic promoters. A large group of mycobacteriophages code for repressors that are unusual in also terminating transcription elongation at numerous binding sites (stoperators) distributed across the phage genome. Here we provide the X-ray crystal structure of a mycobacteriophage immunity repressor bound to DNA, which reveals the binding of a monomer to an asymmetric DNA sequence using two independent DNA binding domains. The structure is supported by small-angle X-ray scattering, DNA binding, molecular dynamics, and in vivo immunity assays. We propose a model for how dual DNA binding domains facilitate regulation of both transcription initiation and elongation, while enabling evolution of other superinfection immune specificities.