Small molecule ligands exhibit a diverse range of conformations in solution. Upon binding to a target protein, this conformational diversity is reduced. However, ligands can retain some degree of conformational flexibility even when bound to a receptor. In the Protein Data Bank, a small number of ligands have been modeled with distinct alternative conformations that are supported by macromolecular X-ray crystallography density maps. However, the vast majority of structural models are fit to a single-ligand conformation, potentially ignoring the underlying conformational heterogeneity present in the sample. We previously developed qFit-ligand to sample diverse ligand conformations and to select a parsimonious ensemble consistent with the density. While this approach indicated that many ligands populate alternative conformations, limitations in our sampling procedures often resulted in non-physical conformations and could not model complex ligands like macrocycles. Here, we introduce several improvements to qFit-ligand, including integrating RDKit for stochastic conformational sampling. This new sampling method greatly enriches low-energy conformations of small molecules and macrocycles. We further extended qFit-ligand to identify alternative conformations in PanDDA-modified density maps from high-throughput X-ray fragment screening experiments, as well as single-particle cryo-electron microscopy density maps. The new version of qFit-ligand improves fit to electron density and reduces torsional strain relative to deposited single-conformer models and our prior version of qFit-ligand. These advances enhance the analysis of residual conformational heterogeneity present in ligand-bound structures, which can provide important insights for the rational design of therapeutic agents.
Biomolecules exchange between multiple conformational states in their folded state, crucial for their function. Traditional structural biology methods, such as X-ray crystallography and cryogenic electron microscopy (cryo-EM), produce density maps that are ensemble averages, capturing molecules in various conformations. Yet, most models derived from these maps represent only a single conformation, overlooking the complexity of biomolecular structures. The pressing need to accurately reflect the diversity of biomolecular forms by shifting towards modeling structural ensembles that mirror experimental data is complicated by the challenge of distinguishing signal from noise in manual model creation efforts. We have developed qFit, automatic multiconformer modeling software, to model multiconformer models in high-resolution X-ray and cryo-EM structures. Importantly, unlike ensemble models, the multiconformer models produced by qFit can be manually modified in most major model-building software (e.g. Coot) and fit can be further improved by refinement using standard pipelines (e.g. Phenix, Refmac, Buster). The advancement of automated multiconformer modeling is poised to transform the interpretation of structural biology data, fostering new hypotheses about the relationship between macromolecular conformational dynamics and their functions, and marking a new era in the prediction of protein structural ensembles.
TGF-β, essential for development and immunity, is expressed as a latent complex (L-TGF-β) non-covalently associated with its prodomain and presented on immune cell surfaces by covalent association with GARP. Binding to integrin αvβ8 activates L-TGF-β1/GARP. The dogma is that mature TGF-β must physically dissociate from L-TGF-β1 for signaling to occur. Our previous studies discovered that αvβ8-mediated TGF-β autocrine signaling can occur without TGF-β1 release from its latent form. Here, we show that mice engineered to express TGF-β1 that cannot release from L-TGF-β1 survive without early lethal tissue inflammation, unlike those with TGF-β1 deficiency. Combining cryogenic electron microscopy with cell-based assays, we reveal a dynamic allosteric mechanism of autocrine TGF-β1 signaling without release where αvβ8 binding redistributes the intrinsic flexibility of L-TGF-β1 to expose TGF-β1 to its receptors. Dynamic allostery explains the TGF-β3 latency/activation mechanism and why TGF-β3 functions distinctly from TGF-β1, suggesting that it broadly applies to other flexible cell surface receptor/ligand systems.
During protein folding, proteins transition from a disordered polymer into a globular structure, markedly decreasing their conformational degrees of freedom and consequently leading to a substantial reduction in entropy. Nonetheless, folded proteins still retain significant entropy as they fluctuate between the conformations that make up their native state. This residual entropy contributes to crucial functions like binding or catalysis. Here, we outline three major ways that protein use conformational entropy to perform their functions; first, pre-paying entropic cost through ordering of the ground state; second, spatially redistributing entropy, where an decrease in entropy in one area is reciprocated by an increase in entropy elsewhere; third, populating catalytically-competent ensembles, where conformational entropy within the enzymatic scaffold aids in lowering transition state barriers. Given the growing evidence of the biological significance of conformational entropy, emerging largely from NMR and simulation studies, solving the current challenge of structurally defining the ensembles encoding conformational entropy will open new paths for control of binding, catalysis, and allostery.
In their folded state, biomolecules exchange between multiple conformational states, crucial for their function. However, most structural models derived from experiments and computational predictions only encode a single state. To represent biomolecules more accurately, we must move towards modeling and predicting structural ensembles. Information about structural ensembles exists within experimental data from X-ray crystallography and cryo electron microscopy (cryoEM). While new tools are available to detect conformational and compositional heterogeneity that exist within these ensembles, the legacy PDB data structure does not robustly encapsulate this complexity. We propose modifications to the Macromolecular Crystallographic Information File (mmCIF) to improve the representation and interrelation of conformational and compositional heterogeneity. These modifications will enable improved tools to capture macromolecular ensembles in a way that is human and machine interpretable, potentially catalyzing breakthroughs for ensemble-function predictions, analogous to AlphaFold's achievements with single structure prediction.
Despite advances and social progress, the exclusion of diverse groups in academia, especially science, technology, engineering, and mathematics (STEM) fields, across the US and Europe persists, resulting in the underrepresentation of diverse people in higher education. There is extensive literature about theory, observation, and evidence-based practices that can help create a more equitable, inclusive, and diverse learning environment. In this article, we propose the implementation of a Diversity, Equity, Inclusion, and Justice (DEIJ) journal club as a strategic initiative to foster education and promote action towards making academia a more equitable institution. By creating a space for people to engage with DEIJ theories* and strategize ways to improve their learning environment, we hope to normalize the practice and importance of analyzing academia through an equity lens. Guided by restorative justice principles, we offer 10 recommendations for fostering community cohesion through education and mutual understanding. This approach underscores the importance of appropriate action and self-education in the journey toward a more diverse, equitable, inclusive, and just academic environment. *Authors’ note: We understand that “DEIJ” is a multidisciplinary organizational framework that relies on numerous fields of study, including history, sociology, philosophy, and more. We use this term to refer to these different fields of study for brevity purposes.
In their folded state, biomolecules exchange between multiple conformational states that are crucial for their function. Traditional structural biology methods, such as X-ray crystallography and cryogenic electron microscopy (cryo-EM), produce density maps that are ensemble averages, reflecting molecules in various conformations. Yet, most models derived from these maps explicitly represent only a single conformation, overlooking the complexity of biomolecular structures. To accurately reflect the diversity of biomolecular forms, there is a pressing need to shift toward modeling structural ensembles that mirror the experimental data. However, the challenge of distinguishing signal from noise complicates manual efforts to create these models. In response, we introduce the latest enhancements to qFit, an automated computational strategy designed to incorporate protein conformational heterogeneity into models built into density maps. These algorithmic improvements in qFit are substantiated by superior R free and geometry metrics across a wide range of proteins. Importantly, unlike more complex multicopy ensemble models, the multiconformer models produced by qFit can be manually modified in most major model building software (e.g., Coot) and fit can be further improved by refinement using standard pipelines (e.g., Phenix, Refmac, Buster). By reducing the barrier of creating multiconformer models, qFit can foster the development of new hypotheses about the relationship between macromolecular conformational dynamics and function.
By converting a disordered polymer into a globular structure, protein folding reduces many conformational degrees of freedom, resulting in a significant conformational entropy penalty. Nonetheless, residual entropy persists in the protein's native state as it fluctuates between thermally accessible conformations. Here, we review biophysical evidence, primarily from NMR studies, for how conformational entropy modulates the free energy of ligand binding and catalysis. The major theme that emerges is that selection based on free energy has converged on mechanisms to mitigate the effects of entropy loss during crucial functions like binding or catalysis. The modulation of conformational entropy occurs primarily via two main mechanisms: pre-paying entropic costs through ordering in the ground state and spatial compensation through increases in conformational entropy in distal regions after binding. In examining these mechanisms, it also becomes clear that conformational entropy is highly intertwined with classic definitions of conformational changes. We argue that given the ample evidence of the biological significance of conformational entropy, structurally defining the ensembles encoding conformational entropy will open new paths for control of binding, catalysis, and allostery.
Supplemental Table 2: Clinical Characteristics Supplemental Table 3: Sequencing Characteristics Supplemental Table 4: IHC Staining Results
Includes: Table 1S. Baseline patient and disease characteristics. Table 2S. Tumor mutation burden status. Table 3S. PDL1 expression data. Figure 1S. Overall survival for the total cohort.
Conformational ensembles underlie all protein functions. Thus, acquiring atomic-level ensemble models that accurately represent conformational heterogeneity is vital to deepen our understanding of how proteins work. Modeling ensemble information from X-ray diffraction data has been challenging, as traditional cryo-crystallography restricts conformational variability while minimizing radiation damage. Recent advances have enabled the collection of high quality diffraction data at ambient temperatures, revealing innate conformational heterogeneity and temperature-driven changes. Here, we used diffraction datasets for Proteinase K collected at temperatures ranging from 313 to 363K to provide a tutorial for the refinement of multiconformer ensemble models. Integrating automated sampling and refinement tools with manual adjustments, we obtained multiconformer models that describe alternative backbone and sidechain conformations, their relative occupancies, and interconnections between conformers. Our models revealed extensive and diverse conformational changes across temperature, including increased bound peptide ligand occupancies, different Ca2+ binding site configurations and altered rotameric distributions. These insights emphasize the value and need for multiconformer model refinement to extract ensemble information from diffraction data and to understand ensemble-function relationships.
While protein conformational heterogeneity plays an important role in many aspects of biological function, including ligand binding, its impact has been difficult to quantify. Macromolecular X-ray diffraction is commonly interpreted with a static structure, but it can provide information on both the anharmonic and harmonic contributions to conformational heterogeneity. Here, through multiconformer modeling of time- and space-averaged electron density, we measure conformational heterogeneity of 743 stringently matched pairs of crystallographic datasets that reflect unbound/apo and ligand-bound/holo states. When comparing the conformational heterogeneity of side chains, we observe that when binding site residues become more rigid upon ligand binding, distant residues tend to become more flexible, especially in non-solvent-exposed regions. Among ligand properties, we observe increased protein flexibility as the number of hydrogen bonds decreases and relative hydrophobicity increases. Across a series of 13 inhibitor-bound structures of CDK2, we find that conformational heterogeneity is correlated with inhibitor features and identify how conformational changes propagate differences in conformational heterogeneity away from the binding site. Collectively, our findings agree with models emerging from nuclear magnetic resonance studies suggesting that residual side-chain entropy can modulate affinity and point to the need to integrate both static conformational changes and conformational heterogeneity in models of ligand binding.
Molecular profiling studies have enabled discoveries for metastatic prostate cancer (MPC) but have predominantly occurred in academic medical institutions and involved non-representative patient populations. We established the Metastatic Prostate Cancer Project (MPCproject, mpcproject.org), a patient-partnered initiative to involve patients with MPC living anywhere in the US and Canada in molecular research. Here, we present results from our partnership with the first 706 MPCproject participants. While 41% of patient partners live in rural, physician-shortage, or medically underserved areas, the MPCproject has not yet achieved racial diversity, a disparity that demands new initiatives detailed herein. Among molecular data from 333 patient partners (572 samples), exome sequencing of 63 tumor and 19 cell-free DNA (cfDNA) samples recapitulated known findings in MPC, while inexpensive ultra-low-coverage sequencing of 318 cfDNA samples revealed clinically relevant AR amplifications. This study illustrates the power of a growing, longitudinal partnership with patients to generate a more representative understanding of MPC.
This paper describes outcomes of the 2019 Cryo-EM Model Challenge. The goals were to (1) assess the quality of models that can be produced from cryogenic electron microscopy (cryo-EM) maps using current modeling software, (2) evaluate reproducibility of modeling results from different software developers and users and (3) compare performance of current metrics used for model evaluation, particularly Fit-to-Map metrics, with focus on near-atomic resolution. Our findings demonstrate the relatively high accuracy and reproducibility of cryo-EM models derived by 13 participating teams from four benchmark maps, including three forming a resolution series (1.8 to 3.1 Å). The results permit specific recommendations to be made about validating near-atomic cryo-EM structures both in the context of individual experiments and structure data archives such as the Protein Data Bank. We recommend the adoption of multiple scoring parameters to provide full and objective annotation and assessment of the model, reflective of the observed cryo-EM map density.
ABSTRACTMolecular profiling studies have enabled numerous discoveries for metastatic prostate cancer (MPC), but they have mostly occurred in academic medical institutions focused on select patient populations. We developed the Metastatic Prostate Cancer Project (MPCproject, mpcproject.org), a patient-partnered initiative to empower MPC patients living anywhere in the U.S. and Canada to participate in molecular research and contribute directly to translational discovery. Here we present clinicogenomic results from our partnership with the first 706 MPCproject participants. We found that a patient-centered and remote research strategy enhanced engagement with patients in rural and medically underserved areas. Furthermore, patient-reported data achieved 90% consistency with abstracted health records for therapies and provided a mechanism for patient-partners to share information about their cancer experience not documented in medical records. Among the molecular profiling data from 333 patient-partners (n = 573 samples), whole exome sequencing of 63 tumor samples obtained from hospitals across the U.S. and Canada and 19 plasma cell-free DNA (cfDNA) samples from blood donated remotely recapitulated known findings in MPC and enabled longitudinal study of prostate cancer evolution. Inexpensive ultra-low coverage whole genome sequencing of 318 cfDNA samples from donated blood revealed clinically relevant genomic changes like AR amplification, even in the context of low tumor burden. Collectively, this study illustrates the power of a longitudinal partnership with patients to generate a more representative clinical and molecular understanding of MPC.NoteTo assist our patient-partners and the wider MPC community interpret the results of this study, we have included a glossary of terms in the Supplementary Materials.
Protein conformations are shaped by cellular environments, but how environmental changes alter the conformational landscapes of specific proteins in vivo remains largely uncharacterized, in part due to the challenge of probing protein structures in living cells. Here, we use deep mutational scanning to investigate how a toxic conformation of α-synuclein, a dynamic protein linked to Parkinson's disease, responds to perturbations of cellular proteostasis. In the context of a course for graduate students in the UCSF Integrative Program in Quantitative Biology, we screened a comprehensive library of α-synuclein missense mutants in yeast cells treated with a variety of small molecules that perturb cellular processes linked to α-synuclein biology and pathobiology. We found that the conformation of α-synuclein previously shown to drive yeast toxicity-an extended, membrane-bound helix-is largely unaffected by these chemical perturbations, underscoring the importance of this conformational state as a driver of cellular toxicity. On the other hand, the chemical perturbations have a significant effect on the ability of mutations to suppress α-synuclein toxicity. Moreover, we find that sequence determinants of α-synuclein toxicity are well described by a simple structural model of the membrane-bound helix. This model predicts that α-synuclein penetrates the membrane to constant depth across its length but that membrane affinity decreases toward the C terminus, which is consistent with orthogonal biophysical measurements. Finally, we discuss how parallelized chemical genetics experiments can provide a robust framework for inquiry-based graduate coursework.