IHMCIF (github.com/ihmwg/IHMCIF) is a data information framework that supports archiving and disseminating macromolecular structures determined by integrative or hybrid modeling (IHM), and making them Findable, Accessible, Interoperable, and Reusable (FAIR). IHMCIF is an extension of the Protein Data Bank Exchange/macromolecular Crystallographic Information Framework (PDBx/mmCIF) that serves as the framework for the Protein Data Bank (PDB) to archive experimentally determined atomic structures of biological macromolecules and their complexes with one another and small molecule ligands (e.g., enzyme cofactors and drugs). IHMCIF serves as the foundational data standard for the PDB-Dev prototype system, developed for archiving and disseminating integrative structures. It utilizes a flexible data representation to describe integrative structures that span multiple spatiotemporal scales and structural states with definitions for restraints from a variety of experimental methods contributing to integrative structural biology. The IHMCIF extension was created with the benefit of considerable community input and recommendations gathered by the Worldwide Protein Data Bank (wwPDB) Task Force for Integrative or Hybrid Methods (wwpdb.org/task/hybrid). Herein, we describe the development of IHMCIF to support evolving methodologies and ongoing advancements in integrative structural biology. Ultimately, IHMCIF will facilitate the unification of PDB-Dev data and tools with the PDB archive so that integrative structures can be archived and disseminated through PDB.
Advances in computational tools for atomic model building are leading to accurate models of large molecular assemblies seen in electron microscopy, often at challenging resolutions of 3-4 Å. We describe new methods in the UCSF ChimeraX molecular modeling package that take advantage of machine-learning structure predictions, provide likelihood-based fitting in maps, and compute per-residue scores to identify modeling errors. Additional model-building tools assist analysis of mutations, post-translational modifications, and interactions with ligands. We present the latest ChimeraX model-building capabilities, including several community-developed extensions. ChimeraX is available free of charge for noncommercial use at https://www.rbvi.ucsf.edu/chimerax.
UCSF ChimeraX is the next-generation interactive visualization program from the Resource for Biocomputing, Visualization, and Informatics (RBVI), following UCSF Chimera. ChimeraX brings (a) significant performance and graphics enhancements; (b) new implementations of Chimera's most highly used tools, many with further improvements; (c) several entirely new analysis features; (d) support for new areas such as virtual reality, light-sheet microscopy, and medical imaging data; (e) major ease-of-use advances, including toolbars with icons to perform actions with a single click, basic "undo" capabilities, and more logical and consistent commands; and (f) an app store for researchers to contribute new tools. ChimeraX includes full user documentation and is free for noncommercial use, with downloads available for Windows, Linux, and macOS from https://www.rbvi.ucsf.edu/chimerax.
Structures of biomolecular systems are increasingly computed by integrative modeling. In this approach, a structural model is constructed by combining information from multiple sources, including varied experimental methods and prior models. In 2019, a Workshop was held as a Biophysical Society Satellite Meeting to assess progress and discuss further requirements for archiving integrative structures. The primary goal of the Workshop was to build consensus for addressing the challenges involved in creating common data standards, building methods for federated data exchange, and developing mechanisms for validating integrative structures. The summary of the Workshop and the recommendations that emerged are presented here.
Chlamydia trachomatis(Ct) is the most common sexually transmitted bacterium with more than 131 million cases occurring annually worldwide.Ctinfections are often asymptomatic, persisting for many years despite treatment.In vitrorecovery from persistence occurs when indole is utilized by the organism’s tryptophan synthase to synthesize tryptophan, an essential amino acid for replication. Ocular but not urogenitalCtstrains contain mutations in the synthase that abrogate tryptophan synthesis. Here, we discovered that the genomes of serial isolates from a woman with recurrent, treatedCtSTIs over many years were identical with a novel synthase mutation. This likely allowed long-termin vivopersistence where active infection resumed only when tryptophan became available. Our findings indicate an emerging adaptive host-pathogen evolutionary strategy for survival in the urogenital tract that will prompt the field to further explore chlamydial persistence, evaluate the genetics of mutantCtstrains and fitness within the host, and their implications for disease pathogenesis.
Can virtual reality be useful for visualizing and analyzing molecular structures and three-dimensional (3D) microscopy? Uses we are exploring include studies of drug binding to proteins and the effects of mutations, building accurate atomic models in electron microscopy and x-ray density maps, understanding how immune system cells move using 3D light microscopy, and teaching schoolchildren about biomolecules that are the machinery of life. Virtual reality (VR) offers immersive display with a wide field of view and head tracking for better perception of molecular architectures and uses 6-degree-of-freedom hand controllers for simple manipulation of 3D data. Conventional computer displays with trackpad, mouse and keyboard excel at two-dimensional tasks such as writing and studying research literature, uses for which VR technology is at present far inferior. Adding VR to the conventional computing environment could improve 3D capabilities if new user-interface problems can be solved. We have developed three VR applications: ChimeraX for analyzing molecular structures and electron and light microscopy data, AltPDB for collaborative discussions around atomic models, and Molecular Zoo for teaching young students characteristics of biomolecules. Investigations over three decades have produced an extensive literature evaluating the potential of VR in research and education. Consumer VR headsets are now affordable to researchers and educators, allowing direct tests of whether the technology is valuable in these areas. We survey here advantages and disadvantages of VR for molecular biology in the context of affordable and dramatically more powerful VR and graphics hardware than has been available in the past.
With ever-increasing amounts of sequence data available in both the primary literature and sequence repositories, there is a bottleneck in annotating molecular function to a sequence. This article describes the biocuration process and methods used in the structure-function linkage database (SFLD) to help address some of the challenges. We discuss how the hierarchy within the SFLD allows us to infer detailed functional properties for functionally diverse enzyme superfamilies in which all members are homologous, conserve an aspect of their chemical function and have associated conserved structural features that enable the chemistry. Also presented is the Enzyme Structure-Function Ontology (ESFO), which has been designed to capture the relationships between enzyme sequence, structure and function that underlie the SFLD and is used to guide the biocuration processes within the SFLD.Database URL:http://sfld.rbvi.ucsf.edu/.
Protein function identification remains a significant problem. Solving this problem at the molecular functional level would allow mechanistic determinant identification-amino acids that distinguish details between functional families within a superfamily. Active site profiling was developed to identify mechanistic determinants. DASP and DASP2 were developed as tools to search sequence databases using active site profiling. Here, TuLIP (Two-Level Iterative clustering Process) is introduced as an iterative, divisive clustering process that utilizes active site profiling to separate structurally characterized superfamily members into functionally relevant clusters. Underlying TuLIP is the observation that functionally relevant families (curated by Structure-Function Linkage Database, SFLD) self-identify in DASP2 searches; clusters containing multiple functional families do not. Each TuLIP iteration produces candidate clusters, each evaluated to determine if it self-identifies using DASP2. If so, it is deemed a functionally relevant group. Divisive clustering continues until each structure is either a functionally relevant group member or a singlet. TuLIP is validated on enolase and glutathione transferase structures, superfamilies well-curated by SFLD. Correlation is strong; small numbers of structures prevent statistically significant analysis. TuLIP-identified enolase clusters are used in DASP2 GenBank searches to identify sequences sharing functional site features. Analysis shows a true positive rate of 96%, false negative rate of 4%, and maximum false positive rate of 4%. F-measure and performance analysis on the enolase search results and comparison to GEMMA and SCI-PHY demonstrate that TuLIP avoids the over-division problem of these methods. Mechanistic determinants for enolase families are evaluated and shown to correlate well with literature results.
Leukocytes and other amoeboid cells change shape as they move, forming highly dynamic, actin-filled pseudopods. Although we understand much about the architecture and dynamics of thin lamellipodia made by slow-moving cells on flat surfaces, conventional light microscopy lacks the spatial and temporal resolution required to track complex pseudopods of cells moving in three dimensions. We therefore employed lattice light sheet microscopy to perform three-dimensional, time-lapse imaging of neutrophil-like HL-60 cells crawling through collagen matrices. To analyze three-dimensional pseudopods we: (i) developed fluorescent probe combinations that distinguish cortical actin from dynamic, pseudopod-forming actin networks, and (ii) adapted molecular visualization tools from structural biology to render and analyze complex cell surfaces. Surprisingly, three-dimensional pseudopods turn out to be composed of thin (<0.75 µm), flat sheets that sometimes interleave to form rosettes. Their laminar nature is not templated by an external surface, but likely reflects a linear arrangement of regulatory molecules. Although we find that Arp2/3-dependent pseudopods are dispensable for three-dimensional locomotion, their elimination dramatically decreases the frequency of cell turning, and pseudopod dynamics increase when cells change direction, highlighting the important role pseudopods play in pathfinding.
Peroxiredoxins (Prxs or Prdxs) are a large protein superfamily of antioxidant enzymes that rapidly detoxify damaging peroxides and/or affect signal transduction and, thus, have roles in proliferation, differentiation, and apoptosis. Prx superfamily members are widespread across phylogeny and multiple methods have been developed to classify them. Here we present an updated atlas of the Prx superfamily identified using a novel method called MISST (Multi-level Iterative Sequence Searching Technique). MISST is an iterative search process developed to be both agglomerative, to add sequences containing similar functional site features, and divisive, to split groups when functional site features suggest distinct functionally-relevant clusters. Superfamily members need not be identified initially—MISST begins with a minimal representative set of known structures and searches GenBank iteratively. Further, the method’s novelty lies in the manner in which isofunctional groups are selected; rather than use a single or shifting threshold to identify clusters, the groups are deemed isofunctional when they pass a self-identification criterion, such that the group identifies itself and nothing else in a search of GenBank. The method was preliminarily validated on the Prxs, as the Prxs presented challenges of both agglomeration and division. For example, previous sequence analysis clustered the Prx functional families Prx1 and Prx6 into one group. Subsequent expert analysis clearly identified Prx6 as a distinct functionally relevant group. The MISST process distinguishes these two closely related, though functionally distinct, families. Through MISST search iterations, over 38,000 Prx sequences were identified, which the method divided into six isofunctional clusters, consistent with previous expert analysis. The results represent the most complete computational functional analysis of proteins comprising the Prx superfamily. The feasibility of this novel method is demonstrated by the Prx superfamily results, laying the foundation for potential functionally relevant clustering of the universe of protein sequences.
UCSF ChimeraX is next-generation software for the visualization and analysis of molecular structures, density maps, 3D microscopy, and associated data. It addresses challenges in the size, scope, and disparate types of data attendant with cutting-edge experimental methods, while providing advanced options for high-quality rendering (interactive ambient occlusion, reliable molecular surface calculations, etc.) and professional approaches to software design and distribution. This article highlights some specific advances in the areas of visualization and usability, performance, and extensibility. ChimeraX is free for noncommercial use and is available from http://www.rbvi.ucsf.edu/chimerax/ for Windows, Mac, and Linux.
Variation in the expression level and activity of genes involved in drug disposition and action (‘pharmacogenes’) can affect drug response and toxicity, especially when in tissues of pharmacological importance. Previous studies have relied primarily on microarrays to understand gene expression differences, or have focused on a single tissue or small number of samples. The goal of this study was to use RNA-sequencing (RNA-seq) to determine the expression levels and alternative splicing of 389 Pharmacogenomics Research Network pharmacogenes across four tissues (liver, kidney, heart and adipose) and lymphoblastoid cell lines, which are used widely in pharmacogenomics studies. Analysis of RNA-seq data from 139 different individuals across the 5 tissues (20–45 individuals per tissue type) revealed substantial variation in both expression levels and splicing across samples and tissue types. Comparison with GTEx data yielded a consistent picture. This in-depth exploration also revealed 183 splicing events in pharmacogenes that were previously not annotated. Overall, this study serves as a rich resource for the research community to inform biomarker and drug discovery and use.
Background: Development of automatable processes for clustering proteins into functionally relevant groups is a critical hurdle as an increasing number of sequences are deposited into databases. Experimental function determination is exceptionally time-consuming and can't keep pace with the identification of protein sequences. A tool, DASP (Deacon Active Site Profiler), was previously developed to identify protein sequences with active site similarity to a query set. Development of two iterative, automatable methods for clustering proteins into functionally relevant groups exposed algorithmic limitations to DASP.Results: The accuracy and efficiency of DASP was significantly improved through six algorithmic enhancements implemented in two stages: DASP2 and DASP3. Validation demonstrated DASP3 provides greater score separation between true positives and false positives than earlier versions. In addition, DASP3 shows similar performance to previous versions in clustering protein structures into isofunctional groups (validated against manual curation), but DASP3 gathers and clusters protein sequences into isofunctional groups more efficiently than DASP and DASP2.Conclusions: DASP algorithmic enhancements resulted in improved efficiency and accuracy of identifying proteins that contain active site features similar to those of the query set. These enhancements provide incremental improvement in structure database searches and initial sequence database searches; however, the enhancements show significant improvement in iterative sequence searches, suggesting DASP3 is an appropriate tool for the iterative processes required for clustering proteins into isofunctional groups.
This file contains additional DASP methods, including the development and validation of DASP2, and supplemental figures with additional validation of TuLIP and MISST. (DOCX 1543Â kb)
Homology modeling predicts protein structures using known structures of related proteins as templates. We developed MULTIDOMAIN ASSEMBLER (MDA) to address the special problems that arise when modeling proteins with large numbers of domains, such as fibronectin with 30 domains, as well as cases with hundreds of templates. These problems include how to spatially arrange nonoverlapping template structures, and how to get the best template coverage when some sequence regions have hundreds of available structures while other regions have a few distant homologs. MDA automates the tasks of template searching, visualization, and selection followed by multidomain model generation, and is part of the widely used molecular graphics package UCSF CHIMERA (University of California, San Francisco). We demonstrate applications and discuss MDA’s benefits and limitations.
MOTIVATION:Contact maps are a convenient method for the structural biologists to identify structural features through two-dimensional simplification. Binary (yes/no) contact maps with a single cutoff distance can be generalized to show continuous distance ranges. We have developed a UCSF Chimera tool, RRDistMaps, to compute such generalized maps in order to analyze pairwise variations in intramolecular contacts. An interactive utility, RRDistMaps, visualizes conformational changes, both local (e.g. binding-site residues) and global (e.g. hinge motion), between unbound and bound proteins through distance patterns. Users can target residue pairs in RRDistMaps for further navigation in Chimera. The interface contains the unique features of identifying long-range residue motion and aligning sequences to simultaneously compare distance maps.AVAILABILITY AND IMPLEMENTATION:RRDistMaps was developed as part of UCSF Chimera release 1.10, which is freely available at http://rbvi.ucsf.edu/chimera/download.html, and operates on Linux, Windows, and Mac OS.CONTACT:conrad@cgl.ucsf.edu.
setsApp (http://apps.cytoscape.org/apps/setsapp) is a relatively simple Cytoscape 3 app for users to handle groups of nodes and/or edges. It supports several important biological workflows and enables various set operations. setsApp provides basic tools to create sets of nodes or edges, import or export sets, and perform standard set operations (union, difference, intersection) on those sets. Automatic set partitioning and layout functions are also provided. The sets functionality is also exposed to users and app developers in the form of a set of commands that can be used for scripting purposes or integrated in other Cytoscape apps.