An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Multi-omics approaches use a diversity of high-throughput technologies to profile the different molecular layers of living cells. Ideally, the integration of this information should result in comprehensive systems models of cellular physiology and regulation. However, most multi-omics projects still include a limited number of molecular assays and there have been very few multi-omic studies that evaluate dynamic processes such as cellular growth, development and adaptation. Hence, we lack formal analysis methods and comprehensive multi-omics datasets that can be leveraged to develop true multi-layered models for dynamic cellular systems. Here we present the STATegra multi-omics dataset that combines measurements from up to 10 different omics technologies applied to the same biological system, namely the well-studied mouse pre-B-cell differentiation. STATegra includes high-throughput measurements of chromatin structure, gene expression, proteomics and metabolomics, and it is complemented with single-cell data. To our knowledge, the STATegra collection is the most diverse multi-omics dataset describing a dynamic biological system.
In this article, the Ponceau staining presented in Fig. 1b (right, bottom) does not follow best practices for figure preparation since itinadvertently included duplications from the Ponceau staining presented in Supplementary Fig. 1b (for which the same preparation ofnucleosomes from HeLa cells had been used). A new Fig. 1b is provided in the Author Correction.
BACKGROUND:Peak calling is a fundamental step in the analysis of data generated by ChIP-seq or similar techniques to acquire epigenetics information. Current peak callers are often hard to parameterise and may therefore be difficult to use for non-bioinformaticians. In this paper, we present the ChIP-seq analysis tool available in CLC Genomics Workbench and CLC Genomics Server (version 7.5 and up), a user-friendly peak-caller designed to be not specific to a particular *-seq protocol.RESULTS:We illustrate the advantages of a shape-based approach and describe the algorithmic principles underlying the implementation. Thanks to the generality of the idea and the fact the algorithm is able to learn the peak shape from the data, the implementation requires only minimal user input, while still being applicable to a range of *-seq protocols. Using independently validated benchmark datasets, we compare our implementation to other state-of-the-art algorithms explicitly designed to analyse ChIP-seq data and provide an evaluation in terms of receiver-operator characteristic (ROC) plots. In order to show the applicability of the method to similar *-seq protocols, we also investigate algorithmic performances on DNase-seq data.CONCLUSIONS:The results show that CLC shape-based peak caller ranks well among popular state-of-the-art peak callers while providing flexibility and ease-of-use.
BACKGROUND:The heat shock protein 90 (Hsp90) is required for the stability of many signalling kinases. As a target for cancer therapy it allows the simultaneous inhibition of several signalling pathways. However, its inhibition in healthy cells could also lead to severe side effects. This is the first comprehensive analysis of the response to Hsp90 inhibition at the kinome level.METHODS:We quantitatively profiled the effects of Hsp90 inhibition by geldanamycin on the kinome of one primary (Hs68) and three tumour cell lines (SW480, U2OS, A549) by affinity proteomics based on immobilized broad spectrum kinase inhibitors ("kinobeads"). To identify affected pathways we used the KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway classification. We combined Hsp90 and proteasome inhibition to identify Hsp90 substrates in Hs68 and SW480 cells. The mutational status of kinases from the used cell lines was determined using next-generation sequencing. A mutation of Hsp90 candidate client RIPK2 was mapped onto its structure.RESULTS:We measured relative abundances of > 140 protein kinases from the four cell lines in response to geldanamycin treatment and identified many new potential Hsp90 substrates. These kinases represent diverse families and cellular functions, with a strong representation of pathways involved in tumour progression like the BMP, MAPK and TGF-beta signalling cascades. Co-treatment with the proteasome inhibitor MG132 enabled us to classify 64 kinases as true Hsp90 clients. Finally, mutations in 7 kinases correlate with an altered response to Hsp90 inhibition. Structural modelling of the candidate client RIPK2 suggests an impact of the mutation on a proposed Hsp90 binding domain.CONCLUSIONS:We propose a high confidence list of Hsp90 kinase clients, which provides new opportunities for targeted and combinatorial cancer treatment and diagnostic applications.
C/EBPs are implied in an amazing number of cellular functions: C/EBPs regulate tissue and cell type specific gene expression, proliferation, and differentiation control. C/EBPs assist in energy metabolism, female reproduction, innate immunity, inflammation, senescence, and the development of neoplasms. How can C/EBPs fulfill so many functions? Here we discuss that C/EBPs are extensively modified by methylation of arginine and lysine side chains and that regulated methylation profoundly affects the activity of C/EBPs.
Summary: Contact maps are a valuable visualization tool in structural biology. They are a convenient way to display proteins in two dimensions and to quickly identify structural features such as domain architecture, secondary structure and contact clusters. We developed a tool called CMView which integrates rich contact map analysis with 3D visualization using PyMol. Our tool provides functions for contact map calculation from structure, basic editing, visualization in contact map and 3D space and structural comparison with different built-in alignment methods. A unique feature is the interactive refinement of structural alignments based on user selected substructures. Availability: CMView is freely available for Linux, Windows and MacOS. The software and a comprehensive manual can be downloaded from http://www.bioinformatics.org/cmview/. The source code is licensed under the GNU General Public License. Contact: lappe@molgen. mpg. de , stehr@molgen. mpg. de
The use of local covalent geometry for quality assessment and refinement of protein structure models is a well-established methodology. The question arises whether information on non-covalent geometry contained within resolved structures can be harnessed to improve structure prediction. Moreover, incorporation of different combinations of priors would pave the way towards multi-body potentials. Existing empirical force-fields do not facilitate an interactive exploration of the parameter space and an assignment of spatial propensities to contacts. Hence, we investigate the possibility of making such propensities available for synergistic modeling. We present an approach that facilitates the extraction and analysis of anisotropic contact potentials for a multitude of parameters describing an amino acid and the conditions within its microenvi-ronment. For this purpose, two novel visualization principles will be introduced. The first visualization illustrates anisotropic residue-dependent contact density potentials in the form of a map projection. A second visualization is overlaid onto this, showing similar local neighborhoods as abstract traces of residues contained within each individual neighborhood. The Contact Geometry Analysis Plugin (CGAP) (for CMView) we developed allows incorporation of geometric orientation propensities into the process of interactive protein modeling and can be used for the generation of improved energy functions. It further supports the analysis of model quality, as it directly illustrates model consistency with known spatial propensities which, in turn, enables users to detect possible structural errors.
BACKGROUND:Current large-scale cancer sequencing projects have identified large numbers of somatic mutations covering an increasing number of different cancer tissues and patients. However, the characterization of these mutations at the structural and functional level remains a challenge.RESULTS:We present results from an analysis of the structural impact of frequent missense cancer mutations using an automated method. We find that inactivation of tumor suppressors in cancer correlates frequently with destabilizing mutations preferably in the core of the protein, while enhanced activity of oncogenes is often linked to specific mutations at functional sites. Furthermore, our results show that this alteration of oncogenic activity is often associated with mutations at ATP or GTP binding sites.CONCLUSIONS:With our findings we can confirm and statistically validate the hypotheses for the gain-of-function and loss-of-function mechanisms of oncogenes and tumor suppressors, respectively. We show that the distinct mutational patterns can potentially be used to pre-classify newly identified cancer-associated genes with yet unknown function.
Background Colorectal cancer (CRC) is with approximately 1 million cases the third most common cancer worldwide. Extensive research is ongoing to decipher the underlying genetic patterns with the hope to improve early cancer diagnosis and treatment. In this direction, the recent progress in next generation sequencing technologies has revolutionized the field of cancer genomics. However, one caveat of these studies remains the large amount of genetic variations identified and their interpretation. Methodology/Principal Findings Here we present the first work on whole exome NGS of primary colon cancers. We performed 454 whole exome pyrosequencing of tumor as well as adjacent not affected normal colonic tissue from microsatellite stable (MSS) and microsatellite instable (MSI) colon cancer patients and identified more than 50,000 small nucleotide variations for each tissue. According to predictions based on MSS and MSI pathomechanisms we identified eight times more somatic non-synonymous variations in MSI cancers than in MSS and we were able to reproduce the result in four additional CRCs. Our bioinformatics filtering approach narrowed down the rate of most significant mutations to 359 for MSI and 45 for MSS CRCs with predicted altered protein functions. In both CRCs, MSI and MSS, we found somatic mutations in the intracellular kinase domain of bone morphogenetic protein receptor 1A, BMPR1A, a gene where so far germline mutations are associated with juvenile polyposis syndrome, and show that the mutations functionally impair the protein function. Conclusions/Significance We conclude that with deep sequencing of tumor exomes one may be able to predict the microsatellite status of CRC and in addition identify potentially clinically relevant mutations.
The success of community projects such as Wikipedia has recently prompted a discussion about the applicability of such tools in the life sciences. Currently, there are several such 'science-wikis' that aim to collect specialist knowledge from the community into centralized resources. However, there is no consensus about how to achieve this goal. For example, it is not clear how to best integrate data from established, centralized databases with that provided by 'community annotation'. We created PDBWiki, a scientific wiki for the community annotation of protein structures. The wiki consists of one structured page for each entry in the the Protein Data Bank (PDB) and allows the user to attach categorized comments to the entries. Additionally, each page includes a user editable list of cross-references to external resources. As in a database, it is possible to produce tabular reports and 'structure galleries' based on user-defined queries or lists of entries. PDBWiki runs in parallel to the PDB, separating original database content from user annotations. PDBWiki demonstrates how collaboration features can be integrated with primary data from a biological database. It can be used as a system for better understanding how to capture community knowledge in the biological sciences. For users of the PDB, PDBWiki provides a bug-tracker, discussion forum and community annotation system. To date, user participation has been modest, but is increasing. The user editable cross-references section has proven popular, with the number of linked resources more than doubling from 17 originally to 39 today. Database URL: http://www.pdbwiki.org.
Contact maps have been extensively used as a simplified representation of protein structures. They capture most important features of a protein's fold, being preferred by a number of researchers for the description and study of protein structures. Inspired by the model's simplicity many groups have dedicated a considerable amount of effort towards contact prediction as a proxy for protein structure prediction. However a contact map's biological interest is subject to the availability of reliable methods for the 3-dimensional reconstruction of the structure.
We show that multiple structure alignment (MStA) using contact maps is equivalent to the problem of f nding a sample mean of contact maps. From this result, we derive a subgradient method for solving the MStA method. Experiments show that the proposed algorithm is a f exible alignment method that provides an excellent tradeoff between accuracy and speed.
Much attention has recently been given to the statistical significance of topological features observed in biological networks. Here, we consider residue interaction graphs (RIGs) as network representations of protein structures with residues as nodes and inter-residue interactions as edges. Degree-preserving randomized models have been widely used for this purpose in biomolecular networks. However, such a single summary statistic of a network may not be detailed enough to capture the complex topological characteristics of protein structures and their network counterparts. Here, we investigate a variety of topological properties of RIGs to find a well fitting network null model for them. The RIGs are derived from a structurally diverse protein data set at various distance cut-offs and for different groups of interacting atoms. We compare the network structure of RIGs to several random graph models. We show that 3-dimensional geometric random graphs, that model spatial relationships between objects, provide the best fit to RIGs. We investigate the relationship between the strength of the fit and various protein structural features. We show that the fit depends on protein size, structural class, and thermostability, but not on quaternary structure. We apply our model to the identification of significantly over-represented structural building blocks, i.e., network motifs, in protein structure networks. As expected, choosing geometric graphs as a null model results in the most specific identification of motifs. Our geometric random graph model may facilitate further graph-based studies of protein conformation space and have important implications for protein structure comparison and prediction. The choice of a well-fitting null model is crucial for finding structural motifs that play an important role in protein folding, stability and function. To our knowledge, this is the first study that addresses the challenge of finding an optimized null model for RIGs, by comparing various RIG definitions against a series of network models.
The network of native non-covalent residue contacts determines the three-dimensional structure of a protein. However, not all contacts are of equal structural significance, and little knowledge exists about a minimal, yet sufficient, subset required to define the global features of a protein. Characterisation of this “structural essence” has remained elusive so far: no algorithmic strategy has been devised to-date that could outperform a random selection in terms of 3D reconstruction accuracy (measured as the Ca RMSD). It is not only of theoretical interest (i.e., for design of advanced statistical potentials) to identify the number and nature of essential native contacts—such a subset of spatial constraints is very useful in a number of novel experimental methods (like EPR) which rely heavily on constraint-based protein modelling. To derive accurate three-dimensional models from distance constraints, we implemented a reconstruction pipeline using distance geometry. We selected a test-set of 12 protein structures from the four major SCOP fold classes and performed our reconstruction analysis. As a reference set, series of random subsets (ranging from 10% to 90% of native contacts) are generated for each protein, and the reconstruction accuracy is computed for each subset. We have developed a rational strategy, termed “cone-peeling” that combines sequence features and network descriptors to select minimal subsets that outperform the reference sets. We present, for the first time, a rational strategy to derive a structural essence of residue contacts and provide an estimate of the size of this minimal subset. Our algorithm computes sparse subsets capable of determining the tertiary structure at approximately 4.8 Å Ca RMSD with as little as 8% of the native contacts (Ca-Ca and Cb-Cb). At the same time, a randomly chosen subset of native contacts needs about twice as many contacts to reach the same level of accuracy. This “structural essence” opens new avenues in the fields of structure prediction, empirical potentials and docking.
Novel high-throughput technologies for directed evolution enable experimental coverage of an impressive number of sequences. Nevertheless, the success of such experiments hinges on the initial sequence libraries. Here we consider the computational design of smart focused libraries and review insights from experimental strategies and theoretic advances in modelling their energy landscapes. In library design as in structure prediction, the applied energy function is the key. Current knowledge-based potentials have proven more successful than purely physics-based ones. Here we summarize novel approaches that extend the classical pairwise treatment of residue contacts towards adaptive knowledge-based multi-body potentials. We suggest that minimal sets of probabilistic constraints will lead to much more efficient sampling of permissible conformations and sequence space.
Background: For over 30 years potentials of mean force have been used to evaluate the relative energy of protein structures. The most commonly used potentials define the energy of residue-residue interactions and are derived from the empirical analysis of the known protein structures. However, single-body residue 'environment' potentials, although widely used in protein structure analysis, have not been rigorously compared to these classical two-body residue-residue interaction potentials. Here we do not try to combine the two different types of residue interaction potential, but rather to assess their independent contribution to scoring protein structures.Results: A data set of nearly three thousand monomers was used to compare pairwise residue-residue 'contact-type' propensities to single-body residue 'contact-count' propensities. Using a large and standard set of protein decoys we performed an in-depth comparison of these two types of residue interaction propensities. The scores derived from the contact-type and contact-count propensities were assessed using two different performance metrics and were compared using 90 different definitions of residue-residue contact. Our findings show that both types of score perform equally well on the task of discriminating between near-native protein decoys. However, in a statistical sense, the contact-count based scores were found to carry more information than the contact-type based scores.Conclusion: Our analysis has shown that the performance of either type of score is very similar on a range of different decoys. This similarity suggests a common underlying biophysical principle for both types of residue interaction propensity. However, several features of the contact-count based propensity suggests that it should be used in preference to the contact-type based propensity. Specifically, it has been shown that contact-counts can be predicted from sequence information alone. In addition, the use of a single-body term allows for efficient alignment strategies using dynamic programming, which is useful for fold recognition, for example. These facts, combined with the relative simplicity of the contact-count propensity, suggests that contact-counts should be studied in more detail in the future.