Natural products play a significant role in drug discovery and development. Many topological pharmacophore patterns are common between natural products and commercial drugs. A better understanding of the specific physicochemical and structural features of natural products is important for corresponding drug development. Several encyclopedias of natural compounds have been composed, but the information remains scattered or not freely available. The first version of the Supernatural database containing ∼ 50,000 compounds was published in 2006 to face these challenges. Here we present a new, updated and expanded version of natural product database, Super Natural II (http://bioinformatics.charite.de/supernatural), comprising ∼ 326,000 molecules. It provides all corresponding 2D structures, the most important structural and physicochemical properties, the predicted toxicity class for ∼ 170,000 compounds and the vendor information for the vast majority of compounds. The new version allows a template-based search for similar compounds as well as a search for compound names, vendors, specific physical properties or any substructures. Super Natural II also provides information about the pathways associated with synthesis and degradation of the natural products, as well as their mechanism of action with respect to structurally similar drugs and their target proteins.
Mass spectrometry (MS) has become the method of choice to identify and quantify proteins, typically by fragmenting peptides and inferring protein identification by reference to sequence databases. Well-established programs have largely solved the problem of identifying peptides in complex mixtures. However, to prevent the search space from becoming prohibitively large, most search engines need a list of expected modifications. Therefore, unexpected modifications limit both the identification of proteins and peptide-based quantification. We developed mass spectrometry-peak shift analysis (MS-PSA) to rapidly identify related spectra in large data sets without reference to databases or specified modifications. Peptide identifications from established tools, such as MASCOT or SEQUEST, may be propagated onto MS-PSA results. Modification of a peptide alters the mass of the precursor ion and some of the fragmentation ions. MS-PSA identifies characteristic fragmentation masses from MS/MS spectra. Related spectra are identified by pattern matching of unchanged and mass-shifted fragment ions. We illustrate the use of MS-PSA with simple and complex mixtures with both high and low mass accuracy data sets. MS-PSA is not limited to the analysis of peptides but can be used for the identification of related groups of spectra in any set of fragmentation patterns.
We present a simple, intuitive, and effective approach for network clustering. It is based on basic concepts of linear algebra such as efficient calculation of spanning trees, and can be implemented in a few lines of code. We introduce the node separation measure spanning tree separation (STS) and the corresponding graph distance measure spanning tree vector similarity distance (STVSD). We demonstrate that the STS is a link salience measure able to identify the backbone of networks. The STVSD is used to reveal the hierarchical community structure of networks. We show that it, together with the clustering quality measure partition density, is on a par with the best graph or network clustering methods known, in terms of both quality and efficiency. In perspective, we note that our approach could also handle weighted and directed networks and could be used for identification of overlapping communities.
Background Gene reshuffling, point mutations and horizontal gene transfer contribute to bacterial genome variation, but require the genome to rewire its transcriptional circuitry to ensure that inserted, mutated or reshuffled genes are transcribed at appropriate levels. The genomes of Epsilonproteobacteria display very low synteny, due to high levels of reshuffling and reorganisation of gene order, but still share a significant number of gene orthologs allowing comparison. Here we present the primary transcriptome of the pathogenic Epsilonproteobacterium Campylobacter jejuni , and have used this for comparative and predictive transcriptomics in the Epsilonproteobacteria. Results Differential RNA-sequencing using 454 sequencing technology was used to determine the primary transcriptome of C. jejuni NCTC 11168, which consists of 992 transcription start sites (TSS), which included 29 putative non-coding and stable RNAs, 266 intragenic (internal) TSS, and 206 antisense TSS. Several previously unknown features were identified in the C. jejuni transcriptional landscape, like leaderless mRNAs and potential leader peptides upstream of amino acid biosynthesis genes. A cross-species comparison of the primary transcriptomes of C. jejuni and the related Epsilonproteobacterium Helicobacter pylori highlighted a lack of conservation of operon organisation, position of intragenic and antisense promoters or leaderless mRNAs. Predictive comparisons using 40 other Epsilonproteobacterial genomes suggests that this lack of conservation of transcriptional features is common to all Epsilonproteobacterial genomes, and is associated with the absence of genome synteny in this subdivision of the Proteobacteria. Conclusions Both the genomes and transcriptomes of Epsilonproteobacteria are highly variable, both at the genome level by combining and division of multicistronic operons, but also on the gene level by generation or deletion of promoter sequences and 5′ untranslated regions. Regulatory features may have evolved after these species split from a common ancestor, with transcriptome rewiring compensating for changes introduced by genomic reshuffling and horizontal gene transfer.
Einleitung: Aktuelle Gen-Chip (Array-CGH) Analysen bieten eine zunehmend höhere Auflösung die prinzipiell eine Untersuchung von kleinen homozygoten bzw. heterozygoten chromosomalen Deletionen ermöglicht. Phänomene wie Verlust der Heterozygotie (LOH), oder homozygote Deletionen führen zur Inaktivierung von Tumorsuppressorgenen und tragen somit grundlegend zur Tumorgenese bei. Die vorliegende Studie untersucht mittels aktueller Chip-Analyse das Genom verschiedener Zelllinien des kolorektalen Karzinoms auf rekurrente LOHs und Deletionen und vergleicht die Ergebnisse mit denen der konventionellen Zytogenetik.
BACKGROUND:Expansion of multi-C2H2 domain zinc finger (ZNF) genes, including the Krüppel-associated box (KRAB) subfamily, paralleled the evolution of tetrapodes, particularly in mammalian lineages. Advances in their cataloging and characterization suggest that the functions of the KRAB-ZNF gene family contributed to mammalian speciation.RESULTS:Here, we characterized the human 8q24.3 ZNF cluster on the genomic, the phylogenetic, the structural and the transcriptome level. Six (ZNF7, ZNF34, ZNF250, ZNF251, ZNF252, ZNF517) of the seven locus members contain exons encoding KRAB domains, one (ZNF16) does not. They form a paralog group in which the encoded KRAB and ZNF protein domains generally share more similarities with each other than with other members of the human ZNF superfamily. The closest relatives with respect to their DNA-binding domain were ZNF7 and ZNF251. The analysis of orthologs in therian mammalian species revealed strong conservation and purifying selection of the KRAB-A and zinc finger domains. These findings underscore structural/functional constraints during evolution. Gene losses in the murine lineage (ZNF16, ZNF34, ZNF252, ZNF517) and potential protein truncations in primates (ZNF252) illustrate ongoing speciation processes. Tissue expression profiling by quantitative real-time PCR showed similar but distinct patterns for all tested ZNF genes with the most prominent expression in fetal brain. Based on accompanying expression signatures in twenty-six other human tissues ZNF34 and ZNF250 revealed the closest expression profiles. Together, the 8q24.3 ZNF genes can be assigned to a cerebellum, a testis or a prostate/thyroid subgroup. These results are consistent with potential functions of the ZNF genes in morphogenesis and differentiation. Promoter regions of the seven 8q24.3 ZNF genes display common characteristics like missing TATA-box, CpG island-association and transcription factor binding site (TFBS) modules. Common TFBS modules partly explain the observed expression pattern similarities.CONCLUSIONS:The ZNF genes at human 8q24.3 form a relatively old mammalian paralog group conserved in eutherian mammals for at least 130 million years. The members persisted after initial duplications by undergoing subfunctionalizations in their expression patterns and target site recognition. KRAB-ZNF mediated repression of transcription might have shaped organogenesis in mammalian ontogeny.
Molecular interaction networks establish all cell biological processes. The networks are under intensive research that is facilitated by new high-throughput measurement techniques for the detection, quantification, and characterization of molecules and their physical interactions. For the common model organism yeast Saccharomyces cerevisiae, public databases store a significant part of the accumulated information and, on the way to better understanding of the cellular processes, there is a need to integrate this information into a consistent reconstruction of the molecular interaction network. This work presents and validates RefRec, the most comprehensive molecular interaction network reconstruction currently available for yeast. The reconstruction integrates protein synthesis pathways, a metabolic network, and a protein-protein interaction network from major biological databases. The core of the reconstruction is based on a reference object approach in which genes, transcripts, and proteins are identified using their primary sequences. This enables their unambiguous identification and non-redundant integration. The obtained total number of different molecular species and their connecting interactions is approximately 67,000. In order to demonstrate the capacity of RefRec for functional predictions, it was used for simulating the gene knockout damage propagation in the molecular interaction network in approximately 590,000 experimentally validated mutant strains. Based on the simulation results, a statistical classifier was subsequently able to correctly predict the viability of most of the strains. The results also showed that the usage of different types of molecular species in the reconstruction is important for accurate phenotype prediction. In general, the findings demonstrate the benefits of global reconstructions of molecular interaction networks. With all the molecular species and their physical interactions explicitly modeled, our reconstruction is able to serve as a valuable resource in additional analyses involving objects from multiple molecular -omes. For that purpose, RefRec is freely available in the Systems Biology Markup Language format.
Many large 'omics' datasets have been published and many more are expected in the near future. New analysis methods are needed for best exploitation. We have developed a graphical user interface (GUI) for easy data analysis. Our discovery of all significant substructures (DASS) approach elucidates the underlying modularity, a typical feature of complex biological data. It is related to biclustering and other data mining approaches. Importantly, DASS-GUI also allows handling of multi-sets and calculation of statistical significances. DASS-GUI contains tools for further analysis of the identified patterns: analysis of the pattern hierarchy, enrichment analysis, module validation, analysis of additional numerical data, easy handling of synonymous names, clustering, filtering and merging. Different export options allow easy usage of additional tools such as Cytoscape.
DiProDB (http://diprodb.fli-leibniz.de) is a database of conformational and thermodynamic dinucleotide properties. It includes datasets both for DNA and RNA, as well as for single and double strands. The data have been shown to be important for understanding different aspects of nucleic acid structure and function, and they can also be used for encoding nucleic acid sequences. The database is intended to facilitate further applications of dinucleotide properties. A number of property datasets is highly correlated. Therefore, the database comes with a correlation analysis facility. Authors having determined new sets of dinucleotide property values are invited to submit these data to DiProDB.
BACKGROUND:Bistability underlies basic biological phenomena, such as cell division, differentiation, cancer onset, and apoptosis. So far biologists identified two necessary conditions for bistability: positive feedback and ultrasensitivity.RESULTS:Biological systems are based upon elementary mono- and bimolecular chemical reactions. In order to definitely clarify all necessary conditions for bistability we here present the corresponding minimal system. According to our definition, it contains the minimal number of (i) reactants, (ii) reactions, and (iii) terms in the corresponding ordinary differential equations (decreasing importance from i-iii). The minimal bistable system contains two reactants and four irreversible reactions (three bimolecular, one monomolecular).We discuss the roles of the reactions with respect to the necessary conditions for bistability: two reactions comprise the positive feedback loop, a third reaction filters out small stimuli thus enabling a stable 'off' state, and the fourth reaction prevents explosions. We argue that prevention of explosion is a third general necessary condition for bistability, which is so far lacking discussion in the literature.Moreover, in addition to proving that in two-component systems three steady states are necessary for bistability (five for tristability, etc.), we also present a simple general method to design such systems: one just needs one production and three different degradation mechanisms (one production, five degradations for tristability, etc.). This helps modelling multistable systems and it is important for corresponding synthetic biology projects.CONCLUSION:The presented minimal bistable system finally clarifies the often discussed question for the necessary conditions for bistability. The three necessary conditions are: positive feedback, a mechanism to filter out small stimuli and a mechanism to prevent explosions. This is important for modelling bistability with simple systems and for synthetically designing new bistable systems. Our simple model system is also well suited for corresponding teaching purposes.
Motivation: DiProGB is an easy to use new genome browser that encodes the primary nucleotide sequence by thermodynamical and geometrical dinucleotide properties. The nucleotide sequence is thus converted into a sequence graph. This visualization, supported by different graph manipulation options, facilitates genome analyses, because the human brain can process visual information better than textual information. Also, DiProGB can identify genomic regions where certain physical properties are more conserved than the nucleotide sequence itself. Most of the DiProGB tools can be applied to both, the primary nucleotide sequence and the sequence graph. They include motif and repeat searches as well as statistical analyses. DiProGB adds a new dimension to the common genome analysis approaches by taking into account the physical properties of DNA and RNA. Availability and Implementation: Source code and binaries are freely available for download at http://diprogb.fli-leibniz.de, implemented in C++ and supported on MS Windows and Linux (using e.g. WineHQ). Contact: maikfr@fli-leibniz.de; thomas.wilhelm@bbsrc.ac.uk
Many papers published in recent years show that real-world graphs G(n,m) (n nodes, m edges) are more or less “complex” in the sense that different topological features deviate from random graphs. Here we narrow the definition of graph complexity and argue that a complex graph contains many different subgraphs. We present different measures that quantify this complexity, for instance C1e, the relative number of non-isomorphic one-edge-deleted subgraphs (i.e. DECK size). However, because these different subgraph measures are computationally demanding, we also study simpler complexity measures focussing on slightly different aspects of graph complexity. We consider heuristically defined “product measures”, the products of two quantities which are zero in the extreme cases of a path and clique, and “entropy measures” quantifying the diversity of different topological features. The previously defined network/graph complexity measures Medium Articulation and Offdiagonal complexity (OdC) belong to these two classes. We study OdC measures in some detail and compare it with our new measures. For all measures, the most complex graph GCmax has a medium number of edges, between the edge numbers of the minimum and the maximum connected graph n−1<m(GCmax)<n(n−1)/2. Interestingly, for some measures C̃ this number scales exactly with the geometric mean of the extremes: m(GC̃max)=n/2(n−1)∼n1.5. All graph complexity measures are characterized with the help of different example graphs. For all measures the corresponding time complexity is given.
We present a new data structure, called a Decomposition Tree (DT), for analysing Boolean functions, and demonstrate a variety of applications. In each node of the DT, appropriate bit-string decomposition fragments are combined by a logical operator. The DT has 2k nodes in the worst case, which implies exponential complexity for problems where the whole tree has to be considered. However, it is important to note that many problems are simpler. We show that these can be handled in an efficient way using the DT. Nevertheless, many problems are of exponential complexity and cannot be made any simpler: for example, the calculation of prime implicants. Using our general DT structure, we present a new worst case algorithm to compute all prime implicants. This algorithm has a lower time complexity than the well-known Quine–McCluskey algorithm and is the fastest corresponding worst case algorithm so far.
We present a simple new method to systematically identify all topological structures (e.g., positive feedback loops) potentially leading to locally unstable steady states: ICSA-The instability causing structure analysis. Systems without any instability causing structure (i.e., not fulfilling the necessary topological condition for instabilities) cannot have unstable steady states. It follows that common bistability or multistability and Hopf bifurcations are excluded and sustained oscillations and deterministic chaos are most unlikely. The ICSA leads to new insights into the topological organization of chemical and biochemical systems, such as metabolic, gene regulatory, and signal transduction networks.
Genome-wide gene expression was comparatively investigated in early-passage rheumatoid arthritis (RA) and osteoarthritis (OA) synovial fibroblasts (SFBs; n = 6 each) using oligonucleotide microarrays; mRNA/protein data were validated by quantitative PCR (qPCR) and western blotting and immunohistochemistry, respectively. Gene set enrichment analysis (GSEA) of the microarray data suggested constitutive upregulation of components of the transforming growth factor (TGF)-β pathway in RA SFBs, with 2 hits in the top 30 regulated pathways. The growth factor TGF-β1, its receptor TGFBR1, the TGF-β binding proteins LTBP1/2, the TGF-β-releasing thrombospondin 1 (THBS1), the negative effector SkiL, and the smad-associated molecule SARA were upregulated in RA SFBs compared to OA SFBs, whereas TGF-β2 was downregulated. Upregulation of TGF-β1 and THBS1 mRNA (both positively correlated with clinical markers of disease activity/severity) and downregulation of TGF-β2 mRNA in RA SFBs were confirmed by qPCR. TGFBR1 mRNA (only numerically upregulated in RA SFBs) and SkiL mRNA were not differentially expressed. At the protein level, TGF-β1 showed a slightly higher expression, and the signal-transducing TGFBR1 and the TGF-β-activating THBS1 a significantly higher expression in RA SFBs than in OA SFBs. Consistent with the upregulated TGF-β pathway in RA SFBs, stimulation with TGF-β1 resulted in a significantly enhanced expression of matrix-metalloproteinase (MMP)-11 mRNA and protein in RA SFBs, but not in OA SFBs. In conclusion, RA SFBs show broad, constitutive alterations of the TGF-β pathway. The abundance of TGF-β, in conjunction with an augmented mRNA and/or protein expression of TGF-β-releasing THBS1 and TGFBR1, suggests a pathogenetic role of TGF-β-induced effects on SFBs in RA, for example, the augmentation of MMP-mediated matrix degradation/remodeling.
Recent analyses indicate that differences in protein concentrations are only 20%-40% attributable to variable mRNA levels, underlining the importance of posttranscriptional regulation. Generally, protein concentrations depend on the translation rate (which is proportional to the translational activity, TA) and the degradation rate. By integrating 12 publicly available large-scale datasets and additional database information of the yeast Saccharomyces cerevisiae, we systematically analyzed five factors contributing to TA: mRNA concentration, ribosome density, ribosome occupancy, the codon adaptation index, and a newly developed "tRNA adaptation index.'' Our analysis of the functional relationship between the TA and measured protein concentrations suggests that the TA follows Michaelis-Menten kinetics. The calculated TA, together with measured protein concentrations, allowed us to estimate degradation rates for 4,125 proteins under standard conditions. A significant correlation to recently published degradation rates supports our approach. Moreover, based on a newly developed scoring system, we identified and analyzed genes subjected to the posttranscriptional regulation mechanism, translation on demand. Next we applied these findings to publicly available data of protein and mRNA concentrations under four stress conditions. The integration of these measurements allowed us to compare the condition-specific responses at the posttranscriptional level. Our analysis of all 62 proteins that have been measured under all four conditions revealed proteins with very specific posttranscriptional stress response, in contrast to more generic responders, which were nonspecifically regulated under several conditions. The concept of specific and generic responders is known for transcriptional regulation. Here we show that it also holds true at the posttranscriptional level.
Complex cellular processes are accomplished by the concerted action of hierarchically organized functional modules. Protein complexes are major components which act as highly specialized molecular machines. Here we present a statistical procedure to find insightful substructures in protein complexes based on large-scale protein complex data: we identify statistically significant common protein subcomplexes (SCs) contained in different protein complexes. We analyze recently published data of the two model organisms Saccharomyces cerevisiae (four different data sets) and Escherichia coli, as well as human protein complex data. Our method identifies well-characterized protein assemblies with known functions which act as own functional entities in the cell. In addition, we also identified hitherto unknown functional entities that should be studied experimentally in future. We discuss two typical properties of protein subcomplexes: 1) subcomplexes are enriched with essential proteins (which implies that the whole SCs may be strongly conserved) and 2) SCs are functionally and spatially more homogeneous than the experimentally found protein assemblies. The latter property is exploited to propose functions for so far unknown proteins of S. cerevisiae.
We present a generalised framework for analysing structural robustness of metabolic networks, based on the concept of elementary flux modes (EFMs). Extending our earlier study on single knockouts [Wilhelm, T., Behre, J., Schuster, S., 2004. Analysis of structural robustness of metabolic networks. IEE Proc. Syst. Biol. 1(1), 114-120], we are now considering the general case of double and multiple knockouts. The robustness measures are based on the ratio of the number of remaining EFMs after knockout vs. the number of EFMs in the unperturbed situation, averaged over all combinations of knockouts. With the help of simple examples we demonstrate that consideration of multiple knockouts yields additional information going beyond single-knockout results. It is proven that the robustness score decreases as the knockout depth increases. We apply our extended framework to metabolic networks representing amino acid anabolism in Escherichia coli and human hepatocytes, and the central metabolism in human erythrocytes. Moreover, in the E. coli model the two subnetworks synthesising amino acids that are essential and those that are non-essential for humans are studied separately. The results are discussed from an evolutionary viewpoint. We find that E. coli has the most robust metabolism of all the cell types studied here. Considering only the subnetwork of the synthesis of non-essential amino acids, E. coli and the human hepatocyte show about the same robustness.
We present a new information theoretic approach for network characterizations. It is developed to describe the general type of networks with n nodes and L directed and weighted links, i.e., it also works for the simpler undirected and unweighted networks. The new information theoretic measures for network characterizations are based on a transmitter-receiver analogy of effluxes and influxes. Based on these measures, we classify networks as either complex or non-complex and as either democracy or dictatorship networks. Directed networks, in particular, are furthermore classified as either information spreading and information collecting networks.The complexity classification is based on the information theoretic network complexity measure medium articulation (MA). It is proven that special networks with a medium number of links (L similar to n(1.5)) show the theoretical maximum complexity MA = (log n)(2)/2. A network is complex if its MA is larger than the average MA of appropriately randomized networks: MA > MA(r). A network is of the democracy type if its redundancy R < R-r, otherwise it is a dictatorship network. In democracy networks all nodes are, on average, of similar importance, whereas in dictatorship networks some nodes play distinguished roles in network functioning. In other words, democracy networks are characterized by cycling of information (or mass, or energy), while in dictatorship networks there is a straight through-flow from sources to sinks. The classification of directed networks into information spreading and information collecting networks is based on the conditional entropies of the considered networks (H(A/B) = uncertainty of sender node if receiver node is known, H(B/A) = uncertainty of receiver node if sender node is known): if H(A/B) > H(B/A), it is an information collecting network, otherwise an information spreading network.Finally, different real networks (directed and undirected, weighted and unweighted) are classified according to our general scheme. (C) 2007 Elsevier B.V. All rights reserved.
MOTIVATION:Pattern identification in biological sequence data is one of the main objectives of bioinformatics research. However, few methods are available for detecting patterns (substructures) in unordered datasets. Data mining algorithms mainly developed outside the realm of bioinformatics have been adapted for that purpose, but typically do not determine the statistical significance of the identified patterns. Moreover, these algorithms do not exploit the often modular structure of biological data.RESULTS:We present the algorithm DASS (Discovery of All Significant Substructures) that first identifies all substructures in unordered data (DASS(Sub)) in a manner that is especially efficient for modular data. In addition, DASS calculates the statistical significance of the identified substructures, for sets with at most one element of each type (DASS(P(set))), or for sets with multiple occurrence of elements (DASS(P(mset))). The power and versatility of DASS is demonstrated by four examples: combinations of protein domains in multi-domain proteins, combinations of proteins in protein complexes (protein subcomplexes), combinations of transcription factor target sites in promoter regions and evolutionarily conserved protein interaction subnetworks.AVAILABILITY:The program code and additional data are available at http://www.fli-leibniz.de/tsb/DASS