
Figure 1 – Hits to Campylobacter found using BLAST (red) and Kraken (blue) Detection of Campylobacter in chicken faecal samples is possible from 106 CFU/g using BLAST (red bars) and from 104 CFU/g using Kraken (blue bars). For BLAST results hits are number of contigs matching Campylobacter in proportion to the total number of contigs. For Kraken results hits are number of reads assigned to Campylobacter in proportion to the total number of reads.
Our research group is currently developing software for estimating large-scale gene networks from gene expression data. The software, called SiGN, is specifically designed for the Japanese flagship supercomputer "K computer" which is planned to achieve 10 petaflops in 2012, and other high performance computing environments including Human Genome Center (HGC) supercomputer system. SiGN is a collection of gene network estimation software with three different sub-programs: SiGN-BN, SiGN-SSM and SiGN-L1. In these three programs, five different models are available: static and dynamic nonparametric Bayesian networks, state space models, graphical Gaussian models, and vector autoregressive models. All these models require a huge amount of computational resources for estimating large-scale gene networks and therefore are designed to be able to exploit the speed of 10 petaflops. The software will be available freely for "K computer" and HGC supercomputer system users. The estimated networks can be viewed and analyzed by Cell Illustrator Online and SBiP (Systems Biology integrative Pipeline). The software project web site is available at http://sign.hgc.jp/ .
Elucidating protein-RNA interactions (PRIs) is important for understanding many cellular systems. We developed a PRI prediction method by using a rigid-body protein-RNA docking calculation with tertiary structure data. We evaluated this method by using 78 protein-RNA complex structures from the Protein Data Bank. We predicted the interactions for pairs in 78×78 combinations. Of these, 78 original complexes were defined as positive pairs, and the other 6,006 complexes were defined as negative pairs; then an F-measure value of 0.465 was obtained with our prediction system.
When the DNA damage is generated, the tumor suppressor gene p53 is activated and selects the cell fate such as the cell cycle arrest, the DNA repair and the induction of apoptosis. Recently, the p53 oscillation was observed in MCF7 cell line. However, the biological meaning of p53 oscillation was still unclear. Here, we constructed a novel mathematical model of cell cycle regulatory system with p53 signaling network to investigate the relationship between the p53 oscillation and the cell cycle progression. First, the simulated result without DNA damage agreed with the biological findings. Next, the simulations with DNA damage realized both the p53 oscillation and the cell cycle arrest, and indicated that the generation of multiple p53 pulses disrupted the cell cycle progression. Moreover, the simulated results showed that the cell cycle disruption was caused by the catastrophe of M phase in the cell cycle, which resulted from the decline in cyclin A/cyclin-dependent kinase 2. The results in this study suggested that the generation of multiple p53 pulses against DNA damage may be used as a marker of cell cycle disruption.
A wiki-based repository for crude drugs and Kampo medicine is introduced. It provides taxonomic and chemical information for 158 crude drugs and 348 prescriptions of the traditional Kampo medicine in Japan, which is a variation of ancient Chinese medicine. The system is built on MediaWiki with extensions for inline page search and for sending user-input elements to the server. These functions together realize implementation of word checks and data integration at the user-level. In this scheme, any user can participate in creating an integrated database with controlled vocabularies on the wiki system. Our implementation and data are accessible at http://metabolomics.jp/wiki/.
We developed linear regression models which predict strength of transcriptional activity of promoters from their sequences. Intrinsic transcriptional strength data of 451 human promoter sequences in three cell lines (HEK293, MCF7 and 3T3), which were measured by systematic luciferase reporter gene assays, were used to build the models. The models sum up contributions of CG dinucleotide content and transcription factor binding sites (TFBSs) to transcriptional strength. We evaluated prediction accuracies of the models by cross validation tests and found that they have adequate ability for predicting transcriptional strength of promoters in spite of their simple formalization. We also evaluated statistical significance of the contributions and proposed a picture of regulatory code hidden in promoter sequences. That is, CG dinucleotide content and TFBSs mainly determine strength of transcriptional activity under ubiquitous and specific environments, respectively.
Understanding the evolution and dynamics of metabolism in microbial ecosystems is an ongoing challenge in microbiology. A promising approach towards this goal is the extension of genome-scale flux balance models of metabolism to multiple interacting species. However, since the detailed distribution of metabolic functions among ecosystem members is often unknown, it is important to investigate how compartmentalization of metabolites and reactions affects flux balance predictions. Here, as a first step in this direction, we address the importance of compartmentalization in the well characterized metabolic model of the yeast Saccharomyces cerevisiae, which we treat as an "ecosystem of organelles". In addition to addressing the impact that the removal of compartmentalization has on model predictions, we show that by systematically constraining some individual fluxes in a de-compartmentalized version of the model we can significantly reduce the flux prediction errors induced by the removal of compartments. We expect that our analysis will help predict and understand metabolic functions in complex microbial communities. In addition, further study of yeast as an ecosystem of organelles might provide novel insight on the evolution of endosymbiosis and multicellularity.
The set of chemicals producible and usable by metabolic pathways must have evolved in parallel with the enzymes that catalyze them. One implication of this common historical path should be a correspondence between the innovation steps that gradually added new metabolic reactions to the biosphere-level biochemical toolkit, and the gradual sequence changes that must have slowly shaped the corresponding enzyme structures. However, global signatures of a long-term co-evolution have not been identified. Here we search for such signatures by computing correlations between inter-reaction distances on a metabolic network, and sequence distances of the corresponding enzyme proteins. We perform our calculations using the set of all known metabolic reactions, available from the KEGG database. Reaction-reaction distance on the metabolic network is computed as the length of the shortest path on a projection of the metabolic network, in which nodes are reactions and edges indicate whether two reactions share a common metabolite, after removal of cofactors. Estimating the distance between enzyme sequences in a meaningful way requires some special care: for each enzyme commission (EC) number, we select from KEGG a consensus set of protein sequences using the cluster of orthologous groups of proteins (COG) database. We define the evolutionary distance between protein sequences as an asymmetric transition probability between two enzymes, derived from the corresponding pair-wise BLAST scores. By comparing the distances between sequences to the minimal distances on the metabolic reaction graph, we find a small but statistically significant correlation between the two measures. This suggests that the evolutionary walk in enzyme sequence space has locally mirrored, to some extent, the gradual expansion of metabolism.
Although microarray technology has revealed transcriptomic diversities underlining various cancer phenotypes, transcriptional programs controlling them have not been well elucidated. To decode transcriptional programs governing cancer transcriptomes, we have recently developed a computational method termed EEM, which searches for expression modules from prescribed gene sets defined by prior biological knowledge like TF binding motifs. In this paper, we extend our EEM approach to predict cancer transcriptional networks. Starting from functional TF binding motifs and expression modules identified by EEM, we predict cancer transcriptional networks containing regulatory TFs, associated GO terms, and interactions between TF binding motifs. To systematically analyze transcriptional programs in broad types of cancer, we applied our EEM-based network prediction method to 122 microarray datasets collected from public databases. The data sets contain about 15000 experiments for tumor samples of various tissue origins including breast, colon, lung etc. This EEM based meta-analysis successfully revealed a prevailing cancer transcriptional network which functions in a large fraction of cancer transcriptomes; they include cell-cycle and immune related sub-networks. This study demonstrates broad applicability of EEM, and opens a way to comprehensive understanding of transcriptional networks in cancer cells.
MOTIVATION:Methods like FBA and kinetic modeling are widely used to calculate fluxes in metabolic networks. For the analysis and understanding of simulation results and experimentally measured fluxes visualization software within the network context is indispensable.RESULTS:We present Flux Viz, an open-source Cytoscape plug-in for the visualization of flux distributions in molecular interaction networks. FluxViz supports (i) import of networks in a variety of formats (SBML, GML, XGMML, SIF, BioPAX, PSI-MI) (ii) import of flux distributions as CSV, Cytoscape attributes or VAL files (iii) limitation of views to flux carrying reactions (flux subnetwork) or network attributes like localization (iv) export of generated views (SVG, EPS, PDF, BMP, PNG). Though FluxViz was primarily developed as tool for the visualization of fluxes in metabolic networks and the analysis of simulation results from FASIMU, a flexible software for batch flux-balance computation in large metabolic networks, it is not limited to biochemical reaction networks and FBA but can be applied to the visualization of arbitrary fluxes in arbitrary graphs.AVAILABILITY:The platform-independent program is an open-source project, freely available at http://sourceforge.net/projects/fluxvizplugin/ under GNU public license, including manual, tutorial and examples.
Protein-Protein interactions play an important role in many cellular processes. However experimental determination of the protein complex structure is quite difficult and time consuming. Hence, there is need for fast and accurate in silico protein docking methods. These methods generally consist of two stages: (i) a sampling algorithm that generates a large number of candidate complex geometries (decoys), and (ii) a scoring function that ranks these decoys such that nearnative decoys are higher ranked than other decoys. We have recently developed a neural network based scoring function that performed better than other state-of-the-art scoring functions on a benchmark of 65 protein complexes. Here, we use similar ideas to develop a method that is based on linear scoring functions. We compare the linear scoring function of the present study with other knowledge-based scoring functions such as ZDOCK 3.0, ZRANK and the previously developed neural network. Despite its simplicity the linear scoring function performs as good as the compared state-of-the-art methods and predictions are simple and rapid to compute.
Signaling pathways are often represented by networks where each node corresponds to a protein and each edge corresponds to a relationship between nodes such as activation, inhibition and binding. However, such signaling pathways in a cell may be affected by genetic and epigenetic alteration. Some edges may be deleted and some edges may be newly added. The current knowledge about known signaling pathways is available on some public databases, but most of the signaling pathways including changes upon the cell state alterations remain largely unknown. In this paper, we develop an integer programming-based method for inferring such changes by using gene expression data. We test our method on its ability to reconstruct the pathway of colorectal cancer in the KEGG database.
Many cofactors and nucleotides containing sulfur atoms are known to have important functions in a variety of organisms. Recently, the biosynthetic pathways of these sulfur containing compounds have been revealed, where many enzymes relay sulfur atoms. Increasing evidence also suggests that the prokaryotic sulfur-relay enzymes might be the evolutionary origin of ubiquitination and the related systems that control a wide range of physiological processes in eukaryotic cells. However, these sulfur-relay enzymes have been studied in only a small number of organisms. Here we carried out comparative genomic analysis and examined the presence and absence of sulfurtransferases utilized in the biosynthetic pathways of molybdenum cofactor (Moco), 2-thiouridine (S(2)U), and 4-thiouridine (S(4)U), and IscS, a cysteine desulfurase. We found that all eukaryotes and many other organisms lack the intermediate enzymes in S(2)U biosynthesis. It is also found that most genes lack rhodanese homology domain (RHD), a catalytic domain of sulfurtransferase. Some organisms have a conserved sequence composed of about 100 residues in the C terminus of TusA, different from RHD. Host-associated organisms have a tendency to lose Moco biosynthetic enzymes, and some organisms have MoaD-MoaE fusion protein. Our findings suggest that sulfur-relay pathways have been so diversified that some putative sulfurtransferases possibly function in other unknown pathways.
We address an issue of detecting a switching mechanism in gene expression, where two genes are positively correlated for one experimental condition while they are negatively correlated for another. We compare the performance of existing methods for this issue, roughly divided into two types: interaction test (IT) and the difference of correlation coefficients. Interaction test, currently a standard approach for detecting epistasis in genetics, is the log-likelihood ratio test between two logistic regressions with/without an interaction term, resulting in checking the strength of interaction between two genes. On the other hand, two correlation coefficients can be computed for two experimental conditions and the difference of them shows the alteration of expression trends in a more straightforward manner. In our experiments, we tested three different types of correlation coefficients: Pearson, Spearman and a midcorrelation (biweight midcorrelation). The experiment was performed by using ~ 2.3 × 10(9) combinations selected out of the GEO (Gene Expression Omnibus) database. We sorted all combinations according to the p-values of IT or by the absolute values of the difference of correlation coefficients and then visually evaluated the top ranked combinations in terms of the switching mechanism. The result showed that 1) combinations detected by IT included non-switching combinations and 2) Pearson was affected by outliers easily while Spearman and the midcorrelation seemed likely to avoid them.
We develop a general method to identify gene networks from pair-wise correlations between genes in a microarray data set and apply it to a public prostate cancer gene expression data from 69 primary prostate tumors. We define the degree of a node as the number of genes significantly associated with the node and identify hub genes as those with the highest degree. The correlation network was pruned using transcription factor binding information in VisANT (http://visant.bu.edu/) as a biological filter. The reliability of hub genes was determined using a strict permutation test. Separate networks for normal prostate samples, and prostate cancer samples from African Americans (AA) and European Americans (EA) were generated and compared. We found that the same hubs control disease progression in AA and EA networks. Combining AA and EA samples, we generated networks for low low (<7) and high (≥7) Gleason grade tumors. A comparison of their major hubs with those of the network for normal samples identified two types of changes associated with disease: (i) Some hub genes increased their degree in the tumor network compared to their degree in the normal network, suggesting that these genes are associated with gain of regulatory control in cancer (e.g. possible turning on of oncogenes). (ii) Some hubs reduced their degree in the tumor network compared to their degree in the normal network, suggesting that these genes are associated with loss of regulatory control in cancer (e.g. possible loss of tumor suppressor genes). A striking result was that for both AA and EA tumor samples, STAT5a, CEBPB and EGR1 are major hubs that gain neighbors compared to the normal prostate network. Conversely, HIF-lα is a major hub that loses connections in the prostate cancer network compared to the normal prostate network. We also find that the degree of these hubs changes progressively from normal to low grade to high grade disease, suggesting that these hubs are master regulators of prostate cancer and marks disease progression. STAT5a was identified as a central hub, with ~120 neighbors in the prostate cancer network and only 81 neighbors in the normal prostate network. Of the 120 neighbors of STAT5a, 57 are known cancer related genes, known to be involved in functional pathways associated with tumorigenesis. Our method is general and can easily be extended to identify and study networks associated with any two phenotypes.
DNA replication is a fundamental process that is tightly regulated during the cell cycle. In budding yeast it starts from multiple origins of replication and proceeds in a timely fashion according to a reproducible temporal program until the entire DNA is replicated exactly once per cell cycle. In this program an origin seems to have an inherent firing probability at a specific time in S-phase that is conserved over the population. However, what exactly determines the origin initiation time remains obscure. In this work, we analyze the gene content that clusters around replication origins following the assumption that inherent origin properties that determine staggered initiation times could potentially be mirrored in the close origin proximity. We perform a Gene Ontology term enrichment test and find that metabolic genes are significantly over-represented in the regions that are close to the starting points of DNA replication. Furthermore, functional analysis also reveals that catabolic genes cluster around early firing origins, whereas anabolic genes can rather be found in the proximity of late firing origins of replication. We speculate that, in budding yeast, gene function around replication origins correlates with their intrinsic probability to initiate DNA replication at a given point in S-phase.
Several technologies are currently used for gene expression profiling, such as Real Time RT-PCR, microarray and CAGE (Cap Analysis of Gene Expression). CAGE is a recently developed method for constructing transcriptome maps and it has been successfully applied to analyzing gene expressions in diverse biological studies. The principle of CAGE has been developed to address specific issues such as determination of transcriptional starting sites, the study of promoter regions and identification of new transcripts. Here, we present both quantitative and qualitative comparisons among three major gene expression quantification techniques, namely: CAGE, illumina microarray and Real Time RT-PCR, by showing that the quantitative values of each method are not interchangeable, however, each of them has unique characteristics which render all of them essential and complementary. Understanding the advantages and disadvantages of each technology will be useful in selecting the most appropriate technique for a determined purpose.
In healthy individuals, dehydration of the body leads to release of the hormone vasopressin from the pituitary. Via the bloodstream, vasopressin reaches the collecting duct cells in the kidney, where the water channel Aquaporin-2 (AQP2) is expressed. After stimulation of the vasopressin V2 receptor by vasopressin, intracellular AQP2-containing vesicles fuse with the apical plasma membrane of the collecting duct cells. This leads to increased water reabsorption from the pro-urine into the blood and therefore to enhanced retention of water within the body. Using existing biological data we propose a mathematical model of AQP-2 trafficking and regulation in collecting duct cells. Our model includes the vasopressin receptor, adenylate cyclase, protein kinase A, and intracellular as well as membrane located AQP2. To model the chemical reactions we used ordinary differential equations (ODEs) based on mass action kinetics. We employ known protein concentrations and time series data to estimate the kinetic parameters of our model and demonstrate its validity. Through generating, testing and ranking different versions of the model, we show that some model versions can describe the data well as soon as important regulatory parts such as the reduction of the signal by internalization of the vasopressin-receptor or the negative feedback loop representing phosphodiesterase activity are included. We perform time-dependent sensitivity analysis to identify the reactions that have the greatest influence on the cAMP and membrane located AQP2 levels over time. We predict the time courses for membrane located AQP2 at different vasopressin concentrations, compare them with newly generated data and discuss the competencies of the model.