Cryptococcus neoformans is an opportunistic fungal pathogen with a polysaccharide capsule that becomes greatly enlarged in the mammalian host and during in vitro growth under host-like conditions. To understand how individual environmental signals affect capsule size and gene expression, we grew cells in all combinations of five signals implicated in capsule size and systematically measured cell and capsule sizes. We also sampled these cultures over time and performed RNA-Seq in quadruplicate, yielding 881 RNA-Seq samples. Analysis of the resulting data sets showed that capsule induction in tissue culture medium, typically used to represent host-like conditions, requires the presence of either CO2 or exogenous cyclic AMP (cAMP). Surprisingly, adding either of these pushes overall gene expression in the opposite direction from tissue culture media alone, even though both are required for capsule development. Another unexpected finding was that rich medium blocks capsule growth completely. Statistical analysis further revealed many genes whose expression is associated with capsule thickness; deletion of one of these significantly reduced capsule size. Beyond illuminating capsule induction, our massive, uniformly collected dataset will be a significant resource for the research community.
Cryptococcus neoformans is an opportunistic fungal pathogen with a polysaccharide capsule that becomes greatly enlarged in the mammalian host and during in vitro growth under host-like conditions. To understand how individual environmental signals affect capsule size and gene expression, we grew cells in all combinations of five signals implicated in capsule size and systematically measured cell and capsule sizes. We also sampled these cultures over time and performed RNA-Seq in quadruplicate, yielding 881 RNA-Seq samples. Analysis of the resulting data sets showed that capsule induction in tissue culture medium, typically used to represent host-like conditions, requires the presence of either CO2 or exogenous cyclic AMP (cAMP). Surprisingly, adding either of these pushes overall gene expression in the opposite direction from tissue culture media alone, even though both are required for capsule development. Another unexpected finding was that rich medium blocks capsule growth completely. Statistical analysis further revealed many genes whose expression is associated with capsule thickness; deletion of one of these significantly reduced capsule size. Beyond illuminating capsule induction, our massive, uniformly collected dataset will be a significant resource for the research community.### Competing Interest StatementThe authors have declared no competing interest.
Cryptococcus neoformans is a ubiquitous, opportunistic fungal pathogen that kills almost 200,000 people worldwide each year. It is acquired when mammalian hosts inhale the infectious propagules; these are deposited in the lung and, in the context of immunocompromise, may disseminate to the brain and cause lethal meningoencephalitis. Once inside the host, C. neoformans undergoes a variety of adaptive processes, including secretion of virulence factors, expansion of a polysaccharide capsule that impedes phagocytosis, and the production of giant (Titan) cells. The transcription factor Pdr802 is one regulator of these responses to the host environment. Expression of the corresponding gene is highly induced under host-like conditions in vitro and is critical for C. neoformans dissemination and virulence in a mouse model of infection. Direct targets of Pdr802 include the quorum sensing proteins Pqp1, Opt1, and Liv3; the transcription factors Stb4, Zfc3, and Bzp4, which regulate cryptococcal brain infectivity and capsule thickness; the calcineurin targets Had1 and Crz1, important for cell wall remodeling and C. neoformans virulence; and additional genes related to resistance to host temperature and oxidative stress, and to urease activity. Notably, cryptococci engineered to lack Pdr802 showed a dramatic increase in Titan cells, which are not phagocytosed and have diminished ability to directly cross biological barriers. This explains the limited dissemination of pdr802 mutant cells to the central nervous system and the consequently reduced virulence of this strain. The role of Pdr802 as a negative regulator of Titan cell formation is thus critical for cryptococcal pathogenicity.IMPORTANCE The pathogenic yeast Cryptococcus neoformans presents a worldwide threat to human health, especially in the context of immunocompromise, and current antifungal therapy is hindered by cost, limited availability, and inadequate efficacy. After the infectious particle is inhaled, C. neoformans initiates a complex transcriptional program that integrates cellular responses and enables adaptation to the host lung environment. Here, we describe the role of the transcription factor Pdr802 in the response to host conditions and its impact on C. neoformans virulence. We identified direct targets of Pdr802 and also discovered that it regulates cellular features that influence movement of this pathogen from the lung to the brain, where it causes fatal disease. These findings significantly advance our understanding of a serious disease.
Received wisdom in the field of fungal biology holds that the process of editing a genome by transformation and homologous recombination is inherently mutagenic. However, that belief is based on circumstantial evidence. We provide the first direct measurement of the effects of transformation on a fungal genome by sequencing the genomes of 29 transformants and 30 untransformed controls with high coverage. Contrary to the received wisdom, our results show that transformation of DNA segments flanked by long targeting sequences, followed by homologous recombination and selection for a drug marker, is extremely safe. If a transformation deletes a gene, that may create selective pressure for a few compensatory mutations, but even when we deleted a gene, we found fewer than two point mutations per deletion strain, on average. We also tested these strains for changes in gene expression and found only a few genes that were consistently differentially expressed between the wild type and strains modified by genomic insertion of a drug resistance marker. As part of our report, we provide the assembled genome sequence of the commonly used laboratory strain Cryptococcus neoformans var. grubii strain KN99α.
The ability to rationally manipulate the transcriptional states of cells would be of great use in medicine and bioengineering. We have developed an algorithm, NetSurgeon, which uses genome-wide gene-regulatory networks to identify interventions that force a cell toward a desired expression state. We first validated NetSurgeon extensively on existing datasets. Next, we used NetSurgeon to select transcription factor deletions aimed at improving ethanol production in Saccharomyces cerevisiae cultures that are catabolizing xylose. We reasoned that interventions that move the transcriptional state of cells using xylose toward that of cells producing large amounts of ethanol from glucose might improve xylose fermentation. Some of the interventions selected by NetSurgeon successfully promoted a fermentative transcriptional state in the absence of glucose, resulting in strains with a 2.7-fold increase in xylose import rates, a 4-fold improvement in xylose integration into central carbon metabolism, or a 1.3-fold increase in ethanol production rate. We conclude by presenting an integrated model of transcriptional regulation and metabolic flux that will enable future efforts aimed at improving xylose fermentation to prioritize functional regulators of central carbon metabolism.
A critical step in understanding how a genome functions is determining which transcription factors (TFs) regulate each gene. Accordingly, extensive effort has been devoted to mapping TF networks. In Saccharomyces cerevisiae, protein–DNA interactions have been identified for most TFs by ChIP-chip, and expression profiling has been done on strains deleted for most TFs. These studies revealed that there is little overlap between the genes whose promoters are bound by a TF and those whose expression changes when the TF is deleted, leaving us without a definitive TF network for any eukaryote and without an efficient method for mapping functional TF networks. This paper describes NetProphet, a novel algorithm that improves the efficiency of network mapping from gene expression data. NetProphet exploits a fundamental observation about the nature of TF networks: The response to disrupting or overexpressing a TF is strongest on its direct targets and dissipates rapidly as it propagates through the network. Using S. cerevisiae data, we show that NetProphet can predict thousands of direct, functional regulatory interactions, using only gene expression data. The targets that NetProphet predicts for a TF are at least as likely to have sites matching the TF's binding specificity as the targets implicated by ChIP. Unlike most ChIP targets, the NetProphet targets also show evidence of functional regulation. This suggests a surprising conclusion: The best way to begin mapping direct, functional TF-promoter interactions may not be by measuring binding. We also show that NetProphet yields new insights into the functions of several yeast TFs, including a well-studied TF, Cbf1, and a completely unstudied TF, Eds1.
Determining the beginning and end positions of each exon in each protein coding gene within a genome can be difficult because the DNA patterns that signal a gene’s presence have multiple weakly related alternate forms and the DNA fragments that comprise a gene are generally small in comparison to the size of the genome. In response to this challenge, automated gene predictors were created to generate putative gene structures. N SCAN identifies gene structures in a target DNA sequence and can use conservation patterns learned from alignments between a target and one or more informant DNA sequences. N SCAN uses a Bayesian network, generated from a phylogenetic tree, to probabilistically relate the target sequence to the aligned sequence(s). Phylogenetic substitution models are used to estimate substitution likelihood along the branches of the tree. Although N SCAN’s predictive accuracy is already a benchmark for de novo HMM based gene predictors, optimizing its use of substitution models will allow for improved conservation pattern estimates leading to even better accuracy. Selecting optimal substitution models requires avoiding overfitting as more detailed models require more free parameters; unfortunately, the number of parameters is limited by the number of known genes available for parameter estimation (training). In order to optimize substitution model selection, we tested eight Type of Report: Other Department of Computer Science & Engineering Washington University in St. Louis Campus Box 1045 St. Louis, MO 63130 ph: (314) 935-6160 1 Optimization of Gene Prediction via More Accurate Phylogenetic Substitution Models Ezekiel Maier, Randall H Brown, and Michael R Brent Department of Computer Science and Engineering, Washington University, Saint Louis, MO, 63130 Abstract: Determining the beginning and end positions of each exon in each protein coding gene within a genome can be difficult because the DNA patterns that signal a gene’s presence have multiple weakly related alternate forms and the DNA fragments that comprise a gene are generally small in comparison to the size of the genome. In response to this challenge, automated gene predictors were created to generate putative gene structures. N-SCAN identifies gene structures in a target DNA sequence and can use conservation patterns learned from alignments between a target and one or more informant DNA sequences. N-SCAN uses a Bayesian network, generated from a phylogenetic tree, to probabilistically relate the target sequence to the aligned sequence(s). Phylogenetic substitution models are used to estimate substitution likelihood along the branches of the tree. Although N-SCAN’s predictive accuracy is already a benchmark for de novo HMM based gene predictors, optimizing its use of substitution models will allow for improved conservation pattern estimates leading to even better accuracy. Selecting optimal substitution models requires avoiding overfitting as more detailed models require more free parameters; unfortunately, the number of parameters is limited by the number of known genes available for parameter estimation (training). In order to optimize substitution model selection, we tested eight models on the entire genome including General, Reversible, HKY, Jukes-Cantor, and Kimura. In addition to testing models on the entire genome, genome feature based model selection strategies were investigated by assessing the ability of each model to accurately reflex the unique conservation patterns present in each genome region. Context dependency was examined using Determining the beginning and end positions of each exon in each protein coding gene within a genome can be difficult because the DNA patterns that signal a gene’s presence have multiple weakly related alternate forms and the DNA fragments that comprise a gene are generally small in comparison to the size of the genome. In response to this challenge, automated gene predictors were created to generate putative gene structures. N-SCAN identifies gene structures in a target DNA sequence and can use conservation patterns learned from alignments between a target and one or more informant DNA sequences. N-SCAN uses a Bayesian network, generated from a phylogenetic tree, to probabilistically relate the target sequence to the aligned sequence(s). Phylogenetic substitution models are used to estimate substitution likelihood along the branches of the tree. Although N-SCAN’s predictive accuracy is already a benchmark for de novo HMM based gene predictors, optimizing its use of substitution models will allow for improved conservation pattern estimates leading to even better accuracy. Selecting optimal substitution models requires avoiding overfitting as more detailed models require more free parameters; unfortunately, the number of parameters is limited by the number of known genes available for parameter estimation (training). In order to optimize substitution model selection, we tested eight models on the entire genome including General, Reversible, HKY, Jukes-Cantor, and Kimura. In addition to testing models on the entire genome, genome feature based model selection strategies were investigated by assessing the ability of each model to accurately reflex the unique conservation patterns present in each genome region. Context dependency was examined using zeroth, first, and second order models. All models were tested on the human and D. melanogaster genomes. Analysis of the data suggests that the nucleotide equilibrium frequency assumption (denoted as i) is the strongest predictor of a model’s accuracy, followed by reversibility and transition/transversion inequality. Furthermore, second order models are shown to give an average of 0.6% improvement over first order models, which give an 18% improvement over zeroth order models. Finally, by limiting parameter usage by the number of training examples available for each feature, genome feature based model selection better estimates substitution likelihood leading to a significant improvement in N-SCAN’s gene annotation accuracy.
MOTIVATION The most accurate way to determine the intron-exon structures in a genome is to align spliced cDNA sequences to the genome. Thus, cDNA-to-genome alignment programs are a key component of most annotation pipelines. The scoring system used to choose the best alignment is a primary determinant of alignment accuracy, while heuristics that prevent consideration of certain alignments are a primary determinant of runtime and memory usage. Both accuracy and speed are important considerations in choosing an alignment algorithm, but scoring systems have received much less attention than heuristics. RESULTS We present Pairagon, a pair hidden Markov model based cDNA-to-genome alignment program, as the most accurate aligner for sequences with high- and low-identity levels. We conducted a series of experiments testing alignment accuracy with varying sequence identity. We first created 'perfect' simulated cDNA sequences by splicing the sequences of exons in the reference genome sequences of fly and human. The complete reference genome sequences were then mutated to various degrees using a realistic mutation simulator and the perfect cDNAs were aligned to them using Pairagon and 12 other aligners. To validate these results with natural sequences, we performed cross-species alignment using orthologous transcripts from human, mouse and rat. We found that aligner accuracy is heavily dependent on sequence identity. For sequences with 100% identity, Pairagon achieved accuracy levels of >99.6%, with one quarter of the errors of any other aligner. Furthermore, for human/mouse alignments, which are only 85% identical, Pairagon achieved 87% accuracy, higher than any other aligner. AVAILABILITY Pairagon source and executables are freely available at http://mblab.wustl.edu/software/pairagon/
Background This study analyzes the predictions of a number of promoter predictors on the ENCODE regions of the human genome as part of the ENCODE Genome Annotation Assessment Project (EGASP). The systems analyzed operate on various principles and we assessed the effectiveness of different conceptual strategies used to correlate produced promoter predictions with the manually annotated 5' gene ends. Results The predictions were assessed relative to the manual HAVANA annotation of the 5' gene ends. These 5' gene ends were used as the estimated reference transcription start sites. With the maximum allowed distance for predictions of 1,000 nucleotides from the reference transcription start sites, the sensitivity of predictors was in the range 32% to 56%, while the positive predictive value was in the range 79% to 93%. The average distance mismatch of predictions from the reference transcription start sites was in the range 259 to 305 nucleotides. At the same time, using transcription start site estimates from DBTSS and H-Invitational databases as promoter predictions, we obtained a sensitivity of 58%, a positive predictive value of 92%, and an average distance from the annotated transcription start sites of 117 nucleotides. In this experiment, the best performing promoter predictors were those that combined promoter prediction with gene prediction. The main reason for this is the reduced promoter search space that resulted in smaller numbers of false positive predictions. Conclusion The main finding, now supported by comprehensive data, is that the accuracy of human promoter predictors for high-throughput annotation purposes can be significantly improved if promoter prediction is combined with gene prediction. Based on the lessons learned in this experiment, we propose a framework for the preparation of the next similar promoter prediction assessment.
BACKGROUND:This paper describes Pairagon+N-SCAN_EST, a gene annotation pipeline that uses only native alignments. For each expressed sequence it chooses the best genomic alignment. Systems like ENSEMBL and ExoGean rely on trans alignments, in which expressed sequences are aligned to the genomic loci of putative homologs. Trans alignments contain a high proportion of mismatches, gaps, and/or apparently unspliceable introns, compared to alignments of cDNA sequences to their native loci. The Pairagon+N-SCAN_EST pipeline's first stage is Pairagon, a cDNA-to-genome alignment program based on a PairHMM probability model. This model relies on prior knowledge, such as the fact that introns must begin with GT, GC, or AT and end with AG or AC. It produces very precise alignments of high quality cDNA sequences. In the genomic regions between Pairagon's cDNA alignments, the pipeline combines EST alignments with de novo gene prediction by using N-SCAN_EST. N-SCAN_EST is based on a generalized HMM probability model augmented with a phylogenetic conservation model and EST alignments. It can predict complete transcripts by extending or merging EST alignments, but it can also predict genes in regions without EST alignments. Because they are based on probability models, both Pairagon and N-SCAN_EST can be trained automatically for new genomes and data sets. RESULTS:On the ENCODE regions of the human genome, Pairagon+N-SCAN_EST was as accurate as any other system tested in the EGASP assessment, including ENSEMBL and ExoGean. CONCLUSION:With sufficient mRNA/EST evidence, genome annotation without trans alignments can compete successfully with systems like ENSEMBL and ExoGean, which use trans alignments.
The retrainable, comparative gene predictor N-SCAN integrates multigenome modeling and 5' untranslated region (5' UTR) modeling. In this article, we evaluate N-SCAN's transcription-start site (TSS) and first exon predictions both computationally and experimentally. The computational results indicate that N-SCAN is more accurate than any of the other tools we tested at predicting the TSS and the complete first exon. It is the only one of these tools that can predict complete gene structures together with 5' UTRs. Experimental evaluation shows that N-SCAN can be used to validate novel UTR introns in human gene predictions that do not overlap any RefSeq gene and even to correct RefSeq mRNAs by adding validated UTR exons that are missing from RefSeq.
The genomes of clusters of related eukaryotes are now being sequenced at an increasing rate, creating a need for accurate, low-cost annotation of exon-intron structures. In this paper, we demonstrate that reverse transcription-polymerase chain reaction (RT-PCR) and direct sequencing based on predicted gene structures satisfy this need, at least for single-celled eukaryotes. The TWINSCAN gene prediction algorithm was adapted for the fungal pathogen Cryptococcus neoformans by using a precise model of intron lengths in combination with ungapped alignments between the genome sequences of the two closely related Cryptococcus varieties. This approach resulted in approximately 60% of known genes being predicted exactly right at every coding base and splice site. When previously unannotated TWINSCAN predictions were tested by RT-PCR and direct sequencing, 75% of targets spanning two predicted introns were amplified and produced high-quality sequence. When targets spanning the complete predicted open reading frame were tested, 72% of them amplified and produced high-quality sequence. We conclude that sequencing a small number of expressed sequence tags (ESTs) to provide training data, running TWINSCAN on an entire genome, and then performing RT-PCR and direct sequencing on all of its predictions would be a cost-effective method for obtaining an experimentally verified genome annotation.