De novo transcriptome sequencing and analysis provides a way for researchers of non-model organisms to explore the differences between various conditions and species. The results are typically not definitive but will lead to new hypotheses to study. Therefore, it is important that the results be reproducible, extensible, queryable, and easily available to all members of the team. Towards this end, the Transcriptome Computational Workbench (TCW) is a software package to perform basic computations for transcriptome analysis (singleTCW) and comparative analysis (multiTCW). It is a Java-based desktop application that uses MySQL for the TCW database. The input to singleTCW is sequence and optional count files; the computations are sequence similarity, gene ontology (GO), open reading frame (ORF), and differential expression (DE). TCW provides support for searching with the super-fast DIAMOND program against UniProt taxonomic databases, though the user can provide other databases to search against. The ORF finder uses hit information, 5th-order Markov models and ORF length. For DE and GO enrichment, TCW interfaces with the R environment and an R script, where R scripts are provided for popular methods. The input to multiTCW is multiple singleTCW databases; the computations are homologous pair assignment, pairwise analysis (e.g. Ka/Ks) from codon-based alignments, clustering (bidirectional best hit, Closure, Best Hit, OrthoMCL, user-supplied), and cluster analysis and annotation. Both singleTCW and multiTCW provide a graphical interface for extensive query and display of the data and results. Example results are presented from two rhizome and one non-rhizome plant, where one of the rhizome plants has replicate count data from four tissues. The supplement describes how to reproduce all tables and figures. The TCW V4 software is freely available at <https://github.com/csoderlund/TCW>; the package contains the jar files, external software, and demo files. ### Competing Interest Statement The authors have declared no competing interest.
De novo transcriptome sequencing and analysis provides a way for researchers of non-model organisms to explore the differences between various conditions and species. The results are typically not definitive but will lead to new hypotheses to study. Therefore, it is important that the results be reproducible, extensible, queryable, and easily available to all members of the team. Towards this end, the Transcriptome Computational Workbench (TCW) is a software package to perform basic computations for transcriptome analysis (singleTCW) and comparative analysis (multiTCW). It is a Java-based desktop application that uses MySQL for the TCW database. The input to singleTCW is sequence and optional count files; the computations are sequence similarity, gene ontology (GO), open reading frame (ORF), and differential expression (DE). TCW provides support for searching with the super-fast DIAMOND program against UniProt taxonomic databases, though the user can provide other databases to search against. The ORF finder uses hit information, 5th-order Markov models and ORF length. For DE and GO enrichment, TCW interfaces with the R environment and an R script, where R scripts are provided for popular methods. The input to multiTCW is multiple singleTCW databases; the computations are homologous pair assignment, pairwise analysis (e.g. Ka/Ks) from codon-based alignments, clustering (bidirectional best hit, Closure, Best Hit, OrthoMCL, user-supplied), and cluster analysis and annotation. Both singleTCW and multiTCW provide a graphical interface for extensive query and display of the data and results. Example results are presented from two rhizome and one non-rhizome plant, where one of the rhizome plants has replicate count data from four tissues. The supplement describes how to reproduce all tables and figures. The TCW V4 software is freely available at https://github.com/csoderlund/TCW; the package contains the jar files, external software, and demo files.
The rhizome is responsible for the invasiveness and competitiveness of many plants with great economic and agricultural impact worldwide. Besides its value as an invasive organ, the rhizome plays a role in the establishment and massive growth of forage, providing biomass for biofuel production. Despite these features, little is known about the molecular mechanisms that contribute to rhizome growth, development, and function in plants. In this work, we characterized the proteome of rhizome apical tips and elongation zones from different species using a GeLC-MS/MS (one-dimensional electrophoresis in combination with liquid chromatography coupled online with tandem mass spectrometry) spectral-counting proteomics strategy. Five rhizomatous grasses and an ancient species were compared to study the protein regulation in rhizomes. An average of 2200 rhizome proteins per species were confidently identified and quantified. Rhizome-characteristic proteins showed similar functional distributions across all species analyzed. The over-representation of proteins associated with central roles in cellular, metabolic, and developmental processes indicated accelerated metabolism in growing rhizomes. Moreover, 61 rhizome-characteristic proteins appeared to be regulated similarly among analyzed plants. In addition, 36 showed conserved regulation between rhizome apical tips and elongation zones across species. These proteins were preferentially expressed in rhizome tissues regardless of the species analyzed, making them interesting candidates for more detailed investigative studies about their roles in rhizome development.
Research on the distribution and structure of fungal communities in caves is lacking. Kartchner Caverns is a wet and mineralogically diverse carbonate cave located in an escarpment of Mississippian Escabrosa limestone in the Whetstone Mountains, Arizona, USA. Fungal diversity from speleothem and rock wall surfaces was examined with 454 FLX Titanium sequencing technology using the Internal Transcribed Spacer 1 as a fungal barcode marker. Fungal diversity was estimated and compared between speleothem and rock wall surfaces, and its variation with distance from the natural entrance of the cave was quantified. Effects of environmental factors and nutrient concentrations in speleothem drip water at different sample sites on fungal diversity were also examined. Sequencing revealed 2,219 fungal operational taxonomic units (OTUs) at the 95 % similarity level. Speleothems supported a higher fungal richness and diversity than rock walls. However, community membership and the taxonomic distribution of fungal OTUs at the class level did not differ significantly between speleothems and rock walls. Both OTU richness and diversity decreased significantly with increasing distance from the natural cave entrance. Community membership and taxonomic distribution of fungal OTUs also differed significantly between the sampling sites closest to the entrance and those furthest away. There was no significant effect of temperature, CO 2 concentration, or drip water nutrient concentration on fungal community structure on either speleothems or rock walls. Together, these results suggest that proximity to the natural entrance is a critical factor in determining fungal community structure on mineral surfaces in Kartchner Caverns.
The Asian citrus psyllid (ACP) Diaphorina citri Kuwayama (Hemiptera: Psyllidae) is the insect vector of the fastidious bacterium Candidatus Liberibacter asiaticus (CLas), the causal agent of citrus greening disease, or Huanglongbing (HLB). The widespread invasiveness of the psyllid vector and HLB in citrus trees worldwide has underscored the need for non-traditional approaches to manage the disease. One tenable solution is through the deployment of RNA interference technology to silence protein-protein interactions essential for ACP-mediated CLas invasion and transmission. To identify psyllid interactor-bacterial effector combinations associated with psyllid-CLas interactions, cDNA libraries were constructed from CLas-infected and CLas-free ACP adults and nymphs, and analyzed for differential expression. Library assemblies comprised 24,039,255 reads and yielded 45,976 consensus contigs. They were annotated (UniProt), classified using Gene Ontology, and subjected to in silico expression analyses using the Transcriptome Computational Workbench (TCW) (http://www.sohomoptera.org/ACPPoP/). Functional-biological pathway interpretations were carried out using the Kyoto Encyclopedia of Genes and Genomes databases. Differentially expressed contigs in adults and/or nymphs represented genes and/or metabolic/pathogenesis pathways involved in adhesion, biofilm formation, development-related, immunity, nutrition, stress, and virulence. Notably, contigs involved in gene silencing and transposon-related responses were documented in a psyllid for the first time. This is the first comparative transcriptomic analysis of ACP adults and nymphs infected and uninfected with CLas. The results provide key initial insights into host-parasite interactions involving CLas effectors that contribute to invasion-virulence, and to host nutritional exploitation and immune-related responses that appear to be essential for successful ACP-mediated circulative, propagative CLas transmission.
Sequencing the transcriptome can answer various questions such as determining the transcripts expressed in a given species for a specific tissue or condition, evaluating differential expression, discovering variants, and evaluating allele-specific expression. Differential expression evaluates the expression differences between different strains, tissues, and conditions. Allele-specific expression evaluates expression differences between parental alleles. Both differential expression and allele-specific expression have been studied for heterosis (hybrid vigor), where the hybrid has improved performance over the parents for one or more traits. The Allele Workbench software was developed for a heterosis study that evaluated allele-specific expression for a mouse F1 hybrid using libraries from multiple tissues with biological replicates. This software has been made into a distributable package, which includes a pipeline, a Java interface to build the database, and a Java interface for query and display of the results. The required input is a reference genome, annotation file, and one or more RNA-Seq libraries with optional replicates. It evaluates allelic imbalance at the SNP and transcript level and flags transcripts with significant opposite directional allele-specific expression. The Java interface allows the user to view data from libraries, replicates, genes, transcripts, exons, and variants, including queries on allele imbalance for selected libraries. To determine the impact of allele-specific SNPs on protein folding, variants are annotated with their effect (e.g., missense), and the parental protein sequences may be exported for protein folding analysis. The Allele Workbench processing results in transcript files and read counts that can be used as input to the previously published Transcriptome Computational Workbench, which has a new algorithm for determining a trimmed set of gene ontology terms. The software with demo files is available from https://code.google.com/p/alleleworkbench. Additionally, all software is ready for immediate use from an Atmosphere Virtual Machine Image available from the iPlant Collaborative (www.iplantcollaborative.org).
BACKGROUND:The rhizome, the original stem of land plants, enables species to invade new territory and is a critical component of perenniality, especially in grasses. Red rice (Oryza longistaminata) is a perennial wild rice species with many valuable traits that could be used to improve cultivated rice cultivars, including rhizomatousness, disease resistance and drought tolerance. Despite these features, little is known about the molecular mechanisms that contribute to rhizome growth, development and function in this plant.RESULTS:We used an integrated approach to compare the transcriptome, proteome and metabolome of the rhizome to other tissues of red rice. 116 Gb of transcriptome sequence was obtained from various tissues and used to identify rhizome-specific and preferentially expressed genes, including transcription factors and hormone metabolism and stress response-related genes. Proteomics and metabolomics approaches identified 41 proteins and more than 100 primary metabolites and plant hormones with rhizome preferential accumulation. Of particular interest was the identification of a large number of gene transcripts from Magnaportha oryzae, the fungus that causes rice blast disease in cultivated rice, even though the red rice plants showed no sign of disease.CONCLUSIONS:A significant set of genes, proteins and metabolites appear to be specifically or preferentially expressed in the rhizome of O. longistaminata. The presence of M. oryzae gene transcripts at a high level in apparently healthy plants suggests that red rice is resistant to this pathogen, and may be able to provide genes to cultivated rice that will enable resistance to rice blast disease.
The potato psyllid (PoP) Bactericera cockerelli (Sulc) and Asian citrus psyllid (ACP) Diaphorina citri Kuwayama are the insect vectors of the fastidious plant pathogen, Candidatus Liberibacter solanacearum (CLso) and Ca. L. asiaticus (CLas), respectively. CLso causes Zebra chip disease of potato and vein-greening in solanaceous species, whereas, CLas causes citrus greening disease. The reliance on insecticides for vector management to reduce pathogen transmission has increased interest in alternative approaches, including RNA interference to abate expression of genes essential for psyllid-mediated Ca. Liberibacter transmission. To identify genes with significantly altered expression at different life stages and conditions of CLso/CLas infection, cDNA libraries were constructed for CLso-infected and -uninfected PoP adults and nymphal instars. Illumina sequencing produced 199,081,451 reads that were assembled into 82,224 unique transcripts. PoP and the analogous transcripts from ACP adult and nymphs reported elsewhere were annotated, organized into functional gene groups using the Gene Ontology classification system, and analyzed for differential in silico expression. Expression profiles revealed vector life stage differences and differential gene expression associated with Liberibacter infection of the psyllid host, including invasion, immune system modulation, nutrition, and development.
Carbonate caves represent subterranean ecosystems that are largely devoid of phototrophic primary production. In semiarid and arid regions, allochthonous organic carbon inputs entering caves with vadose-zone drip water are minimal, creating highly oligotrophic conditions; however, past research indicates that carbonate speleothem surfaces in these caves support diverse, predominantly heterotrophic prokaryotic communities. The current study applied a metagenomic approach to elucidate the community structure and potential energy dynamics of microbial communities, colonizing speleothem surfaces in Kartchner Caverns, a carbonate cave in semiarid, southeastern Arizona, USA. Manual inspection of a speleothem metagenome revealed a community genetically adapted to low-nutrient conditions with indications that a nitrogen-based primary production strategy is probable, including contributions from both Archaea and Bacteria. Genes for all six known CO2-fixation pathways were detected in the metagenome and RuBisCo genes representative of the Calvin–Benson–Bassham cycle were over-represented in Kartchner speleothem metagenomes relative to bulk soil, rhizosphere soil and deep-ocean communities. Intriguingly, quantitative PCR found Archaea to be significantly more abundant in the cave communities than in soils above the cave. MEtaGenome ANalyzer (MEGAN) analysis of speleothem metagenome sequence reads found Thaumarchaeota to be the third most abundant phylum in the community, and identified taxonomic associations to this phylum for indicator genes representative of multiple CO2-fixation pathways. The results revealed that this oligotrophic subterranean environment supports a unique chemoautotrophic microbial community with potentially novel nutrient cycling strategies. These strategies may provide key insights into other ecosystems dominated by oligotrophy, including aphotic subsurface soils or aquifers and photic systems such as arid deserts.
Background The analysis of transcriptome data involves many steps and various programs, along with organization of large amounts of data and results. Without a methodical approach for storage, analysis and query, the resulting ad hoc analysis can lead to human error, loss of data and results, inefficient use of time, and lack of verifiability, repeatability, and extensibility. Methodology The Transcriptome Computational Workbench (TCW) provides Java graphical interfaces for methodical analysis for both single and comparative transcriptome data without the use of a reference genome (e.g. for non-model organisms). The singleTCW interface steps the user through importing transcript sequences (e.g. Illumina) or assembling long sequences (e.g. Sanger, 454, transcripts), annotating the sequences, and performing differential expression analysis using published statistical programs in R. The data, metadata, and results are stored in a MySQL database. The multiTCW interface builds a comparison database by importing sequence and annotation from one or more single TCW databases, executes the ESTscan program to translate the sequences into proteins, and then incorporates one or more clusterings, where the clustering options are to execute the orthoMCL program, compute transitive closure, or import clusters. Both singleTCW and multiTCW allow extensive query and display of the results, where singleTCW displays the alignment of annotation hits to transcript sequences, and multiTCW displays multiple transcript alignments with MUSCLE or pairwise alignments. The query programs can be executed on the desktop for fastest analysis, or from the web for sharing the results. Conclusion It is now affordable to buy a multi-processor machine, and easy to install Java and MySQL. By simply downloading the TCW, the user can interactively analyze, query and view their data. The TCW allows in-depth data mining of the results, which can lead to a better understanding of the transcriptome. TCW is freely available from www.agcol.arizona.edu/software/tcw.
A draft sequence of the staple crop kabuli chickpea, together with resequencing and analysis of 90 additional lines from 10 countries, provides a resource for breeders. Chickpea (Cicer arietinum) is the second most widely grown legume crop after soybean, accounting for a substantial proportion of human dietary nitrogen intake and playing a crucial role in food security in developing countries. We report the ∼738-Mb draft whole genome shotgun sequence of CDC Frontier, a kabuli chickpea variety, which contains an estimated 28,269 genes. Resequencing and analysis of 90 cultivated and wild genotypes from ten countries identifies targets of both breeding-associated genetic sweeps and breeding-associated balancing selection. Candidate genes for disease resistance and agronomic traits are highlighted, including traits that distinguish the two main market classes of cultivated chickpea—desi and kabuli. These data comprise a resource for chickpea improvement through molecular breeding and provide insights into both genome diversity and domestication.
Rhizomes are underground stems that serve various purposes including vegetative propagation, invasion of new territory, and bioactive compound synthesis and storage. An important rhizomatous plant is sacred lotus (Nelumbo nucifera), which is prized in Asia as a medicine and a food. RNA-seq and total transcriptome analysis of rhizomes and other lotus tissues was applied to identify genes involved in rhizome growth, development and metabolism. Root, petiole, rhizome internode, and leaf tissues were used for single-read RNA-seq analysis. Two whole transcriptome paired-end read libraries from rhizome apical tip and elongation zone tissues were also generated in order to survey gene expression profiles. In this analysis, 22,803 genes were expressed: 20,476 in rhizome apical meristem and elongation zone, 17,171 in rhizome internode, 16,656 in leaf, 19,457 in root, and 16,845 in petiole. Gene ontology (GO) analysis indicated that “other membrane”, “nucleotide binding”, and “other cellular processes” were highly represented in the expressed genes. A total of 231 genes displayed rhizome-specific expression including several transcription factors, protein kinases, cytochromes P450 and a sulfate transporter. GOseq analysis showed that genes in the “molecular function” GO category and several genes related to cell proliferation based on KEGG IDs were preferentially up-regulated in rhizome tissue. In addition, 1,251 possible exon-skipping events were observed in 1,149 gene models. These results provide valuable insight into gene expression profiles and regulation in sacred lotus, and the identified rhizome-specific genes provide insight into important processes involved in the biology and development of sacred lotus rhizomes.
Caves are relatively accessible subterranean habitats ideal for the study of subsurface microbial dynamics and metabolisms under oligotrophic, non-photosynthetic conditions. A 454-pyrotag analysis of the V6 region of the 16S rRNA gene was used to systematically evaluate the bacterial diversity of ten cave surfaces within Kartchner Caverns, a limestone cave. Results showed an average of 1,994 operational taxonomic units (97 % cutoff) per speleothem and a broad taxonomic diversity that included 21 phyla and 12 candidate phyla. Comparative analysis of speleothems within a single room of the cave revealed three distinct bacterial taxonomic profiles dominated by either Actinobacteria , Proteobacteria , or Acidobacteria. A gradient in observed species richness along the sampling transect revealed that the communities with lower diversity corresponded to those dominated by Actinobacteria while the more diverse communities were those dominated by Proteobacteria. A 16S rRNA gene clone library from one of the Actinobacteria- dominated speleothems identified clones with 99 % identity to chemoautotrophs and previously characterized oligotrophs, providing insights into potential energy dynamics supporting these communities. The robust analysis conducted for this study demonstrated a rich bacterial diversity on speleothem surfaces. Further, it was shown that seemingly comparable speleothems supported divergent phylogenetic profiles suggesting that these communities are very sensitive to subtle variations in nutritional inputs and environmental factors typifying speleothem surfaces in Kartchner Caverns.
BACKGROUND:Ginger (Zingiber officinale) and turmeric (Curcuma longa) accumulate important pharmacologically active metabolites at high levels in their rhizomes. Despite their importance, relatively little is known regarding gene expression in the rhizomes of ginger and turmeric.RESULTS:In order to identify rhizome-enriched genes and genes encoding specialized metabolism enzymes and pathway regulators, we evaluated an assembled collection of expressed sequence tags (ESTs) from eight different ginger and turmeric tissues. Comparisons to publicly available sorghum rhizome ESTs revealed a total of 777 gene transcripts expressed in ginger/turmeric and sorghum rhizomes but apparently absent from other tissues. The list of rhizome-specific transcripts was enriched for genes associated with regulation of tissue growth, development, and transcription. In particular, transcripts for ethylene response factors and AUX/IAA proteins appeared to accumulate in patterns mirroring results from previous studies regarding rhizome growth responses to exogenous applications of auxin and ethylene. Thus, these genes may play important roles in defining rhizome growth and development. Additional associations were made for ginger and turmeric rhizome-enriched MADS box transcription factors, their putative rhizome-enriched homologs in sorghum, and rhizomatous QTLs in rice. Additionally, analysis of both primary and specialized metabolism genes indicates that ginger and turmeric rhizomes are primarily devoted to the utilization of leaf supplied sucrose for the production and/or storage of specialized metabolites associated with the phenylpropanoid pathway and putative type III polyketide synthase gene products. This finding reinforces earlier hypotheses predicting roles of this enzyme class in the production of curcuminoids and gingerols.CONCLUSION:A significant set of genes were found to be exclusively or preferentially expressed in the rhizome of ginger and turmeric. Specific transcription factors and other regulatory genes were found that were common to the two species and that are excellent candidates for involvement in rhizome growth, differentiation and development. Large classes of enzymes involved in specialized metabolism were also found to have apparent tissue-specific expression, suggesting that gene expression itself may play an important role in regulating metabolite production in these plants.
Nearly half the earth’s surface is occupied by dryland ecosystems, regions susceptible to reduced states of biological productivity caused by climate fluctuations. Of these regions, arid zones located at the interface between vegetated semiarid regions and biologically unproductive hyperarid zones are considered most vulnerable. The objective of this study was to conduct a deep diversity analysis of bacterial communities in unvegetated arid soils of the Atacama Desert, to characterize community structure and infer the functional potential of these communities based on observed phylogenetic associations. A 454-pyrotag analysis was conducted of three unvegetated arid sites located at the hyperarid–arid margin. The analysis revealed communities with unique bacterial diversity marked by high abundances of novel Actinobacteria and Chloroflexi and low levels of Acidobacteria and Proteobacteria, phyla that are dominant in many biomes. A 16S rRNA gene library of one site revealed the presence of clones with phylogenetic associations to chemoautotrophic taxa able to obtain energy through oxidation of nitrite, carbon monoxide, iron, or sulfur. Thus, soils at the hyperarid margin were found to harbor a wealth of novel bacteria and to support potentially viable communities with phylogenetic associations to non-phototrophic primary producers and bacteria capable of biogeochemical cycling.
Background: Plants can defend themselves against herbivorous insects prior to the onset of larval feeding by responding to the eggs laid on their leaves. In the European field elm (Ulmus minor), egg laying by the elm leaf beetle (Xanthogaleruca luteola) activates the emission of volatiles that attract specialised egg parasitoids, which in turn kill the eggs. Little is known about the transcriptional changes that insect eggs trigger in plants and how such indirect defense mechanisms are orchestrated in the context of other biological processes.Results: Here we present the first large scale study of egg-induced changes in the transcriptional profile of a tree. Five cDNA libraries were generated from leaves of (i) untreated control elms, and elms treated with (ii) egg laying and feeding by elm leaf beetles, (iii) feeding, (iv) artificial transfer of egg clutches, and (v) methyl jasmonate. A total of 361,196 ESTs expressed sequence tags (ESTs) were identified which clustered into 52,823 unique transcripts (Unitrans) and were stored in a database with a public web interface. Among the analyzed Unitrans, 73% could be annotated by homology to known genes in the UniProt (Plant) database, particularly to those from Vitis, Ricinus, Populus and Arabidopsis. Comparative in silico analysis among the different treatments revealed differences in Gene Ontology term abundances. Defense-and stress-related gene transcripts were present in high abundance in leaves after herbivore egg laying, but transcripts involved in photosynthesis showed decreased abundance. Many pathogen-related genes and genes involved in phytohormone signaling were expressed, indicative of jasmonic acid biosynthesis and activation of jasmonic acid responsive genes. Cross-comparisons between different libraries based on expression profiles allowed the identification of genes with a potential relevance in egg-induced defenses, as well as other biological processes, including signal transduction, transport and primary metabolism.Conclusion: Here we present a dataset for a large-scale study of the mechanisms of plant defense against insect eggs in a co-evolved, natural ecological plant-insect system. The EST database analysis provided here is a first step in elucidating the transcriptional responses of elm to elm leaf beetle infestation, and adds further to our knowledge on insect egg-induced transcriptomic changes in plants. The sequences identified in our comparative analysis give many hints about novel defense mechanisms directed towards eggs.