The pace at which genome references are being generated for plants and animal species is rapidly increasing with Next Generation Sequencing technologies. While this is a major step forward for researchers studying species that previously did not have sequenced genomes, it is only the beginning of the process toward defining the biology underlying the genome. As long as a reference is available, DNA variants can be readily identified on a genome wide scale, often producing lists of 100s of thousands or even millions of variants. Often those occurring in expressed genes are of the most interest; however, if annotation defining where genes exist within a genome is not available or poorly defined, identifying which mutations might affect protein coding may not be possible. To address this challenge we will describe a method whereby RNA-Seq can be readily used to identify transcriptionally active regions which creates transcript annotation for un-annotated or enhanced annotation for any organism. This annotation can then be used in conjunction with whole genome sequencing to annotate variants as to whether they fall within transcriptionally active regions thus facilitating the identification of mutations in larger repertoire of expressed regions of a genome.
Objective: To compare full transcriptome expression levels of matched tumor and normal samples from patients with oropharyngeal carcinoma stratified by known tumor etiologic factors.Patients and Methods: Full transcriptome sequencing was analyzed for 10 matched tumor and normal tissue samples from patients with previously untreated oropharyngeal carcinoma. Transcriptomes were analyzed using massively parallel messenger RNA sequencing and validated using the NanoString nCounter system. Global gene expression levels were compared in samples grouped by smoking status and human papillomavirus status. This study was completed between June 10, 2010, and June 30, 2011.Results: Global gene expression analysis indicated tumor tissue from former smokers grouped more closely to the never smokers than the current smokers. Pathway analysis revealed alterations in the expression of genes involved in the p53 DNA damage-repair pathway, including CHEK2 and ATR, which display patterns of increased expression that is associated with human papillomavirus-negative current smokers rather than former or never smokers.Conclusion: These findings support the application of messenger RNA sequencing technology as an important clinical tool for more accurately stratifying patients based on individual tumor biology with the goal of improving our understanding of tumor prognosis and treatment response, ultimately leading to individualized patient care strategies. (C) 2012 Mayo Foundation for Medical Education and Research square Mayo Clin Proc. 2012;87(3):226-232
By the end of 2011 we will likely know the DNA sequences for 30,000 human genomes. However, to truly understand how the variation between these genomes affect phenotype at a molecular level, future research projects need to analyze these genomes in conjunction with data from multiple ultra-high throughput assays obtained from large sample populations. In cancer research, for example, studies that examine 1000s of specific tumors in 1000s of patients are needed to fully characterize the more than 10,000 types and subtypes of cancer and develop diagnostic biomarkers. These studies will use high throughput DNA sequencing to characterize tumor genomes and their transcriptomes. Sequencing results will be validated with nonsequencing technologies and putative biomarkers will be examined in large populations using rapid targeted assay approaches. Geospiza is transforming the above scenario from vision into reality in several ways. The Company's GeneSifter platform utilizes scalable data management technologies based on open-source HDF5 and BioHDF technologies to capture, integrate, and mine raw data and analysis results from DNA, RNA, and other high-throughput assays. Analysis results are integrated and linked to multiple repositories of information that include variation, expression, pathway, and ontology databases to enable discovery process and support verification assays. Using this platform and RNA-Sequencing and Genomic DNA sequencing from matched tumor/normal samples, we were able to characterize differential gene expression, differential splicing, allele specific expression, RNA editing, somatic mutations and genomic rearrangements as well as validate these observations in a set of patients with oral and other cancers.
Bovine spongiform encephalopathy (BSE) is a transmissible, fatal neurodegenerative disorder of cattle produced by prions. The use of excessive parallel sequencing for comparison of gene expression in bovine control and infected tissues may help to elucidate the molecular mechanisms associated with this disease. In this study, tag profiling Solexa sequencing was used for transcriptome analysis of bovine brain tissues. Replicate libraries were prepared from mRNA isolated from control and infected (challenged with 100 g of BSE-infected brain) medulla tissues 45 mo after infection. For each library, 5-6 million sequence reads were generated and approximately 67-70% of the reads were mapped against the Bovine Genome database to approximately 13,700-14,120 transcripts (each having at least one read). About 42-47% of the total reads mapped uniquely. Using the GeneSifter software package, 190 differentially expressed (DE) genes were identified (> 2.0-fold change, p < .01): 73 upregulated and 117 downregulated. Seventy-nine DE genes had functions described in the Gene Ontology (GO) database and 16 DE genes were involved in 38 different pathways described in the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. Digital analysis expression by tag profiling may be a powerful approach to comprehensive transcriptome analysis to identify changes associated with disease progression, leading to a better understanding of the underlying mechanism of pathogenesis of BSE.
Next Generation Sequencing technologies are limited by the lack of standard bioinformatics infrastructures that can reduce data storage, increase data processing performance, and integrate diverse information. HDF technologies address these requirements and have a long history of use in data-intensive science communities. They include general data file formats, libraries, and tools for working with the data. Compared to emerging standards, such as the SAM/BAM formats, HDF5-based systems demonstrate significantly better scalability, can support multiple indexes, store multiple data types, and are self-describing. For these reasons, HDF5 and its BioHDF extension are well suited for implementing data models to support the next generation of bioinformatics applications.
Next generation DNA sequencing (NGS) technologies are increasing in their appeal for studying cancer genomics. High-throughput data and a growing repertoire of applications that quantitatively measure gene expression, splicing, noncoding RNAs, and genomic variation are revealing that cancer is a more complex and heterogeneous disease than previously imagined. Fully characterizing the ~10,000 types and subtypes of cancer that exist to develop biomarkers that can be used to clinically define tumors and target specific treatments requires large studies that examine specific tumors in thousands of patients. This goal will fail without significantly reducing both data production and analysis costs, so that most cancer biologists and clinicians can conduct NGS assays and analyze their data in routine ways. Currently, most cancer biology NGS papers are published either by genome centers or through collaborations with instrument vendors. However, this is going to change rapidly with efforts like the Cancer Genome Anatomy project. In any case, large teams of bioinformaticians are involved in analyzing data through labor-intensive processes. With refinements offered by the Illumina HiSeq 2000, or Life Technologies SOLiD 4, the cost of collecting data for transcriptome analysis and mate-pair genome sequencing is sufficiently inexpensive for small groups and individuals, beyond genome centers, to conduct the required studies. However, current data analysis methods need to be automated with established tools in scalable and adaptable systems that provide standard reports to make results available to enable interactive exploration by biologists and clinicians. In our presentation, we will examine the time and costs required to analyze data that will be collected in future cancer studies. Using data from existing matched tumor and normal transcriptome studies from random oral cancer samples, and samples grouped by drinking and smoking behavior (as a tool to define data analysis requirements), we will compare the costs of conducting large studies using current data analysis approaches with those using integrated software systems to demonstrate how automation reduces costs, while providing comparable results for identifying transcript isoforms, mutations and novel translocations. Geospiza's GeneSifter distributed cloud- based software architecture, including open source tools, like BioHDF, will be described to share insights into high performance computing requirements for scalable data processing.
Dear Sir, We are writing this letter in response to ‘GeneSifter, leading the blind,’ a letter published last fall (Hoek, 2008) about GeneSifter , a software product for analyzing data from microarray and next generation DNA sequencing experiments (http://www.geospiza.com). We wish to address the areas of concern described in the article: the methods used for data normalization and concerns regarding how fold-change cut-offs and statistical tests are applied. Hoek raises valid points regarding both the need for between chip normalization and that the order of filtering steps can have dramatic affects on the control for false positives; however, we maintain that Hoek’s criticisms were based on an incomplete knowledge of GeneSifter features, capabilities and proper operation. We believe this misunderstanding most likely resulted from a failure on the part of VizXLabs (the previous owners of GeneSifter) to adequately respond to Hoek’s concerns. As the new owners of GeneSifter, we would like to address the criticisms that Hoek raised. Hoek pointed out the necessity for normalizing microarray measurements both within a chip and between multiple chips and wrote that GeneSifter fails to account for the differences between chips when normalizing data (Hoek, 2008). It’s probable that Hoek missed seeing GeneSifter’s options for normalization because they are not visible from the pairwise analysis interface. GeneSifter does include methods for normalizing data within chips and between chips, but also gives users a choice to upload normalized data without further processing; this option allows users to circumvent errors that would result from re-processing already normalized data. The normalization options within GeneSifter include methods such as RMA (Robust Microarray Analysis) and GC-RMA (Irizarry et al., 2003; Wu et al., 2004) for normalizing data between multiple Affymetrix chips. It should be noted, however, that these methods must be applied at the time of loading the data rather than at the time of analysis. Consequently, these methods are not available through the pairwise analysis interface and could be missed by someone who is new to the software. We believe these details were not communicated to Hoek, leaving him with the impression that these more robust normalization methods are missing from the GeneSifter package. Hoek’s second concern was with the order in which a fold-change cut-off is applied and the affects on corrections for multiple testing. Hoek stated that the order of statistical operations in GeneSifter was incorrect because users could apply a fold-change filter to averaged data before performing statistical tests and that users were left unable to change the options or perform steps in the appropriate manner, resulting in weak control for false positives. While it is true that this ordering can be used in GeneSifter, the product has always allowed the statistical test and corrections to be performed prior to applying a fold-change cut-off. As with the normalization concerns, we believe this misunderstanding arose from a failure to communicate the options available in GeneSifter. The default settings for pairwise statistics did perform a threshold cut-off first and then apply a correction based on the number of genes that passed the initial cut-off. These settings were used to facilitate discovery by decreasing the possibility of false negatives. In contrast to the statement in the article; however, GeneSifter users have always had the option to change the order of steps by using the preference settings in the GeneSifter program. The last point we wish to address is the comparison between data analyzed with GeneSifter and the same data analyzed with GeneSpring (Agilent Technologies). Hoek analyzed a publicly available colon cancer data set (GEO accession no. GDS756) with both GeneSifter and GeneSpring. The GeneSpring analysis yielded 449 genes with significant differences in expression where analyzing the data with GeneSifter produced a list of 1556 genes. This difference however, did not result from a true difference in the software platforms, the difference in gene number resulted because different methods were used for the analyses. When we use the same analysis procedure with GeneSifter, that Hoek used with GeneSpring, we obtain similar numbers of significantly expressed genes. Since the publication of Hoek’s letter; we have made changes in the default settings in the pairwise analysis interface to reduce the likelihood that users would inadvertently lower their control for false positives. Users still have the option however, to change their preferences settings and reverse the order should the need arise.
Transcription profiling with microarrays has become a standard procedure for comparing the levels of gene expression between pairs of samples, or multiple samples following different experimental treatments. New technologies, collectively known as next-generation DNA sequencing methods, are also starting to be used for transcriptome analysis. These technologies, with their low background, large capacity for data collection, and dynamic range, provide a powerful and complementary tool to the assays that formerly relied on microarrays. In this chapter, we describe two protocols for working with microarray data from pairs of samples and samples treated with multiple conditions, and discuss alternative protocols for carrying out similar analyses with next-generation DNA sequencing data from two different instrument platforms (Illumina GA and Applied Biosystems SOLiD).
To determine specific molecular features of endothelial cells (ECs) relevant to the physiological process of penile erection we compared gene expression of human EC derived from corpus cavernosum of men with and without erectile dysfunction (HCCEC) to coronary artery (HCAEC) and umbilical vein (HUVEC) using Affymetrix GeneChip microarrays and GeneSifter software. Genes differentially expressed across samples were partitioned around medoids to identify genes with highest expression in HCCEC. A total of 190 genes/transcripts were highly expressed only in HCCEC. Gene Ontology classification indicated cavernosal enrichment in genes related to cell adhesion, extracellular matrix, pattern specification and organogenesis. Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis showed high expression of genes relating to ECM-receptor interaction, focal adhesions, and cytokine-cytokine receptor interaction. Real-time PCR confirmed expression differences in cadherins 2 and 11, claudin 11 (CLDN11), desmoplakin, and versican. CLDN11, a component of tight junctions not previously described in ECs, was highly expressed only in HCCEC and its knockdown by siRNA significantly reduced transendothelial electrical resistance in HCCEC. Overall, cavernosal ECs exhibited a transcriptional profile encoding matrix and adhesion proteins that regulate structural and functional characteristics of blood vessels. Contribution of the tight junction protein CLDN11 to barrier function in endothelial cells is novel and may reflect hemodynamic requirements of the corpus cavernosum.
Next generation sequencing technology is rapidly changing the way laboratories and researchers approach the management and analysis of biological information. The sheer volume of data produced by emerging technologies threatens to overwhelm many researchers, 1, 2 and labs increasingly find themselves in the business of trying to build and run data centers rather than devoting resources to scientific discovery. GeneSifter R Lab Edition provides labs and researchers with next generation tools specifically designed to address the evolving needs of next generation sequencing, including: . A cloud computing model that frees labs from the burden of building and maintaining expensive data centers. . The ability to track experimentally relevant information throughout an entire project within a single unified system. . A framework for bioinformaticians to build automated analysis pipelines that can be shared across the research community. . A platform which allows scientists to run complex application-specific data analysis pipelines on large data sets with the press of a button. . A universally accessible, OS-independent system that can be used to distribute and visualize data from anywhere.