Background Reproducible detection of inherited variants with whole genome sequencing (WGS) is vital for the implementation of precision medicine and is a complicated process in which each step affects variant call quality. Systematically assessing reproducibility of inherited variants with WGS and impact of each step in the process is needed for understanding and improving quality of inherited variants from WGS. Results To dissect the impact of factors involved in detection of inherited variants with WGS, we sequence triplicates of eight DNA samples representing two populations on three short-read sequencing platforms using three library kits in six labs and call variants with 56 combinations of aligners and callers. We find that bioinformatics pipelines (callers and aligners) have a larger impact on variant reproducibility than WGS platform or library preparation. Single-nucleotide variants (SNVs), particularly outside difficult-to-map regions, are more reproducible than small insertions and deletions (indels), which are least reproducible when > 5 bp. Increasing sequencing coverage improves indel reproducibility but has limited impact on SNVs above 30×. Conclusions Our findings highlight sources of variability in variant detection and the need for improvement of bioinformatics pipelines in the era of precision medicine with WGS.
Additional file 12: Table S11. Summary of technical reproducibility.
Additional file 10: Table S9. Factor contribution analysis by boosted tree.
AbstractCOVID-19 is characterised by dysregulated immune responses, metabolic dysfunction and adverse effects on the function of multiple organs. To understand how host responses contribute to COVID-19 pathophysiology, we used a multi-omics approach to identify molecular markers in peripheral blood and plasma samples that distinguish COVID-19 patients experiencing a range of disease severities. A large number of expressed genes, proteins, metabolites and extracellular RNAs (exRNAs) were identified that exhibited strong associations with various clinical parameters. Multiple sets of tissue-specific proteins and exRNAs varied significantly in both mild and severe patients, indicative of multi-organ damage. The continuous activation of IFN-I signalling and neutrophils, as well as a high level of inflammatory cytokines, were observed in severe disease patients. In contrast, COVID-19 in mild patients was characterised by robust T cell responses. Finally, we show that some of expressed genes, proteins and exRNAs can be used as biomarkers to predict the clinical outcomes of SARS-CoV-2 infection. These data refine our understanding of the pathophysiology and clinical progress of COVID-19 and will help guide future studies in this area.
COVID‐19 is characterized by dysregulated immune responses, metabolic dysfunction and adverse effects on the function of multiple organs. To understand host responses to COVID‐19 pathophysiology, we combined transcriptomics, proteomics, and metabolomics to identify molecular markers in peripheral blood and plasma samples of 66 COVID‐19‐infected patients experiencing a range of disease severities and 17 healthy controls. A large number of expressed genes, proteins, metabolites, and extracellular RNAs (exRNAs) exhibit strong associations with various clinical parameters. Multiple sets of tissue‐specific proteins and exRNAs varied significantly in both mild and severe patients suggesting a potential impact on tissue function. Chronic activation of neutrophils, IFN‐I signaling, and a high level of inflammatory cytokines were observed in patients with severe disease progression. In contrast, COVID‐19‐infected patients experiencing milder disease symptoms showed robust T‐cell responses. Finally, we identified genes, proteins, and exRNAs as potential biomarkers that might assist in predicting the prognosis of SARS‐CoV‐2 infection. These data refine our understanding of the pathophysiology and clinical progress of COVID‐19.
Measurement of circulating insulin‐like growth factors ( IGF s), in particular IGF ‐binding protein ( IGFBP )‐2, at the time of diagnosis, is independently prognostic in many cancers, but its clinical performance against other routinely determined prognosticators has not been examined. We measured IGF ‐I, IGF ‐ II , pro‐ IGF ‐ II , IGF bioactivity, IGFBP ‐2, ‐3, and pregnancy‐associated plasma protein A ( PAPP ‐A), an IGFBP regulator, in baseline samples of 301 women with breast cancer treated on four protocols (Odense, Denmark: 1993–1998). We evaluated performance characteristics (expressed as area under the curve, AUC ) using Cox regression models to derive hazard ratios ( HR ) with 95% confidence intervals ( CI s) for 10‐year recurrence‐free survival ( RFS ) and overall survival ( OS ), and compared those against the clinically used Nottingham Prognostic Index ( NPI ). We measured the same biomarkers in 531 noncancer individuals to assess multidimensional relationships ( MDR ), and evaluated additional prognostic models using survival artificial neural network ( SANN ) and survival support vector machines ( SSVM ), as these enhance capture of MDR s. For RFS , increasing concentrations of circulating IGFBP ‐2 and PAPP ‐A were independently prognostic [ HR biomarker doubling : 1.474 (95% CI s: 1.160, 1.875, P = 0.002) and 1.952 (95% CI s: 1.364, 2.792, P < 0.001), respectively]. The AUC RFS for NPI was 0.626 (Cox model), improving to 0.694 ( P = 0.012) with the addition of IGFBP ‐2 plus PAPP ‐A. Derived AUC RFS using SANN and SSVM did not perform superiorly. Similar patterns were observed for OS . These findings illustrate an important principle in biomarker qualification—measured circulating biomarkers may demonstrate independent prognostication, but this does not necessarily translate into substantial improvement in clinical performance.
Background: Aristolochic Acid (AA), a natural component of Aristolochia plants that is found in a variety of herbal remedies and health supplements, is classified as a Group 1 carcinogen by the International Agency for Research on Cancer. Given that microRNAs (miRNAs) are involved in cancer initiation and progression and their role remains unknown in AA-induced carcinogenesis, we examined genome-wide AA-induced dysregulation of miRNAs as well as the regulation of miRNAs on their target gene expression in rat kidney.Results: We treated rats with 10 mg/kg AA and vehicle control for 12 weeks and eight kidney samples (4 for the treatment and 4 for the control) were used for examining miRNA and mRNA expression by deep sequencing, and protein expression by proteomics. AA treatment resulted in significant differential expression of miRNAs, mRNAs and proteins as measured by both principal component analysis (PCA) and hierarchical clustering analysis (HCA). Specially, 63 miRNAs (adjusted p value < 0.05 and fold change > 1.5), 6,794 mRNAs (adjusted p value < 0.05 and fold change > 2.0), and 800 proteins (fold change > 2.0) were significantly altered by AA treatment. The expression of 6 selected miRNAs was validated by quantitative real-time PCR analysis. Ingenuity Pathways Analysis (IPA) showed that cancer is the top network and disease associated with those dysregulated miRNAs. To further investigate the influence of miRNAs on kidney mRNA and protein expression, we combined proteomic and transcriptomic data in conjunction with miRNA target selection as confirmed and reported in miRTarBase. In addition to translational repression and transcriptional destabilization, we also found that miRNAs and their target genes were expressed in the same direction at levels of transcription (169) or translation (227). Furthermore, we identified that up-regulation of 13 oncogenic miRNAs was associated with translational activation of 45 out of 54 cancer-related targets.Conclusions: Our findings suggest that dysregulated miRNA expression plays an important role in AA-induced carcinogenesis in rat kidney, and that the integrated approach of multiple profiling provides a new insight into a post-transcriptional regulation of miRNAs on their target repression and activation in a genome-wide scale.
Single-nucleotide polymorphisms (SNPs) determined based on SNP arrays from the international HapMap consortium (HapMap) and the genetic variants detected in the 1000 genomes project (1KGP) can serve as two references for genomewide association studies (GWAS). We conducted comparative analyses to provide a means for assessing concerns regarding SNP array-based GWAS findings as well as for realistically bounding expectations for next generation sequencing (NGS)-based GWAS. We calculated and compared base composition, transitions to transversions ratio, minor allele frequency and heterozygous rate for SNPs from HapMap and 1KGP for the 622 common individuals. We analysed the genotype discordance between HapMap and 1KGP to assess consistency in the SNPs from the two references. In 1KGP, 90.58% of 36,817,799 SNPs detected were not measured in HapMap. More SNPs with minor allele frequencies less than 0.01 were found in 1KGP than HapMap. The two references have low disc ordance (generally smaller than 0.02) in genotypes of common SNPs, with most discordance from heterozygous SNPs. Our study demonstrated that SNP array-based GWAS findings were reliable and useful, although only a small portion of genetic variances were explained. NGS can detect not only common but also rare variants, supporting the expectation that NGS-based GWAS will be able to incorporate a much larger portion of genetic variance than SNP arrays-based GWAS.
Background Gene expression profiling is being widely applied in cancer research to identify biomarkers for clinical endpoint prediction. Since RNA-seq provides a powerful tool for transcriptome-based applications beyond the limitations of microarrays, we sought to systematically evaluate the performance of RNA-seq-based and microarray-based classifiers in this MAQC-III/SEQC study for clinical endpoint prediction using neuroblastoma as a model. Results We generate gene expression profiles from 498 primary neuroblastomas using both RNA-seq and 44 k microarrays. Characterization of the neuroblastoma transcriptome by RNA-seq reveals that more than 48,000 genes and 200,000 transcripts are being expressed in this malignancy. We also find that RNA-seq provides much more detailed information on specific transcript expression patterns in clinico-genetic neuroblastoma subgroups than microarrays. To systematically compare the power of RNA-seq and microarray-based models in predicting clinical endpoints, we divide the cohort randomly into training and validation sets and develop 360 predictive models on six clinical endpoints of varying predictability. Evaluation of factors potentially affecting model performances reveals that prediction accuracies are most strongly influenced by the nature of the clinical endpoint, whereas technological platforms (RNA-seq vs. microarrays), RNA-seq data analysis pipelines, and feature levels (gene vs. transcript vs. exon-junction level) do not significantly affect performances of the models. Conclusions We demonstrate that RNA-seq outperforms microarrays in determining the transcriptomic characteristics of cancer, while RNA-seq and microarray-based models perform similarly in clinical endpoint prediction. Our findings may be valuable to guide future studies on the development of gene expression-based predictive models and their implementation in clinical practice.
BACKGROUND:Due to a significant decline in the costs associated with next-generation sequencing, it has become possible to decipher the genetic architecture of a population by sequencing a large number of individuals to a deep coverage. The Korean Personal Genomes Project (KPGP) recently sequenced 35 Korean genomes at high coverage using the Illumina Hiseq platform and made the deep sequencing data publicly available, providing the scientific community opportunities to decipher the genetic architecture of the Korean population.METHODS:In this study, we used two single nucleotide variant (SNV) calling pipelines: mapping the raw reads obtained from whole genome sequencing of 35 Korean individuals in KPGP using BWA and SOAP2 followed by SNV calling using SAMtools and SOAPsnp, respectively. The consensus SNVs obtained from the two SNV pipelines were used to represent the SNVs of the Korean population. We compared these SNVs to those from 17 other populations provided by the HapMap consortium and the 1000 Genomes Project (1KGP) and identified SNVs that were only present in the Korean population. We studied the mutation spectrum and analyzed the genes of non-synonymous SNVs only detected in the Korean population.RESULTS:We detected a total of 8,555,726 SNVs in the 35 Korean individuals and identified 1,213,613 SNVs detected in at least one Korean individual (SNV-1) and 12,640 in all of 35 Korean individuals (SNV-35) but not in 17 other populations. In contrast with the SNVs common to other populations in HapMap and 1KGP, the Korean only SNVs had high percentages of non-silent variants, emphasizing the unique roles of these Korean only SNVs in the Korean population. Specifically, we identified 8,361 non-synonymous Korean only SNVs, of which 58 SNVs existed in all 35 Korean individuals. The 5,754 genes of non-synonymous Korean only SNVs were highly enriched in some metabolic pathways. We found adhesion is the top disease term associated with SNV-1 and Nelson syndrome is the only disease term associated with SNV-35. We found that a significant number of Korean only SNVs are in genes that are associated with the drug term of adenosine.CONCLUSION:We identified the SNVs that were found in the Korean population but not seen in other populations, and explored the corresponding genes and pathways as well as the associated disease terms and drug terms. The results expand our knowledge of the genetic architecture of the Korean population, which will benefit the implementation of personalized medicine for the Korean population.
BACKGROUND:Gene expression microarray has been the primary biomarker platform ubiquitously applied in biomedical research, resulting in enormous data, predictive models, and biomarkers accrued. Recently, RNA-seq has looked likely to replace microarrays, but there will be a period where both technologies co-exist. This raises two important questions: Can microarray-based models and biomarkers be directly applied to RNA-seq data? Can future RNA-seq-based predictive models and biomarkers be applied to microarray data to leverage past investment?RESULTS:We systematically evaluated the transferability of predictive models and signature genes between microarray and RNA-seq using two large clinical data sets. The complexity of cross-platform sequence correspondence was considered in the analysis and examined using three human and two rat data sets, and three levels of mapping complexity were revealed. Three algorithms representing different modeling complexity were applied to the three levels of mappings for each of the eight binary endpoints and Cox regression was used to model survival times with expression data. In total, 240,096 predictive models were examined.CONCLUSIONS:Signature genes of predictive models are reciprocally transferable between microarray and RNA-seq data for model development, and microarray-based models can accurately predict RNA-seq-profiled samples; while RNA-seq-based models are less accurate in predicting microarray-profiled samples and are affected both by the choice of modeling algorithm and the gene mapping complexity. The results suggest continued usefulness of legacy microarray data and established microarray biomarkers and predictive models in the forthcoming RNA-seq era.
We present primary results from the Sequencing Quality Control (SEQC) project, coordinated by the US Food and Drug Administration. Examining Illumina HiSeq, Life Technologies SOLiD and Roche 454 platforms at multiple laboratory sites using reference RNA samples with built-in controls, we assess RNA sequencing (RNA-seq) performance for junction discovery and differential expression profiling and compare it to microarray and quantitative PCR (qPCR) data using complementary metrics. At all sequencing depths, we discover unannotated exon-exon junctions, with >80% validated by qPCR. We find that measurements of relative expression are accurate and reproducible across sites and platforms if specific filters are used. In contrast, RNA-seq and microarrays do not provide accurate absolute measurements, and gene-specific biases are observed for all examined platforms, including qPCR. Measurement performance depends on the platform and data analysis pipeline, and variation is large for transcript-level profiling. The complete SEQC data sets, comprising >100 billion reads (10Tb), provide unique resources for evaluating RNA-seq analyses for clinical and regulatory settings.
The rat has been used extensively as a model for evaluating chemical toxicities and for understanding drug mechanisms. However, its transcriptome across multiple organs, or developmental stages, has not yet been reported. Here we show, as part of the SEQC consortium efforts, a comprehensive rat transcriptomic BodyMap created by performing RNA-Seq on 320 samples from 11 organs of both sexes of juvenile, adolescent, adult and aged Fischer 344 rats. We catalogue the expression profiles of 40,064 genes, 65,167 transcripts, 31,909 alternatively spliced transcript variants and 2,367 non-coding genes/non-coding RNAs (ncRNAs) annotated in AceView. We find that organ-enriched, differentially expressed genes reflect the known organ-specific biological activities. A large number of transcripts show organ-specific, age-dependent or sex-specific differential expression patterns. We create a web-based, open-access rat BodyMap database of expression profiles with crosslinks to other widely used databases, anticipating that it will serve as a primary resource for biomedical research using the rat model.
Although many quality control (QC) methods have been developed to improve the quality of single-nucleotide variants (SNVs) in SNV-calling, QC methods for use subsequent to single-nucleotide polymorphism-calling have not been reported. We developed five QC metrics to improve the quality of SNVs using the whole-genome-sequencing data of a monozygotic twin pair from the Korean Personal Genome Project. The QC metrics improved both repeatability between the monozygotic twin pair and reproducibility between SNV-calling pipelines. We demonstrated the QC metrics improve reproducibility of SNVs derived from not only whole-genome-sequencing data but also whole-exome-sequencing data. The QC metrics are calculated based on the reference genome used in the alignment without accessing the raw and intermediate data or knowing the SNV-calling details. Therefore, the QC metrics can be easily adopted in downstream association analysis.
A comparison of RNA-seq and microarray data from samples treated with diverse drugs highlights a dependency of cross-platform concordance on treatment effect. The concordance of RNA-sequencing (RNA-seq) with microarrays for genome-wide analysis of differential gene expression has not been rigorously assessed using a range of chemical treatment conditions. Here we use a comprehensive study design to generate Illumina RNA-seq and Affymetrix microarray data from the same liver samples of rats exposed in triplicate to varying degrees of perturbation by 27 chemicals representing multiple modes of action (MOAs). The cross-platform concordance in terms of differentially expressed genes (DEGs) or enriched pathways is linearly correlated with treatment effect size (R2≈0.8). Furthermore, the concordance is also affected by transcript abundance and biological complexity of the MOA. RNA-seq outperforms microarray (93% versus 75%) in DEG verification as assessed by quantitative PCR, with the gain mainly due to its improved accuracy for low-abundance transcripts. Nonetheless, classifiers to predict MOAs perform similarly when developed using data from either platform. Therefore, the endpoint studied and its biological complexity, transcript abundance and the genomic application are important factors in transcriptomic research and for clinical and regulatory decision making.
RNA-seq facilitates unbiased genome-wide gene-expression profiling. However, its concordance with the well-established microarray platform must be rigorously assessed for confident uses in clinical and regulatory application. Here we use a comprehensive study design to generate Illumina RNA-seq and Affymetrix microarray data from the same set of liver samples of rats under varying degrees of perturbation by 27 chemicals representing multiple modes of action (MOA). The cross-platform concordance in terms of differentially expressed genes (DEGs) or enriched pathways is highly correlated with treatment effect size, gene-expression abundance and the biological complexity of the MOA. RNA-seq outperforms microarray (90% versus 76%) in DEG verification by quantitative PCR and the main gain is its improved accuracy for low expressed genes. Nonetheless, predictive classifiers derived from both platforms performed similarly. Therefore, the endpoint studied and its biological complexity, transcript abundance, and intended application are important factors in transcriptomic research and for decision-making. Emerging technologies facilitate basic science research, but their value in clinical and regulatory settings requires rigorous assessment and consensus within the research community. The U.S. Food and Drug Administration’s (FDA’s) initiative on advancing regulatory science embraces collaborations among various stakeholders to expedite translation of advancement in basic science to regulatory application1. In the past decade, microarrays have been a principal technology for analyzing transcriptomes to support drug development and safety evaluation2. The FDA launched the community-wide MicroArray Quality Control (MAQC) consortium to investigate the reliability and utility of microarrays in identifying differentially expressed genes (DEGs) and predicting patient/toxicity outcomes based on gene-expression data in the first (MAQC-I)3, 4 and second (MAQCII)5, 6 phases of the project, respectively. MAQC-I and MAQC-II demonstrated the critical Wang et al. Page 2 Nat Biotechnol. Author manuscript; available in PMC 2014 November 25. N IH -P A A uhor M anscript N IH -P A A uhor M anscript N IH -P A A uhor M anscript roles of a comprehensive study design and crowd sourcing model to reach community-wide consensus on the fit-for-purpose use of emerging technologies. High-throughput sequencing technologies provide new methods for whole-transcriptome analyses of gene expression7. Recently published studies have compared data obtained from microarrays and RNA-seq in terms of technical reproducibility, variance structure, absolute expression and detection of DEGs or gene isoforms8–20 (Supplementary Table 1). Some of these studies suggested that RNA-seq exhibits lower precision for weakly expressed genes owing to the nature of sampling21, 22, whereas others found higher sensitivity of RNA-seq for gene detection23, 24. The varied conclusions can be attributed to the fact that they used few treatment conditions and hence they do not cover a wide range of biologic complexity. Furthermore, the question has not been adequately addressed about whether predicting toxicity outcomes based on gene-expression data could be enhanced with RNA-seq over microarray. Under the umbrella of the third phase of the MAQC consortium3–6, also known as the SEquencing Quality Control (SEQC) project, we conducted a comprehensive study to evaluate RNA-seq in its differences and similarities to microarrays in terms of identifying DEGs and developing predictive models. In contrast to data generated as part of the SEQC project using reference RNA samples25, our study design provides a comparison of the transcription response for rat livers that each platform detects in terms of extensive chemical treatments, biologic replication and breath of shared mode of action (MOA) of the chemicals beyond simply monitoring performance metrics. Specifically, we report the results of a comparative analysis of gene expression responses profiled by Affymetrix microarray and Illumina RNA-seq in liver tissue from rats exposed to diverse chemicals. We used either microarray or RNA-seq data to generate DEGs and predictive models of MOA of each chemical. This allowed us to assess the influence of the chemical (referred hereafter as the ‘treatment effect’) on the concordance between RNA-seq and microarrays and on the performance of predictive models generated using each technology. Treatment effect is characterized by the number of DEGs and the overexpressed pathways underlying MOA of the chemical. We found that (i) the concordance between array and sequencing platforms for detecting the number of DEGs was positively correlated with the extensive perturbation elicited by the treatment, (ii) RNA-seq performed better than microarrays at detecting weakly expressed genes, and (iii) gene expression–based predictive models generated from RNA-seq and microarray data were similar. The experimental design also allowed us to identify positive correlations in differentially expressed RNA elements (mRNA, splice variants, non-coding RNA and exon-exon junction) with the extensive perturbation elicited by the treatment, and to examine treatment-induced alternative splicing and shortening of 3’ untranslated regions (UTRs).
RNA-Seq provides the capability to characterize the entire transcriptome in multiple levels including gene expression, allele specific expression, alternative splicing, fusion gene detection, and etc. The US FDA-led SEQC (i.e., MAQC-III) project conducted a comprehensive study focused on the transcriptome profiling of rat liver samples treated with 27 chemicals to evaluate the utility of RNA-Seq in safety assessment and toxicity mechanism elucidation. The chemicals represented multiple chemogenomic modes of action (MOA) and exhibited varying degrees of transcriptional response. The paired-end 100 bp sequencing data were generated using Illumina HiScanSQ and/or HiSeq 2000. In addition to the core study, six animals (i.e., three aflatoxin B1 treated rats and three vehicle control rats) were sequenced three times, with two separate library preparations on two sequencing machines. This large toxicogenomics dataset can serve as a resource to characterize various aspects of transcriptomic changes (e.g., alternative splicing) that are byproduct of chemical perturbation.
The aim of this review is to comprehensively summarize the recent achievements in the field of toxicogenomics and cancer research regarding genetic-environmental interactions in carcinogenesis and detection of genetic aberrations in cancer genomes by next-generation sequencing technology. Cancer is primarily a genetic disease in which genetic factors and environmental stimuli interact to cause genetic and epigenetic aberrations in human cells. Mutations in the germline act as either high-penetrance alleles that strongly increase the risk of cancer development, or as low-penetrance alleles that mildly change an individual's susceptibility to cancer. Somatic mutations, resulting from either DNA damage induced by exposure to environmental mutagens or from spontaneous errors in DNA replication or repair are involved in the development or progression of the cancer. Induced or spontaneous changes in the epigenome may also drive carcinogenesis. Advances in next-generation sequencing technology provide us opportunities to accurately, economically, and rapidly identify genetic variants, somatic mutations, gene expression profiles, and epigenetic alterations with single-base resolution. Whole genome sequencing, whole exome sequencing, and RNA sequencing of paired cancer and adjacent normal tissue present a comprehensive picture of the cancer genome. These new findings should benefit public health by providing insights in understanding cancer biology, and in improving cancer diagnosis and therapy.
The rat is used extensively by the pharmaceutical, regulatory, and academic communities for safety assessment of drugs and chemicals and for studying human diseases; however, its transcriptome has not been well studied. As part of the SEQC (i.e., MAQC-III) consortium efforts, a comprehensive RNA-Seq data set was constructed using 320 RNA samples isolated from 10 organs (adrenal gland, brain, heart, kidney, liver, lung, muscle, spleen, thymus, and testes or uterus) from both sexes of Fischer 344 rats across four ages (2-, 6-, 21-, and 104-week-old) with four biological replicates for each of the 80 sample groups (organ-sex- age). With the Ribo-Zero rRNA removal and Illumina RNA-Seq protocols, 41 million 50 bp single-end reads were generated per sample, yielding a total of 13.4 billion reads. This data set could be used to identify and validate new rat genes and transcripts, develop a more comprehensive rat transcriptome annotation system, identify novel gene regulatory networks related to tissue specific gene expression and development, and discover genes responsible for disease and drug toxicity and efficacy.
Whole-transcriptome sequencing (‘RNA-Seq’) has been drastically changing the scale and scope of genomic research. In order to fully understand the power and limitations of this technology, the US Food and Drug Administration (FDA) launched the third phase of the MicroArray Quality Control (MAQC-III) project, also known as the SEquencing Quality Control (SEQC) project. Using two well-established human reference RNA samples from the first phase of the MAQC project, three sequencing platforms were tested across more than ten sites with built-in truths including spike-in of external RNA controls (ERCC), titration data and qPCR verification. The SEQC project generated over 30 billion sequence reads representing the largest RNA-Seq data ever generated by a single project on individual RNA samples. This extraordinarily ultradeep transcriptomic data set and the known truths built into the study design provide many opportunities for further research and development to advance the improvement and application of RNA-Seq.