RNA sequencing (RNA-seq) is gaining popularity as a complementary assay to genome sequencing for precisely identifying the molecular causes of rare disorders. A powerful approach is to identify aberrant gene expression levels as potential pathogenic events. However, existing methods for detecting aberrant read counts in RNA-seq data either lack assessments of statistical significance, so that establishing cutoffs is arbitrary, or rely on subjective manual corrections for confounders. Here, we describe OUTRIDER (Outlier in RNA-Seq Finder), an algorithm developed to address these issues. The algorithm uses an autoencoder to model read-count expectations according to the gene covariation resulting from technical, environmental, or common genetic variations. Given these expectations, the RNA-seq read counts are assumed to follow a negative binomial distribution with a gene-specific dispersion. Outliers are then identified as read counts that significantly deviate from this distribution. The model is automatically fitted to achieve the best recall of artificially corrupted data. Precision-recall analyses using simulated outlier read counts demonstrated the importance of controlling for covariation and significance-based thresholds. OUTRIDER is open source and includes functions for filtering out genes not expressed in a dataset, for identifying outlier samples with too many aberrantly expressed genes, and for detecting aberrant gene expression on the basis of false-discovery-rate-adjusted p values. Overall, OUTRIDER provides an end-to-end solution for identifying aberrantly expressed genes and is suitable for use by rare-disease diagnostic platforms.
Across a variety of Mendelian disorders, ∼50–75% of patients do not receive a genetic diagnosis by exome sequencing indicating disease-causing variants in non-coding regions. Although genome sequencing in principle reveals all genetic variants, their sizeable number and poorer annotation make prioritization challenging. Here, we demonstrate the power of transcriptome sequencing to molecularly diagnose 10% (5 of 48) of mitochondriopathy patients and identify candidate genes for the remainder. We find a median of one aberrantly expressed gene, five aberrant splicing events and six mono-allelically expressed rare variants in patient-derived fibroblasts and establish disease-causing roles for each kind. Private exons often arise from cryptic splice sites providing an important clue for variant prioritization. One such event is found in the complex I assembly factor TIMMDC1 establishing a novel disease-associated gene. In conclusion, our study expands the diagnostic tools for detecting non-exonic variants and provides examples of intronic loss-of-function variants with pathological relevance.
Microarray technologies are established approaches for high throughput gene expression, methylation and genotyping analysis. An accurate mapping of the array probes is essential to generate reliable biological findings. However, manufacturers of the microarray platforms typically provide incomplete and outdated annotation tables, which often rely on older genome and transcriptome versions that differ substantially from up-to-date sequence databases. Here, we present the Re-Annotator, a re-annotation pipeline for microarray probe sequences. It is primarily designed for gene expression microarrays but can also be adapted to other types of microarrays. The Re-Annotator uses a custom-built mRNA reference database to identify the positions of gene expression array probe sequences. We applied Re-Annotator to the Illumina Human-HT12 v4 microarray platform and found that about one quarter (25%) of the probes differed from the manufacturer's annotation. In further computational experiments on experimental gene expression data, we compared Re-Annotator to another probe re-annotation tool, ReMOAT, and found that Re-Annotator provided an improved re-annotation of microarray probes. A thorough re-annotation of probe information is crucial to any microarray analysis. The Re-Annotator pipeline is freely available at http://sourceforge.net/projects/reannotator along with re-annotated files for Illumina microarrays HumanHT-12 v3/v4 and MouseRef-8 v2.
Standfirst text and highlights: Local trans regulation, mainly due to negative feedback, buffers effects of cis-regulatory variants by about 15%. This buffering is stronger for essential genes and genes with low to middle expression levels, for which tight regulation matters most. • Novel experimental design using expression of a diploid hybrid and its haploid spores allows systematic dissection of cis and local trans regulation • Local trans effects buffer effects of cis-regulatory variants in yeast by typically 15% • Local trans buffering is primarily due to negative feedback • Negative feedback as robustness strategy for genes with low to medium expression level
Acute liver failure (ALF) in infancy and childhood is a life-threatening emergency. Few conditions are known to cause recurrent acute liver failure (RALF), and in about 50% of cases, the underlying molecular cause remains unresolved. Exome sequencing in five unrelated individuals with fever-dependent RALF revealed biallelic mutations in NBAS. Subsequent Sanger sequencing of NBAS in 15 additional unrelated individuals with RALF or ALF identified compound heterozygous mutations in an additional six individuals from five families. Immunoblot analysis of mutant fibroblasts showed reduced protein levels of NBAS and its proposed interaction partner p31, both involved in retrograde transport between endoplasmic reticulum and Golgi. We recommend NBAS analysis in individuals with acute infantile liver failure, especially if triggered by fever.
Mechanisms conferring robustness against regulatory variants have been controversial. Previous studies suggested widespread buffering of RNA misexpression on protein levels during translation. We do not find evidence that translational buffering is common. Instead, we find extensive buffering at the level of RNA expression, exerted through negative feedback regulation acting in trans, which reduces the effect of regulatory variants on gene expression. Our approach is based on a novel experimental design in which allelic differential expression in a yeast hybrid strain is compared to allelic differential expression in a pool of its spores. Allelic differential expression in the hybrid is due to cis-regulatory differences only. Instead, in the pool of spores allelic differential expression is not only due to cis-regulatory differences but also due to local trans effects that include negative feedback. We found that buffering through such local trans regulation is widespread, typically compensating for about 15% of cis-regulatory effects on individual genes. Negative feedback is stronger not only for essential genes, indicating its functional relevance, but also for genes with low to middle levels of expression, for which tight regulation matters most. We suggest that negative feedback is one mechanism of Waddington's canalization, facilitating the accumulation of genetic variants that might give selective advantage in different environments.
High-throughput DNA sequencing (HTS) is of increasing importance in the life sciences. One of its most prominent applications is the sequencing of whole genomes or targeted regions of the genome such as all exonic regions (i.e., the exome). Here, the objective is the identification of genetic variants such as single nucleotide polymorphisms (SNPs). The extraction of SNPs from the raw genetic sequences involves many processing steps and the application of a diverse set of tools. We review the essential building blocks for a pipeline that calls SNPs from raw HTS data. The pipeline includes quality control, mapping of short reads to the reference genome, visualization and post-processing of the alignment including base quality recalibration. The final steps of the pipeline include the SNP calling procedure along with filtering of SNP candidates. The steps of this pipeline are accompanied by an analysis of a publicly available whole-exome sequencing dataset. To this end, we employ several alignment programs and SNP calling routines for highlighting the fact that the choice of the tools significantly affects the final results.
(Note: With the exception of the correction of typographical or spelling errors that could be a source of ambiguity, letters and reports are not edited. The original formatting of letters and referee reports may not be reflected in this compilation.) Thank you again for submitting your work to Molecular Systems Biology. We have now heard back from the three referees who accepted to evaluate the study. As you will see, the referees find the topic of your study of potential interest and are supportive. They raise however a series of concerns and make suggestions for modifications, which we would ask you to carefully address in a revision of the present work. The recommendations provided by the reviewers are very clear in this regard and refer mainly to additional clarifications and the need to word some of the conclusions more carefully. The authors use an elegant and efficient experimental design using RNA-seq in a hybrid diploid yeast and pools of halpoid segregants to distinguish between genetic variants close ('local') to a gene that influence expression in cis and trans. This allows them to reach the conclusion that many genes regulate there own expression by feedback ('tho the mechanisms are obscure). They also re-analyse ribosome profiling data that was previously used to suggest buffering of changes in RNA expression via changes in translation. They find no evidence for this but rather show the previous result is likely due to a technical artefact in the data processing. Overall I found this is a very interesting study that systematically addresses a basic and important question in gene regulation genome-wide. Minor suggestions: In the analysis with respect to gene importance the authors (need to) control for expression level. It would be reassuring to see this elsewhere. e.g. is there a bias in the detection of local-trans or local-cis eQTLs depending upon expression levels in the pools and does this effect the results? Molecular mechanisms-(for good reason) the authors speculate little about possible molecular mechanisms. But it would be interesting to know a bit more about the properties/functions of the genes they do detect feedback for-what functional categories/age of proteins do they cover etc? Fig 1B-I don't understand the choice of example shown here. If I understand correctly in this example the example is not any of their classes in fig 1C (looks like the expression levels are swapped between the diploids and haploids?) Abstract-delete the …