MOTIVATION:Sequencing coverage is among key determinants considered in the design of omics studies. To help estimate cost-effective sequencing coverage for specific downstream analysis, downsampling, a technique to sample subsets of reads with a specific size, is routinely used. However, as the size of sequencing becomes larger and larger, downsampling becomes computationally challenging. RESULTS:Here, we developed an approximate downsampling method called s-leaping that was designed to efficiently and accurately process large-size data. We compared the performance of s-leaping with state-of-the-art downsampling methods in a range of practical omics-study downsampling settings and found s-leaping to be up to 39% faster than the second-fastest method, with comparable accuracy to the exact downsampling methods. To apply s-leaping on FASTQ data, we developed a light-weight tool called fadso in C. Using whole-genome sequencing data with 208 million reads, we compared fadso's performance with that of a commonly used FASTQ tool with the same downsampling feature and found fadso to be up to 12% faster with 21% lower memory usage, suggesting fadso to have up to 40% higher throughput in a parallel computing setting. AVAILABILITY AND IMPLEMENTATION:The C source code for s-leaping, as well as the fadso package is freely available at https://github.com/hkuwahara/sleaping.
Autozygosity is associated with rare Mendelian disorders and clinically relevant quantitative traits. We investigated associations between the fraction of the genome in runs of homozygosity (FROH) and common diseases in Genes & Health (n = 23,978 British South Asians), UK Biobank (n = 397,184), and 23andMe. We show that restricting analysis to offspring of first cousins is an effective way of reducing confounding due to social/environmental correlates of FROH. Within this group in G&H+UK Biobank, we found experiment-wide significant associations between FROH and twelve common diseases. We replicated associations with type 2 diabetes (T2D) and post-traumatic stress disorder via within-sibling analysis in 23andMe (median n = 480,282). We estimated that autozygosity due to consanguinity accounts for 5%-18% of T2D cases among British Pakistanis. Our work highlights the possibility of widespread non-additive genetic effects on common diseases and has important implications for global populations with high rates of consanguinity.
Male damselfish typically demonstrate uniparental egg-guarding care in nature. Potential plasticity in sexual behavior has recently been reported in various teleost fish. To examine behavioral plasticity in parental care, we conducted aquarium experiments to explore the potential for egg-guarding care in the female damselfish, Dascyllus reticulatus. After initial caretaking, males were removed from the mating nests, and cohabiting females frequently exhibited egg predation on the same day. However, we confirmed that females showed significantly decreased egg-predation frequencies on the following day and showed egg-caring behaviors. All experimental females guarded their eggs until they hatched. Females subsequently spawned eggs as females even after performing parental care behaviors, indicating no progression of sex change into males. Molecular analysis of select pituitary gland hormones indicated that egg-caring females and males showed high expression levels of prolactin, suggesting its involvement in the development of parental care behaviors. The cryptic possession of caretaking ability in females may be a tactical response to the need for temporary replacement of the care roles in cases where caretaking males are removed, for example, through predation, in damselfish species living in sexually cohabiting groups.
Arabs represent 5% of the world population and have a high prevalence of common disease, yet remain greatly underrepresented in genome-wide association studies, where only 1 in 600 individuals are Arab. We highlight the persistent and unaddressed underrepresentation of Arabs in genomic databases and discuss its impact on public health genomics and missed opportunities for biological discovery.
Arabs account for 5% of the world population and have a high burden of cardiometabolic disease, yet clinical utility of polygenic risk prediction in Arabs remains understudied. Among 5399 Arab patients, we optimize polygenic scores for 10 cardiometabolic traits, achieving a performance that is better than published scores and on par with performance in European-ancestry individuals. Odds ratio per standard deviation (OR per SD) for a type 2 diabetes score was 1.83 (95% CI 1.74–1.92), and each SD of body mass index (BMI) score was associated with 1.18 kg/m 2 difference in BMI. Polygenic scores associated with disease independent of conventional risk factors, and also associated with disease severity—OR per SD for coronary artery disease (CAD) was 1.78 (95% CI 1.66–1.90) for three-vessel CAD and 1.41 (95% CI 1.29–1.53) for one-vessel CAD. We propose a pragmatic framework leveraging public data as one way to advance equitable clinical implementation of polygenic scores in non-European populations.
Introduction: Saudi Arabia is the largest country in the Middle East and has an alarming rise in cardiometabolic disease. While conventional risk factors are highly prevalent, it is not clear if and how genomic risk might be playing a role. Methods: In 5,399 Saudi patients (age 54.8±14.8 years, 35.7% female) referred for cardiac catheterization, we collected detailed cardiometabolic phenotype data and performed whole-genome genotyping. We implemented a computational framework to develop Arab-specific polygenic scores for 9 traits (Figure A, B). We studied the association of the polygenic scores with cardiometabolic disease, and their interplay with conventional risk factors using regression models adjusted for age, sex, and genetic ancestry. Results: Within this disease cohort with high prevalence of conventional risk factors - 65.0% coronary artery disease (CAD), 55.6% type 2 diabetes (DM-2), 41.5% obesity, 48.8% hypercholesterolemia, 80.6% hypertension, and 38.3% smoking - polygenic scores were strongly associated with cardiometabolic traits (Figure C). The effect sizes of polygenic scores for CAD and DM-2 remained unchanged with adjustment for clinical risk factors. Genomic risk was also associated with the severity of CAD, as measured by the number of vessels with severe stenosis on cardiac catheterization (Figure D). CAD patients with high genomic risk - defined as top quintile of the polygenic score - (N=568) had earlier onset (54.8 vs. 57.9 years, p=0.001) and more severe disease (41.3% vs. 29.9% with three-vessel CAD, p<0.001) despite similar clinical risk factors, compared to CAD patients without high genomic risk (N=1144). Conclusions: Polygenic scores were strongly associated with cardiometabolic disease even in a non-European population with high prevalence of conventional risk factors. Genomic risk for CAD appears independent of conventional risk factors and identifies patients with earlier onset and more severe disease.
Recombination is one of the essential genetic processes for sexually reproducing organisms, which can happen more frequently in some regions, called recombination hotspots. Although several factors, such as PRDM9 binding motifs, are known to be related to the hotspots, their contributions to the recombination hotspots have not been quantified, and other determinants are yet to be elucidated. Here, we develop a computational method, RHSNet, based on deep learning and signal processing, to identify and quantify the hotspot determinants in a purely data-driven manner, utilizing datasets from various studies, populations, sexes, and species. In addition to being able to identify hotspot regions and the well-known determinants accurately, RHSNet is sensitive to the difference between different PRDM9 alleles and different sexes, and can generalize to PRDM9-lacking species. The cross-sex, cross-population, and cross-species studies suggest that the proposed method has the potential to identify and quantify the evolutionary determinant motifs. Teaser RHSNet can accurately identify and quantify recombination hotspot determinants across different studies, sexes, populations, and species.
Two-dimensional (2D) chemical fingerprints are widely used as binary features for the quantification of structural similarity of chemical compounds, which is an important step in similarity-based virtual screening (VS). Here, using an eigenvalue-based entropy approach, we identified 2D fingerprints with little to no contribution to shaping the eigenvalue distribution of the feature matrix as related ones and examined the degree to which these related 2D fingerprints influenced molecular similarity scores calculated with the Tanimoto coefficient. Our analysis identified many related fingerprints in publicly available fingerprint schemes and showed that their presence in the feature set could have substantial effects on the similarity scores and bias the outcome of molecular similarity analysis. Our results have implication in the optimal selection of 2D fingerprints for compound similarity analysis and the identification of potential hits for compounds with target biological activity in VS.
Ribonucleic acid (RNA) secondary structures are the determining factor for the many roles RNA plays in life. They can be obtained experimentally by techniques such as X-ray diffraction and NMR imaging. However, given that experimental methods are laborious and expensive, they are not fit for high throughput analysis. Computational prediction algorithms are complementary in predicting the RNA secondary structures for the multitude of known RNA sequences lacking structural information. Here, we introduce NNfold, a novel sequence-based deep neural network method to predict RNA secondary structures. The predictions are made by combining a local and a global model: first, we construct a matrix with the pairing likelihood of each nucleotide by predicting all potential interactions using a convolutional deep learning model. Next, we modify the list of base pairs obtained from the matrix using a second model whose output is used to ensure the contextual validity of the predicted secondary structure. Within the RNA Strand database, NNfold performed much better than thermodynamics-based methods on a diverse set of RNA sequences, improving the average F1 score by 0.20. It is capable of predicting pseudoknots which is a challenging task for other approaches. We also extracted the learned thermodynamic features within the model, which can help advance the construction of new biological models to predict RNA secondary structures. Our developed method is available as a service online at http://www.cbrc.kaust.edu.sa/NNfold/, or as an installable package at https://github.com/ramzan1990/NNfold.
Background Molecular autopsy refers to DNA-based identification of the cause of death. Despite recent attempts to broaden its scope, the term remains typically reserved to sudden unexplained death in young adults. In this study, we aim to showcase the utility of molecular autopsy in defining lethal variants in humans. Methods We describe our experience with a cohort of 481 cases in whom the cause of premature death was investigated using DNA from the index or relatives (molecular autopsy by proxy). Molecular autopsy tool was typically exome sequencing although some were investigated using targeted approaches in the earlier stages of the study; these include positional mapping, targeted gene sequencing, chromosomal microarray, and gene panels. Results The study includes 449 cases from consanguineous families and 141 lacked family history (simplex). The age range was embryos to 18 years. A likely causal variant (pathogenic/likely pathogenic) was identified in 63.8% (307/481), a much higher yield compared to the general diagnostic yield (43%) from the same population. The predominance of recessive lethal alleles allowed us to implement molecular autopsy by proxy in 55 couples, and the yield was similarly high (63.6%). We also note the occurrence of biallelic lethal forms of typically non-lethal dominant disorders, sometimes representing a novel bona fide biallelic recessive disease trait. Forty-six disease genes with no OMIM phenotype were identified in the course of this study. The presented data support the candidacy of two other previously reported novel disease genes (FAAH2 and MSN). The focus on lethal phenotypes revealed many examples of interesting phenotypic expansion as well as remarkable variability in clinical presentation. Furthermore, important insights into population genetics and variant interpretation are highlighted based on the results. Conclusions Molecular autopsy, broadly defined, proved to be a helpful clinical approach that provides unique insights into lethal variants and the clinical annotation of the human genome.
In response to severe genetic and environmental perturbations, wild-type organisms can express hidden alternative phenotypes adaptive to such adverse conditions. While our theoretical understanding of the population-level fitness advantage and evolution of phenotypic switching under variable environments has grown, the mechanism by which these organisms maintain phenotypic switching capabilities under static environments remains to be elucidated. Here, using computational simulations, we analyzed the evolution of gene circuits under natural selection and found that different strategies evolved to increase the gene expression stability near the optimum level. In a population comprising bistable individuals, a strategy of maintaining bistability and raising the potential barrier separating the bistable regimes was consistently taken. Our results serve as evidence that hidden bistable switches can be stably maintained during environmental stasis-an essential property enabling the timely release of adaptive alternatives with small genetic changes in the event of substantial perturbations.
BACKGROUND:At least 50% of patients with suspected Mendelian disorders remain undiagnosed after whole-exome sequencing (WES), and the extent to which non-coding variants that are not captured by WES contribute to this fraction is unclear. Whole transcriptome sequencing is a promising supplement to WES, although empirical data on the contribution of RNA analysis to the diagnosis of Mendelian diseases on a large scale are scarce.RESULTS:Here, we describe our experience with transcript-deleterious variants (TDVs) based on a cohort of 5647 families with suspected Mendelian diseases. We first interrogate all families for which the respective Mendelian phenotype could be mapped to a single locus to obtain an unbiased estimate of the contribution of TDVs at 18.9%. We examine the entire cohort and find that TDVs account for 15% of all "solved" cases. We compare the results of RT-PCR to in silico prediction. Definitive results from RT-PCR are obtained from blood-derived RNA for the overwhelming majority of variants (84.1%), and only a small minority (2.6%) fail analysis on all available RNA sources (blood-, skin fibroblast-, and urine renal epithelial cells-derived), which has important implications for the clinical application of RNA-seq. We also show that RNA analysis can establish the diagnosis in 13.5% of 155 patients who had received "negative" clinical WES reports. Finally, our data suggest a role for TDVs in modulating penetrance even in otherwise highly penetrant Mendelian disorders.CONCLUSIONS:Our results provide much needed empirical data for the impending implementation of diagnostic RNA-seq in conjunction with genome sequencing.
We have previously described a heart-, eye-, and brain-malformation syndrome caused by homozygous loss-of-function variants in SMG9, which encodes a critical component of the nonsense-mediated decay (NMD) machinery. Here, we describe four consanguineous families with four different likely deleterious homozygous variants in SMG8, encoding a binding partner of SMG9. The observed phenotype greatly resembles that linked to SMG9 and comprises severe global developmental delay, microcephaly, facial dysmorphism, and variable congenital heart and eye malformations. RNA-seq analysis revealed a general increase in mRNA expression levels with significant overrepresentation of core NMD substrates. We also identified increased phosphorylation of UPF1, a key SMG1-dependent step in NMD, which most likely represents the loss of SMG8-mediated inhibition of SMG1 kinase activity. Our data show that SMG8 and SMG9 deficiency results in overlapping developmental disorders that most likely converge mechanistically on impaired NMD.
ABSTRACT Motivation Proper prioritization of candidate genes is essential to the genome-based diagnostics of a range of genetic diseases. However, it is a highly challenging task involving limited and noisy knowledge of genes, diseases and their associations. While a number of computational methods have been developed for the disease gene prioritization task, their performance is largely limited by manually crafted features, network topology, or pre-defined rules of data fusion. Results Here, we propose a novel graph convolutional network-based disease gene prioritization method, PGCN, through the systematic embedding of the heterogeneous network made by genes and diseases, as well as their individual features. The embedding learning model and the association prediction model are trained together in an end-to-end manner. We compared PGCN with five state-of-the-art methods on the Online Mendelian Inheritance in Man (OMIM) dataset for tasks to recover missing associations and discover associations between novel genes and diseases. Results show significant improvements of PGCN over the existing methods. We further demonstrate that our embedding has biological meaning and can capture functional groups of genes. Availability The main program and the data are available at https://github.com/lykaust15/Disease_gene_prioritization_GCN .
MOTIVATION:Computational identification of promoters is notoriously difficult as human genes often have unique promoter sequences that provide regulation of transcription and interaction with transcription initiation complex. While there are many attempts to develop computational promoter identification methods, we have no reliable tool to analyze long genomic sequences. RESULTS:In this work, we further develop our deep learning approach that was relatively successful to discriminate short promoter and non-promoter sequences. Instead of focusing on the classification accuracy, in this work we predict the exact positions of the transcription start site inside the genomic sequences testing every possible location. We studied human promoters to find effective regions for discrimination and built corresponding deep learning models. These models use adaptively constructed negative set, which iteratively improves the model's discriminative ability. Our method significantly outperforms the previously developed promoter prediction programs by considerably reducing the number of false-positive predictions. We have achieved error-per-1000-bp rate of 0.02 and have 0.31 errors per correct prediction, which is significantly better than the results of other human promoter predictors. AVAILABILITY AND IMPLEMENTATION:The developed method is available as a web server at http://www.cbrc.kaust.edu.sa/PromID/.
MOTIVATION:Accurate and wide-ranging prediction of thermodynamic parameters for biochemical reactions can facilitate deeper insights into the workings and the design of metabolic systems. RESULTS:Here, we introduce a machine learning method with chemical fingerprint-based features for the prediction of the Gibbs free energy of biochemical reactions. From a large pool of 2D fingerprint-based features, this method systematically selects a small number of relevant ones and uses them to construct a regularized linear model. Since a manual selection of 2D structure-based features can be a tedious and time-consuming task, requiring expert knowledge about the structure-activity relationship of chemical compounds, the systematic feature selection step in our method offers a convenient means to identify relevant 2D fingerprint-based features. By comparing our method with state-of-the-art linear regression-based methods for the standard Gibbs free energy prediction, we demonstrated that its prediction accuracy and prediction coverage are most favorable. Our results show direct evidence that a number of 2D fingerprints collectively provide useful information about the Gibbs free energy of biochemical reactions and that our systematic feature selection procedure provides a convenient way to identify them. AVAILABILITY AND IMPLEMENTATION:Our software is freely available for download at http://sfb.kaust.edu.sa/Pages/Software.aspx. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Transcriptome level analysis has been shown to have the great potential for clinical utility. Here, we introduce omega, a between-sample RNA-seq quantification to estimate the abundance level of functional mRNAs which is suitable for analysis of transcriptional aberrations and molecular diagnostics of a range of genetic diseases. By using five diagnosed cases of Mendelian diseases as a case study, we show evidence that omega can improve the signal to detect genes with deleterious transcriptional aberrations and drastically reduce the disease-gene search space.
The export option will allow you to export the current search results of the entered query to a file. Different formats are available for download. To export the items, click on the button corresponding with the preferred download format.By default, clicking on the export buttons will result in a download of the allowed maximum amount of items. For anonymous users the allowed maximum amount is 50 search results.
Computational identification of promoters is notoriously difficult as human genes often have unique promoter sequences that provide regulation of transcription and interaction with transcription initiation complex. While there are many attempts to develop computational promoter identification methods, we have no reliable tool to analyze long genomic sequences. In this work we further develop our deep learning approach that was relatively successful to discriminate short promoter and non-promoter sequences. Instead of focusing on the classification accuracy, in this work we predict the exact positions of the TSS inside the genomic sequences testing every possible location. We studied human promoters to find effective regions for discrimination and built corresponding deep learning models. These models use adaptively constructed negative set which iteratively improves the models discriminative ability. The developed promoter identification models significantly outperform the previously developed promoter prediction programs by considerably reducing the number of false positive predictions. The best model we have built has recall 0.76, precision 0.77 and MCC 0.76, while the next best tool FPROM achieved precision 0.48 and MCC 0.60 for the recall of 0.75. Our method is available at http://www.cbrc.kaust.edu.sa/PromID/.
The Synthetic Biology Open Language (SBOL) is a community-driven open language to promote standardization in synthetic biology. To support the use of SBOL in metabolic engineering, we developed SBOLme, the first open-access repository of SBOL 2-compliant biochemical parts for a wide range of metabolic engineering applications. The URL of our repository is http://www.cbrc.kaust.edu.sa/sbolme .