The biological process of RNA translation is fundamental to cellular life and has wide-ranging implications for human disease. Accurate delineation of RNA translation variation represents a significant challenge due to the complexity of the process and technical limitations. Here, we introduce RiboTIE, a transformer model-based approach designed to enhance the analysis of ribosome profiling data. Unlike existing methods, RiboTIE leverages raw ribosome profiling counts directly to robustly detect translated open reading frames (ORFs) with high precision and sensitivity, evaluated on a diverse set of datasets. We demonstrate that RiboTIE successfully recapitulates known findings and provides novel insights into the regulation of RNA translation in both normal brain and medulloblastoma cancer samples. Our results suggest that RiboTIE is a versatile tool that can significantly improve the accuracy and depth of Ribo-Seq data analysis, thereby advancing our understanding of protein synthesis and its implications in disease.
Background Many cardiovascular diseases are associated with chronic kidney disease (CKD) and their burden as comorbidities is high. There is a pressing need to understand the molecular mechanisms by which cardiovascular diseases drive the development of CKD in multimorbidity. Methods and Results We performed an analysis of the risk of CKD onset across the full spectrum of cardiovascular disease phenotypes using UK Biobank primary and secondary care data (N = 129,824). We identified 17,347 cases of new onset CKD stage 3–5 within patients with pre-existing type 2 diabetes (T2D) (N = 13,409) and without diabetes (N = 116,415). We identified 8 cardiovascular disease phenotypes significantly associated with CKD onset including heart failure, atrial fibrillation, myocardial infarction, hypertension, ischaemic heart disease, peripheral vascular disease, hypotension, and angina pectoris. We performed a genome-wide association study (GWAS) to identify genetic variants enriched within identified high-risk CKD cardiovascular disease backgrounds and found 9 novel significant (p < 5 x 10-8) genetic loci associated with CKD in varying cardiovascular disease backgrounds. We used epigenomic databases and kidney eQTL data to link identified lead SNPs to likely target genes and found 1 novel CKD-associated gene in a background of hypertension with a clear SNP-to-gene assignment. We corroborated the identified gene-CKD association using single-cell RNA-seq kidney data and found the novel CKD-associated gene to have a highly specific cell-type expression within kidney podocytes. Conclusions We identified 9 novel CKD-associated loci that were enriched within high-risk-CKD, cardiovascular disease backgrounds. We identified a novel CKD-associated gene against a background of hypertension and found this gene to be highly enriched in podocytes in kidneys. This work has identified novel genetic and molecular factors associated with high-risk cardiovascular to CKD trajectories.
AbstractOsteoarthritis (OA) is increasing in prevalence and has a severe impact on patients’ lives. However, our understanding of biomarkers driving OA risk remains limited. We developed a model predicting the five-year risk of OA diagnosis, integrating retrospective clinical, lifestyle and biomarker data from the UK Biobank (19,120 patients with OA, ROC-AUC: 0.72, 95%CI (0.71–0.73)). Higher age, BMI and prescription of non-steroidal anti-inflammatory drugs contributed most to increased OA risk prediction ahead of diagnosis. We identified 14 subgroups of OA risk profiles. These subgroups were validated in an independent set of patients evaluating the 11-year OA risk, with 88% of patients being uniquely assigned to one of the 14 subgroups. Individual OA risk profiles were characterised by personalised biomarkers. Omics integration demonstrated the predictive importance of key OA genes and pathways (e.g., GDF5 and TGF-β signalling) and OA-specific biomarkers (e.g., CRTAC1 and COL9A1). In summary, this work identifies opportunities for personalised OA prevention and insights into its underlying pathogenesis.
The correct mapping of the proteome is an important step towards advancing our understanding of biological systems and cellular mechanisms. Methods that provide better mappings can fuel important processes such as drug discovery and disease understanding. Currently, true determination of translation initiation sites is primarily achieved by in vivo experiments. Here, we propose TIS Transformer, a deep learning model for the determination of translation start sites solely utilizing the information embedded in the transcript nucleotide sequence. The method is built upon deep learning techniques first designed for natural language processing. We prove this approach to be best suited for learning the semantics of translation, outperforming previous approaches by a large margin. We demonstrate that limitations in the model performance are primarily due to the presence of low-quality annotations against which the model is evaluated against. Advantages of the method are its ability to detect key features of the translation process and multiple coding sequences on a transcript. These include micropeptides encoded by short Open Reading Frames, either alongside a canonical coding sequence or within long non-coding RNAs. To demonstrate the use of our methods, we applied TIS Transformer to remap the full human proteome.
A bstract Ribosome profiling is a deep sequencing technique used to chart translation by means of mRNA ribosome occupancy. It has been instrumental in the detection of non-canonical coding sequences. Because of the complex nature of next-generation sequencing data, existing solutions that seek to identify translated open reading frames from the data are still not perfect. We propose RIBO-former, a new approach featuring several innovations for the de novo annotation of translated coding sequences. RIBO-former is built using recent transformer models that have achieved considerable advancements in the field of natural language processing. The presented deep learning approach allows to omit several pre-processing steps as features are automatically extracted from the data. We discuss various steps that improve the detection of coding sequences and show that read length information of all mapped reads can be leveraged to improve the predictive performance of the tool. Our results show RIBO-former to outperform previous methodologies. Additionally, through our study we find support for the existence of translated non-canonical ORFs, present along existing coding sequences or on long non-coding RNAs. Furthermore, several polycistronic mRNAs with multiple translated coding regions were detected.