At the level of secondary structure, circular RNAs (circRNAs) can be understood in terms of base pairing, base-pair stacking, and entropic loop contribution in the same way as linear RNAs and intermolecular RNA-RNA interactions. The folding problem of circular RNAs can thus be solved by dynamic programming algorithms in essentially the same manner. In this chapter, we review the similarities and differences between circular and linear RNAs with a focus on the software tools provided by the ViennaRNA package. Comparative analysis of RNA structures can also be generalized to circular RNA molecules. However, the task of constructing pairwise and multiple alignments of circular sequences is more difficult than those of their linear counterpart, whence fewer and less convenient software solutions are available. This chapter has also touched upon recent developments such as applications of chemical probing to circular RNAs and prediction of secondary structures on the interaction of circular RNAs with other molecules.
Mitochondrial tRNAs have acquired a diverse portfolio of aberrant structures throughout metazoan evolution. With the availability of more than 12,500 mitogenome sequences, it is essential to compile a comprehensive overview of the pattern changes with regard to mitochondrial tRNA repertoire and structural variations. This, of course, requires reanalysis of the sequence data of more than 250,000 mitochondrial tRNAs with a uniform workflow. Here, we report our results on the complete reannotation of all mitogenomes available in the RefSeq database by September 2022 using mitos2. Based on the individual cases of mitochondrial tRNA variants reported throughout the literature, our data pinpoint the respective hotspots of change, i.e. Acanthocephala (Lophotrochozoa), Nematoda, Acariformes, and Araneae (Arthropoda). Less dramatic deviations of mitochondrial tRNAs from the norm are observed throughout many other clades. Loss of arms in animal mitochondrial tRNA clearly is a phenomenon that occurred independently many times, not limited to a small number of specific clades. The summary data here provide a starting point for systematic investigations into the detailed evolutionary processes of structural reduction and loss of mitochondrial tRNAs as well as a resource for further improvements of annotation workflows for mitochondrial tRNA annotation.
Transfer RNA (tRNA) modifications are essential for the temperature adaptation of thermophilic and psychrophilic organisms as they control the rigidity and flexibility of transcripts. To further understand how specific tRNA modifications are adjusted to maintain functionality in response to temperature fluctuations, we investigated whether tRNA modifications represent an adaptation of bacteria to different growth temperatures (minimal, optimal, and maximal), focusing on closely related psychrophilic (P. halocryophilus and E. sibiricum), mesophilic (B. subtilis), and thermophilic (G. stearothermophilus) Bacillales. Utilizing an RNA sequencing approach combined with chemical pre-treatment of tRNA samples, we systematically profiled dihydrouridine (D), 4-thiouridine (s4U), 7-methyl-guanosine (m7G), and pseudouridine (Ψ) modifications at single-nucleotide resolution. Despite their close relationship, each bacterium exhibited a unique tRNA modification profile. Our findings revealed increased tRNA modifications in the thermophilic bacterium at its optimal growth temperature, particularly showing elevated levels of s4U8 and Ψ55 modifications compared to non-thermophilic bacteria, indicating a temperature-dependent regulation that may contribute to thermotolerance. Furthermore, we observed higher levels of D modifications in psychrophilic and mesophilic bacteria, indicating an adaptive strategy for cold environments by enhancing local flexibility in tRNAs. Our method demonstrated high effectiveness in identifying tRNA modifications compared to an established tool, highlighting its potential for precise tRNA profiling studies.
Glomerular-tubular crosstalk within the kidney has been proposed, but the paracrine signals enabling this remain largely unknown. The cold-shock protein Y-box binding protein 1 (YBX1) is known to regulate inflammation and kidney diseases but its role in podocytes remains undetermined. Therefore, we analyzed mice with podocyte specific Ybx1 deletion (Ybx1ΔPod). Albuminuria was increased in unchallenged Ybx1ΔPod mice, which surprisingly was associated with reduced glomerular, but enhanced tubular damage. Tubular toll-like receptor 4 (TLR4) expression, node-like receptor protein 3 (NLRP3) inflammasome activation and kidney inflammatory cell infiltrates were all increased in Ybx1ΔPod mice. In vitro, extracellular YBX1 inhibited NLRP3 inflammasome activation in tubular cells. Co-immunoprecipitation, immunohistochemical analyses, microscale cell-free thermophoresis assays, and blunting of the YBX1-mediated TLR4-inhibition by a unique YBX1-derived decapeptide suggests a direct interaction of YBX1 and TLR4. Since YBX1 can be secreted upon post-translational acetylation, we hypothesized that YBX1 secreted from podocytes can inhibit TLR4 signaling in tubular cells. Indeed, mice expressing a non-secreted YBX1 variant specifically in podocytes (Ybx1PodK2A mice) phenocopied Ybx1ΔPod mice, demonstrating a tubular-protective effect of YBX1 secreted from podocytes. Lipopolysaccharide-induced tubular injury was aggravated in Ybx1ΔPod and Ybx1PodK2A mice, indicating a pathophysiological relevance of this glomerular-tubular crosstalk. Thus, our data show that YBX1 is physiologically secreted from podocytes, thereby negatively modulating sterile inflammation in the tubular compartment, apparently by binding to and inhibiting tubular TLR4 signaling. Hence, we have uncovered an YBX1-dependent molecular mechanism of glomerular-tubular crosstalk.
The in silico prediction of non-coding and protein-coding genetic loci has received considerable attention in comparative genomics aiming in particular at the identification of properties of nucleotide sequences that are informative of their biological role in the cell. We present here a software framework for the alignment-based training, evaluation and application of machine learning models with user-defined parameters. Instead of focusing on the one-size-fits-all approach of pervasive in silico annotation pipelines, we offer a framework for the structured generation and evaluation of models based on arbitrary features and input data, focusing on stable and explainable results. Furthermore, we showcase the usage of our software package in a full-genome screen of Drosophila melanogaster and evaluate our results against the well-known but much less flexible program RNAz.
Circular RNAs (circRNAs) are a regulatory RNA class. While cancer-driving functions have been identified for single circRNAs, how they modulate gene expression in cancer is not well understood. We investigate circRNA expression in the pediatric malignancy, neuroblastoma, through deep whole-transcriptome sequencing in 104 primary neuroblastomas covering all risk groups. We demonstrate that MYCN amplification, which defines a subset of high-risk cases, causes globally suppressed circRNA biogenesis directly dependent on the DHX9 RNA helicase. We detect similar mechanisms in shaping circRNA expression in the pediatric cancer medulloblastoma implying a general MYCN effect. Comparisons to other cancers identify 25 circRNAs that are specifically upregulated in neuroblastoma, including circARID1A. Transcribed from the ARID1A tumor suppressor gene, circARID1A promotes cell growth and survival, mediated by direct interaction with the KHSRP RNA-binding protein. Our study highlights the importance of MYCN regulating circRNAs in cancer and identifies molecular mechanisms, which explain their contribution to neuroblastoma pathogenesis.
Abstract Summary RNA molecules play crucial roles in various biological processes. They mediate their function mainly by interacting with other RNAs or proteins. At present, information about these interactions is distributed over different resources, often providing the data in simple tab-delimited formats that differ between the databases. There is no standardized data format that can capture the nature of all these different interactions in detail. Availability and implementation Here, we propose the RNA interaction format (RIF) for the detailed representation of RNA–RNA and RNA–Protein interactions and provide reference implementations in C/C++, Python, and JavaScript. RIF is released under licence GNU General Public License version 3 (GNU GPLv3) and is available on https://github.com/RNABioInfo/rna-interaction-format.
(1) Background: Cystic fibrosis (CF) is a disease with well-documented clinical differences between female and male patients. However, this gender gap is very poorly studied at the molecular level. (2) Methods: Expression differences in whole blood transcriptomics between female and male CF patients are analyzed in order to determine the pathways related to sex-biased genes and assess their potential influence on sex-specific effects in CF patients. (3) Results: We identify sex-biased genes in female and male CF patients and provide explanations for some sex-specific differences at the molecular level. (4) Conclusion: Genes in key pathways associated with CF are differentially expressed between sexes, and thus may account for the gender gap in morbidity and mortality in CF.
CRISPR-Cas constitutes an adaptive prokaryotic defence system against invasive nucleic acids like viruses and plasmids. Beyond their role in immunity, CRISPR-Cas systems have been shown to closely interact with components of cellular DNA repair pathways, either by regulating their expression or via direct protein-protein contact and enzymatic activity. The integrase Cas1 is usually involved in the adaptation phase of CRISPR-Cas immunity but an additional role in cellular DNA repair pathways has been proposed previously. Here, we analysed the capacity of an archaeal Cas1 from Haloferax volcanii to act upon DNA damage induced by oxidative stress and found that a deletion of the cas1 gene led to reduced survival rates following stress induction. In addition, our results indicate that Cas1 is directly involved in DNA repair as the enzymatically active site of the protein is crucial for growth under oxidative conditions. Based on biochemical assays, we propose a mechanism by which Cas1 plays a similar function to DNA repair protein Fen1 by cleaving branched intermediate structures. The present study broadens our understanding of the functional link between CRISPR-Cas immunity and DNA repair by demonstrating that Cas1 and Fen1 display equivalent roles during archaeal DNA damage repair.
MONSDA runs HTS data analysis from pre- to postprocessing based on a single configuration file. It wraps Snakemake or Nextflow to run workflow steps involving QC, trimming, mapping, deduplication and differential analysis as well as a set of specific workflows for the generation of genome browser tracks and more. Users profit from many of the advantages of two of the most popular workflow management systems without having to learn all the specifics of Snakemake or Nextflow and can interchange and configure tools according to their specific needs. MONSDA is available via Bioconda, pip and https://github.com/jfallmann/MONSDA
Abstract Self-cleaving ribozymes are catalytic RNAs and can be found in all domains of life. They catalyze a site-specific cleavage that results in a 5′ fragment with a 2′,3′ cyclic phosphate (2′,3′ cP) and a 3′ fragment with a 5′ hydroxyl (5′ OH) end. Recently, several strategies to enrich self-cleaving ribozymes by targeted biochemical methods have been introduced by us and others. Here, we develop an alternative strategy in which 5ʹ OH RNAs are specifically ligated by RtcB ligase, which first guanylates the 3′ phosphate of the adapter and then ligates it directly to RNAs with 5′ OH ends. Our results demonstrate that adapter ligation to highly structured ribozyme fragments is much more efficient using the thermostable RtcB ligase from Pyrococcus horikoshii than the broadly applied Escherichia coli enzyme. Moreover, we investigated DNA, RNA and modified RNA adapters for their suitability in RtcB ligation reactions. We used the optimized RtcB-mediated ligation to produce RNA-seq libraries and captured a spiked 3ʹ twister ribozyme fragment from E. coli total RNA. This RNA-seq-based method is applicable to detect ribozyme fragments as well as other cellular RNAs with 5ʹ OH termini from total RNA.
The problem of segmenting linearly ordered data is frequently encountered in time-series analysis, computational biology, and natural language processing. Segmentations obtained independently from replicate data sets or from the same data with different methods or parameter settings pose the problem of computing an aggregate or consensus segmentation. This Segmentation Aggregation problem amounts to finding a segmentation that minimizes the sum of distances to the input segmentations. It is again a segmentation problem and can be solved by dynamic programming. The aim of this contribution is (1) to gain a better mathematical understanding of the Segmentation Aggregation problem and its solutions and (2) to demonstrate that consensus segmentations have useful applications. Extending previously known results we show that for a large class of distance functions only breakpoints present in at least one input segmentation appear in the consensus segmentation. Furthermore, we derive a bound on the size of consensus segments. As show-case applications, we investigate a yeast transcriptome and show that consensus segments provide a robust means of identifying transcriptomic units. This approach is particularly suited for dense transcriptomes with polycistronic transcripts, operons, or a lack of separation between transcripts. As a second application, we demonstrate that consensus segmentations can be used to robustly identify growth regimes from sets of replicate growth curves.
Advances in genome sequencing over the last years have lead to a fundamental paradigm shift in the field. With steadily decreasing sequencing costs, genome projects are no longer limited by the cost of raw sequencing data, but rather by computational problems associated with genome assembly. There is an urgent demand for more efficient and and more accurate methods is particular with regard to the highly complex and often very large genomes of animals and plants. Most recently, “hybrid” methods that integrate short and long read data have been devised to address this need. LazyB is such a hybrid genome assembler. It has been designed specificially with an emphasis on utilizing low-coverage short and long reads. LazyB starts from a bipartite overlap graph between long reads and restrictively filtered short-read unitigs. This graph is translated into a long-read overlap graph G. Instead of the more conventional approach of removing tips, bubbles, and other local features, LazyB stepwisely extracts subgraphs whose global properties approach a disjoint union of paths. First, a consistently oriented subgraph is extracted, which in a second step is reduced to a directed acyclic graph. In the next step, properties of proper interval graphs are used to extract contigs as maximum weight paths. These path are translated into genomic sequences only in the final step. A prototype implementation of LazyB, entirely written in python, not only yields significantly more accurate assemblies of the yeast and fruit fly genomes compared to state-of-the-art pipelines but also requires much less computational effort. LazyB is new low-cost genome assembler that copes well with large genomes and low coverage. It is based on a novel approach for reducing the overlap graph to a collection of paths, thus opening new avenues for future improvements. The LazyB prototype is available at https://github.com/TGatter/LazyB .
Dictyostelium discoideum is a social amoeba, which on starvation develops from a single-cell state to a multicellular fruiting body. This developmental process is accompanied by massive changes in gene expression, which also affect non-coding RNAs. Here, we investigate how tRNAs as key regulators of the translation process are affected by this transition. To this end, we used LOTTE-seq to sequence the tRNA pool of D. discoideum at different developmental time points and analyzed both tRNA composition and tRNA modification patterns. We developed a workflow for the specific detection of modifications from reverse transcriptase signatures in chemically untreated RNA-seq data at single-nucleotide resolution. It avoids the comparison of treated and untreated RNA-seq data using reverse transcription arrest patterns at nucleotides in the neighborhood of a putative modification site as internal control. We find that nucleotide modification sites in D. discoideum tRNAs largely conform to the modification patterns observed throughout the eukaroytes. However, there are also previously undescribed modification sites. We observe substantial dynamic changes of both expression levels and modification patterns of certain tRNA types during fruiting body development. Beyond the specific application to D. discoideum our results demonstrate that the developmental variability of tRNA expression and modification can be traced efficiently with LOTTE-seq.
Machine learning (ML) methods are often used to identify members of non-coding RNA classes such as microRNAs or snoRNAs. However, ML methods have not been successfully used for homology search tasks. A systematic evaluation of ML in homology search requires large, controlled, and known ground truth test sets, and thus, methods to construct large realistic artificial data sets. Here we describe a method for producing sets of arbitrarily large and diverse snoRNA sequences based on artificial evolution. These are then used to evaluate supervised ML methods (Support Vector Machine, Artificial Neural Network, and Random Forest) for snoRNA detection in a chordate genome. Our results indicate that ML approaches can indeed be competitive also for homology search.
Background The microbiome has emerged as an environmental factor contributing to obesity and type 2 diabetes (T2D). Increasing evidence suggests links between circulating bacterial components (i.e., bacterial DNA), cardiometabolic disease, and blunted response to metabolic interventions. In this aspect, thorough next-generation sequencing-based and contaminant-aware approaches are lacking. To address this, we tested whether bacterial DNA could be amplified in the blood of subjects with obesity and high metabolic risk under strict experimental and analytical control and whether a putative bacterial signature is related to metabolic improvement after bariatric surgery. Methods Subjects undergoing bariatric surgery were recruited into sex- and BMI-matched subgroups with (n = 24) or without T2D (n = 24). Bacterial DNA in the blood was quantified and prokaryotic 16S rRNA gene amplicons were sequenced. A contaminant-aware approach was applied to derive a compositional microbial signature from bacterial sequences in all subjects at baseline and at 3 and 12 months after surgery. We modeled associations between bacterial load and composition with host metabolic and anthropometric markers. We further tested whether compositional shifts were related to weight loss response and T2D remission. Lastly, bacteria were visualized in blood samples using catalyzed reporter deposition (CARD)-fluorescence in situ hybridization (FISH). Results The contaminant-aware blood bacterial signature was associated with metabolic health. Based on bacterial phyla and genera detected in the blood samples, a metabolic syndrome classification index score was derived and shown to robustly classify subjects along their actual clinical group. T2D was characterized by decreased bacterial richness and loss of genera associated with improved metabolic health. Weight loss and metabolic improvement following bariatric surgery were associated with an early and stable increase of these genera in parallel with improvements in key cardiometabolic risk parameters. CARD-FISH allowed the detection of living bacteria in blood samples in obesity. Conclusions We show that the circulating bacterial signature reflects metabolic disease and its improvement after bariatric surgery. Our work provides contaminant-aware evidence for the presence of living bacteria in the blood and suggests a putative crosstalk between components of the blood and metabolism in metabolic health regulation.
Self-cleaving ribozymes are catalytically active RNAs that cleave themselves into a 5'-fragment with a 2',3'-cyclic phosphate and a 3'-fragment with a 5'-hydroxyl. They are widely applied for the construction of synthetic RNA devices and RNA-based therapeutics. However, the targeted discovery of self-cleaving ribozymes remains a major challenge. We developed a transcriptome-wide method, called cyPhyRNA-seq, to screen for ribozyme cleavage fragments in total RNA extract. This approach employs the specific ligation-based capture of ribozyme 5'-fragments using a variant of the Arabidopsis thaliana tRNA ligase we engineered. To capture ribozyme 3'-fragments, they are enriched from total RNA by enzymatic treatments. We optimized and enhanced the individual steps of cyPhyRNA-seq in vitro and in spike-in experiments. Then, we applied cyPhyRNA-seq to total RNA isolated from the bacterium Desulfovibrio vulgaris and detected self-cleavage of the three predicted type II hammerhead ribozymes, whose activity had not been examined to date. cyPhyRNA-seq can be used for the global analysis of active self-cleaving ribozymes with the advantage to capture both ribozyme cleavage fragments from total RNA. Especially in organisms harbouring many self-cleaving RNAs, cyPhyRNA-seq facilitates the investigation of cleavage activity. Moreover, this method has the potential to be used to discover novel self-cleaving ribozymes in different organisms. [Figure: see text].
The role of the inflammation-silencing ribonuclease, MCPIP1 (monocyte chemoattractant protein-induced protein 1), in neoplasia continuous to emerge. The ribonuclease can cleave not only inflammation-related transcripts but also some microRNAs (miRNAs) and viral RNAs. The suppressive effect of the protein has been hitherto suggested in breast cancer, clear cell renal cell carcinoma, osteosarcoma, and neuroblastoma. Our previous results have demonstrated a reduced levels of several oncogenes, as well as inhibited growth of neuroblastoma cells upon MCPIP1 overexpression. Here, we investigate the mechanisms underlying the suppression of MYCN proto-oncogene, bHLH transcription factor (MYCN)-amplified neuroblastoma cells overexpressing the MCPIP1 protein. We showed that the levels of several transcripts involved in cell cycle progression decreased in BE(2)-C and KELLY cells overexpressing MCPIP1 in a ribonucleolytic activity-dependent manner. However, RNA immunoprecipitation indicated that only AURKA mRNA (encoding for Aurora A kinase) interacts with the ribonuclease. Furthermore, the application of a luciferase assay suggested MCPIP1-dependent destabilization of the transcript. Further analyses demonstrated that the entire conserved region of AURKA seems to be indispensable for the interaction with the MCPIP1 protein. Additionally, we examined the effect of the ribonuclease overexpression on the miRNA expression profile in MYCN-amplified neuroblastoma cells. However, no significant alterations were observed. Our data indicate a key role of the binding and cleavage of the AURKA transcript in an MCPIP1-dependent suppressive effect on neuroblastoma cells.
Tunicates are the sister group of vertebrates and thus occupy a key position for investigations into vertebrate innovations as well as into the consequences of the vertebrate-specific genome duplications. Nevertheless, tunicate genomes have not been studied extensively in the past, and comparative studies of tunicate genomes have remained scarce. The carpet sea squirt Didemnum vexillum, commonly known as “sea vomit”, is a colonial tunicate considered an invasive species with substantial ecological and economical risk. We report the assembly of the D. vexillum genome using a hybrid approach that combines 28.5 Gb Illumina and 12.35 Gb of PacBio data. The new hybrid scaffolded assembly has a total size of 517.55 Mb that increases contig length about eightfold compared to previous, Illumina-only assembly. As a consequence of an unusually high genetic diversity of the colonies and the moderate length of the PacBio reads, presumably caused by the unusually acidic milieu of the tunic, the assembly is highly fragmented (L50 = 25,284, N50 = 6539). It is sufficient, however, for comprehensive annotations of both protein-coding genes and non-coding RNAs. Despite its shortcomings, the draft assembly of the “sea vomit” genome provides a valuable resource for comparative tunicate genomics and for the study of the specific properties of colonial ascidians.
Long non-coding RNAs (lncRNAs) are widely recognized as important regulators of gene expression. Their molecular functions range from miRNA sponging to chromatin-associated mechanisms, leading to effects in disease progression and establishing them as diagnostic and therapeutic targets. Still, only a few representatives of this diverse class of RNAs are well studied, while the vast majority is poorly described beyond the existence of their transcripts. In this review we survey common in silico approaches for lncRNA annotation. We focus on the well-established sets of features used for classification and discuss their specific advantages and weaknesses. While the available tools perform very well for the task of distinguishing coding sequence from other RNAs, we find that current methods are not well suited to distinguish lncRNAs or parts thereof from other non-protein-coding input sequences. We conclude that the distinction of lncRNAs from intronic sequences and untranslated regions of coding mRNAs remains a pressing research gap.