Expression Atlas (https://www.ebi.ac.uk/gxa/home) is EMBL-EBI's comprehensive knowledgebase for gene and protein expression across tissues, cell types, conditions, and multiple species. Since our last update, Expression Atlas has expanded substantially in both content and functionality, now comprising >4500 studies from 67 species, with increased proteomics coverage and updated Genotype-Tissue Expression (GTEx) tissue profiles. The resource also includes hundreds of single-cell RNA-seq experiments spanning 21 species, among them externally analysed community datasets such as Tabula Sapiens and GTEx single-nucleus profiles, allowing exploration of curated atlases while maintaining their original analytical framework. Key methodological advances include a new marker gene analysis module for bulk baseline experiments, alongside workflow updates that improve reproducibility. Expression Atlas data are integrated into EMBL-EBI resources such as Ensembl, UniProt, and Europe PMC and disseminated through collaboration with model organism communities such as FlyBase and Gramene. The resource also supports translational research through the European Diagnostic Transcriptomic Library and integration with the Open Targets platform. Future directions include modernizing analysis pipelines, enhancing programmatic access, and delivering AI-ready data formats, strengthening Expression Atlas as a findable, accessible, interoperable, and reusable (FAIR) community-driven resource for both fundamental and translational discovery.
MOTIVATION:The exponential growth of non-coding RNA research-with over 230 000 papers published since 2000-has created an urgent knowledge management crisis in molecular biology. Despite their crucial regulatory roles, microRNAs (miRNAs) face a significant curation bottleneck, with only 1400 articles manually curated to the Gene Ontology (GO) knowledgebase over a decade. This highlights the critical need for automated systems that can accelerate biocuration while maintaining high-quality standards. RESULTS:We present GOFlowLLM, an automated curation pipeline powered by reasoning-enabled Large Language Models (LLMs) that follows established GO curation flowcharts to extract and structure miRNA-mediated gene silencing data at scale. When evaluated on existing curation, GOFlowLLM selects the correct GO term in 90% of cases, with curators agreeing with 95% of the system's reasoning steps and 90% of the evidence selected. Applied to 6996 previously uncurated articles using the Qwen QwQ-32B model, our system identified 2538 new candidate GO annotations on 1785 articles in just 58 hours-potentially doubling the available miRNA GO curation. Manual review shows curators agreed with the selected term in 87% of cases, the model's reasoning in 92% of cases, and the extracted evidence in 93%. The integration of reasoning traces provides transparent justification for annotations that can be reviewed by human curators, addressing a key challenge in adopting AI for scientific curation. AVAILABILITY AND IMPLEMENTATION:GOFlowLLM is implemented as an automated pipeline that follows expert-designed reasoning frameworks to maintain curation quality. The system is available on GitHub: https://github.com/RNAcentral/GO_Flow_LLM.
Members of the genus Pestivirus (family Flaviviridae) comprise economically important livestock pathogens like classical swine fever virus (CSFV) and bovine viral diarrhea virus (BVDV). Research over recent years revealed 11 recognized and eight proposed species. The single-stranded, positive-sense RNA genome encodes one large polyprotein that is processed by viral and cellular proteases into 12 mature proteins. In addition to its protein-coding function, the RNA genome contains secondary structures critical for various stages of the viral life cycle. Some of these structures, including the internal ribosome entry site (IRES) and a 3 ' stem-loop, essential for genome replication, have been studied in individual pestiviruses. Here, we present the first genome-wide multiple sequence alignment comprising all known pestivirus species (accepted and tentative) and a comprehensive analysis of phylogenetically conserved RNA secondary structures across the genus. Well-characterized elements, such as a 5 ' stem-loop, the IRES, and the 3 ' stem-loop SL I, were conserved in all pestiviruses, whereas additional 3 ' untranslated region structures were conserved only in subsets of species. We identified 29 novel conserved RNA secondary structures within the protein-coding region, with thus far unresolved functional importance. A miR-17 binding site, previously described in species A, B, and C, was detected in ten additional species but absent in species K, S, Q, and R. We identified a putative long-distance RNA interaction between the IRES and the 3 ' end of the genome. Together, these findings and the comprehensive MSA of all 19 pestivirus species provide a valuable resource for future research and diagnostic applications.
Rfam is a comprehensive database of non-coding RNA (ncRNA) families providing curated sequence alignments, consensus secondary structures, and covariance models for thousands of RNA families. The database is essential for identifying structured non-coding RNAs in newly sequenced genomes and understanding RNA structure-function relationships. Here we present computational protocols for automated ncRNA annotation of viral genomes, and for programmatic interaction with Rfam through its RESTful API. We showcase genome-wide RNA structure visualization from a genome sequence and from a multiple sequence alignment by generating comprehensive 2D structure diagrams using newly developed features in R2DT. We also present practical examples for retrieving family metadata, downloading alignments, accessing secondary structures, and searching user sequences from the Rfam API. These methods enable researchers in virology and RNA biology to integrate Rfam data into custom bioinformatics pipelines, comparative analyses, and machine learning workflows.
Curation of literature in life sciences is a growing challenge. The continued increase in the rate of publication, coupled with the relatively fixed number of curators worldwide, presents a major challenge to developers of biomedical knowledgebases. Very few knowledgebases have resources to scale to the whole relevant literature and all have to prioritize their efforts. In this work, we take a first step to alleviating the lack of curator time in RNA science by generating summaries of literature for noncoding RNAs using large language models (LLMs). We demonstrate that high-quality, factually accurate summaries with accurate references can be automatically generated from the literature using a commercial LLM and a chain of prompts and checks. Manual assessment was carried out for a subset of summaries, with the majority being rated extremely high quality. We apply our tool to a selection of >4600 ncRNAs and make the generated summaries available via the RNAcentral resource. We conclude that automated literature summarization is feasible with the current generation of LLMs, provided that careful prompting and automated checking are applied. Database URL: https://rnacentral.org/
RNA secondary (2D) structure visualization is an essential tool for understanding RNA function. R2DT is a software package designed to visualize RNA 2D structures in consistent, recognizable, and reproducible layouts. The latest release, R2DT 2.0, introduces multiple significant features, including the ability to display position-specific information, such as single nucleotide polymorphisms or SHAPE reactivities. It also offers a new template-free mode allowing visualization of RNAs without pre-existing templates, alongside a constrained folding mode and support for animated visualizations. Users can interactively modify R2DT diagrams, either manually or using natural language prompts, to generate new templates or create publication-quality images. Additionally, R2DT features faster performance, an expanded template library, and a growing collection of compatible tools and utilities. Already integrated into multiple biological databases, R2DT has evolved into a comprehensive platform for RNA 2D visualization, accessible at https://r2dt.bio.
Hepatitis C virus (HCV) is a plus-stranded RNA virus that often chronically infects liver hepatocytes and causes liver cirrhosis and cancer. These viruses replicate their genomes employing error-prone replicases. Thereby, they routinely generate a large 'cloud' of RNA genomes (quasispecies) which-by trial and error-comprehensively explore the sequence space available for functional RNA genomes that maintain the ability for efficient replication and immune escape. In this context, it is important to identify which RNA secondary structures in the sequence space of the HCV genome are conserved, likely due to functional requirements. Here, we provide the first genome-wide multiple sequence alignment (MSA) with the prediction of RNA secondary structures throughout all representative full-length HCV genomes. We selected 57 representative genomes by clustering all complete HCV genomes from the BV-BRC database based on k-mer distributions and dimension reduction and adding RefSeq sequences. We include annotations of previously recognized features for easy comparison to other studies. Our results indicate that mainly the core coding region, the C-terminal NS5A region, and the NS5B region contain secondary structure elements that are conserved beyond coding sequence requirements, indicating functionality on the RNA level. In contrast, the genome regions in between contain less highly conserved structures. The results provide a complete description of all conserved RNA secondary structures and make clear that functionally important RNA secondary structures are present in certain HCV genome regions but are largely absent from other regions. Full-genome alignments of all branches of Hepacivirus C are provided in the supplement.
A-to-I RNA editing is the most common non-transient epitranscriptome modification. It plays several roles in human physiology and has been linked to several disorders. Large-scale deep transcriptome sequencing has fostered the characterization of A-to-I editing at the single nucleotide level and the development of dedicated computational resources. REDIportal is a unique and specialized database collecting similar to 16 million of putative A-to-I editing sites designed to face the current challenges of epitranscriptomics. Its running version has been enriched with sites from the TCGA project (using data from 31 studies). REDIportal provides an accurate, sustainable and accessible tool enriched with interconnections with widespread ELIXIR core resources such as Ensembl, RNAcentral, UniProt and PRIDE. Additionally, REDIportal now includes information regarding RNA editing in putative double-stranded RNAs, relevant for the immune-related roles of editing, as well as an extended catalog of recoding events. Finally, we report a reliability score per site calculated using a deep learning model trained using a huge collection of positive and negative instances. REDIportal is available at http://srv00.recas.ba.infn.it/atlas/.
The Rfam database, a widely-used repository of non-coding RNA (ncRNA) families, has undergone significant updates in release 15.0. This paper introduces major improvements, including the expansion of Rfamseq to 26,106 genomes, a 76% increase, incorporating the latest UniProt reference proteomes and additional viral genomes. Sixty-five RNA families were enhanced using experimentally determined 3D structures, improving the accuracy of consensus secondary structures and annotations. R-scape covariation analysis was used to refine structural predictions in 26 families. Gene Ontology and Sequence Ontology annotations were comprehensively updated, increasing GO term coverage to 75% of families. The release adds 14 new Hepatitis C Virus RNA families and completes microRNA family synchronisation with miRBase, resulting in 1,603 microRNA families. New data types, including FULL alignments, have been implemented. Integration with APICURON for improved curator attribution and multiple website enhancements further improve user experience. These updates significantly expand Rfam's coverage and improve annotation quality, reinforcing its critical role in RNA research, genome annotation, and the development of machine learning models. Rfam is freely available at https://rfam.org.
Motivation: Curation of literature in life sciences is a growing challenge. The continued increase in the rate of publication, coupled with the relatively fixed number of curators worldwide presents a major challenge to developers of biomedical knowledgebases. Very few knowledgebases have resources to scale to the whole relevant literature and all have to prioritise their efforts. Results: In this work, we take a first step to alleviating the lack of curator time in RNA science by generating summaries of literature for non-coding RNAs using large language models (LLMs). We demonstrate that high-quality, factually accurate summaries with accurate references can be automatically generated from the literature using a commercial LLM and a chain of prompts and checks. Manual assessment was carried out for a subset of summaries, with the majority being rated extremely high quality. We also applied the most commonly used automated evaluation approaches, finding that they do not correlate with human assessment. Finally, we apply our tool to a selection of over 4,600 ncRNAs and make the generated summaries available via the RNAcentral resource. We conclude that automated literature summarization is feasible with the current generation of LLMs, provided careful prompting and automated checking are applied. Availability: Code used to produce these summaries can be found here: https://github.com/RNAcentral/litscan-summarization and the dataset of contexts and summaries can be found here: https://huggingface.co/datasets/RNAcentral/litsumm-v1. Summaries are also displayed on the RNA report pages in RNAcentral (https://rnacentral.org/)
AbstractThe protein structure prediction problem has been solved for many types of proteins by AlphaFold. Recently, there has been considerable excitement to build off the success of AlphaFold and predict the 3D structures of RNAs. RNA prediction methods use a variety of techniques, from physics-based to machine learning approaches. We believe that there are challenges preventing the successful development of deep learning-based methods like AlphaFold for RNA in the short term. Broadly speaking, the challenges are the limited number of structures and alignments making data-hungry deep learning methods unlikely to succeed. Additionally, there are several issues with the existing structure and sequence data, as they are often of insufficient quality, highly biased and missing key information. Here, we discuss these challenges in detail and suggest some steps to remedy the situation. We believe that it is possible to create an accurate RNA structure prediction method, but it will require solving several data quality and volume issues, usage of data beyond simple sequence alignments, or the development of new less data-hungry machine learning methods.
Sequence variation in a widespread, recurrent, structured RNA 3D motif, the Sarcin/Ricin (S/R), was studied to address three related questions: First, how do the stabilities of structured RNA 3D motifs, composed of non-Watson-Crick (non-WC) basepairs, compare to WC-paired helices of similar length and sequence? Second, what are the effects on the stabilities of such motifs of isosteric and non-isosteric base substitutions in the non-WC pairs? And third, is there selection for particular base combinations in non-WC basepairs, depending on the temperature regime to which an organism adapts? A survey of large and small subunit rRNAs from organisms adapted to different temperatures revealed the presence of systematic sequence variations at many non-WC paired sites of S/R motifs. UV melting analysis and enzymatic digestion assays of oligonucleotides containing the motif suggest that more stable motifs tend to be more rigid. We further found that the base substitutions at non-Watson-Crick pairing sites can significantly affect the thermodynamic stabilities of S/R motifs and these effects are highly context specific indicating the importance of base-stacking and base-phosphate interactions on motif stability. This study highlights the significance of non-canonical base pairs and their contributions to modulating the stability and flexibility of RNA molecules.
Non-coding RNAs (ncRNA) are essential for all life, and their functions often depend on their secondary (2D) and tertiary structure. Despite the abundance of software for the visualisation of ncRNAs, few automatically generate consistent and recognisable 2D layouts, which makes it challenging for users to construct, compare and analyse structures. Here, we present R2DT, a method for predicting and visualising a wide range of RNA structures in standardised layouts. R2DT is based on a library of 3,647 templates representing the majority of known structured RNAs. R2DT has been applied to ncRNA sequences from the RNAcentral database and produced >13 million diagrams, creating the world's largest RNA 2D structure dataset. The software is amenable to community expansion, and is freely available at https://github.com/rnacentral/R2DT and a web server is found at https://rnacentral.org/r2dt .
Non-coding RNAs (ncRNA) are essential for all life, and the functions of many ncRNAs depend on their secondary (2D) and tertiary (3D) structure. Despite proliferation of 2D visualisation software, there is a lack of methods for automatically generating 2D representations in consistent, reproducible, and recognisable layouts, making them difficult to construct, compare and analyse. Here we present R2DT, a comprehensive method for visualising a wide range of RNA structures in standardised layouts. R2DT is based on a library of 3,632 templates representing the majority of known structured RNAs, from small RNAs to the large subunit ribosomal RNA. R2DT has been applied to ncRNA sequences from the RNAcentral database and produced >13 million diagrams, creating the world’s largest RNA 2D structure dataset. The software is freely available at https://github.com/rnacentral/R2DT and a web server is found at https://rnacentral.org/r2dt .
Non-coding RNAs are essential for all life and carry out a wide range of functions. Information about these molecules is distributed across dozens of specialized resources. RNAcentral is a database of non-coding RNA sequences that provides a unified access point to non-coding RNA annotations from >40 member databases and helps provide insight into the function of these RNAs. This article describes different ways of accessing the data, including searching the website and retrieving the data programmatically over web APIs and a public database. We also demonstrate an example Galaxy workflow for using RNAcentral for RNA-seq differential expression analysis. RNAcentral is available at https://rnacentral.org. © 2020 The Authors. Basic Protocol 1: Viewing RNAcentral sequence reports Basic Protocol 2: Using RNAcentral text search to explore ncRNA sequences Basic Protocol 3: Using RNAcentral sequence search Basic Protocol 4: Using RNAcentral FTP archive Support Protocol 1: Using web APIs for programmatic data access Support Protocol 2: Using public Postgres database to export large datasets Support Protocol 3: Analyze non-coding RNA in RNA-seq datasets using RNAcentral and Galaxy.
RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and >18 million ncRNA sequences from a wide range of organisms and RNA types. RNAcentral now also includes secondary (2D) structure information for >13 million sequences, making RNAcentral the world's largest RNA 2D structure database. The 2D diagrams are displayed using R2DT, a new 2D structure visualization method that uses consistent, reproducible and recognizable layouts for related RNAs. The sequence similarity search has been updated with a faster interface featuring facets for filtering search results by RNA type, organism, source database or any keyword. This sequence search tool is available as a reusable web component, and has been integrated into several RNAcentral member databases, including Rfam, miRBase and snoDB. To allow for a more fine-grained assignment of RNA types and subtypes, all RNAcentral sequences have been annotated with Sequence Ontology terms. The RNAcentral database continues to grow and provide a central data resource for the RNA community. RNAcentral is freely available at https://rnacentral.org.
RNA 3D motifs occupy places in structured RNA molecules that correspond to the hairpin, internal and multi-helix junction "loops" of their secondary structure representations. As many as 40% of the nucleotides of an RNA molecule can belong to these structural elements, which are distinct from the regular double helical regions formed by contiguous AU, GC, and GU Watson-Crick basepairs. With the large number of atomic- or near atomic-resolution 3D structures appearing in a steady stream in the PDB/NDB structure databases, the automated identification, extraction, comparison, clustering and visualization of these structural elements presents an opportunity to enhance RNA science. Three broad applications are: (1) identification of modular, autonomous structural units for RNA nanotechnology, nanobiology and synthetic biology applications; (2) bioinformatic analysis to improve RNA 3D structure prediction from sequence; and (3) creation of searchable databases for exploring the binding specificities, structural flexibility, and dynamics of these RNA elements. In this contribution, we review methods developed for computational extraction of hairpin and internal loop motifs from a non-redundant set of high-quality RNA 3D structures. We provide a statistical summary of the extracted hairpin and internal loop motifs in the most recent version of the RNA 3D Motif Atlas. We also explore the reliability and accuracy of the extraction process by examining its performance in clustering recurrent motifs from homologous ribosomal RNA (rRNA) structures. We conclude with a summary of remaining challenges, especially with regard to extraction of multi-helix junction motifs.
Many non-coding RNAs have been identified and may function by forming 2D and 3D structures. RNA hairpin and internal loops are often represented as unstructured on secondary structure diagrams, but RNA 3D structures show that most such loops are structured by non-Watson-Crick basepairs and base stacking. Moreover, different RNA sequences can form the same RNA 3D motif. JAR3D finds possible 3D geometries for hairpin and internal loops by matching loop sequences to motif groups from the RNA 3D Motif Atlas, by exact sequence match when possible, and by probabilistic scoring and edit distance for novel sequences. The scoring gauges the ability of the sequences to form the same pattern of interactions observed in 3D structures of the motif. The JAR3D webserver at http://rna.bgsu.edu/jar3d/ takes one or many sequences of a single loop as input, or else one or many sequences of longer RNAs with multiple loops. Each sequence is scored against all current motif groups. The output shows the ten best-matching motif groups. Users can align input sequences to each of the motif groups found by JAR3D. JAR3D will be updated with every release of the RNA 3D Motif Atlas, and so its performance is expected to improve over time.
Predicting RNA 3D structure from sequence is a major challenge in biophysics. An important sub-goal is accurately identifying recurrent 3D motifs from RNA internal and hairpin loop sequences extracted from secondary structure (2D) diagrams. We have developed and validated new probabilistic models for 3D motif sequences based on hybrid Stochastic Context-Free Grammars and Markov Random Fields (SCFG/MRF). The SCFG/MRF models are constructed using atomic-resolution RNA 3D structures. To parameterize each model, we use all instances of each motif found in the RNA 3D Motif Atlas and annotations of pairwise nucleotide interactions generated by the FR3D software. Isostericity relations between non-Watson-Crick basepairs are used in scoring sequence variants. SCFG techniques model nested pairs and insertions, while MRF ideas handle crossing interactions and base triples. We use test sets of randomly-generated sequences to set acceptance and rejection thresholds for each motif group and thus control the false positive rate. Validation was carried out by comparing results for four motif groups to RMDetect. The software developed for sequence scoring (JAR3D) is structured to automatically incorporate new motifs as they accumulate in the RNA 3D Motif Atlas when new structures are solved and is available free for download.