The Spanish National Bioinformatics Institute (INB), founded in 2003 as a distributed network, is the ELIXIR Node in Spain and has two objectives: 1) deepen its involvement and leadership within ELIXIR and broaden the resources provided as part of ELIXIR infrastructure to the Life Sciences community; and 2) increase its impact within the Spanish National Health System . INB/ELIXIR-ES continues to strengthen its technological capabilities in federated data infrastructures, interoperability, and FAIR data management within ELIXIR. The Node is actively involved in several ELIXIR-driven projects and commissioned services (CoS) within the 2024-28 Work Programme. From the ELIXIR perspective, the Service Delivery Plan (SDP) maintains 40 resources offered by 24 groups belonging to 12 institutions. Regarding national activities, the INB/ELIXIR-ES leads the Translational Bioinformatics Network (TransBioNet) and serves as proxy between IMPaCT-Data activities, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine, and European efforts. These activities align with major European data projects such as the Genomic Data Infrastructure (GDI), EUCAIM, and the Federated European Genome-phenome Archive (FEGA), while implementing Global Alliance for Genomics and Health (GA4GH) standards in its technological developments. The Node strongly engages within the 2024-28 Work Programme: TechnologyTier: co-leadership of Data, Tools and Training Platforms, with contributions across all Platforms. During this period, the co-led ELIXIR Beacon Network Infrastructure Service secured funding for this service, strengthening federated data discovery capabilities. ScienceTier : co-leadership of CMR and HDTR Science priority areas; co-leadership of Rare Diseases, FHD, Cancer Data and Biodiversity Communities; and Pathogens Data, and RNA Data Focus Group; leading and participating in several CoS in the HDTR and CMR areas. PeopleTier : co-leadership of the ELEAD2.0 leadership programme, a CoS built on the experiences of Bioinfo4Women, and active role in the PeoplePulse CoS. Active role in the NodeTier NSCS, together with 4 Platforms, 13 Communities, and 8 Focus Groups. Regarding the INB/ELIXIR-ES portfolio, EGA, an ELIXIR CDR, is co-developed and maintained by CRG and EMBL-EBI with BSC’s infrastructure support, with the current focus on its extension through Federated EGA. Canada joined the Federated EGA, marking the first major expansion beyond Europe and reinforcing its global dimension. Additionally, four resources are recognised as ELIXIR RIRs: 3DBIONOTES-API, FAIRtracks, FAIRCookbook and OpenEBench. Various ELIXIR Communities have adopted OpenEBench as their community-driven benchmarking platform. The INB/ELIXIR-ES continued its training activities and organised key meetings within ELIXIR. It gathered its national community in the XV Symposium on Bioinformatics (JBI2025) jointly organised with ELIXIR-PT and INSTRUCT-ES. Other ELIXIR events organised were the ELIXIR 3DBioinfo Community Annual General Meeting with the 3D-SIG Community, and the Biodiversity and Microbiome Community meetings. https://inb-elixir.es https://inb-elixir.es/resources
Abstract Long-read sequencing enables transcriptome-wide isoform discovery. However, it generates substantial technical and structural ambiguity that complicates transcript interpretation. Here, we present SQANTI-browser, a classification-aware visualization framework that converts SQANTI3 outputs into interactive UCSC Genome Browser Track Hubs, preserving full transcript structural metadata. By integrating SQANTI classifications directly within the UCSC ecosystem, SQANTI-browser enables dynamic filtering and evidence-guided curation alongside public resource tracks. Furthermore, its adaptive architecture natively supports non-reference genomes, orthogonal data, and custom metadata fields. Applied to clinical, noisy, and synthetic datasets, SQANTI-browser resolves alignment artifacts and rescues actionable novel isoforms, providing a robust framework for long-read transcriptome curation.
Long-read sequencing (LRS) platforms, such as Oxford Nanopore and Pacific Biosciences, enable comprehensive transcriptome analysis but face challenges such as sequencing errors, sample quality variability, and library preparation biases. Current benchmarking approaches address these issues insufficiently: BUSCO assesses transcriptome completeness using conserved single-copy orthologous genes but can misinterpret alternative splicing as gene duplications, while SIRV spike-ins and ERCCs oversimplify real sample complexity, neglecting RNA degradation and RNA-extraction artifacts, thus inflating performance metrics. Simulation algorithms are limited in their ability to recapitulate the complexity of real samples. To overcome these limitations, we introduce the Transcriptome Universal Single-isoform COntrol (TUSCO) benchmarking framework, centered on a curated TUSCO gene set of genes lacking alternative isoforms that can be confidently treated as an internal ground truth. The TUSCO evaluation quantifies precision by identifying reconstructed transcripts that deviate from reference annotations and quantifies sensitivity by verifying detection completeness in human and mouse samples. Masking TUSCO gene set transcripts and replacing them with modified splice variants in the annotation creates a TUSCO-novel challenge that assesses reconstruction of the true, now-unannotated isoforms. Our validation demonstrates that TUSCO metrics provide accurate and reliable benchmarking without external controls, significantly improving quality control standards for transcriptome reconstruction using LRS.
Colorectal cancer (CRC) arises through distinct molecular and histological routes that may be shaped by interactions between the mucosa-associated microbiota and host epigenetic regulation. We analyzed paired tumor and adjacent non-tumor colorectal mucosa from CRC patients stratified by histological subtype (conventional vs serrated-pathway groups). Microbiota composition was profiled by Illumina sequencing, and host DNA methylation was assessed using genome-wide CpG arrays with targeted validation. Alpha diversity showed modest tumor-non-tumor differences, with variation by anatomical location driven by distal tumors. Differential abundance testing identified tumor-associated genera, and linear discriminant modeling highlighted taxa with high discriminatory power, including Bacteroides, Eubacterium, Fusobacterium, and Acinetobacter. Methylome ordination separated tumor from non-tumor samples and revealed prominent contributions of zinc finger (ZNF) loci. Integrative latent-variable modeling (PLS/sparse partial least squares and multi-block sparse partial least squares discriminant analysis) supported coordinated microbiome-methylome variation distinguishing tumor from non-tumor mucosa, prioritized a limited set of bacterial and methylation signals, and suggested partial discrimination between conventional and serrated tumor profiles, particularly in the methylome block. Focusing on Fusobacterium, methylation changes in selected host loci were associated with its abundance, and qPCR-based quantification of Fusobacterium nucleatum correlated with CpG methylation at ZNF788. F. nucleatum levels were also associated with serrated/microsatellite instability-related CRC features and more advanced disease. These findings identify subtype-aware microbiome-methylome signatures in CRC and highlight ZNF788 methylation as a candidate epigenetic correlate of intratumoral F. nucleatum burden.IMPORTANCEColorectal cancer does not develop in a single way, and different tumor types may interact differently with bacteria living on the bowel lining. This study examined both tissue-associated bacteria and DNA methylation, an epigenetic mark that helps regulate genes, in paired tumor and nearby non-tumor tissue from patients with colorectal cancer. By analyzing these two layers together, we identified microbial and host methylation patterns linked to tumor tissue and to serrated-pathway cancers. In particular, Fusobacterium nucleatum was associated with serrated/microsatellite instability-related features, more advanced disease, and methylation of the host gene ZNF788. These findings suggest that combining microbiome and epigenetic information may help explain why colorectal cancer subtypes behave differently and may support future subtype-aware biomarkers.
As humans continue the manned exploration of space, it is critical to understand the impact of this harsh environment on the beneficial microbes that interact with their bodies. Here, we explore whether the onset of symbiotic associations between microbes and animals are impacted during spaceflight. We used the association between the bobtail squid Euprymna scolopes and its beneficial bacterium Vibrio fischeri as an animal model system to examine how spaceflight affects symbiotic interactions at the transcriptomic, metabolomic, and lipidomic levels over time. Our results suggest that in the spaceflight environment, symbiotic microbes can mitigate molecular stress responses of the host animal and accelerate normal developmental pathways, such as neurogenesis and tissue morphogenesis. Overall, this work provides evidence that beneficial microbes can effectively colonize nascent host epithelial tissues in microgravity and play a critical role in shaping the host tissue environment to promote stability of symbiosis during spaceflight.
While isoform identification from long-read RNA sequencing (lrRNA-seq) data has received significant attention, the handling of biologically replicated lrRNA-seq datasets remains less explored. This study defines two strategies for obtaining consensus transcriptomes from multi-sample lrRNA-seq data: Join & Call, where reads from all samples are combined before transcript reconstruction, and Call & Join, where transcript reconstruction is performed on individual samples before combining the resulting annotations. We apply these strategies to mouse brain and kidney tissue datasets, using PacBio and ONT technologies, across six transcript reconstruction tools. Our results indicate that the optimal strategy depends on the tool and research objective. We find that Join & Call is generally preferable for discovering novel isoforms, while Call & Join is often preferable for highly replicated datasets when the discovery of novelty is secondary. Our findings provide a conceptual and practical framework for multi-sample transcript reconstruction, guiding best practices for increasingly large-scale lrRNA-seq studies. When combining long-read RNA sequencing data from multiple samples, transcripts can be identified either before or after merging. This study shows that neither approach is universally optimal and provides guidance based on the software tool and study goals.
While the production of a draft genome has become more accessible due to long-read sequencing, the annotation of these new genomes has not been developed at the same pace. Long-read RNA sequencing offers a promising solution for enhancing gene annotation. In this study, we explore how sequencing platforms, Oxford Nanopore R9.4.1 chemistry or Pacific Biosciences (PacBio) Sequel II CCS, and data processing methods influence evidence-driven genome annotation using long reads. Incorporating PacBio transcripts into our annotation pipeline significantly outperformed traditional methods, such as ab initio predictions and short-read-based annotations. We applied this strategy to a nonmodel species, the Florida manatee, and compared our results to existing short-read-based annotation. At the loci level, both annotations were highly concordant, with 90% agreement. However, at the transcript level, the agreement was only 35%. We identified 4906 novel loci, represented by 5707 isoforms, with 64% of these isoforms matching known sequences in other mammalian species. Overall, our findings underscore the importance of using high-quality curated transcript models in combination with ab initio methods for effective genome annotation.
Common bean (Phaseolus vulgaris), a staple food in Latin America and Africa, serves as a vital source of energy, protein, and essential minerals for millions of people. However, genomics knowledge that breeders could leverage for improvement of this crop is scarce. We have developed and validated a comparative genomics approach to predict conserved transcription factor binding sites (TFBS) in common bean and studied gene regulatory networks. We analyzed promoter regions and identified TFBS for 12,631 bean genes with an average of 6 conserved motifs per gene. Moreover, we discovered a statistically significant relationship between the number of conserved motifs and amount of available experimental evidence of gene regulation. Notably, ERF, MYB, and bHLH transcription factor families dominated conserved motifs, with implications for starch biosynthesis regulation. Furthermore, we provide gene regulatory data as a resource that can be interrogated for the regulatory landscape of any set of genes. Our results underscore the significance of TFBS conservation in legumes and aligns with the notion that core genes often exhibit a more conserved regulatory makeup. The study demonstrates the effectiveness of a comparative genomics approach for addressing genome information gaps in non-model organisms and provides valuable insights into the regulatory networks governing starch biosynthesis genes that can support crop improvement programs.
The Yeast Metabolic Cycle (YMC) is a molecular system that serves as a model to study the internal clock that maintains homeostasis in complex organisms. Traditionally, this ultradian rhythm has been studied in the three phases where mature mRNA transcripts show peak accumulation. However, recent studies have shown that the YMC can be interpreted as a two-phase cycle based on altered redox states, known as the high (HOC) and low oxygen consumption (LOC) phases. The length of the HOC phase is fixed and its frequency is nutrient dependent but the nature of the HOC to LOC transition is poorly defined. Here, we use multivariate statistics to integrate metabolic, chromatin and transcriptional changes across the YMC to study the levels of organization that connect them. Our model reveals that both the HOC-LOC and LOC-HOC phase transitions in the YMC are coordinated by accumulating metabolites, reflecting cellular energetics and redox state. We propose that the cycling behavior of chromatin states, transcription and transcripts is a consequence of accumulating metabolites at phase transitions, which function by modulating protein activity and coordinating biochemical pathways to maintain cellular homeostasis.
The Yeast Metabolic Cycle (YMC) is a molecular system that serves as a model to study internal clocks maintaining homeostasis in complex organisms. Traditionally, this ultradian rhythm has been studied in the three phases where mature mRNA transcripts show accumulation. However, recent studies have shown that the YMC can be interpreted as a two-phase cycle based on altered redox states, known as the high (HOC) and low oxygen consumption (LOC) phases. The length of the HOC phase is fixed and its frequency is nutrient-dependent but the nature of the HOC to LOC transition is poorly defined. Here, we use a multimodal latent-space regression framework to study the cross-talk among metabolic, chromatin and transcriptional changes during the YMC. Our model reveals that both the HOC-LOC and LOC-HOC phase transitions in the YMC are coordinated by accumulating metabolites, reflecting cellular energetics and redox state. We propose that the cycling behavior of chromatin states, transcription and transcript levels is a consequence of metabolites accumulation at phase transitions, which modulate protein activity and biochemical pathways to maintain cellular homeostasis.
SQANTI-reads leverages SQANTI3, a tool for the analysis of the quality of transcript models, to develop a read-level quality control framework for replicated long-read RNA-seq experiments. The number and distribution of reads, as well as the number and distribution of unique junction chains (transcript splicing patterns), in SQANTI3 structural categories are informative of raw data quality. Multisample visualizations of QC metrics are presented by experimental design factors to identify outliers. We introduce new metrics for (1) the identification of potentially under-annotated genes and putative novel transcripts and for (2) quantifying variation in junction donors and acceptors. We applied SQANTI-reads to two different data sets, a Drosophila developmental experiment and a multiplatform data set from the LRGASP project and demonstrate that the tool effectively reveals the impact of read coverage on data quality, and readily identifies strong and weak splicing sites.
Long-read sequencing (LRS) technologies have revolutionized transcriptomic research by enabling the comprehensive sequencing of full-length transcripts. Using these technologies, researchers have reported tens of thousands of novel transcripts, even in well-annotated genomes, while developing new algorithms and experimental approaches to handle the noisy data. The Long-read RNA-seq Genome Annotation Assessment Project community effort benchmarked LRS methods in transcriptomics and validated many novel, lowly expressed, often times sample-specific transcripts identified by long reads. These molecules represent deviations of the major transcriptional program that were overlooked by short-read sequencing methods but are now captured by the full-length, single-molecule approach. This Perspective discusses the challenges and opportunities associated with LRS' capacity to unravel this fraction of the transcriptome, in terms of both transcriptome biology and genome annotation. For transcriptome biology, we need to develop novel experimental and computational methods to effectively differentiate technology errors from rare but real molecules. For genome annotation, we must agree on the strategy to capture molecular variability while still defining reference annotations that are useful for the genomics community.
Long-read sequencing (LRS) platforms, such as Oxford Nanopore and Pacific Biosciences, enable comprehensive transcriptome analysis but face challenges such as sequencing errors, sample quality variability, and library preparation biases. Current benchmarking approaches address these issues insufficiently: BUSCO assesses transcriptome completeness using conserved single-copy orthologs but can misinterpret alternative splicing as gene duplications, while spike-ins (SIRVs, ERCCs) oversimplify real- sample complexity, neglecting RNA degradation and RNA extraction artifacts, thus inflating performance metrics. Simulation algorithms are limited to recapitulate this complexity. To overcome these limitations, we introduce the Transcriptome Universal Single-isoform Control (TUSCO), a curated internal reference set of genes lacking alternative isoforms. TUSCO evaluates precision by identifying transcripts deviating from reference annotations and assesses sensitivity by verifying detection completeness in human and mouse samples. Masking TUSCO transcripts—and optionally inserting decoy splice variants—creates a ‘novel- isoform’ challenge that assesses recovery of the true, now-unannotated isoforms. Our validation demonstrates that TUSCO provides accurate and reliable benchmarking without external controls, significantly improving quality control standards for transcriptome reconstruction using LRS. ### Competing Interest Statement AC has received in-kind funding from Pacific Biosciences AC and TL are partners with Oxford Nanopore in MSCA-DN LongTREC project. Maria Curie-Skłodowska Actions, GA 101072892 Spanish Ministry of Science, PID2023-152976NB-I00 Spanish Ministry of Universities, FPU21/01597
Transcriptome sequencing revolutionized the analysis of gene expression, providing an unbiased approach to gene detection and quantification that enabled the discovery of novel isoforms, alternative splicing events and fusion transcripts. However, although short-read sequencing technologies have surpassed the limited dynamic range of previous technologies such as microarrays, they have limitations, for example, in resolving full-length transcripts and complex isoforms. Over the past 5 years, long-read sequencing technologies have matured considerably, with improvements in instrumentation and analytical methods, enabling their application to RNA sequencing (RNA-seq). Benchmarking studies are beginning to identify the strengths and limitations of long-read RNA-seq, although there remains a need for comprehensive resources to guide newcomers through the intricacies of this approach. In this Review, we provide a comprehensive overview of the long-read RNA-seq workflow, from library preparation and sequencing challenges to core data processing, downstream analyses and emerging developments. We present an extensive inventory of experimental and analytical methods and discuss current challenges and prospects. Advances in long-read sequencing are driving the implementation of these technologies for transcriptome profiling. The authors provide a comprehensive guide to long-read RNA sequencing, including experimental and computational tools, current applications, challenges and opportunities.
As multi-omics sequencing technologies advance, the need for simulation tools capable of generating realistic and diverse (bulk and single-cell) multi-omics datasets for method testing and benchmarking becomes increasingly important. We present MOSim, an R package that simulates both bulk (via mosim function) and single-cell (via sc_mosim function) multi-omics data. The mosim function generates bulk transcriptomics data (RNA-seq) and additional regulatory omics layers (ATAC-seq, miRNA-seq, ChIP-seq, Methyl-seq, and transcription factors), while sc_mosim simulates single-cell transcriptomics data (scRNA-seq) with scATAC-seq and transcription factors as regulatory layers. The tool supports various experimental designs, including simulation of gene co-expression patterns, biological replicates, and differential expression between conditions. MOSim enables users to generate quantification matrices for each simulated omics data type, capturing the heterogeneity and complexity of bulk and single-cell multi-omics datasets. Furthermore, MOSim provides differentially abundant features within each omics layer and elucidates the active regulatory relationships between regulatory omics and gene expression data at both bulk and single-cell levels. By leveraging MOSim, researchers will be able to generate realistic and customizable bulk and single-cell multi-omics datasets to benchmark and validate analytical methods specifically designed for the integrative analysis of diverse regulatory omics data.
Long-read RNA sequencing (lrRNA-seq) has revolutionized transcriptomics facilitating the study of alternative splicing and resulting in identification of thousands of novel transcripts. While isoform identification has received significant attention, the handling of biologically replicated lrRNA-seq datasets remains less explored. However, how multiple samples are combined in a lrRNA-seq study may strongly impact transcript identification. This study defines and evaluates two strategies for obtaining consensus transcriptomes from multi-sample lrRNA-seq data: "Join & Call", where reads from all samples are combined before transcript identification, and "Call & Join", where transcript identification is performed on individual samples before combining the resulting annotations. We applied these strategies to a highly replicated dataset of mouse brain and kidney tissues, using both PacBio and ONT technologies, across six widely used transcript reconstruction tools. Our results indicate that the optimal strategy depends on the chosen computational tool and research objective. We found that Join & Call is generally more suitable for discovering rarely occurring, novel isoforms, as pooling evidence increases confidence in calling lowly-expressed transcripts. Conversely, Call & Join is computationally more efficient and often preferable for highly replicated datasets when the investigation of rare novel transcripts is not the primary objective. Our findings provide a conceptual and practical framework for multi-sample transcriptome reconstruction, guiding best practices in the context of increasingly large-scale lrRNA-seq studies.
The advent of single-molecule, long-read sequencing (LRS) technologies by Oxford Nanopore Technologies and Pacific Biosciences has revolutionized genomics, transcriptomics and, more recently, epigenomics research. These technologies offer distinct advantages, including the direct detection of methylated DNA and simultaneous assessment of DNA sequences spanning multiple kilobases along with their modifications at the single-molecule level. This has enabled the development of new assays for analyzing chromatin states and made it possible to integrate data for DNA methylation, chromatin accessibility, transcription factor binding and histone modifications, thereby facilitating comprehensive epigenomic profiling. Owing to recent advancements, alternative, nascent and translating transcripts can be detected using LRS approaches. This Review discusses LRS-based experimental and computational strategies for characterizing chromatin states and highlights their advantages over short-read sequencing methods. Furthermore, we demonstrate how various long-read methods can be integrated to design multi-omics studies to investigate the relationship between chromatin states and transcriptional dynamics. Long-read sequencing technologies have revolutionized genomics and transcriptomics, and more recently enabled comprehensive epigenomic profiling. These advances now also allow investigation of the relationship between chromatin states and transcriptional dynamics.