During the SARS-CoV-2 pandemic, wastewater-based genomic surveillance (WWGS) emerged as an efficient viral surveillance tool that takes into account asymptomatic cases and can identify known and novel mutations and offers the opportunity to assign known virus lineages based on the detected mutations profiles. WWGS can also hint towards novel or cryptic lineages, but it is difficult to clearly identify and define novel lineages from wastewater (WW) alone. While WWGS has significant advantages in monitoring SARS-CoV-2 viral spread, technical challenges remain, including poor sequencing coverage and quality due to viral RNA degradation. As a result, the viral RNAs in wastewater have low concentrations and are often fragmented, making sequencing difficult. WWGS analysis requires advanced computational tools that are yet to be developed and benchmarked. The existing bioinformatics tools used to analyze wastewater sequencing data are often based on previously developed methods for quantifying the expression of transcripts or viral diversity. Those methods were not developed for wastewater sequencing data specifically, and are not optimized to address unique challenges associated with wastewater. While specialized tools for analysis of wastewater sequencing data have also been developed recently, it remains to be seen how they will perform given the ongoing evolution of SARS-CoV-2 and the decline in testing and patient-based genomic surveillance. Here, we discuss opportunities and challenges associated with WWGS, including sample preparation, sequencing technology, and bioinformatics methods.
Cardiomyopathies comprise a heterogeneous group of myocardial disorders characterized by intrinsic structural and functional abnormalities that are not explained by secondary cardiovascular or systemic conditions. Although genetically determined cardiomyopathies have traditionally been interpreted within a Mendelian framework, this paradigm does not fully account for the marked variability in penetrance, expressivity, and clinical outcomes observed in affected individuals. Increasing evidence indicates that disease manifestation arises from a complex interplay between rare pathogenic variants, common genetic variation, epigenetic regulation, environmental factors, and stochastic molecular processes. This review focuses on hypertrophic and dilated cardiomyopathies, the most prevalent and extensively studied forms, and critically examines how epigenetic mechanisms, genetic modifiers, and molecular noise challenge classical pathophysiology concepts. We discuss how these factors contribute to phenotypic heterogeneity and influence disease severity, progression, and therapeutic response. Recognition of this multilayered genetic architecture has important clinical implications, supporting more refined risk stratification, improved genetic counseling, and the development of personalized and potentially variant-agnostic therapeutic strategies.
RNA viruses represent an integral component of human-associated environments and human health. However, the ecology of environmental RNA viruses remains largely unexplored. Here, we analyzed 2922 metatranscriptomic samples collected from urban and surrounding environments-including human-dense settings (e.g., transit hubs, hospitals, banks), alongside peri-urban settings - across 102 cities in 31 countries and constructed the Urban & Peri-urban RNA Virus Atlas (UPVAtlas), comprising 54,945 RNA viruses, 77% of which had not been previously observed. Phylogenetic reconstruction based on RNA-dependent RNA polymerases from UPVAtlas greatly expanded the evolutionary diversity of RNA viruses, leading to the identification of two potential candidate phyla, one candidate class, and several unclassified clades. Host association analyses further revealed the ecological complexity of environmental RNA viruses, with the diversity of vertebrate-related and ESKAPE pathogen-related viruses underscoring the importance of continued monitoring of urban environments for tracking RNA viral prevalence and dynamics, with direct relevance to future public health.
Disparities in formal bioinformatics training exacerbate the global skills gap, impeding the democratized application of advanced genomic technologies. To bridge this divide, we introduce a scalable, hybrid training framework designed to rapidly accelerate regional bioinformatics capacity. We exemplify this approach through the Eastern European Bioinformatics and Genomics (EEBG) workshop series — a cross-disciplinary initiative that pairs international faculty with local institutions to deliver modular, hands-on curricula. Functioning as a structured knowledge-transfer pipeline, the series has catalyzed a sustainable educational ecosystem, evidenced by the establishment of multiple independent summer schools across the region. The assessment of the 2025 EEBG workshop in Kraków, Poland, validates the model's viability; participant metrics confirm high efficacy in skill acquisition (mean satisfaction: 4.4/5.0) and community building. Crucially, the hybrid delivery mode dismantled geographic barriers, serving as a vital mechanism for maintaining scientific continuity for researchers facing displacement and crisis. Synthesizing these outcomes, we define the core features of a replicable blueprint for scientific readiness in resource-constrained environments. We conclude by presenting a strategic roadmap — organized around infrastructure standardization, governance sustainability, and geographical expansion — for adapting this regional proof-of-concept into a global export-ready model, offering a critical path toward ensuring universal access to genomic innovation. ### Competing Interest Statement The authors have declared no competing interest. Ministry of Research, Innovation and Digitization, under the Romanias National Recovery and Resilience Plan Funded by EU NextGenerationEU program, number 760286/27.03.2024, code 167/31.07.2023, within Pillar III, Component C9, Investment 8 BMFTR (Federal Ministry of Research, Technology and Space) in Germany, GENOMOLD project (01DK26005) that aims at strengthening integration of Eastern Partnership countries into the European Research Area (Bridge2ERA-EaP) Swedish Ministry for Foreign Affairs at the Kyiv School of Economics Ministry of Education and Research, CCCDI UEFISCDI, project number PN-IV-PCB-RO-MD-20240555, within PNCDI IV
To ensure that the benefits of biomedical research are shared globally, we must overcome the barriers preventing the establishment of national genomic projects in underrepresented regions by navigating the ethical, legal and financial complexities of launching these initiatives. We identify seven major challenges — ranging from policy awareness to data privacy regulations — and conclude that successful implementation requires strategic local capacity building and a commitment to open data standards that ensure genomic discoveries are accessible to all.
Large-scale biological data repositories, particularly when deployed as single, non-mirrored instances or governed within narrow contexts, face structural vulnerabilities, from cyberattacks to funding disruptions. We propose a hybrid framework that integrates federated and decentralized models to ensure the resilience, sustainability and FAIR/CARE stewardship of scientific data as a global public good.
Metagenomics has revolutionized our understanding of microbial communities, offering unprecedented insights into their genetic and functional diversity across Earth’s diverse ecosystems. Beyond their roles as environmental constituents, microbiomes act as symbionts, profoundly influencing the health and function of their host organisms. Given the inherent complexity of these communities and the diverse environments where they reside, the components of a metagenomics study must be carefully tailored to yield accurate results that are representative of the populations of interest. This Primer examines the methodological advancements and current practices that have shaped the field, from initial stages of sample collection and DNA extraction to the advanced bioinformatics tools employed for data analysis, with a particular focus on the profound impact of next-generation sequencing on the scale and accuracy of metagenomics studies. We critically assess the challenges and limitations inherent in metagenomics experimentation, available technologies and computational analysis methods. Beyond technical methodologies, we explore the application of metagenomics across various domains, including human health, agriculture and environmental monitoring. Looking ahead, we advocate for the development of more robust computational frameworks and enhanced interdisciplinary collaborations. This Primer serves as a comprehensive guide for advancing the precision and applicability of metagenomic studies, positioning them to address the complexities of microbial ecology and their broader implications for human health and environmental sustainability. Metagenomics describes the analysis of genetic information in a microbial community to provide taxonomic or functional information on the constituent microorganisms. This Primer describes suitable sample types, sampling handling and processing workflows for metagenomics, and gives a detailed discussion of the various analysis techniques to generate meaningful information from metagenomic data.
Accurately characterizing the structure and variability of microbial pangenomes is essential for understanding genome evolution, adaptation, and functional diversity. Traditional descriptive measures such as genomic fluidity and the openness coefficient (α) have provided valuable insights into gene content diversity and pangenome expansion dynamics. Still, they often lose resolution in transitional genome states where conserved and accessory genes coexist. Here, we introduce the pangenome variability index (PVI), a frequency-aware sampling-independent measure that captures gene presence-absence variability across microbial genomes. Unlike the openness coefficient (α) or genomic fluidity, PVI exhibits a unimodal response across the core genome continuum and peaks in intermediate regimes, reflecting maximal compositional heterogeneity. We validate PVI across simulated pangenomes ranging from fully open to fully closed states and show that it captures dimensions of gene content structure orthogonal to conventional measures. Unlike classical measures that saturate in transitional regimes, PVI retains discriminative power and offers a robust, interpretable summary of genomic variability. We recommend integrating PVI into pangenome analysis pipelines as a complementary measure to guide comparative analyses, especially in studies targeting transitional genome architectures and genome evolution in dynamic environments. Future directions include extending PVI to strain-resolved metagenomics, functional annotation layers, and longitudinal analyses of microbial communities.
RNA sequencing (RNA-seq) has emerged as an exemplary technology in biology and clinical applications, offering a crucial complement to other transcriptomic profiling protocols due to its high sensitivity, precision, and accuracy in characterizing transcriptomes. However, the rapid proliferation of RNA-seq tools necessitates the adoption of robust software development practices. Such development underscores the critical need to examine how RNA-seq tools are developed, maintained, and distributed; and whether the data they generate is reproducible as all of these factors are essential for ensuring software reliability, transparency, and trust in scientific findings. We conducted a comprehensive assessment of 434 RNA-seq tools developed between 2008 and 2024, categorizing them based on the type of analysis they perform. Our evaluation encompassed their software development and distribution methodologies, as well as the attributes contributing to their widespread adoption and dependability within the biomedical community, which were quantified by factors such as package manager availability, containerization, multithreading support, documentation quality, and inclusion of example datasets. Our findings establish the first documented positive association between rigorous software development practices and their adoption of published RNA-seq tools as measured by citations (Mann-Whitney U test, p-value = 4.9e-26). By identifying key characteristics of widely adopted software, our findings guide developing robust and user-friendly RNA-seq tools, thereby reinforcing the call for rigorous community-wide standards.
Wastewater-based epidemiology (WBE) has proven to be a valuable tool for monitoring the evolution and spread of global health threats, from pathogens to antimicrobial resistances. Throughout the COVID-19 pandemic, multiple wastewater surveillance programmes have advanced statistical and machine learning methods for detecting pathogens from wastewater sequencing data and correlating measured targets with the represented population to infer meaningful conclusions for public health. Integrating contextual data can account for measurement uncertainties across the WBE workflow that affect the reliability of analyses. However, the broader availability and harmonization of data are major obstacles to method development. Here we review the benefits and limitations of wastewater-related data streams, highlighting the potential of machine learning to leverage these streams for normalization and other WBE applications. We emphasize the relevance of developing global frameworks for integrating WBE with other health surveillance systems and discuss next steps to address current and foreseeable challenges for robust and interpretable machine learning-enhanced WBE. Wastewater-based epidemiology has already proven to be a powerful tool to monitor the spread of a number of diseases. This Perspective discusses the integration with machine learning, highlighting its potential in a number of wastewater-based epidemiology applications.
The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.
BACKGROUND:Recent advances in high-throughput sequencing technologies have enabled the collection and sharing of a massive amount of omics data, along with its associated metadata-descriptive information that contextualizes the data, including phenotypic traits and experimental design. Enhancing metadata availability is critical to ensure data reusability and reproducibility and to facilitate novel biomedical discoveries through effective data reuse. Yet, incomplete metadata accompanying public omics data may hinder reproducibility and reusability and limit secondary analyses. RESULTS:Our study assesses the completeness of metadata in over 253 scientific studies, covering more than 164,000 samples from both human and non-human mammalian studies. We find that over 25% of critical metadata are omitted, with only 74.8% of relevant phenotypes available in publications or public repositories. Notably, public repositories alone contain 62% of the phenotypes, surpassing the textual content of publications by 3.5%. Only 11.5% of studies completely shared all phenotypes, while 37.9% shared less than 40% of the phenotypes. Additionally, studies with non-human samples are more likely to include complete metadata compared to human studies. Similar trends are observed in an extended dataset comprising 61,000 studies and 2.1 million samples from the Gene Expression Omnibus (GEO) data repository. CONCLUSIONS:These findings highlight significant gaps in metadata sharing, underscoring the need for standardized practices to improve metadata availability. Enhanced metadata reporting would foster data reusability, support better-informed decision-making, and promote reproducible research across the biomedical field.
Metadata, or "data about data," is essential for organizing, understanding, and managing large-scale omics datasets. It enhances data discovery, integration, and interpretation, enabling reproducibility, reusability, and secondary analysis. However, metadata sharing remains hindered by perceptual and technical barriers, including the lack of uniform standards, privacy concerns, study design limitations, insufficient incentives, inadequate infrastructure, and a shortage of trained personnel. These challenges compromise data reliability and obstruct integrative meta-analyses. Addressing these issues requires standardization, education, stronger roles for journals and funding agencies, and improved incentives and infrastructure. Looking ahead, emerging technologies such as artificial intelligence and machine learning may offer promising solutions to automate metadata processes, increasing accuracy and scalability. Fostering a collaborative culture of metadata sharing will maximize the value of omics data, accelerating innovation and scientific discovery.
In light of the continuous transmission and evolution of SARS-CoV-2 coupled with a significant decline in clinical testing, there is a pressing need for scalable, cost-effective, long-term, passive surveillance tools to effectively monitor viral variants circulating in the population. Wastewater genomic surveillance of SARS-CoV-2 has arrived as an alternative to clinical genomic surveillance, allowing to continuously monitor the prevalence of viral lineages in communities of various size at a fraction of the time, cost, and logistic effort and serving as an early warning system for emerging variants, critical for developed communities and especially for underserved ones. Importantly, lineage prevalence estimates obtained with this approach aren't distorted by biases related to clinical testing accessibility and participation. However, the relative performance of bioinformatics methods used to measure relative lineage abundances from wastewater sequencing data is unknown, preventing both the research community and public health authorities from making informed decisions regarding computational tool selection. Here, we perform comprehensive benchmarking of 18 bioinformatics methods for estimating the relative abundance of SARS-CoV-2 (sub)lineages in wastewater by using data from 36 in vitro mixtures of synthetic lineage and sublineage genomes. In addition, we use simulated data from 78 mixtures of lineages and sublineages co-occurring in the clinical setting with proportions mirroring their prevalence ratios observed in real data. Importantly, we investigate how the accuracy of the evaluated methods is impacted by the sequencing technology used, the associated error rate, the read length, read depth, but also by the exposure of the synthetic RNA mixtures to wastewater, with the goal of capturing the effects induced by the wastewater matrix, including RNA fragmentation and degradation.
Data-driven computational analysis is becoming increasingly important in biomedical research, as the amount of data being generated continues to grow. However, the lack of practices of sharing research outputs, such as data, source code and methods, affects transparency and reproducibility of studies, which are critical to the advancement of science. Many published studies are not reproducible due to insufficient documentation, code, and data being shared. We conducted a comprehensive analysis of 453 manuscripts published between 2016–2021 and found that 50.1% of them fail to share the analytical code. Even among those that did disclose their code, a vast majority failed to offer additional research outputs, such as data. Furthermore, only one in ten articles organized their code in a structured and reproducible manner. We discovered a significant association between the presence of code availability statements and increased code availability. Additionally, a greater proportion of studies conducting secondary analyses were inclined to share their code compared to those conducting primary analyses. In light of our findings, we propose raising awareness of code sharing practices and taking immediate steps to enhance code availability to improve reproducibility in biomedical research. By increasing transparency and reproducibility, we can promote scientific rigor, encourage collaboration, and accelerate scientific discoveries. We must prioritize open science practices, including sharing code, data, and other research products, to ensure that biomedical research can be replicated and built upon by others in the scientific community.
During the SARS-CoV-2 pandemic, wastewater-based genomic surveillance (WWGS) emerged as an efficient viral surveillance tool that takes into account asymptomatic cases and can identify known and novel mutations and offers the opportunity to assign known virus lineages based on the detected mutations profiles. WWGS can also hint towards novel or cryptic lineages, but it is difficult to clearly identify and define novel lineages from wastewater (WW) alone. While WWGS has significant advantages in monitoring SARS-CoV-2 viral spread, technical challenges remain, including poor sequencing coverage and quality due to viral RNA degradation. As a result, the viral RNAs in wastewater have low concentrations and are often fragmented, making sequencing difficult. WWGS analysis requires advanced computational tools that are yet to be developed and benchmarked. The existing bioinformatics tools used to analyze wastewater sequencing data are often based on previously developed methods for quantifying the expression of transcripts or viral diversity. Those methods were not developed for wastewater sequencing data specifically, and are not optimized to address unique challenges associated with wastewater. While specialized tools for analysis of wastewater sequencing data have also been developed recently, it remains to be seen how they will perform given the ongoing evolution of SARS-CoV-2 and the decline in testing and patient-based genomic surveillance. Here, we discuss opportunities and challenges associated with WWGS, including sample preparation, sequencing technology, and bioinformatics methods.
RNA sequencing (RNA-seq) has become an exemplary technology in modern biology and clinical science. Its immense popularity is due in large part to the continuous efforts of the bioinformatics community to develop accurate and scalable computational tools to analyze the enormous amounts of transcriptomic data that it produces. RNA-seq analysis enables genes and their corresponding transcripts to be probed for a variety of purposes, such as detecting novel exons or whole transcripts, assessing expression of genes and alternative transcripts, and studying alternative splicing structure. It can be a challenge, however, to obtain meaningful biological signals from raw RNA-seq data because of the enormous scale of the data as well as the inherent limitations of different sequencing technologies, such as amplification bias or biases of library preparation. The need to overcome these technical challenges has pushed the rapid development of novel computational tools, which have evolved and diversified in accordance with technological advancements, leading to the current myriad of RNA-seq tools. These tools, combined with the diverse computational skill sets of biomedical researchers, help to unlock the full potential of RNA-seq. The purpose of this review is to explain basic concepts in the computational analysis of RNA-seq data and define discipline-specific jargon.
In light of the continuous transmission and evolution of SARS-CoV-2 coupled with a significant decline in clinical testing, there is a pressing need for scalable, cost-effective, long-term, passive surveillance tools to effectively monitor viral variants circulating in the population. Wastewater genomic surveillance of SARS-CoV-2 has arrived as an alternative to clinical genomic surveillance, allowing to continuously monitor the prevalence of viral lineages in communities of various size at a fraction of the time, cost, and logistic effort and serving as an early warning system for emerging variants, critical for developed communities and especially for underserved ones. Importantly, lineage prevalence estimates obtained with this approach aren't distorted by biases related to clinical testing accessibility and participation. However, the relative performance of bioinformatics methods used to measure relative lineage abundances from wastewater sequencing data is unknown, preventing both the research community and public health authorities from making informed decisions regarding computational tool selection. Here, we perform comprehensive benchmarking of 18 bioinformatics methods for estimating the relative abundance of SARS-CoV-2 (sub)lineages in wastewater by using data from 36 in vitro mixtures of synthetic lineage and sublineage genomes. In addition, we use simulated data from 78 mixtures of lineages and sublineages co-occurring in the clinical setting with proportions mirroring their prevalence ratios observed in real data. Importantly, we investigate how the accuracy of the evaluated methods is impacted by the sequencing technology used, the associated error rate, the read length, read depth, but also by the exposure of the synthetic RNA mixtures to wastewater, with the goal of capturing the effects induced by the wastewater matrix, including RNA fragmentation and degradation.