Endometriosis is a chronic inflammatory condition with significant diagnostic delays impacting one in ten reproductive age women worldwide. While machine learning (ML) models trained on transcriptomic data show promise for disease prediction, limited generalizability across independent patient cohorts has hindered clinical translation. Foundations models (FMs) pretrained on large-scale transcriptomic data offer promise to learn transferrable, biologically meaningful representations that could support cross-cohort predictions. We assembled a 12-cohort bulk RNA-seq benchmark (334 samples) and developed a computationally efficient pipeline to test whether FMs improve endometriosis classification, an approach not previously applied to this disease. Using AutoXAI4Omics with cohort-aware validation, we compared embeddings derived from five state-of-the-art RNA FMs against TPM baselines. In cross-cohort prediction, FM embeddings significantly improved performance, achieving a weighted F1-score of 0.83 vs. 0.68 for the baseline. To allow gene-level interpretation of FM embedding models, we introduce classified-aligned integrated gradients (CA-IG), an interpretability approach aligning gene-level attributions to the downstream classifier without end-to-end finetuning. CA-IG revealed a conserved set of predictive genes from FM embeddings across cohort-validation regimes, contrasting with unstable baseline explainability, suggesting that FM embeddings prioritized transferable disease-related signal over cohort-specific effects. These genes include novel candidates that converge on biologically plausible pathways for endometriosis.
High-throughput omics data offer unprecedented opportunities to advance our understanding of human health and disease. However, their analysis remains challenging due to high dimensionality, noise, heterogeneity, and sparsity-factors that severely limit both the predictive accuracy and the biological interpretability of conventional machine learning approaches, particularly in patient stratification tasks. Microbiome data exemplify these challenges, as their extreme dimensionality and complex structure often obscure mechanistic insights into disease processes. We present a novel computational framework integrating optimal transport theory with meta-learning for omics data analysis. Our approach models each sample as a probability distribution in a learned latent feature space, enabling capture of distributional properties typically lost in vector-based representations. By employing the Wasserstein distance, we quantify sample dissimilarities that capture subtle but biologically meaningful differences often missed by standard metrics. The latent space is optimized through a composite loss function balancing intra-class cohesion with interclass separation. Within this, we systematically investigate different ground cost metrics to identify optimal distance measures between the features themselves. The meta-learning architecture enhances generalizability across diverse datasets, addressing common sample size limitations in biomedical studies. This framework provides both methodological advances in computational biology and practical tools for improved omics analysis, with applications in precision medicine and systems-level disease understanding. Our method improves patient stratification by enhancing classifier performance, while also offering greater interpretability, robustness to small or imbalanced datasets, and applicability across diverse data modalities. This novel methodology is evaluated on synthetic and clinical datasets and can be used to uncover disease mechanisms and guide precision medicine strategies.
Graph-based machine learning methods are valuable tools for identifying and predicting variation in genetic data. In particular, understanding phenotypic effects at the cellular level is an accelerating area in pharmacogenomics. Insight into how drugs or disease affect bio-networks could aid drug development and precision medicine. This article proposes a novel graph-theoretic approach to infer a co-occurrence network from 16S microbiome data, designed specifically for smallsample datasets. Such datasets pose challenges due to sparsity, compositionality, and complex interactions. The methodology includes steps to enrich and statistically filter the inferred networks. The approach extracts informative, feature-rich, biologically meaningful, and statistically significant networks from limited data. While tailored for small datasets, it is broadly applicable and can be extended to multi-omics integration. The method is tested on data from chickens vaccinated and challenged with Eimeria tenella. Genetic reads are processed, and networks inferred to characterize intestinal ecosystems at three disease progression stages. Analysis of network features yields biologically intuitive conclusions using statistical methods. Notably, the distribution of node features evolves with disease progression, and distributions reveal mutualistic and parasitic species clusters. A sub-network consistently appears across all conditions, suggesting a 'persistent microbiome'. A clustering algorithm is also applied to demonstrate the methods utility for downstream analysis.
Metagenomics can provide insight into the microbial taxa present in a sample and, through gene identification, the functional potential of the community. However, taxonomic and functional information are typically considered separately in downstream analyses. We develop interpretable machine learning (ML) approaches for modelling metagenomic data, combining the biological representation of species with their associated genetically encoded functions within models. We apply our methods to investigate soil organic carbon (SOC) stocks. First, we combine a diverse global set of soil microbiome samples with environmental data, improving the predictive performance of classic ML and providing new insights into the role of soil microbiomes in global carbon cycling. Our network analysis of predictive taxa identified by classical ML models provides context for their ecological significance, extending the focus beyond just the most predictive taxa to ‘hidden’ features within the model that might be considered less predictive using standard methods for explainability. We next develop unique graph representations for individual microbiomes, linking microbial taxa to their associated functions directly, enabling predictions of SOC via deep graph convolutional neural networks (DGCNNs). Interpretation of the DGCNNs distinguished between the importance of functions of key individual species, providing genome sequence differences, e.g., gene loss/acquisition, that associate with SOC. These approaches identify several members of the Verrucomicrobiaceae family and a range of genetically encoded functions, e.g., related to carbohydrate metabolism, as important for SOC stocks and effective global SOC predictors. These relatively understudied but widespread organisms could play an important role in SOC dynamics globally.
The identification and prediction of variation in genetic data can be explored using graph-based machine learning methods. In particular, the comprehension of phenotypic effects at the cellular level is an accelerating research area in pharmacogenomics. Understanding the effect of drugs or disease on the underlying bionetwork could facilitate future drug development and improvement of precision medicine. In this article, a novel graph theoretic approach is proposed to infer a co-occurrence network from 16S microbiome data, specialised to handle small sampled data sets. Small data sets exacerbate the significant challenges faced by biological data, which exhibit properties such as sparsity, compositionality, and complexity of interactions. Methodologies are also proposed to statistically enrich and filter the inferred networks. The method is specialised for small data sets, which are abundant, but it can be generally applied to any 16S data set, and can also be extended to be integrated with other multi-omics data. The proposed methodology is tested on a data set of chickens vaccinated against and challenged by the protozoan parasite Eimeria tenella. Analysis of the expression of network features under three different stages of disease progression derive biologically intuitive conclusions from purely statistical methods. The distributions reveal clusters of species interacting mutualistically and parasitically, as expected. Moreover, a specific subnetwork is found to persist through all experimental conditions, representative of a 'core microbiome'.
We present a proof of concept implementation of the in-memory computing paradigm that we use to facilitate the analysis of metagenomic sequencing reads. In doing so we compare the performance of POSIX™file systems and key-value storage for omics data, and we show the potential for integrating high-performance computing (HPC) and cloud native technologies. We show that in-memory key-value storage offers possibilities for improved handling of omics data through more flexible and faster data processing. We envision fully containerized workflows and their deployment in portable micro-pipelines with multiple instances working concurrently with the same distributed in-memory storage. To highlight the potential usage of this technology for event driven and real-time data processing, we use a biological case study focused on the growing threat of antimicrobial resistance (AMR). We develop a workflow encompassing bioinformatics and explainable machine learning (ML) to predict life expectancy of a population based on the microbiome of its sewage while providing a description of AMR contribution to the prediction. We propose that in future, performing such analyses in 'real-time' would allow us to assess the potential risk to the population based on changes in the AMR profile of the community.
Achieving food security requires resilient agricultural systems with improved nutrient-use efficiency, optimized water and nutrient storage in soils, and reduced gaseous emissions. Success relies on understanding coupled nitrogen and carbon metabolism in soils, their associated influences on soil structure and the processes controlling nitrogen transformations at scales relevant to microbial activity. Here we show that the influence of organic matter on arable soil nitrogen transformations can be decoded by integrating metagenomic data with soil structural parameters. Our approach provides a mechanistic explanation of why organic matter is effective in reducing nitrous oxide losses while supporting system resilience. The relationship between organic carbon, soil-connected porosity and flow rates at scales relevant to microbes suggests that important increases in nutrient-use efficiency could be achieved at lower organic carbon stocks than currently envisaged.
The assembly of divergent haplotypes using noisy long-read data presents a challenge to the reconstruction of haploid genome assemblies, due to overlapping distributions of technical sequencing error, intralocus genetic variation, and interlocus similarity within these data. Here, we present a comparative analysis of assembly algorithms representing overlap-layout-consensus, repeat graph, and de Bruijn graph methods. We examine how postprocessing strategies attempting to reduce redundant heterozygosity interact with the choice of initial assembly algorithm and ultimately produce a series of chromosome-level assemblies for an agricultural pest, the diamondback moth, Plutella xylostella (L.). We compare evaluation methods and show that BUSCO analyses may overestimate haplotig removal processing in long-read draft genomes, in comparison to a k-mer method. We discuss the trade-offs inherent in assembly algorithm and curation choices and suggest that "best practice" is research question dependent. We demonstrate a link between allelic divergence and allele-derived contig redundancy in final genome assemblies and document the patterns of coding and noncoding diversity between redundant sequences. We also document a link between an excess of nonsynonymous polymorphism and haplotigs that are unresolved by assembly or postassembly algorithms. Finally, we discuss how this phenomenon may have relevance for the usage of noisy long-read genome assemblies in comparative genomics.
Additional file 8. Sample GO sequence database.
In a changing climate where future food security is a growing concern, researchers are exploring new methods and technologies in the effort to meet ambitious crop yield targets. The application of Artificial Intelligence (AI) including Machine Learning (ML) methods in this area has been proposed as a potential mechanism to support this. This review explores current research in the area to convey the state-of-the-art as to how AI/ML have been used to advance research, gain insights, and generally enable progress in this area. We address the question-Can AI improve crops and plant health? We further discriminate the bluster from the lustre by identifying the key challenges that AI has been shown to address, balanced with the potential issues with its usage, and the key requisites for its success. Overall, we hope to raise awareness and, as a result, promote usage, of AI related approaches where they can have appropriate impact to improve practices in agricultural and plant sciences.
The circadian clock is an important adaptation to life on Earth. Here, we use machine learning to predict complex, temporal, and circadian gene expression patterns in Arabidopsis Most significantly, we classify circadian genes using DNA sequence features generated de novo from public, genomic resources, facilitating downstream application of our methods with no experimental work or prior knowledge needed. We use local model explanation that is transcript specific to rank DNA sequence features, providing a detailed profile of the potential circadian regulatory mechanisms for each transcript. Furthermore, we can discriminate the temporal phase of transcript expression using the local, explanation-derived, and ranked DNA sequence features, revealing hidden subclasses within the circadian class. Model interpretation/explanation provides the backbone of our methodological advances, giving insight into biological processes and experimental design. Next, we use model interpretation to optimize sampling strategies when we predict circadian transcripts using reduced numbers of transcriptomic timepoints. Finally, we predict the circadian time from a single, transcriptomic timepoint, deriving marker transcripts that are most impactful for accurate prediction; this could facilitate the identification of altered clock function from existing datasets.
Genomics is both a data- and compute-intensive discipline. The success of genomics depends on an adequate informatics infrastructure that can address growing data demands and enable a diverse range of resource-intensive computational activities. Designing a suitable infrastructure is a challenging task, and its success largely depends on its adoption by users. In this article, we take a user-centric view of the genomics, where users are bioinformaticians, computational biologists, and data scientists. We try to take their point of view on how traditional computational activities for genomics are expanding due to data growth, as well as the introduction of big data and cloud technologies. The changing landscape of computational activities and new user requirements will influence the design of future genomics infrastructures.
Additional file 4. GO reference functional hierarchy tree structure.
Background Widespread bioinformatic resource development generates a constantly evolving and abundant landscape of workflows and software. For analysis of the microbiome, workflows typically begin with taxonomic classification of the microorganisms that are present in a given environment. Additional investigation is then required to uncover the functionality of the microbial community, in order to characterize its currently or potentially active biological processes. Such functional analysis of metagenomic data can be computationally demanding for high-throughput sequencing experiments. Instead, we can directly compare sequencing reads to a functionally annotated database. However, since reads frequently match multiple sequences equally well, analyses benefit from a hierarchical annotation tree, e.g. for taxonomic classification where reads are assigned to the lowest taxonomic unit. Results To facilitate functional microbiome analysis, we re-purpose well-known taxonomic classification tools to allow us to perform direct functional sequencing read classification with the added benefit of a functional hierarchy. To enable this, we develop and present a tree-shaped functional hierarchy representing the molecular function subset of the Gene Ontology annotation structure. We use this functional hierarchy to replace the standard phylogenetic taxonomy used by the classification tools and assign query sequences accurately to the lowest possible molecular function in the tree. We demonstrate this with simulated and experimental datasets, where we reveal new biological insights. Conclusions We demonstrate that improved functional classification of metagenomic sequencing reads is possible by re-purposing a range of taxonomic classification tools that are already well-established, in conjunction with either protein or nucleotide reference databases. We leverage the advances in speed, accuracy and efficiency that have been made for taxonomic classification and translate these benefits for the rapid functional classification of microbiomes. While we focus on a specific set of commonly used methods, the functional annotation approach has broad applicability across other sequence classification tools. We hope that re-purposing becomes a routine consideration during bioinformatic resource development.
The majority of bacterial genomes have high coding efficiencies, but there are some genomes of intracellular bacteria that have low gene density. The genome of the endosymbiont Sodalis glossinidius contains almost 50% pseudogenes containing mutations that putatively silence them at the genomic level. We have applied multiple ‘omic strategies, combining: Illumina and Pacific Biosciences Single-Molecule Real Time DNA-sequencing and annotation; stranded RNA-sequencing; and proteome analysis to better understand the transcriptional and translational landscape of Sodalis pseudogenes, and potential mechanisms for their control. Between 53% and 74% of the Sodalis transcriptome remains active in cell-free culture. Mean sense transcription from Coding Domain Sequences (CDS) is four-times greater than that from pseudogenes. Comparative genomic analysis of six Illumina-sequenced Sodalis isolates from different host Glossina species shows pseudogenes make up ~40% of the 2,729 genes in the core genome, suggesting that they are stable and/or Sodalis is a recent introduction across the Glossina genus as a facultative symbiont. These data further shed light on the importance of transcriptional and translational control in deciphering host-microbe interactions. The combination of genomics, transcriptomics and proteomics give a multidimensional perspective for studying prokaryotic genomes with a view to elucidating evolutionary adaptation to novel environmental niches. Importance Bacterial genes are generally 1Kb in length, organized efficiently (i.e. with few gaps between genes or operons), and few open reading frames (ORFs) lack any predicted function. Intracellular bacteria have been removed from extracellular selection pressures acting on pathways of declining importance to fitness and thus, these bacteria tend to delete redundant genes in favour of smaller functional repertoires. In the genomes of endosymbionts with a recent evolutionary relationship with their host, however, this process of genome reduction is not complete; Genes and pathways may be at an intermediate stage, undergoing mutation linked to reduced selection and small population numbers being vertically transmitted from mother to offspring in their hosts, resulting in an increase in abundance of pseudogenes and reduced coding capacities. A greater knowledge of the genomic architecture of persistent pseudogenes, with respect to their DNA structure, mRNA transcription and even putative translation to protein products, will lead to a better understanding of the evolutionary trajectory of endosymbiont genomes, many of which have important roles in arthropod ecology.
During the development of new drugs or compounds there is a requirement for preclinical trials, commonly involving animal tests, to ascertain the safety of the compound prior to human trials. Machine learning techniques could provide an in-silico alternative to animal models for assessing drug toxicity, thus reducing expensive and invasive animal testing during clinical trials, for drugs that are most likely to fail safety tests. Here we present a machine learning model to predict kidney dysfunction, as a proxy for drug induced renal toxicity, in rats. To achieve this, we use inexpensive transcriptomic profiles derived from human cell lines after chemical compound treatment to train our models combined with compound chemical structure information. Genomics data due to its sparse, high-dimensional and noisy nature presents significant challenges in building trustworthy and transparent machine learning models. Here we address these issues by judiciously building feature sets from heterogenous sources and coupling them with measures of model uncertainty achieved through Gaussian Process based Bayesian models. We combine the use of insight into the feature-wise contributions to our predictions with the use of predictive uncertainties recovered from the Gaussian Process to improve the transparency and trustworthiness of the model.
Increasingly available microbial reference data allow interpreting the composition and function of previously uncharacterized microbial communities in detail, via high-throughput sequencing analysis. However, efficient methods for read classification are required when the best database matches for short sequence reads are often shared among multiple reference sequences. Here, we take advantage of the fact that microbial sequences can be annotated relative to established tree structures, and we develop a highly scalable read classifier, PRROMenade, by enhancing the generalized Burrows-Wheeler transform with a labeling step to directly assign reads to the corresponding lowest taxonomic unit in an annotation tree. PRROMenade solves the multi-matching problem while allowing fast variable-size sequence classification for phylogenetic or functional annotation. Our simulations with 5% added differences from reference indicated only 1.5% error rate for PRROMenade functional classification. On metatranscriptomic data PRROMenade highlighted biologically relevant functional pathways related to diet-induced changes in the human gut microbiome.
Biomedical data, particularly in the field of genomics, has characteristics which make it challenging for machine learning applications - it can be sparse, high dimensional and noisy. Biomedical applications also present challenges to model selection - whilst powerful, accurate predictions are necessary, they alone are not sufficient for a model to be deemed useful. Due to the nature of the predictions, a model must also be trustworthy and transparent, empowering a practitioner with confidence that its use is appropriate and reliable. In this paper, we propose that this can be achieved through the use of judiciously built feature sets coupled with Bayesian models, specifically Gaussian processes. We apply Gaussian processes to drug discovery, using inexpensive transcriptomic profiles from human cell lines to predict animal kidney and liver toxicity after treatment with specific chemical compounds. This approach has the potential to reduce invasive and expensive animal testing during clinical trials if in vitro human cell line analysis can accurately predict model animal phenotypes. We compare results across a range of feature sets and models, to highlight model importance for medical applications.
Genomics and related technologies, collectively known as Omics, have transformed life sciences research. These technologies produce mountain of data that needs to be managed and analysed. Rapid developments in the Next Generation Sequencing technologies have helped genomics become mainstream, but the compute support systems, meant to enable genomics, have lagged behind. As genomics is making inroads into personalised health care and clinical settings, it is paramount that a robust compute infrastructure be designed to meet the growing needs of the field. Infrastructure design to deal with omics datasets is an active area of research and a critical one, for omics to be adopted in industrial healthcare and clinical settings. In this paper, we propose a blueprint for an as-a service compute infrastructure for fast and scalable processing of omics datasets. We explain our approach with help of a well-known bioinformatics workflow and a compute environment that can be tailored to achieve portability, reproducibility and scalability using modern High Performance Computing systems.
Background Aedes aegypti is the principal vector of several important arboviruses. Among the methods of vector control to limit transmission of disease are genetic strategies that involve the release of sterile or genetically modified non-biting males, which has generated interest in manipulating mosquito sex ratios. Sex determination in Ae. aegypti is controlled by a non-recombining Y chromosome-like region called the M locus, yet characterisation of this locus has been thwarted by the repetitive nature of the genome. In 2015, an M locus gene named Nix was identified that displays the qualities of a sex determination switch. Results With the use of a whole-genome bacterial artificial chromosome (BAC) library, we amplified and sequenced a ~200 kb region containing the male-determining gene Nix . In this study, we show that Nix is comprised of two exons separated by a 99 kb intron primarily composed of repetitive DNA, especially transposable elements. Conclusions Nix , an unusually large and highly repetitive gene, exhibits features in common with Y chromosome genes in other organisms. We speculate that the lack of recombination at the M locus has allowed the expansion of repeats in a manner characteristic of a sex-limited chromosome, in accordance with proposed models of sex chromosome evolution in insects.