Conventional annotation of single-cell RNA-sequencing (scRNA-seq) data relies heavily on manual, marker-based thresholding, an approach that can obscure subtle transcriptomic gradients and collapse functionally distinct cell states into broad, heterogeneous populations. Here we apply the Gaussian multi-Graphical Model (GmGM) framework, which jointly infers cell-cell and gene-gene dependency structure from a single scRNA-seq data matrix, to a 10x Genomics PBMC dataset. Ten independent GMGM-Leiden clustering runs were integrated into a robust consensus partition using a soft cluster ensemble approach and benchmarked against reference cell-type annotations. This strategy yielded stable cluster partitions that resolve biologically meaningful sub-populations not distinguished by the reference annotation. In parallel, for each cluster, gene co-expression modules were extracted from the fitted model via consensus Leiden clustering across resolutions, evaluated using standard network metrics, and validated functionally with the Network Enrichment Analysis Test (NEAT), which confirmed non-random enrichment signal. A module-scoring procedure linked network topology to per-cell, per-cluster expression signatures, and a novel extension of GmGM, recovering a shared cell-cell network together with population-specific gene networks in a single model run, was demonstrated in a case study on the CD4+ T-cell population. These results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.
Abstract Background Access to omics datasets, such as single-cell RNA-sequencing, enables us to estimate the regulatory networks governing the differentiation, proliferation, and interaction of cells in our body. Knowledge of such networks can give us valuable insight on the structure and progression of diseases; the difficulty is in estimating them. Most methods that estimate these networks either make an independence assumption (‘the cells in your body do not interact’), or ignore the directional nature of gene regulation. Methodology In this paper, we introduce the Cartesian Linear Gaussian Additive Noise Model to learn both cell-cell and gene-gene interactions. Our method is a statistical method; it is fit with maximum likelihood estimation, and we prove that a unique optimum always exists (under certain assumptions) using tools from high-dimensional statistics. Results Our method differs from prior work in its lack of an independence assumption; we show that this leads to a real improvement in gene regulatory network and cell network reconstructions relative to analogous independence-assuming methods. Conclusions We have developed and proved viable a novel method that learns directed gene regulatory networks, without assuming independence of cells. Our method is also extensible to more complicated omics datasets, such as longitudinal bulk RNA-sequencing datasets, through its ability to handle ‘tensor-variate’ datasets. Trial Registration Clinical trial number: not applicable.
MOTIVATION:Networks underlie the generation and interpretation of many biological datasets: gene networks shed light on the regulatory structure of the genome, and cell networks can capture structure of the tumor micro-environment. However, most methods that learn such networks make the faulty "independence assumption"; to learn the gene network, they assume that no cell network exists. "Multi-axis" methods, which do not make this assumption, fail to scale beyond a few thousand cells or genes. This limits their applicability to only the smallest datasets. RESULTS:We develop a multi-axis method, which learns conditional dependency networks, capable of processing million-cell datasets within minutes. This was previously impossible, and unlocks the use of such methods on modern scRNA-seq datasets, as well as more complex datasets. We apply the method to a new scRNA-seq dataset for neuronal cell development, and compare the result to an existing state of the art method, hdWGCNA. We demonstrate that the new method yields gene networks that have a more focused biological interpretation and that the simultaneously learned cell network has advantages over a conventional kNN-based clustering. Further, our method yields novel biological insights by identifying long non-coding RNAs that potentially have a role in neuronal development. AVAILABILITY AND IMPLEMENTATION:Our methodology is available as a Python package GmGM on PyPI (https://pypi.org/project/GmGM/0.5.3/). The code for all experiments performed in this article is available on GitHub (https://github.com/BaileyAndrew/GmGM-Bioinformatics) and Zenodo (10.5281/zenodo.20384566).
In this paper we develop a graph-learning algorithm, MED-MAGMA, to fit multi-axis (Kronecker-sum-structured) models corrupted by multiplicative noise. This type of noise is natural in many application domains, such as that of single-cell RNA sequencing, in which it naturally captures technical biases of RNA sequencing platforms. Our work is evaluated against prior work on each and every public dataset in the Single Cell Expression Atlas under a certain size, demonstrating that our methodology learns networks with better local and global structure. MED-MAGMA is made available as a Python package (MED-MAGMA).
Motivation:The rapid expansion of single-cell RNA sequencing (scRNA-seq) technologies has increased the need for robust and scalable clustering evaluation methods. To address these challenges, we developed robin2, an optimized version of our R package robin. It introduces enhanced computational efficiency, support for high-dimensional datasets, and harmonious integration with R's base functionalities for robust network analysis. Results:robin2 offers improved functionality for clustering stability validation and enables systematic evaluation of community detection algorithms across various resolutions and pipelines. The application to Tabula Muris and PBMC scRNA-seq datasets confirmed its ability to identify biologically meaningful cell subpopulations with high statistical significance. The new version reduces computational time by 9-fold on large-scale datasets using parallel processing. Availability and implementation:The robin2 package is freely available on CRAN at https://CRAN.R-project.org/package=robin. Comprehensive documentation and a detailed analysis vignette are available on GitHub at https://drighelli.github.io/scrobinv2/index.html.
The independence assumption between random variables is a useful tool to increase the tractability of a modelling framework. However, this assumption can be too simplistic; failing to take dependencies into account can cause models to fail dramatically. The field of multi-axis graphical modelling (also called multi-way modelling, Kronecker-separable modelling) has seen growth over the past decade, but these models require that the data have zero mean. In the multi-axis case, inference is typically done in the single sample scenario, making mean inference impossible. In this paper, we demonstrate how the zero-mean assumption can cause egregious modelling errors for Kronecker-sum-decomposable Gaussian graphical models, as well as propose a relaxation to the zero-mean assumption that allows the avoidance of such errors. Specifically, we propose the "Kronecker-sum-structured mean" assumption, which leads to models with nonconvex-but-unimodal log-likelihoods that can be solved efficiently with coordinate descent.
Motivation: Networks underlie the generation and interpretation of many biological datasets: gene networks shed light on the regulatory structure of the genome, and cell networks can capture structure of the tumor micro-environment. However, most methods that learn such networks make the faulty 'independence assumption'; to learn the gene network, they assume that no cell network exists. 'Multi-axis' methods, which do not make this assumption, fail to scale beyond a few thousand cells or genes. This limits their applicability to only the smallest datasets. Results: We develop a multi-axis method capable of processing million-cell datasets within minutes. This was previously impossible, and unlocks the use of such methods on modern scRNA-seq datasets, as well as more complex datasets. We show that our method yields novel biological insights from real single-cell data, and compares favorably to the existing hdWGCNA methodology. In particular, it identifies long non-coding RNA genes that potentially have a regulatory or functional role in neuronal development. Availability and implementation: Our methodology is available as a Python package GmGM on PyPI (https://pypi.org/project/GmGM/0.5.3/). The code for all experiments performed in this paper is available on GitHub (https://github.com/BaileyAndrew/GmGM-Bioinformatics). Contact: sceba@leeds.ac.uk Supplementary information: Our proofs, and some additional experiments, are available in the supplementary material. Keywords: gaussian graphical models, multi-axis models, transcriptomics, multi-omics, scalability
Quiescence, a reversible state of cell-cycle arrest, is an important state during both normal development and cancer progression. For example, in glioblastoma (GBM) quiescent glioblastoma stem cells (GSCs) play an important role in re-establishing the tumour, leading to relapse. While most studies have focused on identifying differentially expressed genes between proliferative and quiescent cells as potential drivers of this transition, recent studies have shown the importance of protein oscillations in controlling the exit from quiescence of neural stem cells. Here, we have undertaken a genome-wide bioinformatic inference approach to identify genes whose expression oscillates and which may be good candidates for controlling the transition to and from the quiescent cell state in GBM. Our analysis identified, among others, a list of important transcription regulators as potential oscillators, including the stemness gene SOX2 , which we verified to oscillate in quiescent GSCs. These findings expand on the way we think about gene regulation and introduce new candidate genes as key regulators of quiescence.
This paper introduces the Gaussian multi-Graphical Model, a model to construct sparse graph representations of matrix- and tensor-variate data. We generalize prior work in this area by simultaneously learning this representation across several tensors that share axes, which is necessary to allow the analysis of multimodal datasets such as those encountered in multi-omics. Our algorithm uses only a single eigendecomposition per axis, achieving an order of magnitude speedup over prior work in the ungeneralized case. This allows the use of our methodology on large multi-modal datasets such as single-cell multi-omics data, which was challenging with previous approaches. We validate our model on synthetic data and five real-world datasets.
Background Glioblastoma (GBM) brain tumors lacking IDH1 mutations (IDHwt) have the worst prognosis of all brain neoplasms. Patients receive surgery and chemoradiotherapy but tumors almost always fatally recur. Results Using RNA sequencing data from 107 pairs of pre- and post-standard treatment locally recurrent IDHwt GBM tumors, we identify two responder subtypes based on longitudinal changes in gene expression. In two thirds of patients, a specific subset of genes is upregulated from primary to recurrence (Up responders), and in one third, the same genes are downregulated (Down responders), specifically in neoplastic cells. Characterization of the responder subtypes indicates subtype-specific adaptive treatment resistance mechanisms that are associated with distinct changes in the tumor microenvironment. In Up responders, recurrent tumors are enriched in quiescent proneural GBM stem cells and differentiated neoplastic cells, with increased interaction with the surrounding normal brain and neurotransmitter signaling, whereas Down responders commonly undergo mesenchymal transition. ChIP-sequencing data from longitudinal GBM tumors suggests that the observed transcriptional reprogramming could be driven by Polycomb-based chromatin remodeling rather than DNA methylation. Conclusions We show that the responder subtype is cancer-cell intrinsic, recapitulated in in vitro GBM cell models, and influenced by the presence of the tumor microenvironment. Stratifying GBM tumors by responder subtype may lead to more effective treatment.
Transcriptomes and translatomes measure genome-wide levels of total and ribosome-associated RNAs. A few hundred translatomes were reported over >250,000 transcriptomes highlighting the challenges of identifying translating RNAs. Here, we used a human isogenic inducible model of TDP-43-linked amyotrophic lateral sclerosis, which exhibits altered expression of thousands of transcripts, as a paradigm for the direct comparison of whole-cell, cytoplasmic and translating RNAs, showing broad uncoupling and poor correlation between disease-altered transcripts. Moreover, based on precipitation of endogenous ribosomes, we developed GRASPS (Genome-wide RNA Analysis of Stalled Protein Synthesis), a simple-to-operate translatome technology. Remarkably, GRASPS identified three times more differentially-expressed transcripts with higher fold changes and statistical significance, providing unprecedented opportunities for data modeling at stringent filtering and discovery of previously omics-missed disease-relevant pathways, which functionally map on dense gene regulatory networks of protein-protein interactions. Based on its simplicity and robustness, GRASPS is widely applicable across disciplines in the biotechnologies and biomedical sciences.### Competing Interest StatementProfs Guillaume M. Hautbergue and Pamela J. Shaw are Founders of Crucible Therapeutics Limited, a start-up company developing gene therapeutics to treat neurological disorders. The authors declare having no other competing interests.
This paper introduces the Gaussian multi-Graphical Model, a model to construct sparse graph representations of matrix- and tensor-variate data. We generalize prior work in this area by simultaneously learning this representation across several tensors that share axes, which is necessary to allow the analysis of multimodal datasets such as those encountered in multi-omics. Our algorithm uses only a single eigendecomposition per axis, achieving an order of magnitude speedup over prior work in the ungeneralized case. This allows the use of our methodology on large multi-modal datasets such as single-cell multi-omics data, which was challenging with previous approaches. We validate our model on synthetic data and five real-world datasets.
Cases of laryngeal cancer are rising, with diagnosis often involving invasive biopsy procedures.An alternate approach is to identify high-risk patients by analysis of voice recordings which can alert clinical teams to those patients that need prioritisation.We propose a pipeline for evaluating speech classifier performance in the presence of noise.We perform experiments using the pipeline with several classifiers and denoising techniques.Random forest classifier performed best with an accuracy of 81.2% on clean data dropping to 63.8% when noise was added to recordings.The accuracy of all classifiers was reduced by added noise, signal denoising improved classifier accuracy but could not fully reverse the effects of noise.The effects of noise on classification is a complex issue which must be resolved for these detection systems to be implemented in clinical practice.We show that the proposed pipeline allows for the evaluation of classifier performance in the presence of noise.
In this paper we propose a new approach to detect clusters in undirected graphs with attributed vertices. We incorporate structural and attribute similarities between the vertices in an augmented graph by creating additional vertices and edges as proposed in [1, 2]. The augmented graph is then embedded in a Euclidean space associated to its Laplacian and we cluster vertices via a modified K-means algorithm, using a new vector-valued distance in the embedding space. Main novelty of our method, which can be classified as an early fusion method, i.e., a method in which additional information on vertices are fused to the structure information before applying clustering, is the interpretation of attributes as new realizations of graph vertices, which can be dealt with as coordinate vectors in a related Euclidean space. This allows us to extend a scalable generalized spectral clustering procedure which substitutes graph Laplacian eigenvectors with some vectors, named algebraically smooth vectors, obtained by a linear-time complexity Algebraic MultiGrid (AMG) method. We discuss the performance of our proposed clustering method by comparison with recent literature approaches and public available results. Extensive experiments on different types of synthetic datasets and real-world attributed graphs show that our new algorithm, embedding attributes information in the clustering, outperforms structure-only-based methods, when the attributed network has an ambiguous structure. Furthermore, our new method largely outperforms the method which originally proposed the graph augmentation, showing that our embedding strategy and vector-valued distance are very effective in taking advantages from the augmented-graph representation.
Classically, statistical datasets have a larger number of data points than features (n > p). The standard model of classical statistics caters for the case where data points are considered conditionally independent given the parameters. However, for n approximate to p or p > n such models are poorly determined. Kalaitzis et al. (2013) introduced the Bigraphical Lasso, an estimator for sparse precision matrices based on the Cartesian product of graphs. Unfortunately, the original Bigraphical Lasso algorithm is not applicable in case of large p and n due to memory requirements. We exploit eigenvalue decomposition of the Cartesian product graph to present a more efficient version of the algorithm which reduces memory requirements from O(n(2)p(2)) to O(n(2) + p(2)). Many datasets in different application fields, such as biology, medicine and social science, come with count data, for which Gaussian based models are not applicable. Our multiway network inference approach can be used for discrete data. Our methodology accounts for the dependencies across both instances and features, reduces the computational complexity for high dimensional data and enables to deal with both discrete and continuous data. Numerical studies on both synthetic and real datasets are presented to showcase the performance of our method.
Classically, statistical datasets have a larger number of data points than features ($n > p$). The standard model of classical statistics caters for the case where data points are considered conditionally independent given the parameters. However, for $n\approx p$ or $p > n$ such models are poorly determined. Kalaitzis et al. (2013) introduced the Bigraphical Lasso, an estimator for sparse precision matrices based on the Cartesian product of graphs. Unfortunately, the original Bigraphical Lasso algorithm is not applicable in case of large p and n due to memory requirements. We exploit eigenvalue decomposition of the Cartesian product graph to present a more efficient version of the algorithm which reduces memory requirements from $O(n^2p^2)$ to $O(n^2 + p^2)$. Many datasets in different application fields, such as biology, medicine and social science, come with count data, for which Gaussian based models are not applicable. Our multi-way network inference approach can be used for discrete data. Our methodology accounts for the dependencies across both instances and features, reduces the computational complexity for high dimensional data and enables to deal with both discrete and continuous data. Numerical studies on both synthetic and real datasets are presented to showcase the performance of our method.
In this paper we propose a new approach to detect clusters in undirected graphs with attributed vertices. The aim is to group vertices which are similar not only in terms of structural connectivity but also in terms of attribute values. We incorporate structural and attribute similarities between the vertices in an augmented graph by creating additional vertices and edges as proposed in [5, 27]. The augmented graph is embedded in a Euclidean space associated to its Laplacian and apply a modified K-means algorithm to identify clusters. The modified K-means uses a vector distance measure where to each original vertex is assigned a vector-valued set of coordinates depending on both structural connectivity and attribute similarities. To define the coordinate vectors we employ an adaptive AMG (Algebraic MultiGrid) method to identify the coordinate directions in the embedding Euclidean space extending our previous result for graphs without attributes. We demonstrate the effectiveness of our proposed clustering method on both synthetic and real-world attributed graphs.
In network analysis, many community detection algorithms have been developed.However, their implementation leaves unaddressed the question of the statistical validation of the results.Here, we present robin (ROBustness In Network), an R package to assess the robustness of the community structure of a network found by one or more methods to give indications about their reliability.The procedure initially detects if the community structure found by a set of algorithms is statistically significant and then compares two selected detection algorithms on the same graph to choose the one that better fits the network of interest.We demonstrate the use of our package on the American College Football benchmark dataset.
Background Loss of motor neurons in amyotrophic lateral sclerosis (ALS) leads to progressive paralysis and death. Dysregulation of thousands of RNA molecules with roles in multiple cellular pathways hinders the identification of ALS-causing alterations over downstream changes secondary to the neurodegenerative process. How many and which of these pathological gene expression changes require therapeutic normalisation remains a fundamental question. Methods Here, we investigated genome-wide RNA changes in C9ORF72-ALS patient-derived neurons and Drosophila , as well as upon neuroprotection taking advantage of our gene therapy approach which specifically inhibits the SRSF1-dependent nuclear export of pathological C9ORF72 -repeat transcripts. This is a critical study to evaluate (i) the overall safety and efficacy of the partial depletion of SRSF1, a member of a protein family involved itself in gene expression, and (ii) a unique opportunity to identify neuroprotective RNA changes. Results Our study shows that manipulation of 362 transcripts out of 2257 pathological changes, in addition to inhibiting the nuclear export of repeat transcripts, is sufficient to confer neuroprotection in C9ORF72-ALS patient-derived neurons. In particular, expression of 90 disease-altered transcripts is fully reverted upon neuroprotection leading to the characterisation of a human C9ORF72-ALS disease-modifying gene expression signature. These findings were further investigated in vivo in diseased and neuroprotected Drosophila transcriptomes, highlighting a list of 21 neuroprotective changes conserved with 16 human orthologues in patient-derived neurons. We also functionally validated the high neuroprotective potential of one of these disease-modifying transcripts, demonstrating that inhibition of ALS-upregulated human KCNN1–3 ( Drosophila SK) voltage-gated potassium channel orthologs mitigates degeneration of human motor neurons and Drosophila motor deficits. Conclusions Strikingly, the partial depletion of SRSF1 leads to expression changes in only a small proportion of disease-altered transcripts, indicating that not all RNA alterations need normalization and that the gene therapeutic approach is safe in the above preclinical models as it does not disrupt globally gene expression. The efficacy of this intervention is also validated at genome-wide level with transcripts modulated in the vast majority of biological processes affected in C9ORF72-ALS. Finally, the identification of a characteristic signature with key RNA changes modified in both the disease state and upon neuroprotection also provides potential new therapeutic targets and biomarkers.