Mutational signatures are characteristic patterns of mutation frequencies assumed to be generated by specific mutagenic processes. A growing catalog of mutational signatures exists, but tools to systematically infer relationships between them remain limited. A mutational signature can be viewed as the combined outcome of two processes: DNA damage and DNA repair. Since cancer therapies often target DNA repair, inferring DNA repair pathways is important for treatment design, even when the mutagenic process is unknown. Here, we model the DNA repair step as a transformation, called RePrint, from damaged nucleotides to repair-related mutation patterns conditioned on the damage. We demonstrate that RePrint similarity is indicative of shared DNA repair mechanisms, enabling guilt-by-association prediction of DNA repair pathways. Using experimentally annotated signatures from environmental exposures and CRISPR gene knockouts as gold standards, we demonstrate that RePrint-based clustering consistently outperforms signature-based clustering across multiple evaluation metrics. We validate several guilt-by-association predictions with literature evidence, demonstrating RePrint's ability to identify shared repair mechanisms even among signatures with divergent mutational profiles. RePrint provides the first approach to systematically transfer DNA repair information between signatures, opening doors to understanding signatures of unknown origin and informing therapeutic strategies. An open-source implementation is available at https://github.com/wojtowicz-lab/RePrintPy.
Mutational processes shape cancer genomes, leaving characteristic marks that are termed signatures. The level of activity of each such process, or its signature exposure, provides important information on the disease, improving patient stratification and the prediction of drug response. Thus, there is growing interest in developing refitting methods that accurately decipher those exposures. Previous work in this domain was unsupervised in nature, employing algebraic decomposition and probabilistic inference methods. We present SuRe, a supervised approach to signature refitting that demonstrates superiority over current methods. SuRe leverages a neural network model to capture correlations between signature exposures in real data. We show that SuRe outperforms previous methods on sparse mutation data from both tumor-type-specific and pan-cancer data sets, with an increasing performance advantage as the data become sparser. We further demonstrate the model’s utility in clinical settings by predicting homologous recombination deficiency in breast cancer from sparse data. Furthermore, SuRe outperforms standard methods in the unsupervised stratification of over 13,000 patients from large-scale panel sequencing cohorts, highlighting its potential for analyzing targeted sequencing data.
Motivation Hierarchical clustering is a fundamental problem in computational biology, with popular greedy heuristics such as average linkage dating back to the 1950s but no well-defined objective. Recently, a combinatorial optimization criterion for the problem was suggested by Dasgupta. While minimizing this criterion is NP-hard, the popular average linkage method serves as a strong baseline. Nevertheless, its myopic, greedy nature frequently leads to structurally suboptimal hierarchies.Results To remedy this, we introduce a novel average-linkage-based clustering approach that combines local and global considerations by generating multiple views of the input data and learning how to blend them into an integrated similarity measure. We demonstrate that our method, DOMUS, consistently outperforms strong baselines, including a beam search heuristic, on a wide range of synthetic and classic benchmark datasets. Furthermore, we validate its real-world applicability through a rigorous benchmark on single-cell RNA sequencing data, where it compares favorably with the state-of-the-art HiDeF algorithm.Availability and implementation The DOMUS framework is implemented in Python and freely available at https://github.com/GalGilad/DOMUS.
MOTIVATION:Protein-protein interactions (PPIs) provide the skeleton for signaling pathways in the cell. Their experimental measurement, however, reveals only the existence of an interaction without any information on its functional roles. A key step in developing a working logical model of cell signaling is annotating activation/repression (sign) of an interaction. RESULTS:Here, we develop SIGN Annotation aLgorithm (SIGNAL), a method for annotating PPI networks with signs based on cause-effect data. The approach is based on a multiplicative model in which the effect of a pathway is assumed to be the product of the signs along its edges. The algorithm uses network propagation techniques to quantify the influence of each edge on gene expression changes, and the resulting features are fed to a classifier for sign prediction. We validate our method using known annotations and demonstrate the utility of SIGNAL for predicting the effect of a knockout on gene expression and on telomere length. AVAILABILITY AND IMPLEMENTATION:SIGNAL code is available at https://github.com/L-F-S/PPI_Network_Signer.
SUMMARY:Pathway enrichment analysis is a fundamental technique in bioinformatics for interpreting gene expression data to pinpoint biological pathways associated with specific conditions or diseases. We introduce Pathway Enrichment Analysis through Network UTilization (PEANUT), a web-based tool for pathway enrichment analysis that enhances traditional pipelines by integrating network propagation computations within a network of protein-protein interactions (PPIs). By diffusing gene expression scores through the PPI network, PEANUT amplifies the signals of connected sets of genes, thereby improving the detection of relevant pathways. AVAILABILITY AND IMPLEMENTATION:The tool is accessible as an open-source web application at https://peanut.cs.tau.ac.il/. The source code is available at https://github.com/Yapibe/PEANUT with a permanent identifier (DOI: https://doi.org/10.5281/zenodo.15184862).
To begin deciphering the hierarchical structure of the cell, we need to integrate multiple types of data of different scales on subcellular organization. To this end, we developed MIRAGE, a multi-modal generative model for integrating protein sequence, protein-protein interaction, and protein localization data. Our adversarial approach successfully learns a joint embedding space that captures the complex relationships among these diverse modalities and allows us to generate missing modalities. We evaluate our model’s performance against existing methods, obtaining superior performance in protein function prediction and protein complex detection. We apply MIRAGE to construct a hierarchical map of subcellular organization in HEK293T cells, recovering known protein assemblies across multiple scales.
Summary:Proteomics has developed many approaches to inform the subcellular organization of proteins, each with differing coverage and sensitivity to distinct scales. Here, we develop a self-supervised deep learning framework, ProteinProjector, that flexibly integrates all available data for a protein from any number of modalities, resulting in a unified map of protein position. As initial proof-of-concept we integrate four proteome-wide characterizations of HEK293 human embryonic kidney cells, including protein affinity purification, proximity ligation, and size-exclusion-chromatography mass spectrometry (AP-MS, PL-MS, SEC-MS), as well as protein fluorescent imaging. Map coverage and accuracy grow substantially as new data modes are added, with maximal recovery of known complexes observed when using all four proteomic datasets. We find that ProteinProjector outperforms individual modalities and other integration methods in recovery of orthogonal functional and physical associations not used during training. ProteinProjector provides a foundation for integration of diverse modalities that characterize subcellular structure. Availability and implementation:ProteinProjector is available as part of the Cell Mapping Toolkit at https://github.com/idekerlab/cellmaps_coembedding.
Recent advances in single-cell RNA sequencing (scRNA-seq) techniques have provided unprecedented insights into the heterogeneity of various tissues. However, gene expression data alone often fails to capture and identify changes in cellular pathways and complexes, as they are more discernible at the protein level. Moreover, analyzing scRNA-seq data presents further challenges due to inherent characteristics such as high noise levels and zero inflation. In this study, we propose an approach to address these limitations by integrating scRNA-seq datasets with a protein-protein interaction network. Our method utilizes a unique dual-view architecture based on graph neural networks, enabling joint representation of gene expression and protein-protein interaction network data. This approach models gene-to-gene relationships under specific biological contexts and refines cell-cell relations using an attention mechanism. Next, through comprehensive evaluations, we demonstrate that scNET better captures gene annotation, pathway characterization and gene-gene relationship identification, while improving cell clustering and pathway analysis across diverse cell types and biological conditions.
Tumor heterogeneity drives drug resistance and relapse, influencing immune evasion and tumor progression. While intratumor heterogeneity has been extensively studied at the genomic level, its functional outcomes and interactions with the tumor microenvironment remain underexplored. In contrast, the functional outcome of heterogeneity and the interplay with the tumor microenvironment have not been addressed. In this study, we integrate multi-region spatial MS-based proteomics of 280 tumor regions, exome sequencing, and imaging to investigate spatial proteomic heterogeneity in breast cancer. Our findings reveal increased proteomic heterogeneity with tumor progression, independent of genomic heterogeneity but closely associated with microenvironmental differences. Integration with immune and stromal imaging highlighted a dynamic interplay where low-grade tumors exhibit constrained immune infiltration, and upon progression to higher grades, macrophages and T cells infiltrate. However, anti-inflammatory pathways involving kynurenine and prostaglandins are more highly expressed in infiltrated regions, suggesting that anti-tumorigenic activities are inhibited. Integration with the global protein network provides potential targetable mediators of immune evasion in breast cancer that can serve as the basis for future development of personalized breast cancer therapies.
Autism spectrum disorder (ASD) is a highly heritable complex disease that affects 1% of the population, yet its underlying molecular mechanisms are largely unknown. Here we study the problem of predicting causal genes for ASD by combining genome-scale data with a network propagation approach. We construct a predictor that integrates multiple omic data sets that assess genomic, transcriptomic, proteomic, and phosphoproteomic associations with ASD. In cross validation our predictor yields mean area under the ROC curve of 0.87 and area under the precision-recall curve of 0.89. We further show that it outperforms previous gene-level predictors of autism association. Finally, we show that we can use the model to predict genes associated with Schizophrenia which is known to share genetic components with ASD.
Network biology, an interdisciplinary field at the intersection of computational and biological sciences, is critical for deepening understanding of cellular functioning and disease. While the field has existed for about two decades now, it is still relatively young. There have been rapid changes to it and new computational challenges have arisen. This is caused by many factors, including increasing data complexity, such as multiple types of data becoming available at different levels of biological organization, as well as growing data size. This means that the research directions in the field need to evolve as well. Hence, a workshop on Future Directions in Network Biology was organized and held at the University of Notre Dame in 2022, which brought together active researchers in various computational and in particular algorithmic aspects of network biology to identify pressing challenges in this field. Topics that were discussed during the workshop include: inference and comparison of biological networks, multimodal data integration and heterogeneous networks, higher-order network analysis, machine learning on networks, and network-based personalized medicine. Video recordings of the workshop presentations are publicly available on YouTube. For even broader impact of the workshop, this paper, co-authored mostly by the workshop participants, summarizes the discussion from the workshop. As such, it is expected to help shape short- and long-term vision for future computational and algorithmic research in network biology.
The data deluge in biology calls for computational approaches that can integrate multiple datasets of different types to build a holistic view of biological processes or structures of interest. An emerging paradigm in this domain is the unsupervised learning of data embeddings that can be used for downstream clustering and classification tasks. While such approaches for integrating data of similar types are becoming common, there is scarcer work on consolidating different data modalities such as network and image information. Here, we introduce DICE (Data Integration through Contrastive Embedding), a contrastive learning model for multi-modal data integration. We apply this model to study the subcellular organization of proteins by integrating protein-protein interaction data and protein image data measured in HEK293 cells. We demonstrate the advantage of data integration over any single modality and show that our framework outperforms previous integration approaches. Availability: https://github.com/raminass/protein-contrastive Contact: raminass@gmail.com.
It is becoming clear that bulk gene expression measurements represent an average over very different cells. Elucidating the expression and abundance of each of the encompassed cells is key to disease understanding and precision medicine approaches. A first step in any such deconvolution is the inference of cell type abundances in the given mixture. Numerous approaches to cell-type deconvolution have been proposed, yet very few take advantage of the emerging discipline of deep learning and most approaches are limited to input data regarding the expression profiles of the cell types in question. Here we present DECODE, a deep learning method for the task that is data-driven and does not depend on input expression profiles. DECODE builds on a deep unfolded non-negative matrix factorization technique. It is shown to outperform previous approaches on a range of synthetic and real data sets, producing abundance estimates that are closer to and better correlated with the real values.
Motivation:Protein-protein interactions (PPIs) play essential roles in the buildup of cellular machinery and provide the skeleton for cellular signaling. However, these biochemical roles are context dependent and interactions may change across cell type, time, and space. In contrast, PPI detection assays are run in a single condition that may not even be an endogenous condition of the organism, resulting in static networks that do not reflect full cellular complexity. Thus, there is a need for computational methods to predict cell-type-specific interactions. Results:Here we present SPIDER (Supervised Protein Interaction DEtectoR), a graph attention-based model for predicting cell-type-specific PPI networks. In contrast to previous attempts at this problem, which were unsupervised in nature, our model's training is guided by experimentally measured cell-type-specific networks, enhancing its performance. We evaluate our method using experimental data of cell-type-specific networks from both humans and mice, and show that it outperforms current approaches by a large margin. We further demonstrate the ability of our method to generalize the predictions to datasets of tissues lacking prior PPI experimental data. We leverage the networks predicted by the model to facilitate the identification of tissue-specific disease genes. Availability and implementation:Our code and data are available at https://github.com/Kuper994/SPIDER.
Motivation:Technical differences between gene expression sequencing experiments can cause variations in the data in the form of batch effect biases. These do not represent true biological variations between samples and can lead to false conclusions or hinder the ability to integrate multiple datasets. Since there is a growing need for the joint analysis of single-cell sequencing datasets from different sources, there is also a need to correct the resulting batch effects while maintaining the true biological variations in the data.Results:We developed a semi-supervised deep learning architecture called Autoencoder-based Batch Correction (ABC) for integrating single-cell sequencing datasets. Our method removes batch effects through a guided process of data compression using supervised cell type classifier branches for biological signal retention. It aligns the different batches using an adversarial training approach. We comprehensively evaluate the performance of our method using four single-cell sequencing datasets and multiple measures for batch effect removal and biological variation conservation. ABC outperforms 10 state-of-the-art methods for this task including Seurat, scGen, ComBat, scanorama, scVI, scANVI, AutoClass, Harmony, scDREAMER, and CLEAR, correcting various types of batch effects while preserving intricate biological variations.
Executable models of biological circuits offer the ability to simulate their behavior under different settings with important biomedical applications. In particular, Boolean network models have been a prime research focus and dozens of manually curated Boolean models are available in public databases. A key challenge in studying the dynamics of these models is determining their asymptotic behavior, that is the state-sets or attractors they converge to. This is particularly challenging for large networks, as the state space size grows exponentially. Here we introduce a novel method for identifying stable components within attractors under an asynchronous update scheme. Our method leverages the observation that the majority of cellular functions in current models can be described as linear threshold functions, facilitating an efficient integer programming formulation for the problem. We conduct simulations on both synthetic and real biological networks, demonstrating that our proposed method is highly efficient and outperforms previous methods.
We hypothesized that via extracellular vesicles (EVs), chronic lymphocytic leukemia (CLL) cells turn endothelial cells into CLL-supportive cells. To test this, we treated vein-derived (HUVECs) and artery-derived (HAOECs) endothelial cells with EVs isolated from the peripheral blood of 45 treatment-naïve patients. Endothelial cells took up CLL-EVs in a dose- and time-dependent manner. To test whether CLL-EVs turn endothelial cells into IL-6-producing cells, we exposed them to CLL-EVs and found a 50% increase in IL-6 levels. Subsequently, we filtered out the endothelial cells and added CLL cells to this IL-6-enriched medium. After 15 min, STAT3 became phosphorylated, and there was a 40% decrease in apoptosis rate, indicating that IL-6 activated the STAT3-dependent anti-apoptotic pathway. Phospho-proteomics analysis of CLL-EV-exposed endothelial cells revealed 23 phospho-proteins that were upregulated, and network analysis unraveled the central role of phospho-β-catenin. We transfected HUVECs with a β-catenin-containing plasmid and found by ELISA a 30% increase in the levels of IL-6 in the culture medium. By chromatin immunoprecipitation assay, we observed an increased binding of three transcription factors to the IL-6 promoter. Importantly, patients with CLL possess significantly higher levels of peripheral blood IL-6 compared to normal individuals, suggesting that the inducers of endothelial IL-6 are the neoplastic EVs derived from the CLL cells versus those of healthy people. Taken together, we found that CLL cells communicate with endothelial cells through EVs that they release. Once they are taken up by endothelial cells, they turn them into IL-6-producing cells.
Clustering is a fundamental problem in data science with diverse applications in biology. The problem has many combinatorial and statistical variants, yet few allow clusters to overlap which is common in the biological domain. Recently, Bonchi et al. defined a new variant of the clustering problem, termed overlapping correlation clustering, which calls for multi-label cluster assignments that correlate with an input similarity between elements as much as possible. This variant is NP-hard and was solved by Bonchi et al. using a local search heuristic. We revisit this heuristic and develop exact integer-programming based variants for it. We show that these variants perform well across several datasets and evaluation measures.
Cell-cell crosstalk involves simultaneous interactions of multiple receptors and ligands, followed by downstream signaling cascades working through receptors converging at dominant transcription factors, which then integrate and propagate multiple signals into a cellular response. Single-cell RNAseq of multiple cell subsets isolated from a defined microenvironment provides us with a unique opportunity to learn about such interactions reflected in their gene expression levels. We developed the interFLOW framework to map the potential ligand-receptor interactions between different cell subsets based on a maximum flow computation in a network of protein-protein interactions (PPIs). The maximum flow approach further allows characterization of the intracellular downstream signal transduction from differentially expressed receptors towards dominant transcription factors, therefore, enabling the association between a set of receptors and their downstream activated pathways. Importantly, we were able to identify key transcription factors toward which the convergence of multiple receptor signaling pathways occurs. These identified factors have a unique role in the integration and propagation of signaling following specific cell-cell interactions.