Multi-modal data-based algorithms have gained attention in their capabilities in prediction tasks in cancer-related research. This paper introduces SPACT, a multi-modal capable of predicting cancer survival probability based on a deep-learning application on the histopathological patch features, whole-slide histopathology (WSI) in particular. What SPACT improves upon compared to the existing survival-prediction models is aiming to improve the robustness/resilience and the prediction accuracy, both overall and in subcategories. It does that by instead of relying only on the TCGA-based datasets, it benefits from the use of an external dataset collected from the Başkent Hospital. By doing cross-comparison on the different encoders’ accuracy on both datasets, an optimal encoder that performs well in both the conventional TCGA and newly gathered Başkent Hospital dataset is chosen. Using a combination of multiple encoders on different models, we find that encoders that perform well on both datasets outperform the encoders that work well on a single one. Additionally, in-depth analyses on the ablation studies, risk stratification, attention maps, and on how integrated gradients may have an effect on the performance and reasoning of SPACT. SPACT matches or outperforms all the other state-of-the-art multi-modal prediction algorithms in 5 of the 7 different cancer types, particularly in the case of ovarian cancer, by a wide margin, with a c-index score of 0.77. It has the highest robustness, having the best overall performance on all 3 encoders the models have been tested with. These findings solidify SPACT’s position as a contemporary and improved multi-modal based deep learning model targeting individual prediction accuracy, overall prediction accuracy, and robustness under diverse conditions. The code for SPACT is available at https://github.com/ezgiogulmus/SPACT.
Quantum annealing (QA) has great potential to solve combinatorial optimization problems efficiently. However, the effectiveness of QA algorithms is heavily based on the embedding of problem instances, represented as logical graphs, into the quantum processing unit (QPU) whose topology is in the form of a limited connectivity graph, known as the minor embedding problem. Because the minor embedding problem is an NP-hard problem [11], existing methods for the minor embedding problem suffer from scalability issues when faced with larger problem sizes. In this article, we propose a novel approach utilizing Reinforcement Learning (RL) techniques to address the minor embedding problem, named CHARME. CHARME includes three key components: a Graph Neural Network (GNN) architecture for policy modeling, a state transition algorithm that ensures solution validity, and an order exploration strategy for effective training. Through comprehensive experiments on synthetic and real-world instances, we demonstrate the efficiency of our proposed order exploration strategy as well as our proposed RL framework, CHARME. In particular, CHARME yields superior solutions in terms of qubit usage compared to fast embedding methods such as Minorminer and ATOM. Moreover, our method surpasses the OCT-based approach, known for its slower runtime but high-quality solutions, in several cases. In addition, our proposed exploration enhances the efficiency of the training of the CHARME framework by providing better solutions compared to the greedy strategy.
Quantum computing has the potential to revolutionize fields like quantum optimization and quantum machine learning. However, current quantum devices are hindered by noise, reducing their reliability. A key challenge in gate-based quantum computing is improving the reliability of quantum circuits, measured by process fidelity, during the transpilation process, particularly in the routing stage. In this article, we address the Fidelity Maximization in Routing Stage (FMRS) problem by introducing FIDDLE, a novel learning framework comprising two modules: a Gaussian Process-based surrogate model to estimate process fidelity with limited training samples and a reinforcement learning module to optimize routing. Our approach is the first to directly maximize process fidelity, outperforming traditional methods that rely on indirect metrics such as circuit depth or gate count. We rigorously evaluate FIDDLE by comparing it with state-of-the-art fidelity estimation techniques and routing optimization methods. The results demonstrate that our proposed surrogate model is able to provide abetter estimation on the process fidelity compared to existing learning techniques, and our end-to-end framework significantly improves the process fidelity of quantum circuits across various noise models.
Biological networks, characterized by complex interactions among genes, proteins, and metabolites, are often modeled as graphs to study their organizational principles and dynamics. Network motifs-recurring, statistically significant subgraphs provide critical insights into the functional properties and structural organization of these networks. Existing research has explored various facets of motif detection, including enumeration, edge and node independence, dynamic updates, multi-layered networks, and stochastic settings. Although there have been studies in exploring the functionality of individual motif instances, a significant gap remains in understanding how collections of motif instances act as a group to influence the overall functionality of the network. In this paper, we aim to fill this gap. We model the influence of collections of motifs as a novel problem, which we call the Closest $k$-Motif Set Selection problem. We prove that this problem is NP-hard and develop MOSAIC (MOtif Set with mAximal InfluenCe), a novel greedy algorithm to address this problem. MOSAIC operates in two phases: an initialization phase for computing distances between motifs and nodes, and an update phase that incrementally selects motif instances to optimize their collective impact on the network. We prove that MOSAIC is efficient with a low degree polynomial time complexity. Our experimental results demonstrate that MOSAIC achieves optimal or near-optimal results, and scales to the entire human network efficiently. Our experiments on the human transcriptional regulatory network demonstrate that MOSAIC can identify Alzheimer's genes effectively, and select Alzheimer's genes that are missed by state-of-the-art node-based selection methods. This work advances our understanding of motif-based network analysis and opens new avenues for exploring the functional implications of network motifs.
Classic RNA sequencing dissociates cells from their native tissue architecture, discarding spatial information that critically shapes transcriptional programs in development, homeostasis, and cancer. However, current ST platforms often produce incomplete and noisy profiles due to technical limitations and tissue variability. These limitations obscure biologically meaningful spatial patterns and hinder downstream interpretation. Here, we introduce STORM (spatial transcriptomics optimization by resolution via matrix factorization), a machine learning framework that improves the fidelity of spatial transcriptomics data under severe sparsity. STORM formulates spatial transcriptomics recovery as a low-rank tensor decomposition problem and integrates multimodal biological priors through a principled regularization strategy. Specifically, the model jointly captures spatial continuity, tissue morphology derived from whole-slide histology images, and gene-gene interaction structure informed by protein-protein interaction networks. This method enables accurate reconstruction at unobserved locations while preserving biologically meaningful spatial structure. Across diverse lung tissue profiles, including both healthy and malignant samples, STORM consistently outperforms existing state-of-the-art methods in recovering spatial gene-expression patterns and remains robust even when a majority of spatial measurements are missing. By explicitly embedding biological structure into the reconstruction process, STORM provides a reliable foundation for high-resolution spatial transcriptomic analysis in settings where experimental data are sparse or incomplete. Availability: The source code developed in this study is publicly available at https://github.com/denizgurarslan/STORM.
Crop models are widely used as decision support tools in agriculture and natural resource management. However, the current practices in crop modeling and evaluation vary significantly, making it challenging to compare uncertainties across different models and studies. This study aims to quantify and compare the uncertainties associated with four different crop modeling practices using a standardized evaluation framework. The four modeling practices are: 1) using a single model, considering only model bias for uncertainty; 2) using a single model, accounting for both model bias and parameter uncertainty; 3) employing a multi-model ensemble to account for model bias, parameter uncertainty, and structural uncertainties; and 4) using a multi-model ensemble to consider uncertainties induced by model bias, parameters, structures, and inputs. We developed a framework that integrates Markov Chain Monte Carlo (MCMC) and Bayesian Model Averaging (BMA) to consistently quantify prediction uncertainty across these practices. The framework was applied to an Asian rice dataset. The results revealed that common model evaluation approach (Practice 1) tends to underestimate the uncertainty of model predictions. Relying on a single process-based model presents a substantial risk in critical decision-making situations. In contrast, the BMA ensemble predictor (e-BMA) demonstrated higher reliability, making it a preferable choice for future decision support. Our Bayesian framework provides a more robust and adaptable approach for project-specific decision-making, with promising applications in digital agriculture.
Biological networks are dynamic structures, continuously evolving by rewiring their interactions. These rewirings happen at different rates for different cells, and the rates can change over time, yet we can only observe the cell at a limited number of stages of their evolution, limiting the number of possible observed gene networks. In this paper, we consider the problem of predicting entire gene networks of dynamic biological networks. We develop a novel algorithm PLATO (Predicting Longitudinally-Aligned Time Observations), which utilizes dynamic network alignment that maps multiple systems of networks to improve the prediction accuracy of the matrix factorization model. We evaluate our method on gene-gene interaction networks using a mouse model with evolutionary patterns caused by chronic myeloid leukemia (CML) and compare it to four existing state of the art methods, including two deep learning and one matrix factorization techniques. Our experimental results demonstrate that PLATO outperforms both traditional matrix factorization and other competing methods in terms of gene-gene interaction prediction accuracy.
Drug resistance is one of the fundamental challenges in modern medicine. Using combinations of drugs is an effective solution to counter drug resistance as is harder to develop resistance to multiple drugs simultaneously. Finding the correct dosage for each drug in the combination remains to be a challenging task. Testing all possible drug-drug combinations on various cell lines for different dosages in wet-lab experiments is infeasible since there are many combinations of drugs as well as their dosages yet the drugs and the cell lines are limited in availability and each wet-lab test is costly and time-consuming. Efficient and accurate in silico prediction methods are surely needed. Here we present a novel computational method, PartialFibers, to address this challenge. Unlike existing prediction methods PartialFibers takes advantage of the distribution of the missing drug-drug interactions and effectively predicts the dosage of a drug in the combination. Our results on real datasets demonstrate that PartialFibers is more flexible, scalable, and achieves higher accuracy in less time than the state of the art algorithms.
Sex differences appear in healthy and pathological conditions and may influence sex-specific therapeutic responses. Understanding such differences is a key activity for developing precision medicine strategies. This study investigates sex differences in gene expression across 40 human tissues by applying a Differential Causal Network (DCN) analysis using data from the Genotype-Tissue Expression project. We identified sex-based DCNs that highlight distinct molecular mechanisms influencing both health and disease in men and women. For example, in pancreas tissue, genes associated with immune system show significant differences in their regulatory patterns between sexes, demonstrating a possible different response to diseases such as diabetes mellitus and cancer. Our findings provide valuable information on the biological underpinnings of sex differences, offering potential pathways for the development of precision medicine strategies.
Lung cancer is one of the most common and deadly cancers worldwide. Accurate survival prediction is critical for guiding treatment, yet existing deep learning approaches often struggle with capturing the complexity of histological and tabular features and fusing them effectively. We address these challenges by introducing a novel Kolmogorov-Arnold Network for tabular and fusion tasks, combined with advanced vision models for histology image processing. Experiments show that our method achieves superior survival prediction accuracy compared to unimodal predictors. Furthermore, it provides explainable predictions, as 10 of the top 20 genes identified as most influential are known to play roles in cancer survival and progression.
MOTIVATION:Targeted enrichment via capture probes, also known as baits, is a promising complementary procedure for next-generation sequencing methods. This technique uses short biotinylated oligonucleotide probes that hybridize with complementary genetic material in a sample. Following hybridization, the target fragments can be easily isolated and processed with minimal contamination from irrelevant material. Designing an efficient set of baits for a set of target sequences, however, is an NP-hard problem. RESULTS:We develop a novel heuristic algorithm that leverages the similarities between the characteristics of the Minimum Bait Cover and the Closest String problems to reduce the number of baits to cover a given target sequence. Our results on real and synthetic datasets demonstrate that our algorithm, OLTA produces fewest baits for nearly all experimental settings and datasets. On average, it produces 6% and 11% fewer baits than the next best state-of-the-art methods for two major real datasets, AIV and MEGARES. Also, its bait set has the highest utilization and the minimum redundancy. AVAILABILITY AND IMPLEMENTATION:Our algorithm is available at github.com/FuelTheBurn/OLTA-Optimizing-bait-seLection-for-TArgeted-sequencing. Test data and other software are archived at doi.org/10.5281/zenodo.15086636.
Drug resistance, the decrease in the effectiveness of a medication over time, is a major global threat for public health as it makes it harder and more expensive to fight against diseases, harmful microbial species such as bacteria and viruses. It is established that genes play a significant role in sensitivity to drugs. In this paper, we address the problem of establishing causality between transcription patterns of genes and drug resistance. Class separation based models can be used to provide an explainable solution for the causality definition for drug resistance. However, for $m$ samples and $n$ genes, the time and space complexities of the class separation problem are, respectively $O\left(m^{2} n^{2}\right)$ and $O\left(n^{2}\right)$ making it too costly to study this problem at whole genome scale. We develop an efficient implementation of the class separation model, named Hierarchical Class Separation Transformation (HCST), which solves this problem in $O\left(h n m^{2} k\right)$ time, where $k$ and $h$ are user controlled parameters indicating the partition size for the gene set and gene set mixing limit, with $h k \ll n$, and space $O\left(k m+k^{2}\right)$. HCST allows solving the class separation problem at entire human genome scale in an efficient way, scaling in an efficient way (i.e., less than 2 minutes of running time). Our results demonstrate that HCST is scalable, robust, and can accurately identify genes which affect drug resistance. Code developed in this paper is available at https://github.com/richiebailey74/HCST.
Biological networks are dynamic structures. They continuously evolve by rewiring their interactions. These rewirings happen at different rates for different cells, and the rates can change over time, yet we can only observe the cell at a limited number of stages of their evolution. In this paper, we consider the problem of determining evolutionary trajectories of dynamic biological networks. We develop a novel algorithm DANTE (Determining Adaptation trajectories in biological Networks Through Evolutionary mapping), which maps multiple cellular network evolution patterns in accordance with the greatest possible similarity. We evaluate our method on protein-protein interaction (PPI) networks using a mouse model with evolutionary patterns caused by chronic myeloid leukemia (CML) and compare it to four alternative strategies. Our experimental results demonstrate that DANTE outperforms competing methods in terms of trajectory similarity, and the advantages of DANTE over competing methods grow when network trajectories are incomplete.
Understanding the genetic components of Alzheimer's disease (AD) via transcriptome analysis often necessitates the use of invasive methods. This work focuses on overcoming the difficulties associated with the invasive process of collecting brain tissue samples in order to measure and investigate the transcriptome behavior of AD. Our approach called IDEEA (Information D iffusion model for integrating gene E xpression and E EG data in identifying A lzheimer's disease markers) involves systematically linking two different but complementary modalities: transcriptomics and electroencephalogram (EEG) data. We preprocess these two data types by calculating the spectral and transcriptional sample distances, over 11 brain regions encompassing 6 distinct frequency bands. Subsequently, we employ a genetic algorithm approach to integrate the distinct features of the preprocessed data. Our experimental results show that IDEEA converges rapidly to local optima gene subsets, in fewer than 250 iterations. Our algorithm identifies novel genes along with genes that have previously been linked to AD. It is also capable of detecting genes with transcription patterns specific to individual EEG bands as well as those with common patterns among bands. In particular, the alpha2 (10-13 Hz) frequency band yielded 8 AD-associated genes out of the top 100 most frequently selected genes by our algorithm, with a p-value of 0.05. Our method not only identifies AD-related genes but also genes that interact with AD genes in terms of transcription regulation. We evaluated various aspects of our approach, including the genetic algorithm performance, band-pair association and gene interaction topology. Our approach reveals AD-relevant genes with transcription patterns inferred from EEG alone, across various frequency bands, avoiding the risky brain tissue collection process. This is a significant advancement toward the early identification of AD using non-invasive EEG recordings.
Morphological profiling is a valuable tool in phenotypic drug discovery. The advent of high-throughput automated imaging has enabled the capturing of a wide range of morphological features of cells or organisms in response to perturbations at the single-cell resolution. Concurrently, significant advances in machine learning and deep learning, especially in computer vision, have led to substantial improvements in analyzing large-scale high-content images at high-throughput. These efforts have facilitated understanding of compound mechanism-of-action (MOA), drug repurposing, characterization of cell morphodynamics under perturbation, and ultimately contributing to the development of novel therapeutics. In this review, we provide a comprehensive overview of the recent advances in the field of morphological profiling. We summarize the image profiling analysis workflow, survey a broad spectrum of analysis strategies encompassing feature engineering- and deep learning-based approaches, and introduce publicly available benchmark datasets. We place a particular emphasis on the application of deep learning in this pipeline, covering cell segmentation, image representation learning, and multimodal learning. Additionally, we illuminate the application of morphological profiling in phenotypic drug discovery and highlight potential challenges and opportunities in this field.
Network motif identification problem aims to find topological patterns in biological networks. Identifying non-overlapping motifs is a computationally challenging problem using classical computers. Quantum computers enable solving high complexity problems which do not scale using classical computers. In this paper, we develop the first quantum solution, called QOMIC (Quantum Optimization for Motif IdentifiCation), to the motif identification problem. QOMIC transforms the motif identification problem using a integer model, which serves as the foundation to develop our quantum solution. We develop and implement the quantum circuit to find motif locations in the given network using this model. Our experiments demonstrate that QOMIC outperforms the existing solutions developed for the classical computer, in term of motif counts. We also observe that QOMIC can efficiently find motifs in human regulatory networks associated with five neurodegenerative diseases: Alzheimers, Parkinsons, Huntingtons, Amyotrophic Lateral Sclerosis (ALS), and Motor Neurone Disease (MND).
Drug resistance is the drop in the effectiveness of a medication over time, and has severe consequences for cancer patients, as the use of incorrect drugs from the onset of the disorder not only wastes valuable treatment time for the rapidly advancing cancer types but also weakens the natural defense mechanism of the patients. The high cost of wet-lab experiments necessitates cheap computational methods for drug resistance prediction. Although many different approaches have been developed over the years to perform drug response prediction, few studies focus on the reliability of these predictions. In this paper, we develop the first framework, named VultuRe (VULnerabilities in impuTing drUg REsistance), to identify the vulnerabilities in drug resistance prediction. Our results demonstrate that the success of drug resistance imputation might be different for each drug and VultuRe efficiently detects the vulnerabilities in drug response imputation, suggesting alternatives to overcome wrong predictions.
Target Identification by Enzymes (TIE) problem aims to identify the set of enzymes in a given metabolic network, such that their inhibition eliminates a given set of target compounds associated with a disease while incurring minimum damage to the rest of the compounds. This is an NP-complete problem, and thus optimal solutions using classical computers fail to scale to large metabolic networks. In this paper, we consider the TIE problem for identifying drug targets in metabolic networks. We develop the first quantum optimization solution, called QuTIE (Quantum optimization for Target Identification by Enzymes), to this NP-complete problem. We do that by developing an equivalent formulation of the TIE problem in Quadratic Unconstrained Binary Optimization (QUBO) form, then mapping it to a logical graph, which is then embedded on a hardware graph on a quantum computer. Our experimental results on 27 metabolic networks from Escherichia coli, Homo sapiens, and Mus musculus show that QuTIE yields solutions which are optimal or almost optimal. Our experiments also demonstrate that QuTIE can successfully identify enzyme targets already verified in wet-lab experiments for 14 major disease classes.
Alin Dobra合作论文数Department of Computer & Information Science & Engineering
University of Florida18
Martin Theobald合作论文数Institut fur Datenbanken und Informationssysteme12