Diagnostic tools for rare diseases typically rely on curated gene-phenotype associations and static disease models, limiting their effectiveness in cases with atypical presentations or previously uncharacterized disorders. To address these limitations, we present SimPheny, a phenotype-first algorithm for gene prioritization that operates independently of documented gene-phenotype associations. SimPheny identifies phenotypically similar diagnosed patients by comparing an undiagnosed patient's disease presentation to a reference cohort of diagnosed cases, and returns gene hypotheses by matching the undiagnosed patient's candidate gene list to the causative genes of similar patients using a statistical scoring model. Evaluated in diagnosed probands from the Undiagnosed Diseases Network (UDN) with the true diagnostic gene blinded, SimPheny consistently ranked the diagnostic gene among the top five candidates, outperforming existing tools, particularly for genes with limited gene-phenotype association data. When applied to previously unsolved UDN cases, clinical review confirmed that SimPheny's high-confidence causative gene predictions were diagnostic in nearly half of the analyzed cases. As the size of the diagnosed reference cohort increases, SimPheny's diagnostic reach expands without sacrificing ranking performance. By leveraging real patient data rather than curated models, SimPheny provides a generalizable, scalable framework for improving diagnostic yield in rare disease cohorts.
BACKGROUND:Exome sequencing (ES) and genome sequencing (GS) are increasingly used as standard genetic tests to identify diagnostic variants in rare disease cases. However, prioritizing these variants to reduce the time and burden of manual interpretation by clinical teams remains a significant challenge. The Exomiser/Genomiser software suite is the most widely adopted open-source software for prioritizing coding and noncoding variants. Despite its ubiquitous use, limited data-driven guidelines currently exist to optimize its performance for diagnostic variant prioritization. Based on detailed analyses of Undiagnosed Diseases Network (UDN) probands, this study presents optimized parameters and practical recommendations for deploying the Exomiser and Genomiser tools. We also highlight scenarios where diagnostic variants may be missed and propose alternative workflows to improve diagnostic success in such complex cases. METHODS:We analyzed 386 diagnosed probands from the UDN, including cases with coding and noncoding diagnostic variants. We systematically evaluated how tool performance was affected by key parameters, including gene:phenotype association data, variant pathogenicity predictors, phenotype term quality and quantity, and the inclusion and accuracy of family variant data. RESULTS:Parameter optimization significantly improved Exomiser's performance over default parameters. For GS data, the percentage of coding diagnostic variants ranked within the top 10 candidates increased from 49.7% to 85.5%, and for ES, from 67.3% to 88.2%. For noncoding variants prioritized with Genomiser, the top 10 rankings improved from 15.0% to 40.0%. We also explored refinement strategies for Exomiser outputs, including using p-value thresholds and flagging genes that are frequently ranked in the top 30 candidates but rarely associated with diagnoses. CONCLUSION:This study provides an evidence-based framework for variant prioritization in ES and GS data using Exomiser and Genomiser. These recommendations have been implemented in the Mosaic platform to support the ongoing analysis of undiagnosed UDN participants and provide efficient, scalable reanalysis to improve diagnostic yield. Our work also highlights the importance of tracking solved cases and diagnostic variants that can be used to benchmark bioinformatics tools. Exomiser and Genomiser are available at https://github.com/exomiser/Exomiser/ .
Rapid genomic diagnostics in the Neonatal Intensive Care Unit represents a paradigm shift in medicine with increasing evidence of the utility of early diagnosis, impacting management. The goal of the Utah NeoSeq Project was to implement and evaluate a multidisciplinary and longitudinal rapid sequencing program while transitioning to CLIA-certified sequencing. Enrollment of 65 infants resulted in 26 (40%) with a diagnostic variant(s) and 7 (11%) harboring a strong candidate. This includes re-analyses resulting in four additional diagnoses. Parental surveys indicated that 7% (4/59) of parents had a decisional conflict after consent, and 3% (2/59) experienced decisional regret after the results. Fifty-two provider surveys were conducted. Seventy-nine percent (41/52) of results and 86% (19/22) of diagnostic results were “very useful” or “useful” and associated with management changes. The NeoSeq Project demonstrates that a multidisciplinary collaborative approach to diagnosis is feasible. We have developed a generalizable, collaborative protocol that addresses the need for expedited genetic evaluation with emerging technologies.
The enteric nervous system (ENS) is a complex network of neurons and glial cells. Hirschsprung's disease (HSCR) is a congenital condition characterized by the absence of ganglion cells in the distal colon, leading to functional bowel obstruction. In this study, we used single-cell RNA sequencing (scRNA-seq) and whole genome sequencing (WGS) to analyze healthy and aganglionic colon segments from HSCR patients. Using scRNA-seq, we identified 13 major cell types in patient samples and observed that neural progenitor cells were present in both healthy and aganglionic colon regions, while mature neurons were absent from aganglionic colon. In these progenitor cells, critical differentiation pathway genes displayed reduced expression in the aganglionic colon, suggesting a disruption in their transition to mature neuronal cell types. Furthermore, transcriptomic analysis revealed significant alterations in gene expression across several stromal cell types. These transcriptomic shifts, particularly in mast cells, support the hypothesis that altered gene expression in the microenvironment of neural progenitor cells contributes to impaired differentiation. Our findings support the hypothesis that neural precursors in HSCR are capable of migration, but they are defective in their differentiation to mature cell types. Our analysis provides insights into potential therapeutic targets to stimulate neurogenesis in the aganglionic colon.
Bruton's tyrosine kinase (BTK) inhibitors are effective for the treatment of chronic lymphocytic leukemia (CLL) due to BTK's role in B cell survival and proliferation. Treatment resistance is most commonly caused by the emergence of the hallmark BTKC481S mutation that inhibits drug binding. In this study, we aimed to investigate whether the presence of additional CLL driver mutations in cancer subclones harboring a BTKC481S mutation accelerates subclone expansion. In addition, we sought to determine whether BTK-mutated subclones exhibit distinct transcriptomic behavior when compared to other cancer subclones. To achieve these goals, we employ our recently published method (Qiao et al. 2024) that combines bulk DNA sequencing and single-cell RNA sequencing (scRNA-seq) data to genotype individual cells for the presence or absence of subclone-defining mutations. While the most common approach for scRNA-seq includes short-read sequencing, transcript coverage is limited due to the vast majority of the reads being concentrated at the priming end of the transcript. Here, we utilized MAS-seq, a long-read scRNAseq technology, to substantially increase transcript coverage across the entire length of the transcripts and expand the set of informative mutations to link cells to cancer subclones in six CLL patients who acquired BTKC481S mutations during BTK inhibitor treatment. We found that BTK-mutated subclones often acquire additional mutations in CLL driver genes, leading to faster subclone proliferation. When examining subclone-specific gene expression, we found that in one patient, BTK-mutated subclones are transcriptionally distinct from the rest of the malignant B cell population with an overexpression of CLL-relevant genes.
Genetic and gene expression heterogeneity is an essential hallmark of many tumors, allowing the cancer to evolve and to develop resistance to treatment. Currently, the most commonly used data types for studying such heterogeneity are bulk tumor/normal whole-genome or whole-exome sequencing (WGS, WES); and single-cell RNA sequencing (scRNA-seq), respectively. However, tools are currently lacking to link genomic tumor subclonality with transcriptomic heterogeneity by integrating genomic and single-cell transcriptomic data collected from the same tumor. To address this gap, we developed scBayes, a Bayesian probabilistic framework that uses tumor subclonal structure inferred from bulk DNA sequencing data to determine the subclonal identity of cells from single-cell gene expression (scRNA-seq) measurements. Grouping together cells representing the same genetically defined tumor subclones allows comparison of gene expression across different subclones, or investigation of gene expression changes within the same subclone across time (i.e., progression, treatment response, or relapse) or space (i.e., at multiple metastatic sites and organs). We used simulated data sets, in silico synthetic data sets, as well as biological data sets generated from cancer samples to extensively characterize and validate the performance of our method, as well as to show improvements over existing methods. We show the validity and utility of our approach by applying it to published data sets and recapitulating the findings, as well as arriving at novel insights into cancer subclonal expression behavior in our own data sets. We further show that our method is applicable to a wide range of single-cell sequencing technologies including single-cell DNA sequencing as well as Smart-seq and 10x Genomics scRNA-seq protocols.
Precision oncology matches tumors to targeted therapies based on the presence of actionable molecular alterations. However, most tumors lack actionable alterations, restricting treatment options to cytotoxic chemotherapies for which few data-driven prioritization strategies currently exist. Here, we report an integrated computational/experimental treatment selection approach applicable for both chemotherapies and targeted agents irrespective of actionable alterations. We generated functional drug response data on a large collection of patient-derived tumor models and used it to train ScreenDL, a novel deep learning-based cancer drug response prediction model. ScreenDL leverages the combination of tumor omic and functional drug screening data to predict the most efficacious treatments. We show that ScreenDL accurately predicts response to drugs with diverse mechanisms, outperforming existing methods and approved biomarkers. In our preclinical study, this approach achieved superior clinical benefit and objective response rates in breast cancer patient-derived xenografts, suggesting that testing ScreenDL in clinical trials may be warranted.
Abstract Precision oncology hinges on accurate prediction of patient-specific treatment response from tumor molecular inputs. While deep learning (DL) models achieve state-of-the-art response prediction in cell lines, existing methods do not readily translate to the clinic where training data is limited. Here, we have developed and validated ScreenDL, a novel DL framework designed explicitly for use in clinical precision oncology applications. The underlying architecture of ScreenDL consists of two fully connected branches dedicated to learning rich embeddings of a drug’s chemical structure and a tumor’s transcriptomic profile respectively. These tumor omic and drug molecular embeddings are fed into a shared response prediction subnetwork which outputs predicted ln(IC50) values. The ScreenDL training schema incorporates three phases, each designed to improve clinical response prediction: (1) an initial pretraining phase leverages large-scale cell line pharmacogenomic datasets to extract patterns/relationships linking tumor omics and drug response; (2) subsequent transfer learning integrates pharmacogenomic data from patient-derived models of cancer, adapting the pretrained model to a more patient-relevant context; (3) patient-level omics and drug screening data (e.g., from matched patient-derived organoid (PDO) models) is integrated through patient-specific fine-tuning, generating a patient-specific response prediction model. Importantly, biomaterial from surgical tumor resection is often available for organoid establishment and patient-specific drug screening in practical clinical scenarios. To assess the utility of our ScreenDL framework for clinical response prediction, we applied ScreenDL to predict treatment response in 50 advanced/metastatic triple negative breast cancer (TNBC) patient-derived xenograft models (PDX). We leveraged tumor omics profiles for each PDX and drug screening data from matched PDX-derived organoids (PDxOs), mirroring the combination of patient-level tumor omic characterization and functional drug screening in matched PDOs, a protocol currently being piloted in Functional Precision Oncology trials at the University of Utah Huntsman Cancer Institute. After cell line pretraining, ScreenDL achieves a modest median Pearson correlation per drug of 0.15 relative to 0.03 for the leading compared model. However, transfer learning and subsequent PDX-specific fine-tuning significantly improve performance, producing median Pearson correlations per drug of 0.39 and 0.51, respectively. These results demonstrate the ability of our ScreenDL framework to dramatically improve response prediction in clinically relevant domains, bringing DL-based precision treatment selection closer to clinical application. Citation Format: Casey Sederman, Tonya Di Sera, Yi Qiao, Xiaomeng Huang, Bryan E. Welm, Alana L. Welm, Gabor Marth. ScreenDL: A transfer learning framework integrating tumor omics and functional drug screening for personalized clinical drug response prediction [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 907.
A mechanistic understanding of the biological and technical factors that impact transcript measurements is essential to designing and analyzing single-cell and single-nucleus RNA sequencing experiments. Nuclei contain the same pre-mRNA population as cells, but they contain a small subset of the mRNAs. Nonetheless, early studies argued that single-nucleus analysis yielded results comparable to cellular samples if pre-mRNA measurements were included. However, typical workflows do not distinguish between pre-mRNA and mRNA when estimating gene expression, and variation in their relative abundances across cell types has received limited attention. These gaps are especially important given that incorporating pre-mRNA has become commonplace for both assays, despite known gene length bias in pre-mRNA capture. Here, we reanalyze public data sets from mouse and human to describe the mechanisms and contrasting effects of mRNA and pre-mRNA sampling on gene expression and marker gene selection in single-cell and single-nucleus RNA-seq. We show that pre-mRNA levels vary considerably among cell types, which mediates the degree of gene length bias and limits the generalizability of a recently published normalization method intended to correct for this bias. As an alternative, we repurpose an existing post hoc gene length-based correction method from conventional RNA-seq gene set enrichment analysis. Finally, we show that inclusion of pre-mRNA in bioinformatic processing can impart a larger effect than assay choice itself, which is pivotal to the effective reuse of existing data. These analyses advance our understanding of the sources of variation in single-cell and single-nucleus RNA-seq experiments and provide useful guidance for future studies.
Introduction: BTK inhibition is an effective frontline therapy for patients with CLL. However, patients who develop BTK C481S mutations relapse quickly. We recently published a longitudinal study investigating subclonal evolution in 38 patients receiving BTKi treatment for CLL. Of the 38 patients, 6 developed a subclone containing a BTK C481S mutation identified in the whole-exome sequencing (WES) data at the time of relapse. To understand whether a BTK C481S mutation alone can drive relapse or if an additional CLL-relevant mutation is required, we investigate these subclones on a single-cell level using PacBio Multiplexed Arrays Sequencing (MAS-seq). MAS-seq generates single-cell, long-read RNA sequencing data, which is highly accurate and provides coverage across the entire mRNA transcript. Methods: For each patient, B cells were isolated from a pre-BTKi treatment sample and a sample taken at the time of relapse. Single-cell cDNA library preps were generated using the 10X Genomics Chromium platform. The cDNA was then sequenced using MAS-Seq. After processing the MAS-seq data using PacBio's Isoseq pipeline, Seurat was used to perform a single-cell analysis for each sample. We used scBayes, a Bayesian-statistical approach developed in our lab (Qiao et al., Genome Research, in review), to assign subclone identity to individual cells based on the expression of previously identified somatic mutations. The transcript-wide coverage of MAS-seq increases the likelihood of coverage at the somatic mutation sites compared to traditional single-cell RNA sequencing methods. Results: We have sequenced and analyzed samples from 4 patients in this study (pt 1-4). In the pretreatment sample of pt 1, two subclones containing different TP53 mutations were present. At the time of relapse, a BTK C481S subclone evolved from one of these TP53 subclones, becoming the dominant clonal population. No other CLL-relevant mutations were detected in this patient. Pts 2-4 had an additional CLL-relevant mutation evolve with the BTK C481S mutation in the dominant subclone at the time of relapse. In pt 2, a subclone with an NFKBIE and the BTK mutation evolved from a cell population where BRAF, BCOR, and EGFR2 mutations were present. The BTK C481S subclone in Pt 3 also had new KMT2D and CREBBP mutations. This subclone evolved from a cell population with an ASXL1 and SF3B1 mutation. In pts 1-3, the BTK subclone was only seen at the time of relapse. Pt 4 did not have any CLL-relevant mutations identified before treatment. In a sample taken 6 months before relapse, a BTK C481S subclone is visible at a low frequency with no other CLL-relevant mutations. At the time of relapse, a DICER1 variant, a gene recently implicated in CLL (Knisbacher et al., Nature Genetics, 2022), evolved from the BTK subclone and is seen at the time of relapse as the dominant subclone. Furthermore, the full-transcript sequencing at the single-cell level enabled us to identify a second BTK C481S subclone present at the time of relapse and determine that it was independent of the subclone containing both the BTK C481S and DICER1 mutations. This second BTK subclone had no other CLL-relevant mutations and remained at a very low frequency. Conclusion: Single-cell long-read RNA sequencing via MAS-seq provides transcript-wide coverage, which enables genotyping at somatic mutation sites. Using this method, we found that the dominant relapse clones all contained a BTK C481S mutation, which either co-evolved with additional CLL-relevant mutations (3/4 patients) or occurred on the background of existing TP53 mutations (1/4 patients).
ABSTRACTMotivationIn time-critical clinical settings, such as precision medicine, genomic data needs to be processed as fast as possible to arrive at data-informed treatment decisions in a timely fashion. While sequencing throughput has dramatically increased over the past decade, bioinformatics analysis throughput has not, and consequently has now turned into the primary bottleneck. Modern computational hardware are capable of much higher performance than current genomic informatics algorithms can typically utilize, therefore presenting opportunities for significant improvement of performance. Accessing the raw sequencing data from BAM files, for example, is a necessary and time-consuming step in nearly all sequence analysis tools, however existing programming libraries for BAM access do not take full advantage of the parallel input/output capabilities of storage devices.ResultsIn an effort to stimulate the development of a new generation of faster sequence analysis tools, We developed quickBAM, a software library to accelerate sequencing data access by exploiting the parallelism in commodity storage hardware currently widely available. We demonstrate that analysis software ported to quickBAM consistently outperforms their current versions, in some cases finishing an analysis in under 4 minutes while the original version took 1.5 hours, using the same storage solution.Availability and ImplementationOpen source and freely available athttps://gitlab.com/yiq/quickbam/, we envision that quickBAM will enable a new generation of high performance informatics tools, either directly boosting their performance if they are currently dataaccess bottlenecked, or allow data-access to keep up with further optimizations in algorithms and compute techniques.Contactyi.qiao@genetics.utah.edu.
In many cancer types, the evolution of subclonal malignant cells leads to a diverse tumor population that can affect treatment efficacy. In addition, evolution throughout treatment can lead to resistant subclonal populations and eventual relapse. To effectively treat these cancers, it is essential to understand the subclonal architecture of the tumor, including the somatic mutations that differentiate each population. Many tools have been developed to attempt to cluster mutations or reconstruct the clonal architecture of a given tumor sample. However, few can accommodate multiple samples, which is common in longitudinal or metastasis studies. In addition, these tools typically focus on just one of the steps necessary for a full analysis (e.g., variant clustering, subclone hierarchical tree reconstruction, and tree visualization). Many require non-standard data formats, necessitating data format conversion between each step. These requirements make it very challenging for investigators without a deep understanding of each tool to perform a full subclonal analysis. To make the identification of temporal and spatial heterogeneity more streamlined and accessible to analysts with basic programming knowledge, we have developed a pipeline that takes somatic mutations in standard VCF file format as input and produces meaningful results that are easy to interpret. Variant data is extracted and formatted for the input to PyClone-VI (Gillis and Roth, 2020), which is used to cluster somatic variants based on their allele frequencies at each time point. SuperSeeker, an improvement on the SubcloneSeeker method developed within our lab (Qiao et al., 2014), implements advanced tumor subclone reconstruction algorithms to jointly analyze multiple samples and construct hierarchical trees that account for all subclones observed in a patient. A sample trace, or the cellular fraction of each subclone found in each sample, is also provided. The output of SuperSeeker is an updated VCF file with the tree and sample trace information added to the header. Finally, a GraphViz rendering of the hierarchical tree is made for easy visualization. In addition to this rendering, the final VCF file can also be used for interactive analysis and visualization of the results using our Oncogene.iobio web tool. This pipeline makes identifying temporal and spatial heterogeneity more efficient, requiring only standard file types as input, and basic command line knowledge. In a recently published study, we investigated subclonal evolution in 38 patients with CLL being treated with a BTKi. The initial subclonal architecture analysis for each patient in this study was laborious and time intensive. Using this pipeline, we can now identify subclonal evolution in these patients and create an interactive visualization more efficiently. This approach can streamline analyses, increase efficiency, and lead to a deeper understanding of a tumor’s subclonal evolution. Citation Format: Gage Black, Yi Qiao, Xiaomeng Huang, Gabor Marth. Streamlining the reconstruction of subclonal evolution in DNA sequencing data. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 4295.
INTRODUCTION. Aggressive systemic mastocytosis (ASM) is characterized by the expansion of clonal mast cells (MCs) that cause organ damage. Some ASM cases involve only the mast cell lineage, but most are associated with another myeloid hematologic neoplasm (ASM - AHN), most commonly chronic myelomonocytic leukemia (CMML). The D816V mutation of KIT (KITD816V) is present in almost all ASM cases and thought to be the key disease driver. ASM-AHN patients (pts) have few MCs in the peripheral blood, yet KITD816V can be detectable at high variant allele frequency, consistent with the presence of KITD816V in non-MC lineage cells. To understand how MC and Non-MC compartments of ASM-AHN respond to disruption of KITD816V signaling, we combined single cell RNA sequencing (scRNAseq) with whole genome sequencing (WGS) to map clonal dynamics in ASM-AHN pts treated with avapritinib, a selective KITD816V inhibitor. PATIENTS AND METHODS. White blood cells (WBCs) from 4 pts with ASM-AHN (Table 1) collected prior to and after initiation of avapritinib therapy were subjected to deep (120X) WGS, with skin fibroblasts as controls. WBCs from the same and one additional timepoint were analyzed by scRNAseq (10x Genomics). Mononuclear cells from CMML pts without MC component (N=3) and age-matched healthy donors (N=3) served as controls. scRNAseq data were analyzed using Seurat and scBayes, a Bayesian-statistical approach developed by us for clonal attribution of single cells (Yi et al. Genome Research, in review). MC burden in bone marrow biopsies (BmBx) was analyzed with anti-tryptase antibody. RESULTS. WGS detected an average of 1891 mutations/genome. Pt1 had a focal homozygous deletion in chromosome (Chr) 21 (including RUNX1), and pts2 and 3 had subclonal Chr 7q and 4q LOH, respectively, all of which did not change on therapy. Pt4 had a subclonal Chr 18 deletion (no change on treatment), a subclonal Chr 7q LOH (increased), and a subclonal Chr 21 deletion (became undetectable). Subclone analysis revealed stable clonal structure in pts 1 and 2. In contrast, in pts 3 and 4, the KITD816V containing subclone shrank/disappeared on avapritinib, while founder clones with mutations in TET2, SRSF2, ETNK1, or ASXL1 persisted (Figure 1). BmBx showed reduced MC burden in all 4 pts. scRNAseq revealed abnormal monocytes in pts 1-4 that clustered distinct from each other, CMML monocytes and healthy controls. On avapritinib, monocytes from pts 1 and 2 gradually shifted toward normal controls, while monocytes from pts 3 and 4 continued to cluster closer to normal than CMML monocytes. Pseudo bulking revealed that transcriptomes from ASM-AHN monocytes are distinct from CMML and normal monocytes. Other features distinguishing pt cells from controls included the presence of immature neutrophils and higher levels of CXCR4 in both CD4+ and CD8+ T cells. Pt 4 (ASM-CEL) showed a prominent population of CD34+ cells expressing markers of mature eosinophils (LAIR1, ITGA4, IL3RA) and mast cells (TPSAB1, CPA3), suggesting origin from a transformed granulocyte-macrophage progenitor (GMP) with restricted MC/eosinophil potential (Drissen et al. PMID: 31126997). CONCLUSIONS. (i) Clonal dynamics of ASM-AHN on avapritinib are dependent on the localization of KITD816V in the hierarchy, with truncal upstream mutations predicting AHN responses. (ii) ASM-AHN monocytes are distinct from CMML monocytes. (iii) Eosinophils in ASM-CEL may be derived from a GMP whose lineage potential is limited to eosinophil/MC differentiation. Validation using targeted single cell and colony sequencing is underway and will be presented. Figure 1View largeDownload PPTFigure 1View largeDownload PPT Close modal
Selection of the treatments most likely to combat a patient’s tumor is a central aim of precision oncology. We are currently developing a functional precision oncology program at the University of Utah and Huntsman Cancer Institute that combines the multi-omic characterization of a patient's tumor with the functional screening of relevant drug candidates in patient-derived organoid tumor models to identify which drugs are uniquely capable of combating the patient’s cancer. Here we present our model-based approach for utilizing the multi-omic tumor data to computationally predict the patient’s response to each drug. In this approach, we first identify the genomic and transcriptomic vulnerabilities of the tumor as genes harboring somatic DNA mutations or copy number changes (from paired tumor/normal WGS/WES DNA sequencing data), as well as genes whose expression levels have been significantly altered in the tumor (based on bulk or single-cell RNA sequencing data). Second, using targeting information from drug-gene interaction databases (DGIdb) we compile a list of genes targeted by each drug relevant in the patient’s treatment. Third, we utilize gene interaction database knowledge (KEGG) to construct a gene interaction graph to link each drug’s gene targets with all of the tumor’s vulnerability genes. We then identify all gene interaction paths connecting specific drug target genes with specific vulnerability genes; and score each path and combine all path-specific scores for the target-vulnerability gene pair. Subsequently, we further combine all pairwise scores across all drug target genes and all tumor vulnerability genes and determine the statistical significance of the resulting score by sampling the background distribution of such scores across random drugs, target genes, and vulnerability genes. We have applied our algorithm to predict drug response in advanced breast cancer patients as well as using publicly available tumor cell line data (GDSC2). We have found a high level of concordance between our computational prediction and organoid/cell line screening results, clearly separating sensitive and non-sensitive models. Using Bayesian probability, we compare the drug-specific score distributions of sensitive and non-sensitive models and are able to identify sensitive cell lines/patients with high accuracy. By encapsulating available information on direct gene-gene interactions, the drug’s direct gene targets and the collected omic vulnerabilities of the tumor, our model is not only capable of predicting sensitivity to both targeted and chemotherapy agents, but can also provide a mechanistic understanding for the targeting of the tumor. Our approach will be validated and fine-tuned on a large cohort of breast and brain cancer patients as well as patient PDX models, interrogating gene-target interactions to identify novel relevant target genes/pathways. Citation Format: Szabolcs Tarapcsak, Yi Qiao, Xiaomeng Huang, Tony Di Sera, Matthew H. Bailey, Bryan E. Welm, Alana L. Welm, Gabor T. Marth. Model-based cancer therapy selection by linking tumor vulnerabilities to drug mechanism [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 2723.