Circulating tumor DNA (ctDNA) profiling from liquid biopsies is increasingly adopted as a minimally invasive solution for clinical cancer diagnostic applications. Current methods for inferring gene expression from ctDNA require specialized assays or ultra-deep, targeted sequencing, which preclude transcriptome-wide profiling at single-gene resolution. Herein we jointly introduce Triton, a tool for comprehensive fragmentomic and nucleosome profiling of cell-free DNA (cfDNA), and Proteus, a multi-modal deep learning framework for predicting single gene expression, using standard depth (~30-120x) whole genome sequencing of cfDNA. By synthesizing fragmentation and inferred nucleosome positioning patterns in the promoter and gene body from Triton, Proteus reproduced expression profiles using pure ctDNA from patient-derived xenografts (PDX) with an accuracy similar to RNA-Seq technical replicates. Applying Proteus to cfDNA from four patient cohorts with matched tumor RNA-Seq, we show that the model accurately predicted the expression of specific prognostic and phenotype markers and therapeutic targets. As an analog to RNA-Seq, we further confirmed the immediate applicability of Proteus to existing tools through accurate prediction of gene pathway enrichment scores. Our results demonstrate the potential clinical utility of Triton and Proteus as non-invasive tools for precision oncology applications such as cancer monitoring and therapeutic guidance.
Differential regulation metrics and values for select features Sheet 1: Log2 fold-change, p-value, and q-value between ARPC and NEPC lines for NPS in the 47 phenotype defining gene bodies (two tailed Mann-Whitney U test, Benjamini-Hochberg adjusted). Sheet 2: Differentially expressed list of 514 transcription factor (TF). All statistical comparison and fold change estimation was done against ARPCs. Sheet 3: Log2 fold-change, p-value, and q-value between ARPC and NEPC lines for central mean coverage in 108 TFs overlapping RNA-Seq up/down regulated TFs (two tailed Mann-Whitney U test, Benjamini-Hochberg adjusted). Sheet 4: Paralogous transcription factors (TFs) for each of the 38 TFs with differential RNA expression between ARPC and NEPC and differential TFBS accessibility. Paralogous TFs that are also differentially expressed by RNAseq analysis of PDX tumors are shown in red text. Paralogs were obtained from ensembl biomart human genes GRCh38.p13 (http://uswest.ensembl.org/biomart/martview/64d8bd7fe9851a2501aece9b74b03631) Sheet 5: Mean values in ARPC, NEPC, and HD lines for central and window means for ARPC and NEPC specific open chromatin regions.
Unsupervised model predictions and patient/validation cohort sequencing metrics Sheet 1: DFCI Cohort I: Tumor phenotype by histology and estimates of tumor fraction, subtyping score, and inferred subtype calls using ctdPheno. Sheet 2: DFCI cohort I: Tumor phenotype by histology and estimates of tumor fraction, subtyping score, and inferred phenotypic subtype calls using different ATAC-seq site restricted analysis using ctdPheno. Sheet 3: DFCI cohort I: Tumor phenotype by histology and estimates of tumor fraction, subtyping score, and inferred phenotypic subtype calls using different ATAC-seq site restricted analysis using Keraon. Sheet 4: DFCI Cohort II: Tumor phenotype by histology, summary of clinical correlatives and estimates of tumor fraction, subtyping score, and inferred subtype calls using ctdPheno. Sheet 5: Histology, tumor fraction, subtype score, NEPC fraction, and subtype calls for WGS and ULP patient samples (UW cohort). Sheet 6: Complete sequencing metrics for DFCI cohort I. Sheet 7: Complete sequencing metrics for DFCI cohort II (ULP). Sheet 8: Complete sequencing metrics for DFCI cohort II (deep WGS) Sheet 9: Complete sequencing metrics and ichorCNA estimated tumor fractions for the UW cohort (ULP). Sheet 10: Complete sequencing metrics for the UW cohort (deep WGS) Sheet 11: Complete sequencing metrics for the healthy donor cohort. Sheet 12: Clinical data summary for UW cohort.
The Exon Junction Complex (EJC) decorates RNA exon-exon junctions and modulates mRNA fate at multiple post-transcriptional steps until its disassembly during translation. Our investigation of the EJC disassembly factor PYM1 in human embryonic kidney 293 (HEK293) cells show that the EJC-PYM1 interaction is required for translation-independent EJC destabilization but not for translation-dependent disassembly. Surprisingly, PYM1 interaction deficient EJCs are enriched on locations away from canonical EJC binding site, particularly on transcripts with no or few introns. Such non-canonical EJCs are capable of inducing nonsense-mediated mRNA decay when present downstream of stop codons. Suppression of PYM1 in human cells, including by previously reported PYM1-flavivirus capsid protein interaction, stabilizes mRNAs with fewer and longer exons that localize to endoplasmic reticulum associated TIS-granules. In summary, PYM1 limits non-canonical EJC and thereby acts as an EJC specificity factor that is hijacked by flaviviruses to reshape host cell mRNA regulation.
PTM peak data and phenotype 47 fragment variability Sheet 1: PDX sample representation in 3 histone PTM CUT&RUN nucleosome profiling assays (H3K4me1, H2K27ac and H3K27me3). Sheet 2: Log2 fold-change, p-value, and q-value between ARPC and NEPC lines for coefficient of variation in the 47 phenotype defining gene bodies (two tailed Mann-Whitney U test, Benjamini-Hochberg adjusted). Sheet 3: Log2 fold-change, p-value, and q-value between ARPC and NEPC lines for coefficient of variation in the 47 phenotype defining gene promoters (two tailed Mann-Whitney U test, Benjamini-Hochberg adjusted).
Feature-region combination AUCs and benchmarking Sheet 1: Log2 fold-change, p-value, and q-value between ARPC and NEPC lines for central mean coverage in all queried (338) TFs (two tailed Mann-Whitney U test, Benjamini-Hochberg adjusted). Sheet 2: 100-fold cross-validation AUCs for all region and feature combinations subset by the ‘AR10’ overlapping features (see Methods). Sheet 3: 100-fold cross-validation AUCs for all region and feature combinations subset by the ‘Phenotype 47-defining’ overlapping features. Sheet 4: 100-fold cross-validation AUCs for all region and feature combinations (global). Sheet 5: Predictions scores, tumor fraction, depth of coverage, and subtype for benchmarking admixtures. Sheet 6: AUCs for unsupervised prediction of admixture subtypes grouped by tumor fraction and depth. Sheet 7: NEPC:ARPC ratio, tumor fraction, and ARPC and NEPC fraction predictions for mixed phenotype admixtures using Keraon (see Methods).
Preeclampsia is characterized by placental dysfunction and results in significant morbidity, but reliable early prediction remains challenging. We investigated whether clinically obtained prenatal cell-free DNA (cfDNA) screening (PDNAS) using whole-genome sequencing (WGS) data can be leveraged to predict preeclampsia risk early in pregnancy (≤16 weeks). Using 1,854 routinely collected clinical PDNAS samples (median, 12.1 weeks) with low-coverage (0.5×) WGS data, we developed a framework to quantify maternal and fetal tissue signatures using nucleosome accessibility, revealing early placental and endothelial dysfunction. These signatures informed a prediction model for preeclampsia risk, which achieved a validation performance of 0.85 area under the receiver operating characteristic curve (AUC) (81% sensitivity at 80% specificity) for preterm phenotypes several months prior to disease onset in a separate cohort of 831 consecutively collected samples, and subsequently confirmed in an external cohort of 141 samples (AUC 0.84, 79% sensitivity). We demonstrate that assessment of cfDNA nucleosome accessibility from early-pregnancy cfDNA sequence data enables the detection of early placental and endothelial-tissue aberrations and may aid in the determination of preeclampsia risk. Using 1,854 routinely collected clinical samples from early in pregnancy, with validation in an external cohort, low-coverage cfDNA sequence data identified distinctive features among those who developed preeclampsia.
Abstract Introduction: Metastatic castration-resistant prostate cancer (mCRPC) is a heterogeneous disease which can be classified into clinically relevant subtypes based on the expression of genes, such as the androgen receptor (AR) and neuroendocrine markers. Neuroendocrine prostate cancer (NEPC), characterized by gain of stem-like and neuroendocrine features and lack of AR expression is a clinically aggressive variant. Due to the lack of adequate biomarkers, NEPC is usually detected at a very advanced stage. There is mounting evidence that molecular subtype changes seen in NEPC are enforced by widespread epigenetic alterations, in particular DNA methylation changes. In this study, we aim to devise a novel DNA methylation-based assay for molecular subtyping and disease monitoring from cell-free DNA (cfDNA). Methods: We analyzed genome wide methylation patterns in 56 prostate cancer patient-derived xenograft (PDX) and 128 mCRPC tumors using array- and sequencing-based assays. We integrated DNA methylation at promoters, gene bodies and transcription factor binding site (TFBS) to determine the landscape of methylation alterations at key lineage specific genes. Using whole genome methylation derived from tissue with matched expression data we developed a deep learning framework to predict gene expression directly from tissue or cfDNA. Using key marker genes, the model was used to discern tumor molecular phenotypes from tissue and cfDNA in three independent cohorts of mCRPC patients using whole genome bisulfite sequencing and low-pass Enzymatic Methyl-Seq (EM-seq). Results: We observed a tight association between promoter, gene body and TFBS methylation with gene expression. Inferring gene expression from methylation for lineage specific markers such as AR, KLK3, ASCL1, INSM1, SRRM4 and DLL3 we classified molecular subtypes from both tissue and cfDNA. Additionally, for AR and ASCL1, we identified core sets of TFBSs whose differential methylation allowed for accurate assay-independent molecular subtype quantification. Applying the optimized quantitative model to mCRPC patients who underwent comprehensive tissue sampling by rapid autopsy we observed accurate subtype classification from both tissue samples and cfDNA for all cases. A similar analytical performance was observed in additional clinical mCRPC cohorts with cfDNA. Conclusion: Whole-genome methylation analysis of cfDNA allows for the prediction of gene expression patterns in tumor tissues, enabling non-invasive tumor subclassification and assessment of therapeutic targets. Citation Format: Mohamed Adil, Brian Hanratty, Pallabi Mustafi, Chitvan Mittal, Helen Richards, Ilsa Coleman, Radhika Patel, Anna-Lisa Doebley, Robert Patton, Eden Cruikshank, Patricia Galipeau, Ruth Dumpit, Martine Roudier, Jin-Yih Low, Navonil Sarkar, Robert Montgomery, Eva Corey, Colm Morrissey, Peter Nelson, Gavin Ha, Michael Haffner. Advance prostate cancer detection through epigenomic profiling of cell-free DNA [abstract]. In: Proceedings of the AACR Special Conference: Liquid Biopsy: From Discovery to Clinical Implementation; 2024 Nov 13-16; San Diego, CA. Philadelphia (PA): AACR; Clin Cancer Res 2024;30(21_Suppl):Abstract nr PR011.
Cell-free DNA (cfDNA) has the potential to inform tumor subtype classification and help guide clinical precision oncology. Here we developed Griffin, a new method for profiling nucleosome protection and accessibility from cfDNA to study the phenotype of tumors using as low as 0.1x coverage whole genome sequencing (WGS) data. Griffin employs a novel GC correction procedure tailored to variable cfDNA fragment sizes, which improves the prediction of chromatin accessibility. Griffin achieved excellent performance for detecting tumor cfDNA in early-stage cancer patients (AUC=0.96). Next, we applied Griffin for the first demonstration of estrogen receptor (ER) subtyping in metastatic breast cancer from cfDNA. We analyzed 254 samples from 139 patients and predicted ER subtype with high performance (AUC=0.89), leading to insights about tumor heterogeneity. In summary, Griffin is a framework for accurate clinical subtyping and can be generalizable to other cancer types for precision oncology applications.
Introduction: Metastatic castration-resistant prostate cancer (mCRPC) is a heterogeneous disease which can be classified into clinically relevant subtypes based on the expression of transcription factors (TF), such as the androgen receptor (AR) and neuroendocrine markers. Neuroendocrine prostate cancer (NEPC), characterized by gain of stem-like and neuroendocrine features and lack of AR expression is a clinically aggressive variant. Due to the absence of adequate biomarkers, NEPC is usually detected at a very advanced stage. There is mounting evidence that molecular subtype changes seen in NEPC are enforced by widespread epigenetic alterations, in particular DNA methylation changes. In this study, we aim to devise a novel DNA methylation-based assay for molecular subtyping and disease monitoring from cell-free DNA (cfDNA). Methods: We analyzed genome wide methylation patterns in 60 prostate cancer patient-derived xenograft (PDX) and 133 mCRPC tumors using array- and sequencing-based assays. We integrated DNA methylation with TF cistrome data to determine the landscape of methylation alterations at key lineage TF binding sites (TFBS). A linear regression model was trained on low-pass Enzymatic Methyl-Seq (EM-seq) cfDNA data derived from PDXs to identify molecular subtype specific DNA methylation changes at these TFBS. The model performance was optimized with in silico admixture experiments. This model was then used to discern tumor molecular phenotypes from cfDNA in three independent cohorts of mCRPC patients using low-pass whole genome bisulfite sequencing and EM-seq. Results: We observed a strong association between TFBS methylation and TF expression. For lineage specific TFs such as AR and ASCL1, we identified core sets of TFBSs whose differential methylation allowed for accurate assay-independent molecular subtype classification in tumor tissues. Applying an optimized quantitative model to mCRPC patients who underwent comprehensive tissue sampling by rapid autopsy we observed perfect subtype prediction from both tissue samples and cfDNA (AUC=1). A similar analytical performance was observed in additional clinical mCRPC cohorts with cfDNA. Conclusions: We show that methylation patterns at TFBSs can determine TF activity and can be used to classify molecular subtypes from both tumor tissue and cfDNA. For prostate cancer, we demonstrate that this approach can accurately detect NEPC by cost-effective low-pass EM-seq. More broadly, this study provides a novel analysis framework for robustly assessing molecular tumor phenotypes in cfDNA with applications in solid and liquid tumor diagnostics. Citation Format: Mohamed Adil, Brian Hanratty, Pallabi Mustafi, Ilsa Coleman, Radhika Patel, Anna-Lisa Doebley, Robert Patton, Eden Cruikshank, Patricia Galipeau, Ruth Dumpit, Martine Roudier, Jin-Yih Low, Navonil De Sarkar, Robert Montgomery, Eva Corey, Colm Morrissey, Peter Nelson, Gavin Ha, Michael Haffner. Molecular phenotype classification of metastatic prostate cancer by cell-free DNA methylation analysis [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 2 (Clinical Trials and Late-Breaking Research); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(8_Suppl):Abstract nr LB298.
Abstract Advanced prostate cancers comprise distinct phenotypes, but tumor classification remains clinically challenging. Here, we harnessed circulating tumor DNA (ctDNA) to study tumor phenotypes by ascertaining nucleosome positioning patterns associated with transcription regulation. We sequenced plasma ctDNA whole genomes from patient-derived xenografts representing a spectrum of androgen receptor active (ARPC) and neuroendocrine (NEPC) prostate cancers. Nucleosome patterns associated with transcriptional activity were reflected in ctDNA at regions of genes, promoters, histone modifications, transcription factor binding, and accessible chromatin. We identified the activity of key phenotype-defining transcriptional regulators from ctDNA, including AR, ASCL1, HOXB13, HNF4G, and GATA2. To distinguish NEPC and ARPC in patient plasma samples, we developed prediction models that achieved accuracies of 97% for dominant phenotypes and 87% for mixed clinical phenotypes. Although phenotype classification is typically assessed by IHC or transcriptome profiling from tumor biopsies, we demonstrate that ctDNA provides comparable results with diagnostic advantages for precision oncology. Significance: This study provides insights into the dynamics of nucleosome positioning and gene regulation associated with cancer phenotypes that can be ascertained from ctDNA. New methods for classification in phenotype mixtures extend the utility of ctDNA beyond assessments of somatic DNA alterations with important implications for molecular classification and precision oncology. This article is highlighted in the In This Issue feature, p. 517
Nonsense-mediated mRNA decay (NMD) is governed by the three conserved factors-UPF1, UPF2, and UPF3. While all three are required for NMD in yeast, UPF3B is dispensable for NMD in mammals, and its paralog UPF3A is suggested to only weakly activate or even repress NMD due to its weaker binding to the exon junction complex (EJC). Here, we characterize the UPF3A/B-dependence of NMD in human cell lines deleted of one or both UPF3 paralogs. We show that in human colorectal cancer HCT116 cells, NMD can operate in a UPF3B-dependent and -independent manner. While UPF3A is almost dispensable for NMD in wild-type cells, it strongly activates NMD in cells lacking UPF3B. Notably, NMD remains partially active in cells lacking both UPF3 paralogs. Complementation studies in these cells show that EJC-binding domain of UPF3 paralogs is dispensable for NMD. Instead, the conserved "mid" domain of UPF3 paralogs is consequential for their NMD activity. Altogether, our results demonstrate that the mammalian UPF3 proteins play a more active role in NMD than simply bridging the EJC and the UPF complex.
ABSTRACTNonsense-mediated mRNA decay (NMD) is governed by the three conserved factors - UPF1, UPF2 and UPF3. While all three are required for NMD in yeast, UPF3B is dispensable for NMD in mammals, with its paralog UPF3A suggested to only weakly activate or even repress NMD due to its weaker binding to the exon junction complex (EJC). Here we characterize the UPF3B-dependent and -independent NMD in human cell lines knocked-out of one or bothUPF3paralogs. We show that in human colorectal cancer HCT116 cells, EJC-mediated NMD can operate in UPF3B-dependent and -independent manner. While UPF3A is almost completely dispensable for NMD in wild-type cells, it strongly activates EJC-mediated NMD in cells lacking UPF3B. Surprisingly, this major NMD branch can operate in UPF3-independent manner questioning the idea that UPF3 is needed to bridge UPF proteins to the EJC during NMD. Complementation studies in UPF3 knockout cells further show that EJC-binding domain of UPF3 paralogs is not essential for NMD. Instead, the conserved mid domain of UPF3B, previously shown to engage with ribosome release factors, is required for its full NMD activity. Altogether, UPF3 plays a more active role in NMD than simply being a bridge between the EJC and the UPF complex.
In eukaryotic cells, proteins that associate with RNA regulate its activity to control cellular function. To fully illuminate the basis of RNA function, it is essential to identify such RNA-associated proteins, their mode of action on RNA, and their preferred RNA targets and binding sites. By analyzing catalogs of human RNA-associated proteins defined by ultraviolet light (UV)-dependent and -independent approaches, we classify these proteins into two major groups: (i) the widely recognized RNA binding proteins (RBPs), which bind RNA directly and UV-crosslink efficiently to RNA, and (ii) a new group of RBP-associated factors (RAFs), which bind RNA indirectly via RBPs and UV-crosslink poorly to RNA. As the UV crosslinking and immunoprecipitation followed by sequencing (CLIP-seq) approach will be unsuitable to identify binding sites of RAFs, we show that formaldehyde crosslinking stabilizes RAFs within ribonucleoproteins to allow for their immunoprecipitation under stringent conditions. Using an RBP (CASC3) and an RAF (RNPS1) within the exon junction complex (EJC) as examples, we show that formaldehyde crosslinking combined with RNA immunoprecipitation in tandem followed by sequencing (xRIPiT-seq) far exceeds CLIP-seq to identify binding sites of RNPS1. xRIPiT-seq reveals that RNPS1 occupancy is increased on exons immediately upstream of strong recursively spliced exons, which depend on the EJC for their inclusion.
Many post-transcriptional mechanisms operate via mRNA 3′UTRs to regulate protein expression, and such controls are crucial for development. We show that homozygous mutations in two zebrafish exon junction complex (EJC) core genes rbm8a and magoh leads to muscle disorganization, neural cell death, and motor neuron outgrowth defects, as well as dysregulation of mRNAs subjected to nonsense-mediated mRNA decay (NMD) due to translation termination ≥ 50 nts upstream of the last exon-exon junction. Intriguingly, we find that EJC-dependent NMD also regulates a subset of transcripts that contain 3′UTR introns (3′UI) < 50 nts downstream of a stop codon. Some transcripts containing such stop codon-proximal 3′UI are also NMD-sensitive in cultured human cells and mouse embryonic stem cells. We identify 167 genes that contain a conserved proximal 3′UI in zebrafish, mouse and humans. foxo3b is one such proximal 3′UI-containing gene that is upregulated in zebrafish EJC mutant embryos, at both mRNA and protein levels, and loss of foxo3b function in EJC mutant embryos significantly rescues motor axon growth defects. These data are consistent with EJC-dependent NMD regulating foxo3b mRNA to control protein expression during zebrafish development. Our work shows that the EJC is critical for normal zebrafish development and suggests that proximal 3′UIs may serve gene regulatory function in vertebrates.
The Exon Junction Complex (EJC) regulates many steps in post-transcriptional gene expression and is essential for cellular function and organismal development; however, EJC-regulated genes and genetic pathways during development remain largely unknown. To study EJC function during zebrafish development, we first established that zebrafish EJCs mainly bind ∼24 nucleotides upstream of exon-exon junctions, and are also detected at more distant non-canonical positions. We then generated mutations in two zebrafish EJC core genes, and , and observed that homozygous mutant embryos show paralysis, muscle disorganization, neural cell death, and motor neuron outgrowth defects. Coinciding with developmental defects, mRNAs subjected to Nonsense-Mediated mRNA Decay (NMD) due to translation termination ≥ 50 nts upstream of the last exon-exon junction are upregulated in EJC mutant embryos. Surprisingly, several transcripts containing 3′UTR introns (3′UI) < 50 nts downstream of a stop codon are also upregulated in EJC mutant embryos. These proximal 3′UI-containing transcripts are also upregulated in NMD-compromised zebrafish embryos, cultured human cells, and mouse embryonic stem cells. Loss of function of one of the upregulated proximal 3′UI-containing genes, partially rescues EJC mutant motor neuron outgrowth. In addition to , 166 other genes contain a proximal 3′UI in zebrafish, mouse and humans, and these genes are enriched in nervous system development and RNA binding functions. A proximal 3′UI-containing 3′UTR from one of these genes, , is sufficient to reduce steady state transcript levels when fused to a reporter in HeLa cells. Overall, our work shows that genes with stop codon-proximal 3′UIs encode a new class of EJC-regulated NMD targets with critical roles during vertebrate development.