SUMMARYSpruces (Picea spp.) are coniferous trees widespread in boreal and mountainous forests of the northern hemisphere, with large economic significance and enormous contributions to global carbon sequestration. Spruces harbor very large genomes with high repetitiveness, hampering their comparative analysis. Here, we present and compare the genomes of four different North American spruces: the genome assemblies for Engelmann spruce (Picea engelmannii) and Sitka spruce (Picea sitchensis) together with improved and more contiguous genome assemblies for white spruce (Picea glauca) and for a naturally occurring introgress of these three species known as interior spruce (P. engelmannii × glauca × sitchensis). The genomes were structurally similar, and a large part of scaffolds could be anchored to a genetic map. The composition of the interior spruce genome indicated asymmetric contributions from the three ancestral genomes. Phylogenetic analysis of the nuclear and organelle genomes revealed a topology indicative of ancient reticulation. Different patterns of expansion of gene families among genomes were observed and related with presumed diversifying ecological adaptations. We identified rapidly evolving genes that harbored high rates of non‐synonymous polymorphisms relative to synonymous ones, indicative of positive selection and its hitchhiking effects. These gene sets were mostly distinct between the genomes of ecologically contrasted species, and signatures of convergent balancing selection were detected. Stress and stimulus response was identified as the most frequent function assigned to expanding gene families and rapidly evolving genes. These two aspects of genomic evolution were complementary in their contribution to divergent evolution of presumed adaptive nature. These more contiguous spruce giga‐genome sequences should strengthen our understanding of conifer genome structure and evolution, as their comparison offers clues into the genetic basis of adaptation and ecology of conifers at the genomic level. They will also provide tools to better monitor natural genetic diversity and improve the management of conifer forests. The genomes of four closely related North American spruces indicate that their high similarity at the morphological level is paralleled by the high conservation of their physical genome structure. Yet, the evidence of divergent evolution is apparent in their rapidly evolving genomes, supported by differential expansion of key gene families and large sets of genes under positive selection, largely in relation to stimulus and environmental stress response.
Pediatric AML is characterized by a high rate of relapse of up to ~40% (Im et al. 2016), and nearly half of the patients achieving initial remission experience relapse within 2 years (Alpenc et al. 2016; Rubnitz and Gruber 2018). Due to the lack of recurrently mutated genes associated with relapse as observed through genomic analysis (Farrar et al. 2016; Boluori et al. 2017), we hypothesize that the transcriptome may reveal insights into molecules and mechanisms contributing to treatment resistance. Farrar et al. (2016) provided support for this hypothesis as they observed that somatic mutations across primary and relapse patient samples converged on genes involved in transcriptional regulation. McNeer et al. (2019) showed that a large number of non protein-coding RNAs, such as pseudogenes and long non-coding RNAs, many of which have regulatory roles through interactions with other genes and proteins, were observed to have increased mutational frequency post-induction compared to samples at diagnosis among induction-failure patients. These lines of evidence suggest that regulatory processes and interactions involving RNA molecules may play important roles in treatment resistance. To address our hypothesis, we conducted analysis of rRNA-depleted RNA sequencing data generated from 1325 primary and 396 relapse bone marrow or peripheral blood samples obtained from patients enrolled in the AAML1031 (treatment arms are ADE, ADE+Bortezomib, and ADE +Sorafenib) and AAML0531 (randomized treatment arms chemotherapy with or without Gemtuzumab Ozogamicin) clinical trials. We focused our analysis to malignant cells in 620 primary and 148 relapse samples with >50% blast count. To identify RNAs associated with overall survival, we conducted Cox proportional-hazards regression and generalized linear model via penalized maximum likelihood analyses using RNA expression profiles. We performed transcription factor and regulatory network profiling as adapted from Aibar et al. (2017) to identify significant interactions between RNAs. We used primary samples for 574 patients enrolled in AAML1031 treated with either ADE or ADE+Bortezomib to identify high risk features associated with overall survival. We identified 7 RNAs (AC002401.1, VAV1, RP1-37C10.3, RP11-92C4.3, PRICKLE4, RP11-491H9.3, and NYNRIN) with log hazard ratios >1 (adjusted p-value < 0.000025) and that were assigned positive coefficients as derived from a combinatorial RNA generalized linear model. Such results indicate that these RNAs are significantly associated with low probability of overall survival, and that the expression of these RNAs at diagnosis could be used to identify high-risk patients and to anticipate poor survival outcome. Interestingly, 5 of these genes encode antisense transcripts. RNA expression profiles of 620 primary and 148 relapse samples were compared to identify molecular features and interactions more directly associated with treatment resistance. Though we observed no differences in transcription factor network activities between primary and relapse samples, we identified 14 RNAs that were significantly upregulated at relapse compared to primary samples (log2 fold change >2; BH-adjusted p-value < 0.05), 10 of which were pseudogenes, long intergenic or non protein-coding RNAs. Gene regulatory network inference analysis using the regression tree-based algorithm GENIE3 (Huynh-Thu et al. 2010) identified 1842 genes to have interactions with these 14 genes (3594 interactions total; weight >0.001). Gene set enrichment analysis of the 1856 genes showed the top enriched pathway to be "ribosome, cytoplasmic." These results suggest that RNA processing and translational control could be associated with treatment resistance. Our findings revealed previously uncharacterized molecular features and interactions potentially associated with low probably of overall survival and treatment resistance. Future analyses of these molecules will ideally contribute to deeper insights into the mechanisms driving relapse disease, and extend therapeutic targeting to regulatory RNAs which, through interactions with other molecules, may play important roles in regulating transcription and translation. Disclosures No relevant conflicts of interest to declare.
Spatial heterogeneity of transcriptional and genetic markers between physically isolated biopsies of a single tumor poses major barriers to the identification of biomarkers and the development of targeted therapies that will be effective against the entire tumor. We analyzed the spatial heterogeneity of multiregional biopsies from 35 patients, using a combination of transcriptomic and genomic profiles. Medulloblastomas (MBs), but not high-grade gliomas (HGGs), demonstrated spatially homogeneous transcriptomes, which allowed for accurate subgrouping of tumors from a single biopsy. Conversely, somatic mutations that affect genes suitable for targeted therapeutics demonstrated high levels of spatial heterogeneity in MB, malignant glioma, and renal cell carcinoma (RCC). Actionable targets found in a single MB biopsy were seldom clonal across the entire tumor, which brings the efficacy of monotherapies against a single target into question. Clinical trials of targeted therapies for MB should first ensure the spatially ubiquitous nature of the target mutation.
The development of targeted anti-cancer therapies through the study of cancer genomes is intended to increase survival rates and decrease treatment-related toxicity. We treated a transposon-driven, functional genomic mouse model of medulloblastoma with 'humanized' in vivo therapy (microneurosurgical tumour resection followed by multi-fractionated, image-guided radiotherapy). Genetic events in recurrent murine medulloblastoma exhibit a very poor overlap with those in matched murine diagnostic samples (<5%). Whole-genome sequencing of 33 pairs of human diagnostic and post-therapy medulloblastomas demonstrated substantial genetic divergence of the dominant clone after therapy (<12% diagnostic events were retained at recurrence). In both mice and humans, the dominant clone at recurrence arose through clonal selection of a pre-existing minor clone present at diagnosis. Targeted therapy is unlikely to be effective in the absence of the target, therefore our results offer a simple, proximal, and remediable explanation for the failure of prior clinical trials of targeted therapy.
While significant effort has been dedicated to the characterization of epigenetic changes associated with prenatal differentiation, relatively little is known about the epigenetic changes that accompany post-natal differentiation where fully functional differentiated cell types with limited lifespans arise. Here we sought to address this gap by generating epigenomic and transcriptional profiles from primary human breast cell types isolated from disease-free human subjects. From these data we define a comprehensive human breast transcriptional network, including a set of myoepithelial- and luminal epithelial-specific intronic retention events. Intersection of epigenetic states with RNA expression from distinct breast epithelium lineages demonstrates that mCpG provides a stable record of exonic and intronic usage, whereas H3K36me3 is dynamic. We find a striking asymmetry in epigenomic reprogramming between luminal and myoepithelial cell types, with the genomes of luminal cells harbouring more than twice the number of hypomethylated enhancer elements compared with myoepithelial cells.
BACKGROUND Many mutations that contribute to the pathogenesis of acute myeloid leukemia (AML) are undefined. The relationships between patterns of mutations and epigenetic phenotypes are not yet clear. METHODS We analyzed the genomes of 200 clinically annotated adult cases of de novo AML, using either whole-genome sequencing (50 cases) or whole-exome sequencing (150 cases), along with RNA and microRNA sequencing and DNA-methylation analysis. RESULTS AML genomes have fewer mutations than most other adult cancers, with an average of only 13 mutations found in genes. Of these, an average of 5 are in genes that are recurrently mutated in AML. A total of 23 genes were significantly mutated, and another 237 were mutated in two or more samples. Nearly all samples had at least 1 nonsynonymous mutation in one of nine categories of genes that are almost certainly relevant for pathogenesis, including transcription-factor fusions (18% of cases), the gene encoding nucleophosmin (NPM1) (27%), tumor-suppressor genes (16%), DNA-methylation-related genes (44%), signaling genes (59%), chromatin-modifying genes (30%), myeloid transcription-factor genes (22%), cohesin-complex genes (13%), and spliceosome-complex genes (14%). Patterns of cooperation and mutual exclusivity suggested strong biologic relationships among several of the genes and categories. CONCLUSIONS We identified at least one potential driver mutation in nearly all AML samples and found that a complex interplay of genetic events contributes to AML pathogenesis in individual patients. The databases from this study are widely available to serve as a foundation for further investigations of AML pathogenesis, classification, and risk stratification. (Funded by the National Institutes of Health.).