Supplementary Table 1. List of 516 deep-sequenced cancer-related genes; Supplementary Table 2. Clinicopathological characteristics of screen-detected (n=109) versus interval breast cancers (n=111) in the validation dataset; Supplementary Table 3. Correlation between gene expression clusters and different Characteristics; Supplementary Table 4. Results of gene set enrichment analysis (GSEA Preranked); Supplementary Figure 1. Flow diagram leading to analytical cohort in discovery study; Supplementary Figure 2. Chromosomal plots for copy number aberrations; Supplementary Figure 3: Transcriptome profile of 111 interval breast versus 109 screendetected breast cancers in the validation dataset.
This corrects the article DOI: 10.1038/ncomms5999.
lncRNA in de novo vs. non-de novo AML. Description of data: Differentially expressed lncRNA in de novo vs. non-de novo AML patients. (XLSX 13 kb)
List of genes. Description of data: A list of differentially expressed genes in lncRNA subtypes. (XLSX 2264 kb)
Background: Recent progress in sequencing technologies allows us to explore comprehensive genomic and transcriptomic information to improve the current European LeukemiaNet (ELN) system of acute myeloid leukemia (AML). Methods: We compared the prognostic value of traditional demographic and cytogenetic risk factors, genomic data in the form of somatic aberrations of 25 AML-relevant genes, and whole-transcriptome expression profiling (RNA sequencing) in 267 intensively treated AML patients (Clinseq-AML). Multivariable penalized Cox models (overall survival [OS]) were developed for each data modality (clinical, genomic, transcriptomic), together with an associated prognostic risk score. Results: Of the three data modalities, transcriptomic data provided the best prognostic value, with an integrated area under the curve (iAUC) of a time-dependent receiver operating characteristic (ROC) curve of 0.73. We developed a prognostic risk score (Clinseq-G) from transcriptomic data, which was validated in the independent The Cancer Genome Atlas AML cohort (RNA sequencing, n = 142, iAUC = 0.73, comparing the high-risk group with the low-risk group, hazard ratio [HR] OS = 2.42, 95% confidence interval [CI] = 1.51 to 3.88). Comparison between Clinseq-G and ELN score iAUC estimates indicated strong evidence in favor of the Clinseq-G model (Bayes factor = 26.78). The proposed model remained statistically significant in multivariable analysis including the ELN and other well-known risk factors (HRos = 2.34, 95% CI = 1.30 to 4.22). We further validated the Clinseq-G model in a second independent data set (n = 458, iAUC = 0.66, adjusted HROS = 2.02, 95% CI = 1.33 to 3.08; adjusted HREFS = 2.10, 95% CI = 1.42 to 3.12). Conclusions: Our results indicate that the Clinseq-G prediction model, based on transcriptomic data from RNA sequencing, outperforms traditional clinical parameters and previously reported models based on genomic biomarkers.
Background: Long non-coding RNA (lncRNA) expression has been implicated in a range of molecular mechanisms that are central in cancer. However, lncRNA expression has not yet been comprehensively characterized in acute myeloid leukemia (AML). Here, we assess to what extent lncRNA expression is prognostic of AML patient overall survival (OS) and determine if there are indications of lncRNA-based molecular subtypes of AML. Methods: We performed RNA sequencing of 274 intensively treated AML patients in a Swedish cohort and quantified lncRNA expression. Univariate and multivariate time-to-event analysis was applied to determine association between individual lncRNAs with OS. Unsupervised statistical learning was applied to ascertain if lncRNA-based molecular subtypes exist and are prognostic. Results: Thirty-three individual lncRNAs were found to be associated with OS (adjusted p value < 0.05). We established four distinct molecular subtypes based on lncRNA expression using a consensus clustering approach. LncRNA-based subtypes were found to stratify patients into groups with prognostic information (p value < 0.05). Subsequently, lncRNA expression-based subtypes were validated in an independent patient cohort (TCGA-AML). LncRNA subtypes could not be directly explained by any of the recurrent cytogenetic or mutational aberrations, although associations with some of the established genetic and clinical factors were found, including mutations in NPM1, TP53, and FLT3. Conclusion: LncRNA expression-based four subtypes, discovered in this study, are reproducible and can effectively stratify AML patients. LncRNA expression profiling can provide valuable information for improved risk stratification of AML patients.
Risk stratification of acute myeloid leukemia (AML) patients needs improvement. Several AML risk classification models based on somatic mutations or gene-expression profiling have been proposed. However, systematic and independent validation of these models is required for future clinical implementation. We performed whole-transcriptome RNA-sequencing and panel-based deep DNA sequencing of 23 genes in 274 intensively treated AML patients (Clinseq-AML). We also utilized the The Cancer Genome Atlas (TCGA)-AML study ( N =142) as a second validation cohort. We evaluated six previously proposed molecular-based models for AML risk stratification and two revised risk classification systems combining molecular- and clinical data. Risk groups stratified by five out of six models showed different overall survival in cytogenetic normal-AML patients in the Clinseq-AML cohort ( P -value<0.05; concordance index >0.5). Risk classification systems integrating mutational or gene-expression data were found to add prognostic value to the current European Leukemia Net (ELN) risk classification. The prognostic value varied between models and across cohorts, highlighting the importance of independent validation to establish evidence of efficacy and general applicability. All but one model replicated in the Clinseq-AML cohort, indicating the potential for molecular-based AML risk models. Risk classification based on a combination of molecular and clinical data holds promise for improved AML patient stratification in the future.
Abstract Purpose:Interval breast cancer is of clinical interest as it exhibits an aggressive phenotype and evades detection by screening mammography. A comprehensive picture of somatic changes that drive tumors to become symptomatic in the screening interval can improve understanding of the biology underlying these aggressive tumors. Experimental design:Initiated in April 2013, Clinical Sequencing of Cancer in Sweden (Clinseq) is a scientific and clinical platform for the genomic profiling of cancer. The breast cancer pilot study consisted of women diagnosed with breast cancer between 2001-2012 in the Stockholm/Gotland regions. A subset of 318 breast tumors were sequenced, of which 113 were screen-detected and were 60 interval cancers.We applied targeted deep-sequencing of cancer-related genes, low-pass whole-genome sequencing and RNA-sequencing technology to characterize somatic differences in the genomic and transcriptomic architecture by interval cancer status. Mammographic density and PAM50 molecular subtypes were considered. Results:In the crude analyses, TP53, PPP1R3A, and KMT2B were significantly more frequently mutated in interval cancers than in screen-detected cancers. Acquired somatic copy number aberrations with a frequency difference of at least 15% between the two groups included gains in 17q23-q25.3 and losses in 16q24.2. Gene expression analysis identified 447 significantly differentially expressed genes, of which 120 were replicated in an independent microarray dataset. After adjusting for PAM50, most differences were no longer significant. Conclusions: Molecular differences by interval cancer status were observed, but they were largely explained by PAM50 subtypes. This work offer new insights into the biological differences between the two tumor groups. Translational relevance: Although screen-detected cancers are biologically distinct from interval cancers in terms of somatic mutations, copy number aberrations and gene expression, most of the differences are no longer significant after adjusting for breast cancer intrinsic subtypes (PAM50). We also show that the molecular differences appear to form a spectrum from less aggressive (screen-detected) to more aggressive (interval) manifestations of the disease, which can be characterized by PAM50 subtypes, namely, luminal A, luminal B, HER2-enriched and basal-like, in that order. This work clarifies the picture on what type of breast cancer we are likely to identify through population-based screening, and what type of cancer we are likely to miss. Current knowledge of PAM50 subtype-specific risk factors need to be expanded as our findings might influence how we screen women with a higher risk of basal-like breast cancer for example, beyond known risk groups BRCA1 mutation carriers and women of African-American descent. Citation Format: Czene K, Ivansson E, Klevebring D, Tobin NP, Lindström LS, Holm J, Prochazka G, Hilliges C, Palmgren J, Törnberg S, Humphreys K, Hartman J, Frisell J, Rantalainen M, Lindberg J, Hall P, Bergh J, Grönberg H, Li J. Molecular differences between screen-detected and interval breast cancers are largely explained by PAM50 subtypes [abstract]. In: Proceedings of the 2016 San Antonio Breast Cancer Symposium; 2016 Dec 6-10; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2017;77(4 Suppl):Abstract nr P2-03-03.
AbstractPurpose: Interval breast cancer is of clinical interest, as it exhibits an aggressive phenotype and evades detection by screening mammography. A comprehensive picture of somatic changes that drive tumors to become symptomatic in the screening interval can improve understanding of the biology underlying these aggressive tumors.Experimental Design: Initiated in April 2013, Clinical Sequencing of Cancer in Sweden (Clinseq) is a scientific and clinical platform for the genomic profiling of cancer. The breast cancer pilot study consisted of women diagnosed with breast cancer between 2001 and 2012 in the Stockholm/Gotland regions. A subset of 307 breast tumors was successfully sequenced, of which 113 were screen-detected and 60 were interval cancers. We applied targeted deep sequencing of cancer-related genes; low-pass, whole-genome sequencing; and RNA sequencing technology to characterize somatic differences in the genomic and transcriptomic architecture by interval cancer status. Mammographic density and PAM50 molecular subtypes were considered.Results: In the univariate analyses, TP53, PPP1R3A, and KMT2B were significantly more frequently mutated in interval cancers than in screen-detected cancers. Acquired somatic copy number aberrations with a frequency difference of at least 15% between the two groups included gains in 17q23-q25.3 and losses in 16q24.2. Gene expression analysis identified 447 significantly differentially expressed genes, of which 120 were replicated in an independent microarray dataset. After adjusting for PAM50, most differences were no longer significant.Conclusions: Molecular differences by interval cancer status were observed, but they were largely explained by PAM50 subtypes. This work offers new insights into the biological differences between the two tumor groups. Clin Cancer Res; 23(10); 2584–92. ©2016 AACR.
Sequencing-based breast cancer diagnostics have the potential to replace routine biomarkers and provide molecular characterization that enable personalized precision medicine. Here we investigate the concordance between sequencing-based and routine diagnostic biomarkers and to what extent tumor sequencing contributes clinically actionable information. We applied DNA- and RNA-sequencing to characterize tumors from 307 breast cancer patients with replication in up to 739 patients. We developed models to predict status of routine biomarkers (ER, HER2,Ki-67, histological grade) from sequencing data. Non-routine biomarkers, including mutations in BRCA1 , BRCA2 and ERBB2 (HER2), and additional clinically actionable somatic alterations were also investigated. Concordance with routine diagnostic biomarkers was high for ER status (AUC = 0.95;AUC(replication) = 0.97) and HER2 status (AUC = 0.97;AUC(replication) = 0.92). The transcriptomic grade model enabled classification of histological grade 1 and histological grade 3 tumors with high accuracy (AUC = 0.98;AUC(replication) = 0.94). Clinically actionable mutations in BRCA1 , BRCA2 and ERBB2 (HER2) were detected in 5.5% of patients, while 53% had genomic alterations matching ongoing or concluded breast cancer studies. Sequencing-based molecular profiling can be applied as an alternative to histopathology to determine ER and HER2 status, in addition to providing improved tumor grading and clinically actionable mutations and molecular subtypes. Our results suggest that sequencing-based breast cancer diagnostics in a near future can replace routine biomarkers.
Tumors are composed of multiple cell types besides the tumor cells themselves, including innate immune cells such as macrophages. Tumor-associated macrophages (TAMs) are a heterogeneous population of myeloid cells present in the tumor microenvironment (TME). Here, they contribute to immunosuppression, enabling the establishment and persistence of solid tumors as well as metastatic dissemination. We have found that the pattern recognition scavenger receptor MARCO defines a subtype of suppressive TAMs and is linked to clinical outcome. An anti-MARCO monoclonal antibody was developed, which induces anti-tumor activity in breast and colon carcinoma, as well as in melanoma models through reprogramming TAM populations to a pro-inflammatory phenotype and increasing tumor immunogenicity. This anti-tumor activity is dependent on the inhibitory Fc-receptor, FcγRIIB, and also enhances the efficacy of checkpoint therapy. These results demonstrate that immunotherapies using antibodies designed to modify myeloid cells of the TME represent a promising mode of cancer treatment.
Background: The histologic grade (HG) of breast cancer is an established prognostic factor. The grade is usually reported on a scale ranging from 1 to 3, where grade 3 tumours are the most aggressive. However, grade 2 is associated with an intermediate risk of recurrence, and carries limited information for clinical decision-making. Patients classified as grade 2 are at risk of both under-and over-treatment.Methods: RNA-sequencing analysis was conducted in a cohort of 275 women diagnosed with invasive breast cancer. Multivariate prediction models were developed to classify tumours into high and low transcriptomic grade (TG) based on gene-and isoform-level expression data from RNA-sequencing. HG2 tumours were reclassified according to the prediction model and a recurrence-free survival analysis was performed by the multivariate Cox proportional hazards regression model to assess to what extent the TG model could be used to stratify patients. The prediction model was validated in N = 487 breast cancer cases from the The Cancer Genome Atlas (TCGA) data set. Differentially expressed genes and isoforms associated with HGs were analysed using linear models.Results: The classification of grade 1 and grade 3 tumours based on RNA-sequencing data achieved high accuracy (area under the receiver operating characteristic curve = 0.97). The association between recurrence-free survival rate and HGs was confirmed in the study population (hazard ratio of grade 3 versus 1 was 2.62 with 95 % confidence interval = 1.04-6.61). The TG model enabled us to reclassify grade 2 tumours as high TG and low TG gene or isoform grade. The risk of recurrence in the high TG group of grade 2 tumours was higher than in low TG group (hazard ratio = 2.43, 95 % confidence interval = 1.13-5.20). We found 8200 genes and 13,809 isoforms that were differentially expressed between HG1 and HG3 breast cancer tumours.Conclusions: Gene-and isoform-level expression data from RNA-sequencing could be utilised to differentiate HG1 and HG3 tumours with high accuracy. We identified a large number of novel genes and isoforms associated with HG. Grade 2 tumours could be reclassified as high and low TG, which has the potential to reduce over-and under-treatment if implemented clinically.
Molecular characterization of genome-wide association study (GWAS) loci can uncover key genes and biological mechanisms underpinning complex traits and diseases. Here we present deep, high-throughput characterization of gene regulatory mechanisms underlying prostate cancer risk loci. Our methodology integrates data from 295 prostate cancer chromatin immunoprecipitation and sequencing experiments with genotype and gene expression data from 602 prostate tumor samples. The analysis identifies new gene regulatory mechanisms affected by risk locus SNPs, including widespread disruption of ternary androgen receptor (AR)-FOXA1 and AR-HOXB13 complexes and competitive binding mechanisms. We identify 57 expression quantitative trait loci at 35 risk loci, which we validate through analysis of allele-specific expression. We further validate predicted regulatory SNPs and target genes in prostate cancer cell line models. Finally, our integrated analysis can be accessed through an interactive visualization tool. This analysis elucidates how genome sequence variation affects disease predisposition via gene regulatory mechanisms and identifies relevant genes for downstream biomarker and drug development.
Breast cancer is the most diagnosed malignancy and the second leading cause of cancer mortality in females. Previous association studies have identified variants on 2q35 associated with the risk of breast cancer. To identify functional susceptibility loci for breast cancer, we interrogated the 2q35 gene desert for chromatin architecture and functional variation correlated with gene expression. We report a novel intergenic breast cancer risk locus containing an enhancer copy number variation (enCNV; deletion) located approximately 400Kb upstream to IGFBP5, which overlaps an intergenic ERα-bound enhancer that loops to the IGFBP5 promoter. The enCNV is correlated with modified ERα binding and monoallelic-repression of IGFBP5 following oestrogen treatment. We investigated the association of enCNV genotype with breast cancer in 1,182 cases and 1,362 controls, and replicate our findings in an independent set of 62,533 cases and 60,966 controls from 41 case control studies and 11 GWAS. We report a dose-dependent inverse association of 2q35 enCNV genotype (percopy OR = 0.68 95%CI 0.55-0.83, P = 0.0002; replication OR = 0.77 95% CI 0.73-0.82, P = 2.1 × 10-19) and identify 13 additional linked variants (r2 > 0.8) in the 20Kb linkage block containing the enCNV (P = 3.2 × 10-15 - 5.6 × 10-17). These associations were independent of previously reported 2q35 variants, rs13387042/rs4442975 and rs16857609, and were stronger for ER-positive than ER-negative disease. Together, these results suggest that 2q35 breast cancer risk loci may be mediating their effect through IGFBP5.
Sequencing-based molecular characterization of tumors provides information required for individualized cancer treatment. There are well-defined molecular subtypes of breast cancer that provide improved prognostication compared to routine biomarkers. However, molecular subtyping is not yet implemented in routine breast cancer care. Clinical translation is dependent on subtype prediction models providing high sensitivity and specificity. In this study we evaluate sample size and RNA-sequencing read requirements for breast cancer subtyping to facilitate rational design of translational studies. We applied subsampling to ascertain the effect of training sample size and the number of RNA sequencing reads on classification accuracy of molecular subtype and routine biomarker prediction models (unsupervised and supervised). Subtype classification accuracy improved with increasing sample size up to N = 750 (accuracy = 0.93), although with a modest improvement beyond N = 350 (accuracy = 0.92). Prediction of routine biomarkers achieved accuracy of 0.94 (ER) and 0.92 (Her2) at N = 200. Subtype classification improved with RNA-sequencing library size up to 5 million reads. Development of molecular subtyping models for cancer diagnostics requires well-designed studies. Sample size and the number of RNA sequencing reads directly influence accuracy of molecular subtyping. Results in this study provide key information for rational design of translational studies aiming to bring sequencing-based diagnostics to the clinic.
breast cancer loci at Abstract We recently identi fi ed a novel susceptibility variant, rs865686, for estrogen-receptor positive breast cancer at Here, we report a fi ne-mapping analysis of the 9q31.2 susceptibility locus using 43 160 cases and 42 600 controls of European ancestry ascertained from 52 studies and a further 5795 cases and 6624 controls of Asian ancestry from nine studies. Single nucleotide polymorphism (SNP) rs676256 was most strongly associated with risk in Europeans (odds ratios [OR] = 0.90 [0.88 – 0.92]; P -value = 1.58 × 10 − 25 ). This SNP is one of a cluster of highly correlated variants, including rs865686, that spans ∼ 14.5 kb. We identi fi ed two additional independent association signals demarcated by SNPs rs10816625 (OR = 1.12 [1.08 – 1.17]; P -value = 7.89 × 10 − 09 ) and rs13294895 (OR = 1.09 [1.06 – 1.12]; P -value= 2.97 × 10 − 11 ). SNP rs10816625, but not rs13294895, was also associated with risk of breast cancer in Asian individuals (OR = 1.12 [1.06 – 1.18]; P -value= 2.77× 10 − 05 ). Functional genomic annotation using data derived from breast cancer cell-line models indicates that these SNPs localise to putative enhancer elements that bind known drivers of hormone-dependent breast cancer, including ER- α , FOXA1 and GATA-3. In vitro analyses indicate that rs10816625 and rs13294895 have allele-speci fi c effects on enhancer activity and suggest chromatin interactions with the KLF4 gene locus. These results demonstrate the power of dense genotyping in large studies to identify independent susceptibility variants. Analysis of associations using subjects with different ancestry, combined with bioinformatic and genomic characterisation, can provide strong evidence for the likely causative alleles and their functional basis.
We recently identified a novel susceptibility variant, rs865686, for estrogen-receptor positive breast cancer at 9q31.2. Here, we report a fine-mapping analysis of the 9q31.2 susceptibility locus using 43 160 cases and 42 600 controls of European ancestry ascertained from 52 studies and a further 5795 cases and 6624 controls of Asian ancestry from nine studies. Single nucleotide polymorphism (SNP) rs676256 was most strongly associated with risk in Europeans (odds ratios [OR] = 0.90 [0.88-0.92]; P-value = 1.58 × 10(-25)). This SNP is one of a cluster of highly correlated variants, including rs865686, that spans ∼14.5 kb. We identified two additional independent association signals demarcated by SNPs rs10816625 (OR = 1.12 [1.08-1.17]; P-value = 7.89 × 10(-09)) and rs13294895 (OR = 1.09 [1.06-1.12]; P-value = 2.97 × 10(-11)). SNP rs10816625, but not rs13294895, was also associated with risk of breast cancer in Asian individuals (OR = 1.12 [1.06-1.18]; P-value = 2.77 × 10(-05)). Functional genomic annotation using data derived from breast cancer cell-line models indicates that these SNPs localise to putative enhancer elements that bind known drivers of hormone-dependent breast cancer, including ER-α, FOXA1 and GATA-3. In vitro analyses indicate that rs10816625 and rs13294895 have allele-specific effects on enhancer activity and suggest chromatin interactions with the KLF4 gene locus. These results demonstrate the power of dense genotyping in large studies to identify independent susceptibility variants. Analysis of associations using subjects with different ancestry, combined with bioinformatic and genomic characterisation, can provide strong evidence for the likely causative alleles and their functional basis.
CELL BIOLOGY Correction to Supporting Information for “Majority of differentially expressed genes are down-regulated during malignant transformation in a four-stage model,” by Frida Danielsson, Marie Skogs, Mikael Huss, Elton Rexhepaj, Gillian O’Hurley, Daniel Klevebring, Fredrik Pontén, Annica K. B. Gad, Mathias Uhlén, and Emma Lundberg, which appeared in issue 17, April 23, 2013, of Proc Natl Acad Sci USA (110:6853–6858; first published April 8, 2013; 10.1073/pnas.1216436110). The authors note that Dataset S2 appeared incorrectly. The SI has been corrected online.
Women with contralateral breast cancer (CBC) have significantly worse prognosis compared to women with unilateral cancer. A possible explanation of the poor prognosis of patients with CBC is that in a subset of patients, the second cancer is not a new primary tumor but a metastasis of the first cancer that has potentially obtained aggressive characteristics through selection of treatment. Exome and whole-genome sequencing of solid tumors has previously been used to investigate the clonal relationship between primary tumors and metastases in several diseases. In order to assess the relationship between the first and the second cancer, we performed exome sequencing to identify somatic mutations in both first and second cancers, and compared paired normal tissue of 25 patients with metachronous CBC. For three patients, we identified shared somatic mutations indicating a common clonal origin thereby demonstrating that the second tumor is a metastasis of the first cancer, rather than a new primary cancer. Accordingly, these patients all developed distant metastasis within 3 years of the second diagnosis, compared with 7 out of 22 patients with non-shared somatic profiles. Genomic profiling of both tumors help the clinicians distinguish between true CBCs and subsequent metastases.