Structural variants (SVs) account for over 60% of pediatric cancer driver variants. Pan-cancer analyses on 1,616 pediatric and 2,203 adult whole genomes show that pediatric SV burden varies ∼100-fold across cancer types, is reduced 6- to 16-fold compared to adult brain and solid tumors, but is comparable in hematological malignancies. The top-ranked SV-disrupted genes are drivers in pediatric cancers and fragile sites in adult cancers. Recurrent SV hotspots near RAG recombination signal sequences disrupt immune loci and driver genes in pediatric acute lymphoblastic leukemias, but immune loci exclusively in adult lymphoid cancers. Ten extracted SV signatures implicate RAG-mediated mutagenesis as a potential etiology for COSMIC SV7 in lymphoid cancers, while clustering of spatiotemporally distinct samples from 13 patients reveals the ongoing evolutionary contributions of SVs to intra-tumor heterogeneity and driver selection. Our study expands the known scope of RAG-mediated mutagenesis, while the curated SV dataset can guide future research and clinical testing.
Abstract Recent advances in multi-omics profiling have accelerated the discovery of molecular targets in pediatric cancers. However, clinical interpretation remains constrained by evolving diagnostic standards and limited representation of rare subtypes. To address this, we developed the Cancer Classifications for Kids (CC4K) - a harmonized, molecular classification-driven framework aligned with WHO tumor classification and recent publication standards for pediatric tumors. Using pathogenic variant data hosted on St. Jude Cloud PeCan Knowledge Base (https://pecan.stjude.cloud), we classified 230 subtypes for hematological malignancies (n=70), solid tumors (n=97), and brain tumors (n=63). Most recently, the pathogenic point mutations, CNVs, and gene fusions from ∼1,511 paired tumor-normal samples, profiled by the ongoing NCI’s Childhood Cancer Data Initiative (CCDI), were integrated into PeCan, extending the subtype repertoire by ∼53% (80 new subtypes). Importantly, classification of additional subtypes required aligning molecular data with clinical features, which revealed 16 evidence categories, including “biomarker-confirmed” (n=440) and “rescued” (n=353). Furthermore, the integration of additional multi-modal approaches provided clarity on existing classifications with ambiguous or conflicting data. For example, integrating the data from the Molecular Characterization Initiative (MCI) improved our definition of several previously ambiguous cases including a small round blue cell tumor redefined as Ewing sarcoma following identification of a novel EWSR1::FUS reciprocal fusion event, reclassification of an ependymoma as intracranial mesenchymal tumor, FET::CREB-fusion positive, and validation of an atypical NRAS-positive alveolar rhabdomyosarcoma. Additionally, the integration of CCDI data fine-tuned our knowledgebase on the therapy-relevant molecular drivers such as activation of the Hedgehog signaling pathway in embryonal rhabdomyosarcoma and activation of the PI-3K pathway by recurrent AKT hotspot mutations in multiple cancer types. Distribution of tumor mutation burdens from each cancer subtype revealed hypermutators with distinct etiologies as identified through subsequent mutational signature analyses. Collectively, these results demonstrate the importance of a harmonized framework for systematic cross-cohort integration that advances diagnostic precision through molecular-pathologic consensus, helps unravel the complex landscape of rare pediatric tumor subtypes, and lays the groundwork for future therapeutic and classification refinements across the pediatric oncology ecosystem. Citation Format: Stephanie Sandor, Delaram Rahbarinia, Yuan Feng, Ramzi Alsallaq, Van L. Nguyen, Daniel K. Putnam, David Finkelstein, Jinman Park, Bo Wang, Jobin Sunny, Jian Wang, Sue Qiu, Michael Edmonson, Robert Greenhalgh, Meghann Kirk, Ira Baranova, Stephen Rice, Abbas Shirinifard, Hoaran Chen, Ali F. Pour, Clay McLeod, Lu Wang, Jeffery Klco, Brent Orr, Michael Dyer, Xiang Chen, Xiaotu Ma, Michael Rusch, Jinghui Zhang. Advancing pediatric tumor subtype classification on the pediatric cancer (PeCan) knowledge base by integrating molecular and morphology data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 3488.
Cancer subtype classification is critical for precision therapy and there is a growing trend of augmenting histopathology testing procedures with omics-based machine learning classifiers. However, analytical challenges remain for pediatric cancer on the scope and precision of the current classifiers as well as the evolving subtype standardization. To address these challenges, we built Cancer Identification or CanID, a stacked ensemble machine learning classification scheme, using the transcriptomic features derived from gene-level RNA sequencing count data as the sole input. CanID was developed primarily from 3203 pediatric cancer samples of 13 solid tumor subtypes and 38 hematologic malignancy subtypes with subtype labels curated without the use of RNA-seq data. The accuracies of independent testing in three independent or external data sets for Solid Tumor and Hematologic Malignancy are 99% and 92%–93%, respectively. Notably, CanID was able to classify subtypes challenging for clinical histology evaluation and was robust to both biological and technical challenges, including differences in data collection protocols, class imbalance, potential mislabeled training samples and classes unobserved in training. The high accuracy, robustness, biological interpretability of this transcriptome-based classification scheme represents a valuable approach to advance tumor diagnosis and clinically meaningful stratification of tumor types. CanID can be accessed on GitHub at https://github.com/chenlab-sj/CanID.
Abstract Broad copy number alterations (CNAs) at chromosomal band resolution have been used as cancer diagnostic and prognostic biomarkers for decades. These events have been characterized by cytogenetic imaging, an approach which is powerful for assessing heterogeneity but limited for locus fine mapping. While CNA detection by next-generation sequencing has become a standard analysis, existing methods rarely model CNA heterogeneity in bulk tumor samples and often require a paired normal sample for control. To overcome these limitations, we developed Seq2Karyotype (S2K, https://github.com/chenlab-sj/Seq2Karyotype), a new algorithm for in-silico karyotyping using single-sample whole-genome sequencing (WGS) data. S2K performs joint modeling of read-depth and allelic imbalance (AI) of high-quality heterozygous SNPs to identify reference diploid regions, followed by modeling and segmenting the deviation from the reference. Empirical coverage and AI of segmented regions are fitted to models of single and admixed CNAs to estimate clonality. The final karyotyping considers both the model fitness and minimization of evolution steps.To evaluate S2K’s performance, we analyzed two cell lines commonly used for benchmark test: COLO829 (melanoma) and HCT1395 (breast cancer); both had single-cell (sc) WGS for validation. For COLO829, S2K replicated the four populations detected by scWGS but derived a different clonality estimate. The predominant clone, defined by loss of 1p, 10p, and chr18, was estimated to have 67% cellular fraction (CF) by bulk sample analysis of S2K in contrast to the 10% CF by scWGS analysis. In HCT1395, S2K identified three new CNAs present in >50% of the cells of the matching germline sample which were subsequently validated by FISH, karyotyping and SKY mapping.To demonstrate S2K’s utility, we analyzed three diverse data sets: 17 neuroblastoma cell lines, 24 pediatric AML samples with karyotyping data, and two blood samples from children with myelodysplastic syndromes (MDS) known to harbor mosaic uniparental disomy (UPD). CNA-based intra-tumor heterogeneity was detected in 88% (15/17) of the neuroblastoma cell lines, comprised of 2-4 distinct populations with diverse ranges of CF (~10%-90%) and varying CNA patterns (e.g. admixture of tetraploid and diploid cells), which were validated by cytogenetics or scWGS. In the patient AML samples, S2K detected >95% of the previously reported cytogenetic events and 30% of additional copy-neutral loss-of-heterozygosity events. The mosaic UPD events in MDS patients were detected with projected clonality of 60% and 25%, respectively. These results not only demonstrate the accuracy of in-silico karyotyping performed by S2K but also reveal the dynamic intra-tumor heterogeneity in cancer cell lines, which may impact the design and interpretation of future experiments using these cell lines. Citation Format: Limeng Pu, Karol Szlachta, Virginia Valentine, Xiaolong Chen, Jian Wang, Dennis Kennetz, Daniel Putnam, Sivaraman Natarajan, Li Dong, Thomas Look, Marcin Wlodarski, Lu Wang, Steven Burden, John Easton, Xiang Chen, Jinghui Zhang. Seq2Karyotype (S2K): A method for deconvoluting heterogeneity of copy number alterations using single-sample whole-genome sequencing data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 7419.
Osteosarcoma clinical history and capture validation sequencing data. Supplementary Table S1. Sample and treatment information for relapsed osteosarcoma patients. Supplementary Table S2. SJOS0011101 SNVs detected by capture validation and tumor purity. Supplementary Table S3. SJOS0011105 SNVs detected by capture validation and tumor purity. Supplementary Table S4. SJOS0011107 SNVs detected by capture validation and tumor purity. Supplementary Table S5. SJOS010 SNVs detected by capture validation and tumor purity.
Supplementary Figures S1-S11 show sample histology, quality control data, and additional genomic analysis. S1. Histology of osteosarcoma SJOS001101. S2. Whole-genome sequencing coverage metrics. S3. Deep capture validation coverage metrics. S4. Example genes experiencing copy number alterations. S5. B-allele frequencies confirm homozygous deletions of 3q13.31 and CDKN2A. S6. Identification of cisplatin mutation signature in osteosarcoma. S7. Cisplatin does not induce kataegis. S8. Density plot analysis of SNV clusters reveals unique clones at most sites. S9. SJOS001101 pairwise osteosarcoma sample SNV comparisons reveal cross-seeding. S10. Clonal evolution of SJOS001101 as clarified by MACHINA analysis. S11. Pairwise osteosarcoma sample SNV comparisons in three additional patients.
DNA methylation (DNAm) is a relatively stable regulatory epigenetic mechanism. Although accurate DNAm classifiers have been built for early detection of cancer and subtype classifications, its transcriptional regulatory roles remain unclear. To address this challenge, we developed a deep-learning framework MethylationToActivity (M2A) that accurately predicts individual promoter activity from WGBS methylomes. As array-based methylome is the most common epigenetic data from patients’ tumor samples, we redesigned M2A’s raw features and model topology for array methylomes, which shows an accurate prediction of the promoter activity (R2 = 0.74), approaching its WGBS counterpart (R2 = 0.79). Transfer learning improves the prediction accuracy to 0.77. In a primary rhabdomyosarcoma cohort, M2A prioritized candidate genes with strong subtype-specific alternative promoter usage (APU) and identified distinct isoforms of a subgroup of APU genes including NAV2 expressed in the two major subtypes (aRMS and eRMS). Molecular experiments revealed that PAX3-FOXO1 directly binds to the aRMS-specific promoter and is required for its active transcription. Genetic deletion of PAX3-FOXO1 leads to NAV2 APU. Overexpression of aRMS-specific NAV2 isoform significantly promoted cell proliferation, supporting its oncogenic function. These data indicate that PAX3-FOX1 is critical for APU in aRMS. We conclude that the new M2A provides vast insight into downstream functional interpretation of differential DNAm patterns. Citation Format: Karissa Dieseldorff Jones, Waise Quarni, Daniel Putnam, Shivendra Singh, Qiong Wu, Jun Yang, Xiang Chen. MicroArray-based MethylationToActivity: Advancing biological knowledges from DNA methylation profiles of patient tumors. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 5363.
Gene regulation is critical for cell identity, and its dysregulation is a defining characteristic of common diseases including cancers. Although promoter activity is a strong predictor of gene expression, activities at other regulatory elements, including enhancers and super-enhancers (SE), are major contributors to gene regulation. For example, distinct transcription factor-regulating SEs define neuroblastoma (NB) subtypes with distinct clinical outcomes. However, technical limitations largely prevent ChIP-seq based enhancer/SE activity profiling in primary patient samples. We previously developed MethylationToActivity (M2A) and demonstrated that high-order DNA methylation (DNAm) features are strong predictors of promoter activity measured by ChIP-seq. This tool is important because genomewide DNAm assays are widely used in clinic to classify tumors and stratify patients on clinical trials. However, it is limited by its reliance on known gene position annotations. Here we present MethylationToRegulation (M2R), a method that utilizes a convolutional neural network (CNN)-based deep learning framework to infer intergenic enhancer/SE activities from DNA methylomes. We obtained paired WGBS and H3K27ac ChIP-seq data for 16 pediatric NB samples profiled in the Pediatric Cancer Genome Project. M2R was trained on 6 samples and tested on the remaining 10 samples. It achieved an average relative prediction accuracy of 84% on test samples when compared to the H3K27ac ChIP-seq replicate consistency from the ENCODE project. We adapted the Rank Ordering of Super-Enhancers algorithm to interpret H3K27ac signals inferred from DNA methylomes to identify SEs. M2R faithfully captured subtype-defining SEs with high specificity, including those associated with master NB transcription factors. Our results demonstrate that M2R accurately quantifies enhancer/SE activities and infers critical epigenetic marks from DNA methylomes. Application of M2R will enable the profiling of epigenetic dysregulation in patient tumor samples, seeking to improve clinical outcomes. Citation Format: Daniel K. Putnam, Brian J. Abraham, Xiang Chen. MethylationToRegulation: A deep-learning approach to infer chromatin properties from DNA methylomes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 1936.
Improvements in whole genome amplification (WGA) would enable new types of basic and applied biomedical research, including studies of intratissue genetic diversity that require more accurate single-cell genotyping. Here, we present primary template-directed amplification (PTA), an isothermal WGA method that reproducibly captures >95% of the genomes of single cells in a more uniform and accurate manner than existing approaches, resulting in significantly improved variant calling sensitivity and precision. To illustrate the types of studies that are enabled by PTA, we developed direct measurement of environmental mutagenicity (DMEM), a tool for mapping genome-wide interactions of mutagens with single living human cells at base-pair resolution. In addition, we utilized PTA for genome-wide off-target indel and structural variant detection in cells that had undergone CRISPR-mediated genome editing, establishing the feasibility for performing single-cell evaluations of biopsies from edited tissues. The improved precision and accuracy of variant detection with PTA overcomes the current limitations of accurate WGA, which is the major obstacle to studying genetic diversity and evolution at cellular resolution.
Although genome-wide DNA methylomes have demonstrated their clinical value as reliable biomarkers for tumor detection, subtyping, and classification, their direct biological impacts at the individual gene level remain elusive. Here we present MethylationToActivity (M2A), a machine learning framework that uses convolutional neural networks to infer promoter activities based on H3K4me3 and H3K27ac enrichment, from DNA methylation patterns for individual genes. Using publicly available datasets in real-world test scenarios, we demonstrate that M2A is highly accurate and robust in revealing promoter activity landscapes in various pediatric and adult cancers, including both solid and hematologic malignant neoplasms.
Abstract Transcriptional regulation is fundamental for cell identity and function, and it's deregulation is a defining feature for common diseases, including cancers. Besides directly modifying the canonical promoter activities, tumors frequently utilize alternative promoters to increase the isoform diversity, activate oncogenes when the canonical promoters are repressed, and to evade host immune attacks by immune-editing. However, fresh-tissue availability and technical challenges represent major obstacles in genomewide promoter activity assessment in patient tumors by ChIP-seq experiments. Here, we presented MethylToActivity (M2A), a deep-learning framework that reveals promoter activities (H3K4me3 and H3K27ac) from DNA methylation (DNAm) patterns, a relatively stable epigenetic regulatory mechanism that can be robustly and accurately profiled in various tissues, including retrospective tumor samples archived using the formalin-fixed paraffin-embedded (FFPE) method. Trained from a cohort of neuroblastomas (N = 6), we demonstrate that M2A 1) approaches the accuracies of ChIP-seq experiments in revealing promoter activities, 2) is generalizable to various pediatric and adult cancers, 3) captures changes of promoter activity associated with differentially methylated regions, and 4) faithfully recapitulates differential promoter usages among tumor subtypes. Importantly, M2A uncovers oncogenic activation of alternative promoters in genes critical for tumor survival while the canonical promoter remains inactive. These results substantiate that M2A is capable of accurately measuring promoter activities, which will be of great use not only to functionally interpret differential DNAm patterns, but to unveil alternative promoter usages in patient tumors, which will facilitate precision medicine by tailoring treatments based on both epigenetic deregulations and genetic variants. Citation Format: Justin Williams, Beisi Xu, Daniel Putnam, Xiang Chen. DNA methylation reveals alternative promoter usage in genes critical to pediatric tumors [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 6576.
BACKGROUND:We aimed to systematically evaluate telomere dynamics across a spectrum of pediatric cancers, search for underlying molecular mechanisms, and assess potential prognostic value.METHODS:The fraction of telomeric reads was determined from whole-genome sequencing data for paired tumor and normal samples from 653 patients with 23 cancer types from the Pediatric Cancer Genome Project. Telomere dynamics were characterized as the ratio of telomere fractions between tumor and normal samples. Somatic mutations were gathered, RNA sequencing data for 330 patients were analyzed for gene expression, and Cox regression was used to assess the telomere dynamics on patient survival.RESULTS:Telomere lengthening was observed in 28.7% of solid tumors, 10.5% of brain tumors, and 4.3% of hematological cancers. Among 81 samples with telomere lengthening, 26 had somatic mutations in alpha thalassemia/mental retardation syndrome X-linked gene, corroborated by a low level of the gene expression in the subset of tumors with RNA sequencing. Telomerase reverse transcriptase gene amplification and/or activation was observed in 10 tumors with telomere lengthening, including two leukemias of the E2A-PBX1 subtype. Among hematological cancers, pathway analysis for genes with expressions most negatively correlated with telomere fractions suggests the implication of a gene ontology process of antigen presentation by Major histocompatibility complex class II. A higher ratio of telomere fractions was statistically significantly associated with poorer survival for patients with brain tumors (hazard ratio = 2.18, 95% confidence interval = 1.37 to 3.46).CONCLUSION:Because telomerase inhibitors are currently being explored as potential agents to treat pediatric cancer, these data are valuable because they identify a subpopulation of patients with reactivation of telomerase who are most likely to benefit from this novel therapeutic option.
More than 8,000 genes are turned on or off as progenitor cells produce the 7 classes of retinal cell types during development. Thousands of enhancers are also active in the developing retinae, many having features of cell- and developmental stage-specific activity. We studied dynamic changes in the 3D chromatin landscape important for precisely orchestrated changes in gene expression during retinal development by ultra-deep in situ Hi-C analysis on murine retinae. We identified developmental-stage-specific changes in chromatin compartments and enhancer-promoter interactions. We developed a machine learning-based algorithm to map euchromatin and heterochromatin domains genome-wide and overlaid it with chromatin compartments identified by Hi-C. Single-cell ATAC-seq and RNA-seq were integrated with our Hi-C and previous ChIP-seq data to identify cell- and developmental-stage-specific super-enhancers (SEs). We identified a bipolar neuron-specific core regulatory circuit SE upstream of Vsx2, whose deletion in mice led to the loss of bipolar neurons.
Abstract To investigate the genomic evolution of metastatic pediatric osteosarcoma, we performed whole-genome and targeted deep sequencing on 14 osteosarcoma metastases and two primary tumors from four patients (two to eight samples per patient). All four patients harbored ancestral (truncal) somatic variants resulting in TP53 inactivation and cell-cycle aberrations, followed by divergence into relapse-specific lineages exhibiting a cisplatin-induced mutation signature. In three of the four patients, the cisplatin signature accounted for >40% of mutations detected in the metastatic samples. Mutations potentially acquired during cisplatin treatment included NF1 missense mutations of uncertain significance in two patients and a KIT G565R activating mutation in one patient. Three of four patients demonstrated widespread ploidy differences between samples from the sample patient. Single-cell seeding of metastasis was detected in most metastatic samples. Cross-seeding between metastatic sites was observed in one patient, whereas in another patient a minor clone from the primary tumor seeded both metastases analyzed. These results reveal extensive clonal heterogeneity in metastatic osteosarcoma, much of which is likely cisplatin-induced. Implications: The extent and consequences of chemotherapy-induced damage in pediatric cancers is unknown. We found that cisplatin treatment can potentially double the mutational burden in osteosarcoma, which has implications for optimizing therapy for recurrent, chemotherapy-resistant disease.
VCF2CNA is a tool (Linux commandline or web-interface) for copy-number alteration (CNA) analysis and tumor purity estimation of paired tumor-normal VCF variant file formats. It operates on whole genome and whole exome datasets. To benchmark its performance, we applied it to 46 adult glioblastoma and 146 pediatric neuroblastoma samples sequenced by Illumina and Complete Genomics (CGI) platforms respectively. VCF2CNA was highly consistent with a state-of-the-art algorithm using raw sequencing data (mean F1-score = 0.994) in high-quality whole genome glioblastoma samples and was robust to uneven coverage introduced by library artifacts. In the whole genome neuroblastoma set, VCF2CNA identified MYCN high-level amplifications in 31 of 32 clinically validated samples compared to 15 found by CGI’s HMM-based CNA model. Moreover, VCF2CNA achieved highly consistent CNA profiles between WGS and WXS platforms (mean F1 score 0.97 on a set of 15 rhabdomyosarcoma samples). In addition, VCF2CNA provides accurate tumor purity estimates for samples with sufficient CNAs. These results suggest that VCF2CNA is an accurate, efficient and platform-independent tool for CNA and tumor purity analyses without accessing raw sequence data.
The nuclei of rod photoreceptors in mice and other nocturnal species have an unusual inverted chromatin structure: the heterochromatin is centrally located to help focus light and improve photosensitivity. To better understand this unique nuclear organization, we performed ultra-deep Hi-C analysis on murine retina at 3 stages of development and on purified rod photoreceptors. Predicted looping interactions from the Hi-C data were validated with fluorescence in situ hybridization (FISH). We discovered that a subset of retinal genes that are important for retinal development, cancer, and stress response are localized to the facultative heterochromatin domain. We also used machine learning to develop an algorithm based on our chromatin Hidden Markov Modeling (chromHMM) of retinal development to predict heterochromatin domains and study their dynamics during retinogenesis. FISH data for 264 genomic loci were used to train and validate the algorithm. The integrated data were then used to identify a developmental stage– and cell type-specific core regulatory circuit super-enhancer (CRC-SE) upstream of the Vsx2 gene, which is required for bipolar neuron expression. Deletion of the Vsx2 CRC-SE in mice led to the loss of bipolar neurons in the retina.
Abstract Whole genome sequencing (WGS) is increasingly used in both research and clinical settings. The Variant Call Format (VCF) specification is a widely adopted file format for genetic variation data exchange partially due to its smaller file size compared to raw WGS BAMs. Each variant in a typical VCF file contains its chromosome position, reference/alternative alleles and corresponding allele counts. This makes it possible to identify copy number alterations (CNAs). To this end, we developed VCF2CNA (http://vcf2cna.stjude.org), a web interface tool for CNA analysis from VCF files. A user of VCF2CNA, uploads a VCF file via the provided web interface. The entire analysis runs remotely with an average run time of 23 minutes. Results are emailed to the user as either a downloadable link or file attachments. VCF2CNA also accepts input in the Mutation Annotation Format (MAF) and the variant file format produced by the Bambino program. We analyzed 22 TCGA glioblastoma tumor/normal pairs by Illumina technology to evaluate VCF2CNA’s performance. It achieved high consistency (average F1-score: 0.952 ± 0.082) with CONSERTING, a tool that incorporated read-depth and SV data from raw BAMs for CNA detection. A segment-by-segment comparison between results from CONSERTING and VCF2CNA indicated that the latter was less sensitive to focal CNAs. This is expected because there is less information in the VCF input than in raw BAMs. Further analysis using samples with a “fractured genome” pattern revealed that VCF2CNA was more robust to library artifacts and produced relatively clean CNA profiles (on average 76.2-fold reduction compared to the number of segments reported by CONSERTING). Finally, we analyzed 137 pediatric neuroblastoma samples from the TARGET project, sequenced by Complete Genomics, Inc. (CGI) technology. MYCN amplification has been clinically validated in 33 samples. VCF2CNA identified high amplitude MYCN gains in 32 samples and the remaining sample carried a low-level broad gain covering MYCN. For comparison, CGI’s HMM-based method reported MYCN gains in only 15 out of the 33 samples. VCF2CNA further identified two additional MYCN amplifications among the remaining samples. Collectively, our analysis suggests that VCF2CNA is a platform-independent, efficient, robust and accurate tool for general WGS-based CNA analysis. It further complements CONSERTING, which produces more accurate result in focal CNAs at the cost of significantly higher computational burden. Citation Format: Daniel K. Putnam, Xiaotu Ma, Stephen V. Rice, Yu Liu, Jinghui Zhang, Xiang Chen. VCF2CNA: a tool for efficiently detecting copy number alteration using VCF genotype data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 2587. doi:10.1158/1538-7445.AM2017-2587
Whole genome sequencing (WGS) is increasingly used in both research and clinical settings. The Variant Call Format (VCF) specification is a widely adopted file format for genetic variation data exchange partially due to its smaller file size compared to raw WGS BAMs. Each variant in a typical VCF file contains its chromosome position, reference/alternative alleles and corresponding allele counts. This makes it possible to identify copy number alterations (CNAs). To this end, we developed VCF2CNA (http://vcf2cna.stjude.org), a web interface tool for CNA analysis from VCF files. A user of VCF2CNA, uploads a VCF file via the provided web interface. The entire analysis runs remotely with an average run time of 23 minutes. Results are emailed to the user as either a downloadable link or file attachments. VCF2CNA also accepts input in the Mutation Annotation Format (MAF) and the variant file format produced by the Bambino program. We analyzed 22 TCGA glioblastoma tumor/normal pairs by Illumina technology to evaluate VCF2CNA’s performance. It achieved high consistency (average F1-score: 0.952 ± 0.082) with CONSERTING, a tool that incorporated read-depth and SV data from raw BAMs for CNA detection. A segment-by-segment comparison between results from CONSERTING and VCF2CNA indicated that the latter was less sensitive to focal CNAs. This is expected because there is less information in the VCF input than in raw BAMs. Further analysis using samples with a “fractured genome” pattern revealed that VCF2CNA was more robust to library artifacts and produced relatively clean CNA profiles (on average 76.2-fold reduction compared to the number of segments reported by CONSERTING). Finally, we analyzed 137 pediatric neuroblastoma samples from the TARGET project, sequenced by Complete Genomics, Inc. (CGI) technology. MYCN amplification has been clinically validated in 33 samples. VCF2CNA identified high amplitude MYCN gains in 32 samples and the remaining sample carried a low-level broad gain covering MYCN. For comparison, CGI’s HMM-based method reported MYCN gains in only 15 out of the 33 samples. VCF2CNA further identified two additional MYCN amplifications among the remaining samples. Collectively, our analysis suggests that VCF2CNA is a platform-independent, efficient, robust and accurate tool for general WGS-based CNA analysis. It further complements CONSERTING, which produces more accurate result in focal CNAs at the cost of significantly higher computational burden. Citation Format: Daniel K. Putnam, Xiaotu Ma, Stephen V. Rice, Yu Liu, Jinghui Zhang, Xiang Chen. VCF2CNA: a tool for efficiently detecting copy number alteration using VCF genotype data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 2587. doi:10.1158/1538-7445.AM2017-2587