Ductal Carcinoma In Situ (DCIS) is a precursor lesion of Invasive Ductal Carcinoma (IDC) of the breast. Investigating its temporal progression could provide fundamental new insights for the development of better diagnostic tools to predict which cases of DCIS will progress to IDC. We investigate the problem of reconstructing a plausible progression from single-cell sampled data of an individual with synchronous DCIS and IDC. Specifically, by using a number of assumptions derived from the observation of cellular atypia occurring in IDC, we design a possible predictive model using integer linear programming (ILP). Computational experiments carried out on a preexisting data set of 13 patients with simultaneous DCIS and IDC show that the corresponding predicted progression models are classifiable into categories having specific evolutionary characteristics. The approach provides new insights into mechanisms of clonal progression in breast cancers and helps illustrate the power of the ILP approach for similar problems in reconstructing tumor evolution scenarios under complex sets of constraints.
Abstract We describe computational methods to compute likely evolutionary histories from tumor single-cell copy number data and next generation sequencing data and apply the methods to data collected from diverse types of tumors. Experimental techniques for assessing heterogeneity in tumor cell populations have undergone great advances, but these improvements have created a great need for more sophisticated computer algorithms capable of making sense of these data sources in terms of coherent models of tumor evolution. We have addressed this problem by developing computer algorithms for building phylogenetic trees describing evolution of individual tumors based on copy numbers of fluorescence in situ hybridization (FISH) probes from single cells in these tumors. These algorithms reconstruct evolutionary trees for observed cell populations so as to heuristically minimize the number of mutational events needed to explain the observed combinations of probe counts by evolution from a common diploid ancestral cell. We have extended this work from initial simple evolutionary models of evolution by single copy number changes to account for distinct mechanisms of evolution at the gene, chromosome, or whole-genome scale, with potentially different rates of evolution by mutation type. We have applied these algorithms to several FISH data sets, including cervical cancers probed for four genes (LAMP3, PROX1, PRKAA1 and CCND1) measured for up to 250 cells of paired primary and metastatic samples from 16 patients, head-and-neck cancers probed for four genes (TERC, CCND1, EGFR and TP53) measured on up to 250 cells per patient for 65 patients at four tumor stages, prostate cancers probed for six genes (TBL1XR1, CTTNBP2, MYC, PTEN, MEN1 and PDGFB) measured for up to 407 cells in 6 non-progressive and 7 progressive carcinomas, and breast cancers probed for eight genes (COX-2, MYC, CCND1, HER-2, ZNF217, DBC2, CDH1 and TP53) measured on up to 220 cells of paired of ductal carcinoma in situ and invasive ductal carcinoma samples from 13 patients. We have then applied statistical and machine learning analysis to examine the ability of these trees to classify tumors by stage or potential for progression. The evolutionary tree models reveal robust features of evolutionary processes distinguishing progression stages and predicting future progression that lead to improved classification accuracy relative to predictions from cellular heterogeneity data alone. Our software is freely available at ftp://ftp.ncbi.nlm.nih.gov/pub/FISHtrees. In continuing work, we are exploring extension of these approaches to better modeling and analysis of tumor evolution using single-cell sequencing data and to more detailed models of tumor evolution. Citation Format: Salim A. Chowdhury, Ayshwarya Subramanian, Alejandro A. Schäffer, Stanley E. Shackney, Darawalee Wangsa, Kerstin Heselmeyer-Haddad, Thomas Ried, Russell Schwartz. Inferring evolutionary models of tumor progression from single-cell heterogeneity data. [abstract]. In: Proceedings of the 105th Annual Meeting of the American Association for Cancer Research; 2014 Apr 5-9; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2014;74(19 Suppl):Abstract nr 5338. doi:10.1158/1538-7445.AM2014-5338
Reasoning that overexpression of multiple E2F-responsive genes might be a useful marker for RB1 dysfunction, we compiled a list of E2F-responsive genes from the literature and evaluated their expression in publicly available gene expression microarray data of patients with breast cancer, serous ovarian cancer, and prostate cancer. In breast cancer, a group of tumors was identified, each of which simultaneously overexpressed multiple E2F-responsive genes. Seventy percent of these genes were concerned with cell cycle progression, DNA repair, or mitosis. These E2F-responsive gene overexpressing (ERGO) tumors frequently exhibited additional evidence of Rb/E2F axis dysfunction, were mostly triple negative, and preferentially overexpressed multiple basal cytokeratins, suggesting that they overlapped substantially with the basal-like tumor subset. ERGO tumors were also identified in serous ovarian cancer and prostate cancer. In these cancer types, there was no evidence for a tumor subset comparable to the breast cancer basal-like subset. A core group of about 30 E2F-responsive genes were overexpressed in all three cancer types. Thus, it appears that disorders of the Rb/E2F axis can arise at multiple organ sites and produce tumors that simultaneously overexpress multiple E2F-responsive genes.
We present methods to construct phylogenetic models of tumor progression at the cellular level that include copy number changes at the scale of single genes, entire chromosomes, and the whole genome. The methods are designed for data collected by fluorescence in situ hybridization (FISH), an experimental technique especially well suited to characterizing intratumor heterogeneity using counts of probes to genetic regions frequently gained or lost in tumor development. Here, we develop new provably optimal methods for computing an edit distance between the copy number states of two cells given evolution by copy number changes of single probes, all probes on a chromosome, or all probes in the genome. We then apply this theory to develop a practical heuristic algorithm, implemented in publicly available software, for inferring tumor phylogenies on data from potentially hundreds of single cells by this evolutionary model. We demonstrate and validate the methods on simulated data and published FISH data from cervical cancers and breast cancers. Our computational experiments show that the new model and algorithm lead to more parsimonious trees than prior methods for single-tumor phylogenetics and to improved performance on various classification tasks, such as distinguishing primary tumors from metastases obtained from the same patient population.
Abstract Understanding tumors as evolutionary systems is an important area of study with far-reaching implications in diagnostic and treatment paradigms. Computational phylogenetics is a valuable method for inferring tumor evolution in terms of evolutionary trees, phylogenies, where paths in a tree correspond to possible tumor progression pathways. The location of specific cell-types and patient samples in the tree provide information on tumor sub-types and development of heterogeneity. We previously developed a tumor phylogeny inference pipeline for array comparative genome hybridization (aCGH)-based tumor copy number profiles. Steps in the pipeline included extraction of robust progression markers from the data, which could differentiate stages of tumor evolution or the different paths in the tree, and assigning amplification states to the inferred markers in those stages. We introduced a novel multi-sample model for amplicon identification and calling, HMMCNA, which jointly extracted markers from and assigned amplification states to small sets of tumor aCGH profiles. HMMCNA employs a Hidden Markov Model (HMM), a probabilistic model, to classify data into normal and amplified states based on an underlying distribution for the two copy number states and a hidden state space of possible amplification states. We assumed two possible amplification states per sample: normal (0) or amplified (1). Joint segmentation and calling is performed by identifying a most likely sequence of amplification states across all genomic sites probes and samples. This approach limits in the number of samples the HMM can handle since the number of possible hidden amplification states increases exponentially with the number of samples. Here, we present an extension of the approach to handle large datasets. We incorporate a heuristic prior to the HMM classification to reduce the hidden state space by first screening out amplification states not strongly supported at any individual genome coordinates. The introduction of this heuristic reduces the state space on average by 99%. We further reduce the set of possible amplification states based on the frequency of occurrence of the states by only allowing those states occuring at multiple aCGH probes or array genome coordinate. This step accounts for the presence of random noise in the data and gives a further reduction of 80%. We demonstrate the method on a breast tumor aCGH dataset comprising copy number profiles derived from sectioned biopsy samples (NCBI GEO GSE16672, Navin et al., 2010). Our method was able to quickly segment the data into sets of robust normal and amplified segments suitable for downstream phylogeny building. The amplicons inferred carried several known markers of tumor progression. Further steps include tuning the parameters of the HMM to handle noise-levels across different datasets. Citation Format: Ayshwarya Subramanian, Stanley Shackney, Russell Schwartz. Inference of tumor phylogenetic markers from large copy number datasets. [abstract]. In: Proceedings of the 104th Annual Meeting of the American Association for Cancer Research; 2013 Apr 6-10; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2013;73(8 Suppl):Abstract nr 5133. doi:10.1158/1538-7445.AM2013-5133
Motivation: Development and progression of solid tumors can be attributed to a process of mutations, which typically includes changes in the number of copies of genes or genomic regions. Although comparisons of cells within single tumors show extensive heterogeneity, recurring features of their evolutionary process may be discerned by comparing multiple regions or cells of a tumor. A useful source of data for studying likely progression of individual tumors is fluorescence in situ hybridization (FISH), which allows one to count copy numbers of several genes in hundreds of single cells. Novel algorithms for interpreting such data phylogenetically are needed, however, to reconstruct likely evolutionary trajectories from states of single cells and facilitate analysis of tumor evolution. Results: In this article, we develop phylogenetic methods to infer likely models of tumor progression using FISH copy number data and apply them to a study of FISH data from two cancer types. Statistical analyses of topological characteristics of the tree-based model provide insights into likely tumor progression pathways consistent with the prior literature. Furthermore, tree statistics from the resulting phylogenies can be used as features for prediction methods. This results in improved accuracy, relative to unstructured gene copy number data, at predicting tumor state and future metastasis. Availability: Source code for software that does FISH tree building (FISHtrees) and the data on cervical and breast cancer examined here are available at ftp://ftp.ncbi.nlm.nih.gov/pub/FISHtrees. Contact: sachowdh@andrew.cmu.edu Supplementary information: Supplementary data are available at Bioinformatics online.
Abstract Purpose: Characterizing the common pathways through which tumors progress is critical to understanding the molecular basis of cancer and developing effective treatments. Algorithms for phylogenetics, i.e., evolutionary tree building, can be used to infer progression pathways of single tumors when there is widespread intra-tumor heterogeneity from cell to cell. We describe computational methods to compute likely evolutionary histories from single-cell copy number data and apply the methods to single-cell gene copy number data derived from paired primary and metastatic samples of cervical cancers. Materials and methods: The inputs to our methods are lists of cell types for each patient sample, where a cell type is an array of copy numbers of gene probes measured by FISH. In the cervical cancer data set, copy numbers of four genes (LAMP3, PROX1, PRKAA1 and CCND1) were measured on up to 250 cells each of paired primary and metastatic samples from 16 patients. An example of a cell type is (4, 1, 3, 2). Our algorithms derive maximum parsimony phylogenetic trees explaining how all observed cell types might have evolved from a common diploid ancestor (2,2,2,2), seeking to explain the observed cells with as few copy number changes as possible across the tree. Due to the computational difficulty, we apply a heuristic algorithm to infer an approximately optimal tree by identifying likely unobserved ancestral cell states through a variant of the median joining method and then find the most parsimonious tree connecting these states. Our software also includes a method of computing consensus trees that facilitate the comparison of the tree models for each primary/metastasis sample pair. Results and Conclusion: Our methods allow us to build models of evolution of single tumors at the cellular level that provide insights into possible progression pathways and varying selective pressures in primary and metastatic sites. Analysis of tree geometry shows significant variability in evolutionary patterns between primary and metastatic sites consistent with more specific selective pressures in the metastases. Analysis of specific patterns of gene gain or loss in each environment shows a complicated portrait of tumor evolution with substantial variability from patient to patient. Citation Format: Salim A. Chowdhury, Alejandro A. Schäffer, Stanley E. Shackney, Darawalee Wangsa, Kerstin Heselmeyer-Haddad, Thomas Ried, Russell Schwartz. Phylogenetic models of tumor progression from fluorescence in situ hybridization (FISH) data on many single cells of a solid tumor. [abstract]. In: Proceedings of the 104th Annual Meeting of the American Association for Cancer Research; 2013 Apr 6-10; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2013;73(8 Suppl):Abstract nr 2895. doi:10.1158/1538-7445.AM2013-2895
Tumorigenesis can in principle result from many combinations of mutations, but only a few roughly equivalent sequences of mutations, or “progression pathways,” seem to account for most human tumors. Phylogenetics provides a promising way to identify common progression pathways and markers of those pathways. This approach, however, can be confounded by the high heterogeneity within and between tumors, which makes it difficult to identify conserved progression stages or organize them into robust progression pathways. To tackle this problem, we previously developed methods for inferring progression stages from heterogeneous tumor profiles through computational unmixing. In this paper, we develop a novel pipeline for building trees of tumor evolution from the unmixed tumor data. The pipeline implements a statistical approach for identifying robust progression markers from unmixed tumor data and calling those markers in inferred cell states. The result is a set of phylogenetic characters and their assignments in progression states to which we apply maximum parsimony phylogenetic inference to infer tumor progression pathways. We demonstrate the full pipeline on simulated and real comparative genomic hybridization (CGH) data, validating its effectiveness and making novel predictions of major progression pathways and ancestral cell states in breast cancers.
Computational cancer phylogenetics seeks to enumerate the temporal sequences of aberrations in tumor evolution, thereby delineating the evolution of possible tumor progression pathways, molecular subtypes, and mechanisms of action. We previously developed a pipeline for constructing phylogenies describing evolution between major recurring cell types computationally inferred from whole-genome tumor profiles. The accuracy and detail of the phylogenies, however, depend on the identification of accurate, high-resolution molecular markers of progression, i.e., reproducible regions of aberration that robustly differentiate different subtypes and stages of progression. Here, we present a novel hidden Markov model (HMM) scheme for the problem of inferring such phylogenetically significant markers through joint segmentation and calling of multisample tumor data. Our method classifies sets of genome-wide DNA copy number measurements into a partitioning of samples into normal (diploid) or amplified at each probe. It differs from other similar HMM methods in its design specifically for the needs of tumor phylogenetics, by seeking to identify robust markers of progression conserved across a set of copy number profiles. We show an analysis of our method in comparison to other methods on both synthetic and real tumor data, which confirms its effectiveness for tumor phylogeny inference and suggests avenues for future advances.
Abstract Computational cancer phylogenetics can play an important role in delineating possible tumor progression pathways and identifying molecular subtypes and mechanisms of action. We previously developed a pipeline for constructing tumor phylogenies from recurring cell types computationally inferred from whole genome copy number data. The accuracy and detail of these tumor phylogenies, however, depends on the identification of accurate and high-resolution molecular markers of progression, i.e., reproducible regions of copy number variation that can be used to robustly differentiate different subtypes and stages of progression. Here we present a new method for the problem using hidden Markov models (HMMs) to derive robust, high resolution progression markers from sets of tumor samples. We demonstrate our method on a publicly available array comparative genome hybridization (aCGH) dataset (NCBI GEO GSE16672, Navin et al., 2010) from sectioned primary ductal breast tumors. Our method uses an HMM, a class of probabilistic models, to classify sets of aCGH data into a partitioning of samples into normal (diploid) or amplified at each copy number probe. It differs from other similar HMM methods primarily in seeking a parsimonious set of combinations of amplification states able to explain all aCGH profiles simultaneously in order to identify robust markers of progression across samples. The model learns frequencies with which different combinations of amplifications are observed across the samples by modeling individual probes as Gaussian random variables with either normal or tetraploid means, with data more consistent with tetraploid being classified as amplified and those more consistent with diploid classified as normal. To handle a combinatorial explosion in combinations of amplification states with increasing numbers of samples, the method introduces a Gibbs sampling algorithm to learn a parsimonious model of the most frequently occurring combinations of amplification states. We applied our methods to a previously constructed set of inferred cell types derived from the Navin et al. data (Tolliver et al., 2010) and to a comparison set of 9 random samples from the raw aCGH data. We validated our model relative to manual labeling of amplicons on the same data. In both experiments, the HMM method was able to pick up significantly larger numbers of robustly amplified segments per chromosome than did prior methods or manual analysis. The resulting segments can be directly fed into downstream analysis routines for phylogeny inference or other predictions. In future work, the HMM method may be improved by fine-tuning the underlying model for copy number variation. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr 3964. doi:1538-7445.AM2012-3964
Abstract Computational methods based on phylogenetics, or the inference of evolutionary trees, have been useful for inferring likely pathways of tumor progression, but their use has been limited by our inability to precisely characterize the mutational states of individual cells in heterogeneous tumor samples. We describe a novel approach to tumor phylogenetics using computationally inferred sub-populations of cells derived from whole-genome data on heterogeneous tumor samples. We previously proposed a strategy for computationally separating cell types from heterogeneous tumor samples and inferring genome-wide profiles of RNA expression or DNA copy number for these discrete cell populations from profiles of the full tumor samples. In the present work, we extend this method by showing how one can use these computationally inferred cell states to identify markers informative of progression state and use them in phylogenetic analysis of tumor samples. We demonstrate the approach on a publicly available data set of array comparative genome hybridization (aCGH) data profiles applied to sectioned ductal breast carcinomas (Navin et al., 2010; NCBI GEO GSE16672). Our approach begins by applying our previously described method for cell type separation to generate inferred profiles of six cell types, and computationally inferring virtual aCGH profiles for these types. We then scan over sliding windows of 20 consecutive probes per window, identifying windows for which probe values are significantly aberrant relative to diploid by a modified chi-squared test (p-value 0.001 with Bonferroni correction), and merging overlapping windows to identify phylogenetically informative amplicons. We then identify, for each amplicon, those components that have significantly aberrant copy number by Gaussian test with diploid mean (p-value 0.001). Finally, the identified amplicon states for the inferred cell types, as well as an additional all-diploid normal type, are used as input for a maximum parsimony phylogenetic tree reconstruction using the PAUP package with 300 bootstrap replicates. The method applied to the Navin et al. data identified 15 amplicons phylogenetically informative for breast tumor progression. Bootstrap replicates from the phylogenetic tree reconstruction reveal a robust grouping of cell states into three distinct subgroups and suggest several possible progression pathways by which the identified amplicons may successively accumulate in particular breast cancer sub-types. Continuing work will examine robustness of the inferred markers and pathways to independent data sources. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 102nd Annual Meeting of the American Association for Cancer Research; 2011 Apr 2-6; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2011;71(8 Suppl):Abstract nr 44. doi:10.1158/1538-7445.AM2011-44
The development of cancer is largely driven by the gain or loss of subsets of the genome, promoting uncontrolled growth or disabling defenses against it. Identifying genomic regions whose DNA copy number deviates from the normal is therefore central to understanding cancer evolution. Array-based comparative genomic hybridization (aCGH) is a high-throughput technique for identifying DNA gain or loss by quantifying total amounts of DNA matching defined probes relative to healthy diploid control samples. Due to the high level of noise in microarray data, however, interpretation of aCGH output is a difficult and error-prone task. In this work, we tackle the computational task of inferring the DNA copy number per genomic position from noisy aCGH data. We propose CGHTRIMMER, a novel segmentation method that uses a fast dynamic programming algorithm to solve for a least-squares objective function for copy number assignment. CGHTRIMMER consistently achieves superior precision and recall to leading competitors on benchmarks of synthetic data and real data from the Coriell cell lines. In addition, it finds several novel markers not recorded in the benchmarks but plausibly supported in the oncology literature. Furthermore, CGHTRIMMER achieves superior results with run-times from 1 to 3 orders of magnitude faster than its state-of-art competitors. CGHTRIMMER provides a new alternative for the problem of aCGH discretization that provides superior detection of fine-scale regions of gain or loss yet is fast enough to process very large data sets in seconds. It thus meets an important need for methods capable of handling the vast amounts of data being accumulated in high-throughput studies of tumor genetics.
Motivation: Tumorigenesis is an evolutionary process by which tumor cells acquire sequences of mutations leading to increased growth, invasiveness and eventually metastasis. It is hoped that by identifying the common patterns of mutations underlying major cancer sub-types, we can better understand the molecular basis of tumor development and identify new diagnostics and therapeutic targets. This goal has motivated several attempts to apply evolutionary tree reconstruction methods to assays of tumor state. Inference of tumor evolution is in principle aided by the fact that tumors are heterogeneous, retaining remnant populations of different stages along their development along with contaminating healthy cell populations. In practice, though, this heterogeneity complicates interpretation of tumor data because distinct cell types are conflated by common methods for assaying the tumor state. We previously proposed a method to computationally infer cell populations from measures of tumor-wide gene expression through a geometric interpretation of mixture type separation, but this approach deals poorly with noisy and outlier data. Results: In the present work, we propose a new method to perform tumor mixture separation efficiently and robustly to an experimental error. The method builds on the prior geometric approach but uses a novel objective function allowing for robust fits that greatly reduces the sensitivity to noise and outliers. We further develop an efficient gradient optimization method to optimize this ‘soft geometric unmixing’ objective for measurements of tumor DNA copy numbers assessed by array comparative genomic hybridization (aCGH) data. We show, on a combination of semi-synthetic and real data, that the method yields fast and accurate separation of tumor states. Conclusions: We have shown a novel objective function and optimization method for the robust separation of tumor sub-types from aCGH data and have shown that the method provides fast, accurate reconstruction of tumor states from mixed samples. Better solutions to this problem can be expected to improve our ability to accurately identify genetic abnormalities in primary tumor samples and to infer patterns of tumor evolution. Contact: tolliver@cs.cmu.edu Supplementary information:Supplementary data are available at Bioinformatics online.
Background: While in principle a seemingly infinite variety of combinations of mutations could result in tumor development, in practice it appears that most human cancers fall into a relatively small number of "sub-types," each characterized a roughly equivalent sequence of mutations by which it progresses in different patients. There is currently great interest in identifying the common sub-types and applying them to the development of diagnostics or therapeutics. Phylogenetic methods have shown great promise for inferring common patterns of tumor progression, but suffer from limits of the technologies available for assaying differences between and within tumors. One approach to tumor phylogenetics uses differences between single cells within tumors, gaining valuable information about intra-tumor heterogeneity but allowing only a few markers per cell. An alternative approach uses tissue-wide measures of whole tumors to provide a detailed picture of averaged tumor state but at the cost of losing information about intra-tumor heterogeneity.Results: The present work applies "unmixing" methods, which separate complex data sets into combinations of simpler components, to attempt to gain advantages of both tissue-wide and single-cell approaches to cancer phylogenetics. We develop an unmixing method to infer recurring cell states from microarray measurements of tumor populations and use the inferred mixtures of states in individual tumors to identify possible evolutionary relationships among tumor cells. Validation on simulated data shows the method can accurately separate small numbers of cell states and infer phylogenetic relationships among them. Application to a lung cancer dataset shows that the method can identify cell states corresponding to common lung tumor types and suggest possible evolutionary relationships among them that show good correspondence with our current understanding of lung tumor development.Conclusions: Unmixing methods provide a way to make use of both intra-tumor heterogeneity and large probe sets for tumor phylogeny inference, establishing a new avenue towards the construction of detailed, accurate portraits of common tumor sub-types and the mechanisms by which they develop. These reconstructions are likely to have future value in discovering and diagnosing novel cancer sub-types and in identifying targets for therapeutic development.