BACKGROUND:Pedigree files are ubiquitously used within bioinformatics and genetics studies to convey critical information about relatedness, sex and affected status of study samples. While the text based format of ped files is efficient for computational methods, it is not immediately intuitive to a bioinformatician or geneticist trying to understand family structures, many of which encode the affected status of individuals across multiple generations. The visualization of pedigrees into connected nodes with descriptive shapes and shading provides a far more interpretable format to recognize visual patterns and intuit family structures. Despite these advantages of a visual pedigree, it remains difficult to quickly and accurately visualize a pedigree given a pedigree text file.RESULTS:Here we describe ped_draw a command line and web tool as a simple and easy solution to pedigree visualization. Ped_draw is capable of drawing complex multi-generational pedigrees and conforms to the accepted standards for depicting pedigrees visually. The command line tool can be used as a simple one liner command, utilizing graphviz to generate an image file. The web tool, https://peddraw.github.io , allows the user to either: paste a pedigree file, type to construct a pedigree file in the text box or upload a pedigree file. Users can save the generated image file in various formats.CONCLUSIONS:We believe ped_draw is a useful pedigree drawing tool that improves on current methods due to its ease of use and approachability. Ped_draw allows users with various levels of expertise to quickly and easily visualize pedigrees.
Many types of cancer are found to exhibit a high level of intrasample genetic heterogeneity, as a result of the microevolution process. As genomic aberrations accumulate within tumor cells, a subset of the cells may gain a selective advantage and form an outgrowth within the tumor. When chemotherapeutic agents are administered to the patient, these tumor subclones may respond differently to the treatment and may eventually cause relapse. Identifying resistant subclones, and their genotypes, after each chemotherapy would yield valuable insights into how the tumor population reacted to the previously administered chemo-agent, and how the treatment plan should be adjusted based on the genetic features of surviving subclones. Similarly, understanding the relationships between the primary and metastatic tumor cell populations would provide knowledge about the mechanisms driving, and potentially preventing, metastasis. Our method (SubcloneSeeker) is able to reconstruct subclone structures jointly across all samples (temporal or spatial), using somatic variants jointly clustered with their Cell Prevalence (CP), to infer a unified evolutionary structure. When more than one structure is able to explain the observed somatic variants and their CP values, we can compute a consensus structure that captures the most important features (i.e., which subclones and the corresponding genomic aberrations are responsible for the resistance of a certain chemotherapy drug), as well as suggest a panel of variants to be profiled with single-cell assay in order to reduce computational ambiguity. We applied this method to several datasets, including a breast cancer patient who had received three regimens of chemotherapy and had tumor samples taken at 4 consecutive time points. Our analysis revealed that the tumor population reacted to these treatment in two distinct patterns: 1) although a significant reduction in tumor burden was observed, most of the major subclones survived and expanded again after treatment, indicating that the treatment was effective, but terminated prematurely; 2) a single subclone with extra mutations (e.g., FBXL19, a tumor suppressor by inhibiting E-cadherin downregulation) survived the treatment and became the founding clone for the tumor population post-treatment, and the extra mutations were fixed at very high CP (~100%), indicating that these extra mutations are most likely to be causative for chemoresistance. We also applied our method to breast cancer patient rapid-autopsy datasets with 28 tumor samples. Our method was able to map out subclone evolution across multiple metastatic sites and highlight variants unique to each evolutionary lineage. In conclusion, our method is able to elucidate the dynamics of tumorigenesis, drug response, and metastasis by reconstructing unified subclone evolutionary structures, allowing more informed decisions to be made in adjusting treatment plans in real time and potentially facilitating personalized cancer medicine. Citation Format: Gabor Marth, Yi Qiao, Dillon Lee, Xiaomeng Huang. Real-time spatial and temporal monitoring on tumor subclonal evolution [abstract]. In: Proceedings of the AACR Special Conference on Advancing Precision Medicine Drug Development: Incorporation of Real-World Data and Other Novel Strategies; Jan 9-12, 2020; San Diego, CA. Philadelphia (PA): AACR; Clin Cancer Res 2020;26(12_Suppl_1):Abstract nr 39.
The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs (≥50 bp) per genome. We also discover 156 inversions per genome and 58 of the inversions intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a three to sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The methods and the dataset presented serve as a gold standard for the scientific community allowing us to make recommendations for maximizing structural variation sensitivity for future genome sequencing studies.
Abstract Several state-of-the-art, easy to use tools are available both for short-variant detection (e.g. GATK, FREEBAYES), and structural variant (SV) detection (e.g. LUMPY, MANTRA, DELLY), but these tools often produce divergent variant calls, especially INDELs, and it is very difficult to reconcile such variants into a single, accurate set. Furthermore, while it would be highly desirable to also detect larger, structural variants (SV), existing SV detector packages are typically difficult to integrate, highly resource-intensive to run, and result in call sets that require expert manual review to reduce false positive detection rate. Our algorithm, GRAPHITE (https://github.com/dillonl/graphite) requires as input a collection of variant calls, made by one or more short-variant or SV detection tools. Typically, this starting set is high sensitivity (i.e. inclusive), but low specificity (i.e. have a high false discovery rate). We then apply a novel “variant adjudication” procedure to discard false positives, while keeping true positive calls. This is accomplished by constructing a graph from these variants (the Variant Graph) representing allelic variants as graph branches, in addition to the branches formed by the current, linear genome reference sequence. Using a graph mapping algorithm (GSSW, a graph extension of the Smith-Waterman alignment algorithm) we developed earlier, we re-map all reads from each of the samples contributing to the candidate calls. We retain candidate variants confirmed by mappings to those branches in the graph that represent the corresponding variant allele, and discard those candidates that were not confirmed by such mappings. This procedure results in a highly specific callset that also maintains the high sensitivity of the inclusive starting callset constructed by multiple primary variant calling methods. Because the graph construction and mapping approach works for most types of SVs in addition to all short variants, variants of all different types can be integrated in a single step. Here we present the application of this method for cross-validating structural variants calls from Pacific Biosciences data by remapping deep Illumina WGS read sets to Variant Graphs constructed using the candidate Pacific Biosciences variants, as part of the Human Genome Structural Variation Consortium (HGSVC) data analysis project. We also present GRAPHITE's application to improving the accuracy of allele frequency measurement in tumor sequencing data, which is essential for the accurate reconstruction of subclonal evolution in longitudinal tumor samples. Citation Format: Dillon Lee, Yi Qiao, Gabor Marth. A graph remapping framework for in silico adjudication of SNVs, INDELs, and structural variants from genetic sequencing data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 3272.
Current tumor variant detection software tools are focused on the identification of somatic mutations in a single tumor sample (e.g. the primary tumor) compared to normal control tissue, and are not adequate for studies where multiple samples from the same patient, either at consecutive time points or at different metastatic sites are collected. Here we present a software framework that facilitates simultaneous analysis of such multi-sample datasets. Our framework utilizes our FreeBayes variant caller for jointly calling all tumor samples and normal controls from the patient. This ensures that even very low frequency mutations are called because read evidence is aggregated across all samples; and that allele frequency information for a called somatic variant is provided uniformly for all samples. Short and medium-length INDELs, and structural variants are identified by our RUFUS software, a new reference-free mutation calling algorithm for both germline and somatic mutation detection, with extremely high sensitivity and specificity. We use our Graphite tool, a graph-genome based read realignment algorithm, to eliminate false positive somatic calls i.e. inherited variants masquerading as somatic mutations, and to substantially increase accuracy of allele frequency estimation of the true somatic calls, a feature critical for downstream subclone analysis. In addition, we use the FACETS package for CNV and LOH analysis, also critical for downstream subclone analysis. Finally, we utilize SubcloneSeeker, our algorithm to reconstruct the evolution of the cancer at the subclonal level from the somatic variant allele frequencies, to understand how the tumor evolved over time, in response to multiple courses of treatment, or over space, across multiple metastatic sites. We demonstrate the use of this pipeline in published and ongoing studies involving longitudinal and multi-site tumor sample datasets. Citation Format: Yi Qiao, Xiaomeng Huang, Dillon Lee, Andrew Farrell, Thomas Nicholas, Brent Pederson, Aaron Quinlan, Gabor Marth. Utah somatic variant calling pipeline featuring multi-sample joint calling, variant-graph based accurate allele frequency estimation and subclone analysis [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 3280.