<p>Clonal cluster analysis. Mutations were clustered using DBSCAN based on the VAF and graphed according to timepoint. Panels A-T illustrate possible scenarios of how individual mutations may originate, evolve, and resolve based on VAF.</p>
Gene expression (array based) of recurrently mutated genes. Graph shows normalized expression of each patient and normal bone marrow (BM) controls for 9 of the 10 recurrently mutated genes; DHX30 was not represented on the array.
Clonal evolution of mutations from diagnosis to relapse. These graphs depict the Variant Allele Frequency (VAF) of each patient's mutations at diagnosis and relapse. Note: Mutations absent at diagnosis or relapse are depicted with VAF of zero. Gene names are noted to the left of each variant and are available in Supplemental Table 2.
Characterizing meiotic recombination rates across the genomes of nonhuman primates is important for understanding the genetics of primate populations, performing genetic analyses of phenotypic variation and reconstructing the evolution of human recombination. Rhesus macaques (Macaca mulatta) are the most widely used nonhuman primates in biomedical research. We constructed a high-resolution genetic map of the rhesus genome based on whole genome sequence data from Indian-origin rhesus macaques. The genetic markers used were approximately 18 million SNPs, with marker density 6.93 per kb across the autosomes. We report that the genome-wide recombination rate in rhesus macaques is significantly lower than rates observed in apes or humans, while the distribution of recombination across the macaque genome is more uniform. These observations provide new comparative information regarding the evolution of recombination in primates.
Large-scale, population-based genomic studies have provided a context for modern medical genetics. Among such studies, however, African populations have remained relatively underrepresented. The breadth of genetic diversity across the African continent argues for an exploration of local genomic context to facilitate burgeoning disease mapping studies in Africa. We sought to characterize genetic variation and to assess population substructure within a cohort of HIV-positive children from Botswana-a Southern African country that is regionally underrepresented in genomic databases. Using whole-exome sequencing data from 164 Batswana and comparisons with 150 similarly sequenced HIV-positive Ugandan children, we found that 13%-25% of variation observed among Batswana was not captured by public databases. Uncaptured variants were significantly enriched (p = 2.2 × 10-16) for coding variants with minor allele frequencies between 1% and 5% and included predicted-damaging non-synonymous variants. Among variants found in public databases, corresponding allele frequencies varied widely, with Botswana having significantly higher allele frequencies among rare (<1%) pathogenic and damaging variants. Batswana clustered with other Southern African populations, but distinctly from 1000 Genomes African populations, and had limited evidence for admixture with extra-continental ancestries. We also observed a surprising lack of genetic substructure in Botswana, despite multiple tribal ethnicities and language groups, alongside a higher degree of relatedness than purported founder populations from the 1000 Genomes project. Our observations reveal a complex, but distinct, ancestral history and genomic architecture among Batswana and suggest that disease mapping within similar Southern African populations will require a deeper repository of genetic variation and allelic dependencies than presently exists.
According to the established model of murine innate lymphoid cell (ILC) development, helper ILCs develop separately from natural killer (NK) cells. However, it is unclear how helper ILCs and NK cells develop in humans. Here we elucidated key steps of NK cell, ILC2, and ILC3 development within human tonsils using ex vivo molecular and functional profiling and lineage differentiation assays. We demonstrated that while tonsillar NK cells, ILC2s, and ILC3s originated from a common CD34(+)CD117(+) ILC precursor pool, final steps of ILC2 development deviated independently and became mutually exclusive from those of NK cells and ILC3s, whose developmental pathways overlapped. Moreover, we identified a CD34(-)CD117(+) ILC precursor population that expressed CD56 and gave rise to NK cells and ILC3s but not to ILC2s. These data support a model of human ILC development distinct from the mouse, whereby human NK cells and ILC3s share a common developmental pathway separate from ILC2s.
Humans have a rich awareness of locations and situations that directs how we interpret and interact with our surroundings. The principle aim of this paper is to create ‘Information Spaces' where people will use their awareness to search, browse and learn. In the same way that they navigate in a physical environment, they will navigate through knowledge. An information space is a type of design in which representations of information objects are situated in a principled space. In this chapter we present an architecture based on the principles of electrostatistics, which presents a model for design of information spaces. Our model gives an easy conceptual framework to reason about how information can be represented as well as secure ways of extracting and storing information leading to a design which are easily scalable in virtual team environments.
Whole-genome sequencing (WGS) allows for a comprehensive view of the sequence of the human genome. We present and apply integrated methodologic steps for interrogating WGS data to characterize the genetic architecture of 10 heart- and blood-related traits in a sample of 1,860 African Americans. In order to evaluate the contribution of regulatory and non-protein coding regions of the genome, we conducted aggregate tests of rare variation across the entire genomic landscape using a sliding window, complemented by an annotation-based assessment of the genome using predefined regulatory elements and within the first intron of all genes. These tests were performed treating all variants equally as well as with individual variants weighted by a measure of predicted functional consequence. Significant findings were assessed in 1,705 individuals of European ancestry. After these steps, we identified and replicated components of the genomic landscape significantly associated with heart- and blood-related traits. For two traits, lipoprotein(a) levels and neutrophil count, aggregate tests of low-frequency and rare variation were significantly associated across multiple motifs. For a third trait, cardiac troponin T, investigation of regulatory domains identified a locus on chromosome 9. These practical approaches for WGS analysis led to the identification of informative genomic regions and also showed that defined non-coding regions, such as first introns of genes and regulatory domains, are associated with important risk factor phenotypes. This study illustrates the tractable nature of WGS data and outlines an approach for characterizing the genetic architecture of complex traits.
The cost of Whole Genome Sequencing (WGS) has decreased tremendously in recent years due to advances in next-generation sequencing technologies. Nevertheless, the cost of carrying out large-scale cohort studies using WGS is still daunting. Past simulation studies with coverage at ~2x have shown promise for using low coverage WGS in studies focused on variant discovery, association study replications, and population genomics characterization. However, the performance of low coverage WGS in populations with a complex history and no reference panel remains to be determined.
Hardy Weinberg Equilibrium (HWE) test is widely used as a quality control measure to detect sequencing artifacts like mismapping, allelic dropout and biases. However, in the high throughput sequencing era, where the sample size is beyond a thousand scale, the utility of HWE test in reducing the false positive rate remains unclear. In this paper, we demonstrate that HWE test has limited power in identifying sequencing artifacts when the variant allele frequency is lower than 1% in a variant call set produced from more than five thousand whole genome sequenced samples from two homogeneous populations. We develop a novel strategy of implementing HWE filtering in which we incorporate site frequency spectrum information and determine the p-value cutoff which optimizes the tradeoff between sensitivity and specificity. The novel strategy is shown to outperform the exact test of HWE with an empirical constant p-value cutoff regardless of the sequencing sample size. We also present best practice recommendations for identifying possible sources of false positives from large sequencing datasets based on an analysis of intrinsic biases in the variant calling process. Our novel strategy of determining the HWE test p-value cutoff and applying the test to the common variants provides a practical approach for the variant level quality controls in the upcoming sequencing projects with tens to hundreds of thousand of samples.
Abstract The genomic and clinical information used to develop and implement therapeutic approaches for acute myelogenous leukemia (AML) originated primarily from adult patients and has been generalized to patients with pediatric AML. However, age-specific molecular alterations are becoming more evident and may signify the need to age-stratify treatment regimens. The NCI/COG TARGET-AML initiative used whole exome capture sequencing (WXS) to interrogate the genomic landscape of matched trios representing specimens collected upon diagnosis, remission, and relapse from 20 cases of de novo childhood AML. One hundred forty-five somatic variants at diagnosis (median 6 mutations/patient) and 149 variants at relapse (median 6.5 mutations) were identified and verified by orthogonal methodologies. Recurrent somatic variants [in (greater than or equal to) 2 patients] were identified for 10 genes (FLT3, NRAS, PTPN11, WT1, TET2, DHX15, DHX30, KIT, ETV6, KRAS), with variable persistence at relapse. The variant allele fraction (VAF), used to measure the prevalence of somatic mutations, varied widely at diagnosis. Mutations that persisted from diagnosis to relapse had a significantly higher diagnostic VAF compared with those that resolved at relapse (median VAF 0.43 vs. 0.24, P < 0.001). Further analysis revealed that 90% of the diagnostic variants with VAF >0.4 persisted to relapse compared with 28% with VAF <0.2 (P < 0.001). This study demonstrates significant variability in the mutational profile and clonal evolution of pediatric AML from diagnosis to relapse. Furthermore, mutations with high VAF at diagnosis, representing variants shared across a leukemic clonal structure, may constrain the genomic landscape at relapse and help to define key pathways for therapeutic targeting. Cancer Res; 76(8); 2197–205. ©2016 AACR.
Background: The decreasing costs of sequencing are driving the need for cost effective and real time variant calling of whole genome sequencing data. The scale of these projects are far beyond the capacity of typical computing resources available with most research labs. Other infrastructures like the cloud AWS environment and supercomputers also have limitations due to which large scale joint variant calling becomes infeasible, and infrastructure specific variant calling strategies either fail to scale up to large datasets or abandon joint calling strategies.Results: We present a high throughput framework including multiple variant callers for single nucleotide variant (SNV) calling, which leverages hybrid computing infrastructure consisting of cloud AWS, supercomputers and local high performance computing infrastructures. We present a novel binning approach for large scale joint variant calling and imputation which can scale up to over 10,000 samples while producing SNV callsets with high sensitivity and specificity. As a proof of principle, we present results of analysis on Cohorts for Heart And Aging Research in Genomic Epidemiology (CHARGE) WGS freeze 3 dataset in which joint calling, imputation and phasing of over 5300 whole genome samples was produced in under 6 weeks using four state-of-the-art callers. The callers used were SNPTools, GATK-HaplotypeCaller, GATK-UnifiedGenotyper and GotCloud. We used Amazon AWS, a 4000-core in-house cluster at Baylor College of Medicine, IBM power PC Blue BioU at Rice and Rhea at Oak Ridge National Laboratory (ORNL) for the computation. AWS was used for joint calling of 180 TB of BAM files, and ORNL and Rice supercomputers were used for the imputation and phasing step. All other steps were carried out on the local compute cluster. The entire operation used 5.2 million core hours and only transferred a total of 6 TB of data across the platforms.Conclusions: Even with increasing sizes of whole genome datasets, ensemble joint calling of SNVs for low coverage data can be accomplished in a scalable, cost effective and fast manner by using heterogeneous computing platforms without compromising on the quality of variants.
Background Detection of tandem duplication within coding exons, referred to as internal tandem duplication (ITD), remains challenging due to inefficiencies in alignment of ITD-containing reads to the reference genome. There is a critical need to develop efficient methods to recover these important mutational events. Results In this paper we introduce ITD Assembler, a novel approach that rapidly evaluates all unmapped and partially mapped reads from whole exome NGS data using a De Bruijn graphs approach to select reads that harbor cycles of appropriate length, followed by assembly using overlap-layout-consensus. We tested ITD Assembler on The Cancer Genome Atlas AML dataset as a truth set. ITD Assembler identified the highest percentage of reported FLT3-ITDs when compared to other ITD detection algorithms, and discovered additional ITDs in FLT3 , KIT , CEBPA, WT1 and other genes. Evidence of polymorphic ITDs in 54 genes were also found. Novel ITDs were validated by analyzing the corresponding RNA sequencing data. Conclusions ITD Assembler is a very sensitive tool which can detect partial, large and complex tandem duplications. This study highlights the need to more effectively look for ITD’s in other cancers and Mendelian diseases.
Attacks on the Internet are characterized by several alarming trends: (1) increases in frequency; (2) increases in speed; and (3) increases in severity. Modern computer worms simply propagate too quickly for human detection. Since attacks are now occurring at a speed which prevents direct human intervention, there is a need to develop automated defenses. Since the financial, social and political stakes are so high, we need defenses which are provably good against a worst case attacks and are not too costly to deploy. In this dissertation we present two approaches to tackle these problems.For the first part of the dissertation we consider a game between an alert and a worm over a large network. We show, for this game, that it is possible to design an algorithm for the alerts that can prevent any worm from infecting more than a vanishingly small fraction of the nodes with high probability. Critical to our result is designing a communication network for spreading the alerts that has high expansion. The expansion of the network is related to the gap between the 1st and 2 nd eigenvalues of the adjacency matrix. Intuitively high expansion ensures redundant connectivity. We also present results simulating our algorithm on networks of size up to 225.In the second part of this dissertation we consider the virus inoculation game which models the selfish behavior of the nodes involved. We present a technique for this game which makes it possible to achieve the "windfall of malice" even without the actual presence of malicious players. We also show the limitations of this technique for congestion games that are known to have a windfall of malice.
Consider the following game between a worm and an alert over a network of n nodes. Initially, no nodes are infected or alerted and each node in the network is a special detector node independently with small but constant probability, γ. The game starts with a single node becoming infected. In every round thereafter, every infected node sends out β worms to other nodes in the population for some constant β; in addition, every alerted node sends out α alerts for some constant α. Nodes in the network change state according to the following three rules: 1) If a worm is received by a node that is not a detector and is not alerted, that node becomes infected; 2) If a worm is received by a node that is a detector, that node becomes alerted; 3) If an alert is received by a node that is not infected, that node becomes alerted. We allow an infected node to send worm messages to any other node in the network, but, in contrast, allow the alerts to only be sent over a special precomputed overlay network where every node has O(logn) degree. We assume that the infected nodes collaborate with each other, and know everything except which nodes are detectors, and the alerted nodes’ random coin flips. We show, for this game, that it is possible to design an algorithm that can prevent any worm from infecting more than a vanishingly small fraction of the nodes in logarithmic time. In particular, we describe an algorithm and a network that ensures with high probability that atmost o(n) nodes can be infected in O(logn) time steps by any worm for α a fixed constant depending only on β and γ. In addition, our algorithm ensures that the number of nodes that may receive a spuriously generated “false alert” is polylogarithmic in n. We complement our theoretical analysis with simulations on networks of size up to 2.
We consider a problem at the intersection of distributed computing and game theory, namely: Is it possible to achieve the “windfall of malice” even without the actual presence of malicious players? Our answer to this question is “Yes and No”. Our positive result is that for the virus inoculation game, it is possible to achieve the windfall of malice by use of a mediator. Our negative result is that for symmetric congestion games that are known to have a windfall of malice, it is not possible to design a mediator that achieves this windfall. In proving these two results, we develop novel techniques for mediator design that we believe will be helpful for creating non-trivial mediators to improve social welfare in a large class of games.