When the genetic test for the Huntington’s disease (HD) HTT expansion first became available almost 30 years ago, only 1% of patients tested negative. Since then, the test has become more accessible and the HD phenotype has expanded. More patients are being tested overall, and more negative tests are being received. These patients are deemed “HD phenocopy syndromes” (HDPC). In this study we established a current estimate for the prevalence of these patients. We also surveyed HD clinician experts on what would make them consider an HD test and compared both HD and HDPC patients to these expectations to decide whether they could be distinguished clinically; this proved impossible even when comparing symptom patterns. We re-analysed existing gene panel data for likely and potentially deleterious variants. Furthermore, we determined principles to prioritise patients for whole-genome sequencing (WGS). It was used to probe a 50 patient strong subcohort of HD phenocopy syndromes for known causes of HD-like and other neurodegenerative disease, identifying one ATXN1 expansion using ExpansionHunter ® . This was a small genetic substudy and therefore unsurprisingly no other known deleterious variants could be identified as in these cryptic understudied syndromes. Novel variants in known genes and variants in genes not yet linked to neurodegeneration may play an outsized role.
The allele frequency spectrum of polymorphisms in DNA sequences can be used to test for signatures of natural selection that depart from the expected frequency spectrum under the neutral theory. We observed a significant ( P = 0.001) correlation between the Tajima's D test statistic in full resequencing data and Tajima's D in a dense, genome-wide data set of genotyped polymorphisms for a set of 179 genes. Based on this, we used a sliding window analysis of Tajima's D across the human genome to identify regions putatively subject to strong, recent, selective sweeps. This survey identified seven Contiguous Regions of Tajima's D Reduction (CRTRs) in an African-descent population (AD), 23 in a European-descent population (ED), and 29 in a Chinese-descent population (XD). Only four CRTRs overlapped between populations: three between ED and XD and one between AD and ED. Full resequencing of eight genes within six CRTRs demonstrated frequency spectra inconsistent with neutral expectations for at least one gene within each CRTR. Identification of the functional polymorphism (and/or haplotype) responsible for the selective sweeps within each CRTR may provide interesting insights into the strongest selective pressures experienced by the human genome over recent evolutionary history.
Identifying regions of the human genome that have been targets of natural selection will provide important insights into human evolutionary history and may facilitate the identification of complex disease genes. Although the signature that natural selection imparts on DNA sequence variation is difficult to disentangle from the effects of neutral processes such as population demographic history, selective and demographic forces can be distinguished by analyzing multiple loci dispersed throughout the genome. We studied the molecular evolution of 132 genes by comprehensively resequencing them in 24 African-Americans and 23 European-Americans. We developed a rigorous computational approach for taking into account multiple hypothesis tests and demographic history and found that while many apparent selective events can instead be explained by demography, there is also strong evidence for positive or balancing selection at eight genes in the European-American population, but none in the African-American population. Our results suggest that the migration of modern humans out of Africa into new environments was accompanied by genetic adaptations to emergent selective forces. In addition, a region containing four contiguous genes on Chromosome 7 showed striking evidence of a recent selective sweep in European-Americans. More generally, our results have important implications for mapping genes underlying complex human diseases.
Recent studies have suggested that a significant fraction of the human genome is contained in blocks of strong linkage disequilibrium, ranging from ~5 to >100 kb in length, and that within these blocks a few common haplotypes may account for >90% of the observed haplotypes. Furthermore, previous studies have suggested that common haplotypes in candidate genes are generally shared across populations and represent the majority of chromosomes in each population. The conclusions drawn from these preliminary studies, however, are based on an incomplete knowledge of the variation in the regions examined. To bridge this gap in knowledge, we have completely resequenced 100 candidate genes in a population of African descent and one of European descent. Although these genes have been well studied because of their medical importance, we demonstrate that a large amount of sequence variation has not yet been described. We also report that the average number of inferred haplotypes per gene, when complete data is used, is higher than in previous reports and that the number and proportion of all haplotypes represented by common haplotypes per gene is variable. Furthermore, we demonstrate that haplotypes shared between the two populations constitute only a fraction of the total number of haplotypes observed and that these shared haplotypes represent fewer of the African-descent chromosomes than was expected from previous studies. Finally, we show that restricting variation discovery to coding regions does not adequately describe all common haplotypes or the true haplotype block structure observed when all common variation is used to infer haplotypes. These data, derived from complete knowledge of genetic variation in these genes, suggest that the haplotype architecture of candidate genes across the human genome is more complex than previously suggested, with important implications for candidate gene and genomewide association studies.
The 156 breeds of registered dogs in the United States offer a unique opportunity to map genes important in disease susceptibility, morphology, and behavior. Linkage disequilibrium (LD) is of current interest for its application in whole genome association mapping, since the extent of LD determines the feasibility of such studies. We have measured LD at five genomic intervals, each 5 Mb in length and composed of five clusters of sequence variants spaced 800 kb-1.6 Mb apart. These intervals are located on canine chromosomes 1, 2, 3, 34, and 37, and none is under obvious selective pressure. Approximately 20 unrelated dogs were assayed from each of five breeds: Akita, Bernese Mountain Dog, Golden Retriever, Labrador Retriever, and Pekingese. At each genomic interval, SNPs and indels were discovered and typed by resequencing. Strikingly, LD in canines is much more extensive than in humans: D' falls to 0.5 at 400-700 kb in Golden Retriever and Labrador Retriever, 2.4 Mb in Akita, and 3-3.2 Mb in Bernese Mountain Dog and Pekingese. LD in dog breeds is up to 100x more extensive than in humans, suggesting that a correspondingly smaller number of markers will be required for association mapping studies in dogs compared to humans. We also report low haplotype diversity within regions of high LD, with 80% of chromosomes in a breed carrying two to four haplotypes, as well as a high degree of haplotype sharing among breeds.
Identifying common sequence variations known as single nucleotide polymorphisms (SNPs) in human populations is one of the current objectives of the human genome project. Nearly 3 million SNPs have been identified. Analysis of the relative allele frequency of these markers in human populations and the genetic associations between these markers, known as linkage disequilibrium, is now underway to generate a high-density genetic map. Because of the central role T cells play in immune reactivity, the T-cell receptor (TCR) loci have long been considered important candidates for common disease susceptibility within the immune system (e.g., asthma, atopy and autoimmunity). Over the past two decades, hundreds of SNPs in the TCR loci have been identified. Most studies have focused on defining SNPs in the variable gene segments which are involved in antigenic recognition. On average, the coding sequence of each TCR variable gene segment contains two SNPs, with many more found in the 5', 3' and intronic sequences of these segments. Therefore, a potentially large repertoire of functional variants exists in these loci. Association between SNPs (linkage disequilibrium) extends approximately 30 kb in the TCR loci, although a few larger regions of disequilibrium have been identified. Therefore, the SNPs found in one variable gene segment may or may not be associated with SNPs in other surrounding variable gene segments. This suggests that meaningful association studies in the TCR loci will require the analysis and typing of large marker sets to fully evaluate the role of TCR loci in common disease susceptibility in human populations.
Understanding the pattern of linkage disequilibrium (LD) in the human genome is important both for successful implementation of disease-gene mapping approaches and for inferences about human demographic histories. Previous studies have examined LD between loci within single genes or confined genomic regions, which may not be representative of the genome; between loci separated by large distances, where little LD is seen; or in population groups that differ from one study to the next. We measured LD in a large set of locus pairs distributed throughout the genome, with loci within each pair separated by short distances (average 124 bp). Given current models of the history of the human population, nearly all pairs of loci at such short distances would be expected to show complete LD as a consequence of lack of recombination in the short interval. Contrary to this expectation, a significant fraction of pairs showed incomplete LD. A standard model of recombination applied to these data leads to an estimate of effective human population size of 110,000. This estimate is an order of magnitude higher than most estimates based on nucleotide diversity. The most likely explanation of this discrepancy is that gene conversion increases the apparent rate of recombination between nearby loci.
The T-cell receptor (TCR) plays a central role in the immune system, and > 90% of human T cells present a receptor that consists of the alpha TCR subunit (TCRA) and the beta subunit (TCRB). Here we report an analysis of 63 variable genes (BV), spanning 553 kb of TCRB that yielded 279 single-nucleotide polymorphisms (SNPs). Samples were drawn from 10 individuals and represent four populations-African American, Chinese, Mexican, and Northern European. We found nine variants that produce nonfunctional BV segments, removing those genes from the TCRB genomic repertoire. There was significant heterogeneity among population samples in SNP frequency (including the BV-inactivating sites), indicating the need for multiple-population samples for adequate variant discovery. In addition, we observed considerable linkage disequilibrium (LD) (r(2) > 0.1) over distances of approximately 30 kb in TCRB, and, in general, the distribution of r(2) as a function of physical distance was in close agreement with neutral coalescent simulations. LD in TCRB showed considerable spatial variation across the locus, being concentrated in "blocks" of LD; however, coalescent simulations of the locus illustrated that the heterogeneity of LD we observed in TCRB did not differ markedly from that expected from neutral processes. Finally, examination of the extended genotypes for each subject demonstrated homozygous stretches of >100 kb in the locus of several individuals. These results provide the basis for optimization of locuswide SNP typing in TCRB for studies of genotype-phenotype association.
Strategies for the discovery of single-nucleotide polymorphisms (SNPs) can be characterized by the number of individuals in the discovery sample, and by the minimal required number of observations of each allele. We examine the effect of different strategies on two key properties of the resulting SNP collection: (1) the probability that a SNP with a given population allele frequency is detected; and (2) the allele-frequency distribution of the discovered SNPs. We show that strategies that accept all polymorphic sites lead to collections with a high fraction of SNPs with rare minor alleles, particularly in expanded populations. Such SNPs have a low probability of replication in a second sample. We discuss how to tailor a discovery strategy to the desired properties of a SNP collection.