We performed a two-tiered, whole-genome association study of Parkinson disease (PD). For tier 1, we individually genotyped 198,345 uniformly spaced and informative single-nucleotide polymorphisms (SNPs) in 443 sibling pairs discordant for PD. For tier 2a, we individually genotyped 1,793 PD-associated SNPs (P<.01 in tier 1) and 300 genomic control SNPs in 332 matched case-unrelated control pairs. We identified 11 SNPs that were associated with PD (P<.01) in both tier 1 and tier 2 samples and had the same direction of effect. For these SNPs, we combined data from the case-unaffected sibling pair (tier 1) and case-unrelated control pair (tier 2) samples and employed a liberalization of the sibling transmission/disequilibrium test to calculate odds ratios, 95% confidence intervals, and P values. A SNP within the semaphorin 5A gene (SEMA5A) had the lowest combined P value (P=7.62 x 10(-6)). The protein encoded by this gene plays an important role in neurogenesis and in neuronal apoptosis, which is consistent with existing hypotheses regarding PD pathogenesis. A second SNP tagged the PARK11 late-onset PD susceptibility locus (P=1.70 x 10(-5)). In tier 2b, we also selected for genotyping additional SNPs that were borderline significant (P<.05) in tier 1 but that tested a priori biological and genetic hypotheses regarding susceptibility to PD (n=941 SNPs). In analysis of the combined tier 1 and tier 2b data, the two SNPs with the lowest P values (P=9.07 x 10(-6); P=2.96 x 10(-5)) tagged the PARK10 late-onset PD susceptibility locus. Independent replication across populations will clarify the role of the genomic loci tagged by these SNPs in conferring PD susceptibility.
Allelic variation of gene expression is common in humans, and is of interest because of its potential contribution to variation in heritable traits. To identify human genes with allelic expression differences, we genotype DNA and examine mRNA isolated from the white blood cells of 12 unrelated individuals using oligonucleotide arrays containing 8406 exonic SNPs. Of the exonic SNPs, 1983, located in 1389 genes, are both expressed in the white blood cells and heterozygous in at least one of the 12 individuals, and thus can be examined for differential allelic expression. Of the 1389 genes, 731 (53%) show allele expression differences in at least one individual. To gain insight into the regulatory mechanisms governing allelic expression differences, we analyze a set of 60 genes containing exonic SNPs that are heterozygous in three or more samples, and for which all heterozygotes display differential expression. We find three patterns of allelic expression, suggesting different underlying regulatory mechanisms. Exonic SNPs in three of the 60 genes are monoallelically expressed in the human white blood cells, and when examined in families show expression of only the maternal copy, consistent with regulation by imprinting. Approximately one-third of the genes have the same allele expressed more highly in all heterozygotes, suggesting that their regulation is predominantly influenced by cis-elements in strong linkage disequilibrium with the assayed exonic SNP. The remaining two-thirds of the genes have different alleles expressed more highly in different heterozygotes, suggesting that their expression differences are influenced by factors not in strong linkage disequilibrium with the assayed exonic SNP.
Individual differences in DNA sequence are the genetic basis of human variability. We have characterized whole-genome patterns of common human DNA variation by genotyping 1,586,383 single-nucleotide polymorphisms (SNPs) in 71 Americans of European, African, and Asian ancestry. Our results indicate that these SNPs capture most common genetic variation as a result of linkage disequilibrium, the correlation among common SNP alleles. We observe a strong correlation between extended regions of linkage disequilibrium and functional genomic elements. Our data provide a tool for exploring many questions that remain regarding the causal role of common human DNA variation in complex human traits and for investigating the nature of genetic variation within and between human populations.
Rapid progress in genome research creates a wealth of information on the functional annotation of mammalian genome sequences. However, as we accumulate large amounts of scientific information we are facing problems of how to integrate and relate the data produced by various genomic approaches. Here, we propose the novel concept of an organ atlas where diverse data from expression maps to histological findings to mutant phenotypes can be queried, compared and visualized in the context of a three-dimensional reconstruction of the organ. We will seek proof of concept for the organ atlas by elucidating genetic pathways involved in development and pathophysiology of the kidney. Such a kidney atlas may provide a paradigm for a new systems-biology approach in functional genome research aimed at understanding the genetic bases of organ development, physiology and disease.
High-density SNP screening of panels of inbred mouse strains has been proposed as a method to accelerate the identification of genes associated with complex biomedical phenotypes. To evaluate the potential of these studies, a more detailed understanding of the fine structure of sequence variation across inbred mouse strains is needed. Here, we use high-density oligonucleotide arrays to discover an extremely dense set of SNPs in 13 classical and two wild-derived inbred strains in five genomic intervals totaling 4.6 Mb of DNA sequence, and then analyze the segmental haplotype structure defined by these high-density SNPs. This analysis reveals segments ranging from 12 to 608 kb in length within which the inbred strains have a simple and distinct phylogenetic relationship with typically two or three clades accounting for the 13 classical strains examined. The phylogenetic relationships among strains change abruptly and unpredictably from segment to segment, and are distinct in each of the five genomic regions examined. The data suggest that at least 12 strains would need to be resequenced for exhaustive SNP discovery in every region of the mouse genome, that approximately 97% of the variation among inbred strains is ancestral (between clades) and approximately 3% private (within clades), and provides critical insights into the proposed use of panels of inbred strains to identify genes underlying quantitative trait loci.
Cross-species DNA sequence comparison is a fundamental method for identifying biologically important elements, because functional sequences are evolutionarily conserved, wheres nonfunctional sequences drift. A recent genome-wide comparison of human and mouse DNA discovered over 200,000 conserved noncoding sequences with unknown function. Multispecies DNA comparison has been proposed as a method to prioritize these conserved noncoding sequences for functional analysis based on the hypothesis that elements present in many species are more likely to be functional than elements present in limited numbers of species. Here, we perform a comparative analysis of the single-minded 2 ( SIM2 ) gene interval on human chromosome 21 with horse, cow, pig, dog, cat, and mouse DNA. We classify conserved sequences based on the number of mammals in which they are present, and experimentally test sequences in each class for function. As hypothesized, conserved sequences present in many mammals are frequently functional. Additionally, we demonstrate that sequences conserved in a limited number of mammals are also frequently functional. Examination of genomic deletions in chimpanzee and rhesus macaque DNA showed that several putatively functional conserved noncoding human sequences were absent in these primates. These findings suggest that functional conserved noncoding human sequences can be missing in other mammals, even closely related primate species.
Association studies in populations that are genetically heterogeneous can yield large numbers of spurious associations if population subgroups are unequally represented among cases and controls. This problem is particularly acute for studies involving pooled genotyping of very large numbers of single-nucleotide-polymorphism (SNP) markers, because most methods for analysis of association in structured populations require individual genotyping data. In this study, we present several strategies for matching case and control pools to have similar genetic compositions, based on ancestry information inferred from genotype data for approximately 300 SNPs tiled on an oligonucleotide-based genotyping array. We also discuss methods for measuring the impact of population stratification on an association study. Results for an admixed population and a phenotype strongly confounded with ancestry show that these simple matching strategies can effectively mitigate the impact of population stratification.
Comparative DNA sequence studies between humans and nonhuman primates will be important for understanding the genetic basis of the phenotypic differences between these species. Here we compare approximately 27 Mb of human chromosome 21 with chimpanzee DNA sequences identifying 57 genomic rearrangements (deletions and insertions ranging in size from 0.2 to 8.0 kb) between the two species. These rearrangements are distributed along the entire length of chromosome 21, with approximately 35% found in genomic intervals encoding genes (genic intervals), and have occurred in the genomes of both humans and chimpanzees. Comparison of approximately 9 Mb of human chromosome 21 with orangutan, rhesus macaque, and woolly monkey DNA sequences identified a combined total of 114 genomic rearrangements between humans and nonhuman primates. Analysis of these rearrangements revealed that they are randomly distributed with respect to genic and nongenic intervals and identified one deletion that has likely resulted in the inactivation of a gene (beta1,3-galactosyltransferase) in the woolly monkey. Our data show that genomic rearrangements have occurred frequently during primate genome evolution and significantly contribute to the DNA differences between these species. These DNA rearrangements are commonly found in genic intervals, and thus provide natural starting points for focused investigations of qualitative and quantitative gene expression differences between humans and other primates.