
Introduction Genetic factors underlying complex diseases are difficult to identify: many polymorphisms may contribute, each having a small effect and low penetrance. These factors may be identified by association studies of large populations, an alternative to family-based linkage studies. Allele frequency measurements of pooled DNA selected from population-level DNA repositories can reduce the costs of these studies. We provide guidance for selecting unrelated individuals for pooling and for comparing the power of studies based on pooled measurements to the power of individual genotyping, particularly for studies using single-nucleotide polymorphism (SNP) markers. Materials and methods We used exact numerical calculations to set pooling criteria that maximized the power to detect association as a function of marker frequency, inheritance mode, and additive variance. Analytical approximations are also provided. Results and discussion Power estimates are provided for two pooled DNA designs: the classification of individuals as affected or unaffected, analogous to a case-control design, and the optimized selection of individuals with extreme phenotypic values. Optimized selection is approximately fourfold more efficient than affected/unaffected classification. The optimal design for most markers is to pool the top and bottom 27% of individuals. Neglecting experimental measurement error, this design requires a population only 1.24-fold larger than that required for individual genotyping. When measurement error is included, the pooled DNA association test serves better as a pre-screen to identify candidate markers which then proceed to individual genotyping. This strategy can still provide a 100-fold savings over individual genotyping.
Introduction Monodelphis ( Monodelphis domestica ) shows 20-fold differences, that may be under genetic control, in individual responses to a high-fat challenge diet. Two partially inbred lines of animals were derived by selectively breeding among high responders and low responders and we analysed data from these pedigreed animals and their F 1 progeny. A blood sample was taken from each animal while consuming the basal diet (basal sample). Each animal was then fed a high-fat, high-cholesterol challenge diet for 8 weeks prior to collection of a second blood sample (challenge sample). Lipoprotein measurements included cholesterol concentrations of high-density lipoproteins (HDL-C) and low-density lipoproteins (non-HDL-C) and particle size phenotypes for both. Quantitative genetic analyses indicated strong heritabilities (range, 0.382–0.827) for each of the eight traits. We also tested for single genes with large effects on each trait (major genes). Segregation analyses provided evidence of major genes for three traits: basal non-HDL-C, challenge non-HDL-C, and challenge HDL-C; no major genes were detected for the lipoprotein size traits. Tests for pleiotropy, using bivariate one-locus segregation analyses, showed that the major locus for challenge HDL-C had no effect on the basal HDL-C and that the major locus for challenge non-HDL-C had no effect on basal non-HDL-C. However, the major gene for basal non-HDL-C did significantly influence challenge non-HDL-C. We have found evidence for at least three genes influencing lipoprotein phenotypes under two dietary regimes. Identification of these genes may provide valuable insights into lipoprotein metabolism in other species, including man.
In the HERS trial, hormone therapy did not reduce the risk of coronary events. In post hoc analyses, treatment was associated with early harm and late benefit. According to one hypothesis, a risk factor may well distinguish a susceptible subgroup with early events associated with hormone therapy from a nonsusceptible subgroup who benefit from hormone therapy. In simulation studies, it appeared that only a susceptibility factor with a low prevalence (3–5%) and a high risk ratio (13–25-fold) can produce the pattern of risks seen in HERS. The number of candidate factors is likely to be small.
Studies in unrelated individuals have shown an association between fibrinogen gene polymorphisms and plasma fibrinogen levels, which is itself a risk factor for coronary heart disease. Family-based studies have complementary strengths in the investigation of such hypothesized genetic associations. Genotypes at the β-fibrinogen -455G/A promoter polymorphism, and a neighbouring highly polymorphic microsatellite in the α-fibrinogen gene were examined for evidence of genetic linkage and association with plasma fibrinogen in 568 members of 97 Caucasian families. The heritability of plasma fibrinogen was 0.22 ± 0.08, P = 0.0007. There was no significant evidence of genetic linkage between the fibrinogen locus and plasma fibrinogen with either marker. In contrast, there was evidence for association of plasma fibrinogen with genotype at the β-455G/A polymorphism (χ12 = 11.12; P = 0.0009). Tests examining allelic transmission from heterozygous parents confirmed this association (Monk’s test, T = 2.17; P = 0.03). Genotype at the β-455G/A polymorphism accounted for 2% of the observed variation in fibrinogen. This is equivalent to about 10% of the heritable component, suggesting the presence of other quantitative trait loci (QTL) in unlinked genes. Confirmation of the association of plasma fibrinogen with genotype at the β-455G/A polymorphism in families indicates that the association is due to the physical proximity of this marker to a QTL, although the effect of this QTL was too small to be detected by linkage in this study. These findings are of potential importance for the design of genetic studies of multifactorial quantitative traits.
The ATP-binding cassette (ABC) gene superfamily encodes a series of transporter proteins that move a wide variety of substances across extra- and intracellular membranes. Forty-eight known human ABC genes can be divided into seven phylogenetically distinct subfamilies. The ABCA gene subfamily is found exclusively in multicellular eukaryotes. We report here on a unique tandem array of five ABCA genes on chromosome 17q24 defining a phylogenetically distinct group. This is the largest cluster of mammalian ABC genes described to date. They are arranged head-to-tail and have similar intron/exon organization in both mouse and human. Northern analysis reveals a heterogeneous pattern of expression in human tissues, with ABCA5 and ABCA10 expressed in skeletal muscle, ABCA6 in the liver, ABCA9 in the heart, and ABCA8 in ovaries. This suggests that these proteins have distinct functions.
Introduction Genetic studies to identify linkage or association usually assume participants are sampled from a genetically homogeneous population, so that a single set of marker allele frequencies is appropriate for all individuals in the study. We have developed a method to identify individuals who are population outliers, because the marker allele frequency distributions from which their genotypes arise differ from the distributions of the remaining individuals in the study. Using allele frequencies estimated from an independent sample, the genotype log likelihood (GLL) test statistic calculates the likelihood of each individual’s genotypes across all markers. Extreme values of the statistic indicate that the individual arises from a different population. The distribution of the test statistic is derived and its convergence under the central limit theorem discussed. This method was applied to genome search data from rheumatoid arthritis which identified a single population outlier family. We used allele frequencies from different populations to show that 100 markers provides high power to identify outliers across a range of populations. The GLL test statistic can be used as a screening tool to identify outlier families in any genetic study with genotyping at independent markers.
Introduction Cyclin D1, encoded by the CCND1 gene, is a key regulator of the cell cycle at the G1/S phase checkpoint. A common A/G single nucleotide polymorphism (SNP) at nt870 of the CCND1 gene has been associated with outcome in patients with lung tumours and head and neck cancer. The aim of this study was to ascertain the genotype and allele frequency of the CCND1 polymorphism in five distinct ethnic populations. Polymerase chain reaction–restriction fragment length polymorphism (PCR–RFLP) analysis was carried out on genomic DNA from 505 subjects from five distinct ethnic populations (i.e. Caucasian, South-west Asian, Ghanaian, Kenyan and Chinese subjects). Marked differences in genotype were apparent between the ethnic populations, with homozygosity for the G allele ranging from 13.6% in the Chinese subjects to 62.4% amongst Kenyan individuals ( P < 0.001). Whereas the East and West African populations demonstrated almost identical allele frequencies, both populations differed significantly from each of the remaining populations. the allele frequencies for the South-west Asian population fell between that of the Caucasian and Chinese populations but did not differ significantly from either, while the Caucasian and Chinese subjects displayed significant differences in CCND1 alleles ( P = 0.003). These marked variations in SNP frequencies between ethnic groups may have a significant impact on prognosis of cancer in these populations, because the CCND1 genotype appears to influence prognosis in each tumour type examined to date.
Introduction A causative relationship has been reported between fragile site expression and disease for FRA12A , a rare, folate-sensitive fragile site on chromosome 12q13.1. FRA12A expression has been described in a number of patients with mental retardation, sometimes in combination with clinical abnormalities. In correspondence to the molecular mechanism of previously cloned, rare, fragile sites, it may be expected that FRA12A is caused by repeat expansion, affecting the expression of genes in the region. To identify the repeat and the associated gene, this paper reports the precise mapping of FRA12A on chromosome 12q12–13. Methods Fluorescence in situ hybridization (FISH) techniques were used to map YAC and PAC clones in the neighbourhood of FRA12A . PAC DNA pools and PAC filters were screened to find additional PAC clones spanning the candidate region. Markers in the region were obtained via web searches and used to construct both PAC and YAC contigs. Results and Discussion A single YAC clone that overspans the fragile site was identified and a complete YAC and PAC contig for the FRA12A region was constructed. The region contains several candidate genes, including a calcium ion channel ( CACNLB3 ), a GTP-binding factor ( ARF3 ), a gene involved in brain development ( INT1 ) and two other genes involved in developmental processes ( WNT10B and ALR ). The FXR1 gene, a homologue of the FMR1 gene, that is associated with fragile X syndrome and that maps to chromosome 12q12–13, was ruled out as a possible candidate gene for the FRA12A site.
Introduction The 5q-syndrome is a myelodysplastic syndrome with the 5q deletion as the sole karyotypic abnormality. The MEGF1 gene is the human homologue of the Drosophila fat tumour suppressor gene. Results We have mapped this gene to the 3 Mb critical region of the 5q-syndrome within 5q31–32, using gene dosage analysis. Fine physical mapping of the MEGF1 gene within this genomic interval was then performed by screening YAC and BAC contigs spanning the critical region using PCR amplification. The MEGF1 gene maps between the genes for SPARC and Annexin-6 at 5q32, and is flanked by the genetic markers D5S2146 and D5S2077. We have demonstrated the expression of MEGF1 in a range of haematological tissues using RT-PCR analysis. Discussion Genomic localization, expression and predicted function would suggest that the MEGF1 gene represents a candidate gene for the 5q-syndrome.
Haseman and Elston1 proposed a model-free method for testing linkage between a polymorphic marker and a quantitative trait locus from data on a sample of independent sib pairs. In that method the squared sib-pair trait difference is regressed on the estimated proportion of alleles shared by the sibs at a marker locus, a negative regression coefficient suggesting linkage. It is possible to obtain more power by modelling the sib covariance, as in the variance component method of linkage analysis, and yet retain a method that is computationally fast, involving only linear regression. To do this it is only necessary to change the dependent variable from the squared trait difference to the difference between the squared mean-corrected sum and the squared trait difference. The method can accommodate sibships of arbitrary size by using generalized least squares and can be made more powerful by weighting the two components. The method is robust in large samples in the presence of any trait distribution, and, in the case of ascertained samples, the mean can be chosen to maximize power.
The optimal study design and method of analysis for genetic studies of complex traits have received much attention of late. Most previous works on this topic have assumed that investigators will study a single focal disease trait. In this paper, we approach the question of study design from the perspective of the inherently multivariate study which includes a variety of quantitative risk factors as well as one or more common complex disease traits. We conclude that a sample of randomly ascertained extended pedigrees provides both analytical power and flexibility, permitting profitable investigation of numerous traits using both linkage and linkage-disequilibrium based methods.
Twins provide a useful and powerful tool for identifying genes, by acting as ideally matched sib-pairs, but are also uniquely placed to measure the extent of their action, their expression and the nature of their interaction with the environment. Classical twin studies have provided insight into the relative genetic and environmental contribution to characteristics and diseases in human populations. The search for a more detailed understanding of genetic mechanisms through linkage and association has, however, traditionally been regarded as the province of other family designs. The last few years have seen a resurgence of interest in twin research following an increasing awareness that the study of twins can also provide an important contribution to localizing and understanding the function of specific genes.1 In this brief overview, we focus on these newer developments that are currently being applied in the search to uncover the genetic basis of disease.
The classical twin study is the most popular method for assessing the relative contribution of genes and environment to traits in human populations. Critics argue that several assumptions of the twin method are unjustified, and therefore results from twin studies are misleading. Specifically, it has been suggested that twins differ in important aspects from singletons, that monozygotic (MZ) and dizygotic (DZ) twins are not matched in their degree of environmental similarity, and that MZ twins are neither matched genetically nor in their prenatal environments. These criticisms are addressed and it is suggested that they do not provide serious impediments to the validity of the twin study.
We review the issues involved in the fine-scale mapping of disease loci via allelic association. We argue that it is vital to properly account for uncertainty about the genealogy of the disease locus and founding haplotypes at the surrounding markers: we sketch how this might be done in a Bayesian framework using Markov chain Monte Carlo tecniques, and show that this works well for the much analysed data on the location of the Δ508 mutation for cystic fibrosis.
Introduction Connexins (Cx) comprise a family of homologous proteins that are involved in the intercellular exchange of ions and small metabolites between adjacent cells. So far, mutations in seven different connexins have been found in humans, each resulting in a genetic disease. Methods Based on the sequence alignment of known human Cx genes, we developed degenerate PCR primers that we anticipated would amplify members of the Cx gene family, in order to identify novel Cx genes. Results By subcloning and sequencing the PCR products, we identified a previously unidentified connexin gene that we named Cx59 ( GJA11 ). Using FISH we localized the GJA11 gene to chromosome 1p34, and by the analysis of a YAC contig of this region, we mapped this gene within the linkage interval of an Indonesian family with autosomal dominant nonsyndromic hearing loss (ADNSHL). Because mutations in other connexins, namely GJB3 ( Cx31 ) and GJB6 ( Cx30 ), lead to similar forms of ADNSHL, GJA11 ( Cx59 ) was a very strong candidate gene. Therefore, we performed mutation analysis of the coding region of GJA11 in patients of this Indonesian family, but no disease-causing mutation was found.
Osteoporosis fits well into the category of complex disease with multiple genes and environmental factors likely to be involved in its development. The osteoporosis phenotype itself comprises a number of component parts, aspects of which are captured differently by a range of different clinical and laboratory measurements. This review discusses the nature of the genetic factors that underlie this complex phenotype and considers approaches that take into account this complexity in modelling the effects of specific genes.
Introduction Model-free linkage studies are increasingly used to investigate the genetic factors implicated in complex quantitative traits because they do not require any specification of the underlying genetic model. However, the term model-free does not imply that no assumption is introduced by the corresponding statistical methods. In particular, the widely used variance components approaches assume multivariate normality of the phenotypic distribution and it has been shown that violation of this normality hypothesis could lead to large inflation of the type I error. In this paper, we assess the robustness of the recently developed sibship-oriented Maximum-Likelihood-Binomial (MLB) method for genetic model-free linkage analysis in the context of several types of non-normal phenotypic data using a large simulation study. Simulation study Under the hypothesis of no linkage at the marker locus under study, 20 000 replicates of family samples including 100 or 500 independent sib-pairs were simulated considering four different designs that lead to non-normal phenotypic data: (1) presence of a major gene not linked to the studied marker, (2) gene–environment interaction, (3) analysis of a binary phenotype, and (4) extreme sampling. Further, three levels of residual sib–sib correlation were considered. Results and discussion For each simulation design the empirical type I errors were consistent with their asymptotic expectations showing that the MLB approach is insensitive to non-normal phenotypic distribution whatever the mechanism underlying this non-normality. Therefore, the MLB method should be an attractive alternative method for model-free linkage analysis of QTL, especially for investigators who do not want to worry about the validity of asymptotic thresholds when performing their analyses.
Power calculations for linkage analysis are typically conducted on the assumption of a single locus that affects the trait. Here we report a simple procedure for conducting a power analysis for a genome-wide linkage scan of a quantitative trait under the influence of multiple loci. This procedure is designed for sib pair data analysed by the new Haseman–Elston regression method. The results show that samples as large as 10 000 sib pairs will often not allow quantitative trait loci (QTLs) to be clearly identified. Instead, linkage genome scan using sib pairs must be regarded as a blunt screening tool that will help to focus attention to 10%, or more, of the genome.
Regression analysis is a simple, computationally efficient and often robust method for assessment of genotype–phenotype relationships in quantitative traits. It has long been used in studies of familiality and selection, and recently has been further extended for linkage and association analysis, linkage disequilibrium mapping and population substructure assessment. We review and compare some of the most commonly used regression models and highlight some useful properties of one model via analysis of angiotensin-converting enzyme phenotype and marker data.