Abstract Background The association between genetic variants on the X chromosome to risk of COPD has not been fully explored. We hypothesize that the X chromosome harbors variants important in determining risk of COPD related phenotypes and may drive sex differences in COPD manifestations. Methods Using X chromosome data from three COPD-enriched cohorts of adult smokers, we performed X chromosome specific quality control, imputation, and testing for association with COPD case–control status, lung function, and quantitative emphysema. Analyses were performed among all subjects, then stratified by sex, and subsequently combined in meta-analyses. Results Among 10,193 subjects of non-Hispanic white or European ancestry, a variant near TMSB4X, rs5979771, reached genome-wide significance for association with lung function measured by FEV1/FVC ( $$\beta$$ β 0.020, SE 0.004, p 4.97 × 10–08), with suggestive evidence of association with FEV1 ( $$\beta$$ β 0.092, SE 0.018, p 3.40 × 10–07). Sex-stratified analyses revealed X chromosome variants that were differentially trending in one sex, with significantly different effect sizes or directions. Conclusions This investigation identified loci influencing lung function, COPD, and emphysema in a comprehensive genetic association meta-analysis of X chromosome genetic markers from multiple COPD-related datasets. Sex differences play an important role in the pathobiology of complex lung disease, including X chromosome variants that demonstrate differential effects by sex and variants that may be relevant through escape from X chromosome inactivation. Comprehensive interrogation of the X chromosome to better understand genetic control of COPD and lung function is important to further understanding of disease pathology. Trial registration Genetic Epidemiology of COPD Study (COPDGene) is registered at ClinicalTrials.gov, NCT00608764 (Active since January 28, 2008). Evaluation of COPD Longitudinally to Identify Predictive Surrogate Endpoints Study (ECLIPSE), GlaxoSmithKline study code SCO104960, is registered at ClinicalTrials.gov, NCT00292552 (Active since February 16, 2006). Genetics of COPD in Norway Study (GenKOLS) holds GlaxoSmithKline study code RES11080, Genetics of Chronic Obstructive Lung Disease.
Additional file 4: Table S3. Replication Examination Of Associations From Other Studies In This XWAS Meta-analysis
Additional file 3: Table S2. XWAS Meta-analysis Top ACE2 Xp22.2 Variants
Since the first genome-wide association study (GWAS) identifying variants associated with myocardial infarction was published over 20 years ago, GWASs have emerged as a powerful tool for exploring the genetic basis of complex traits. To date, hundreds of thousands of statistically significant associations have been reported across thousands of human phenotypes. Nevertheless, the design, implementation, and analysis of GWASs remain complex, and the results are easily misinterpreted. Common mistakes include 1) assuming that variants with the strongest statistical associations are causal instead of correlative, 2) believing that associated loci act through nearby genes, and 3) overemphasizing the contribution of individual loci to the total variability of particular traits. Clinical assays have been designed using the results of GWAS that rely on the contribution of such erroneous data interpretations to predict clinical phenotypes, reactions to medications or foods, and/or propensity to develop diseases. The failure to recognize these errors due to fallacies in logical reasoning and statistical inference presents problems for both the scientific community when the wrong targets may be prioritized in future research studies, as well as for communication with the general public when our understanding of the genetic basis of important traits may be misrepresented and overstated. Here, we review statistical data quality, analysis, and meta-analysis, of GWAS results with an emphasis on accurate and reliable interpretation. Placed in the appropriate context, GWASs enable genome-wide discovery of loci associated with diverse traits, but they constitute only a first step towards understanding the biological mechanism(s) underlying the observed associations. Scientific elucidation of these biological mechanisms must be required to establish causality with biochemical and pathophysiological explanations for any putative statistical correlations.
Longitudinal studies play a prominent role in research on growth, change and/or decline in individuals, and in characterising the environmental and social factors which influence change. The essential feature of a longitudinal study is taking repeated measures of an outcome on the same set of individuals at multiple timepoints, thereby allowing investigators to characterise within subject changes during the measurement period. This paper provides an overview of how the basic design features and analysis of longitudinal studies are related to other study designs, including longitudinal clinical trials as well as repeated measures studies. I summarise the use of the linear mixed model as described in Laird and Ware for the analysis of a broad class of designs and present some applications in health and medicine.
Alzheimer’s disease (AD) is a genetically complex disease for which nearly 40 loci have now been identified via genome-wide association studies (GWAS). We attempted to identify groups of rare variants (alternate allele frequency <0.01) associated with AD in a region-based, whole-genome sequencing (WGS) association study (rvGWAS) of two independent AD family datasets (NIMH/NIA; 2247 individuals; 605 families). Employing a sliding window approach across the genome, we identified several regions that achieved association p values <10−6, using the burden test or the SKAT statistic. The genomic region around the dystobrevin beta (DTNB) gene was identified with the burden and SKAT test and replicated in case/control samples from the ADSP study reaching genome-wide significance after meta-analysis (pmeta = 4.74 × 10−8). SKAT analysis also revealed region-based association around the Discs large homolog 2 (DLG2) gene and replicated in case/control samples from the ADSP study (pmeta = 1 × 10−6). In conclusion, in a region-based rvGWAS of AD we identified two novel AD genes, DLG2 and DTNB, based on association with rare variants.
Nan Laird is the Harvey V. Fineberg professor of biostatistics (emerita) at the Harvard T. H. Chan School of Public Health, and the winner of the 2021 International Prize in Statistics. In this interview with Ed Hirschland, she discusses her career and contributions
SARS-CoV-2 mortality has been extensively studied in relation to host susceptibility. How sequence variations in the SARS-CoV-2 genome affect pathogenicity is poorly understood. Starting in October 2020, using the methodology of genome-wide association studies (GWAS), we looked at the association between whole-genome sequencing (WGS) data of the virus and COVID-19 mortality as a potential method of early identification of highly pathogenic strains to target for containment. Although continuously updating our analysis, in December 2020, we analyzed 7548 single-stranded SARS-CoV-2 genomes of COVID-19 patients in the GISAID database and associated variants with mortality using a logistic regression. In total, evaluating 29,891 sequenced loci of the viral genome for association with patient/host mortality, two loci, at 12,053 and 25,088 bp, achieved genome-wide significance (p values of 4.09e-09 and 4.41e-23, respectively), though only 25,088 bp remained significant in follow-up analyses. Our association findings were exclusively driven by the samples that were submitted from Brazil (p value of 4.90e-13 for 25,088 bp). The mutation frequency of 25,088 bp in the Brazilian samples on GISAID has rapidly increased from about 0.4 in October/December 2020 to 0.77 in March 2021. Although GWAS methodology is suitable for samples in which mutation frequencies varies between geographical regions, it cannot account for mutation frequencies that change rapidly overtime, rendering a GWAS follow-up analysis of the GISAID samples that have been submitted after December 2020 as invalid. The locus at 25,088 bp is located in the P.1 strain, which later (April 2021) became one of the distinguishing loci (precisely, substitution V1176F) of the Brazilian strain as defined by the Centers for Disease Control. Specifically, the mutations at 25,088 bp occur in the S2 subunit of the SARS-CoV-2 spike protein, which plays a key role in viral entry of target host cells. Since the mutations alter amino acid coding sequences, they potentially imposing structural changes that could enhance viral infectivity and symptom severity. Our analysis suggests that GWAS methodology can provide suitable analysis tools for the real-time detection of new more transmissible and pathogenic viral strains in databases such as GISAID, though new approaches are needed to accommodate rapidly changing mutation frequencies over time, in the presence of simultaneously changing case/control ratios. Improvements of the associated metadata/patient information in terms of quality and availability will also be important to fully utilize the potential of GWAS methodology in this field.
Noncoding DNA contains gene regulatory elements that alter gene expression, and the function of these elements can be modified by genetic variation. Massively parallel reporter assays (MPRA) enable high-throughput identification and characterization of functional genetic variants, but the statistical methods to identify allelic effects in MPRA data have not been fully developed. In this study, we demonstrate how the baseline allelic imbalance in MPRA libraries can produce biased results, and we propose a novel, nonparametric, adaptive testing method that is robust to this bias. We compare the performance of this method with other commonly used methods, and we demonstrate that our novel adaptive method controls Type I error in a wide range of scenarios while maintaining excellent power. We have implemented these tests along with routines for simulating MPRA data in the Analysis Toolset for MPRA (@MPRA), an R package for the design and analyses of MPRA experiments. It is publicly available at http://github.com/redaq/atMPRA.
With the advent of whole genome-sequencing (WGS) studies, family-based designs enable sex-specific analysis approaches that can be applied to only affected individuals; tests using family-based designs are attractive because they are completely robust against the effects of population substructure. These advantages make family-based association tests (FBATs) that use siblings as well as parents especially suited for the analysis of late-onset diseases such as Alzheimer’s Disease (AD). However, the application of FBATs to assess sex-specific effects can require additional filtering steps, as sensitivity to sequencing errors is amplified in this type of analysis. Here, we illustrate the implementation of robust analysis approaches and additional filtering steps that can minimize the chances of false positive-findings due to sex-specific sequencing errors. We apply this approach to two family-based AD datasets and identify four novel loci ( GRID1 , RIOK3 , MCPH1 , ZBTB7C ) showing sex-specific association with AD risk. Following stringent quality control filtering, the strongest candidate is ZBTB7C (P inter = 1.83 × 10 −7 ), in which the minor allele of rs1944572 confers increased risk for AD in females and protection in males. ZBTB7C encodes the Zinc Finger and BTB Domain Containing 7C, a transcriptional repressor of membrane metalloproteases (MMP). Members of this MMP family were implicated in AD neuropathology.
MOTIVATION:Analysis of rare variants in family-based studies remains a challenge. Transmission-based approaches provide robustness against population stratification, but the evaluation of the significance of test statistics based on asymptotic theory can be imprecise. Also, power will depend heavily on the choice of the test statistic and on the underlying genetic architecture of the locus, which will be generally unknown.RESULTS:In our proposed framework, we utilize the FBAT haplotype algorithm to obtain the conditional offspring genotype distribution under the null hypothesis given the sufficient statistic. Based on this conditional offspring genotype distribution, the significance of virtually any association test statistic can be evaluated based on simulations or exact computations, without the need for asymptotic approximations. Besides standard linear burden-type statistics, this enables our approach to also evaluate other test statistics such as variance components statistics, higher criticism approaches, and maximum-single-variant-statistics, where asymptotic theory might be involved or does not provide accurate approximations for rare variant data. Based on these P-values, combined test statistics such as the aggregated Cauchy association test (ACAT) can also be utilized. In simulation studies, we show that our framework outperforms existing approaches for family-based studies in several scenarios. We also applied our methodology to a TOPMed whole-genome sequencing dataset with 897 asthmatic trios from Costa Rica.AVAILABILITY AND IMPLEMENTATION:FBAT software is available at https://sites.google.com/view/fbatwebpage. Simulation code is available at https://github.com/julianhecker/FBAT_rare_variant_test_simulations. Whole-genome sequencing data for 'NHLBI TOPMed: The Genetic Epidemiology of Asthma in Costa Rica' is available at https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000988.v4.p1.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
SARS-CoV-2 mortality has been extensively studied in relationship to a patient's predisposition to the disease. However, how sequence variations in the SARS-CoV-2 genome affect mortality is not understood. To address this issue, we used a whole-genome sequencing (WGS) association study to directly link death of SARS-CoV-2 patients with sequence variation in the viral genome. Specifically, we analyzed 3,626 single stranded RNA-genomes of SARS-CoV-2 patients in the GISAID database (Elbe and Buckland-Merrett, 2017; Shu and McCauley, 2017) with reported patient’s health status from COVID-19, i.e. deceased versus non-deceased. In total, evaluating 28,492 loci of the viral genome for association with patient/host mortality, two loci, 12,053bp and 25,088bp, achieved genome-wide significance (p-values of 1.24e-12, and 1.24e-26, respectively). Mutations at 25,088bp occur in the S2 subunit of the SARS-CoV-2 spike protein, which plays a key role in viral entry of target host cells. Additionally, mutations at 12,053bp are within the ORF1ab gene, in a region encoding for the protein nsp7, which is necessary to form the RNA polymerase complex responsible for viral replication and transcription. Both mutations altered amino acid coding sequences, potentially imposing structural changes that could enhance viral infectivity and symptom severity, and may be important to consider as targets for therapeutic development.
Protein-coding de novo mutations (DNMs) are significant risk factors in many neurodevelopmental disorders, whereas schizophrenia (SCZ) risk associated with DNMs has thus far been shown to be modest. We analyzed DNMs from 1,695 SCZ-affected trios and 1,077 published SCZ-affected trios to better understand the contribution to SCZ risk. Among 2,772 SCZ probands, exome-wide DNM burden remained modest. Gene set analyses revealed that SCZ DNMs were significantly concentrated in genes that were highly expressed in the brain, that were under strong evolutionary constraint and/or overlapped with genes identified in other neurodevelopmental disorders. No single gene surpassed exome-wide significance; however, 16 genes were recurrently hit by protein-truncating DNMs, corresponding to a 3.15-fold higher rate than the mutation model expectation (permuted 95% confidence interval: 1–10 genes; permuted P = 3 × 10−5). Overall, DNMs explain a small fraction of SCZ risk, and larger samples are needed to identify individual risk genes, as coding variation across many genes confers risk for SCZ in the population. In this study of protein-coding de novo mutations in schizophrenia, researchers found only a small contribution toward overall risk, coming predominantly from genes under negative selection and highly expressed in the brain.
The transmission disequilibrium test (TDT) is the gold standard for testing the association between a genetic variant and disease in samples consisting of affected individuals and their parents. In practice, more complex pedigree structures, that is siblings with no parents, or three-generational pedigrees with possibly missing genotypes, are common. There are several generalizations of the TDT that are suitable for use with arbitrary pedigree structures. We consider three such frequently used generalizations, family-based association test, pedigree disequilibrium test, and generalized disequilibrium test, that have accompanying software and compare them regarding validity and power in the single variant setting. We use simulations to study the effects of population admixture, populations whose genotypes are not in Hardy-Weinberg equilibrium (HWE), different pedigree structures, and the presence of linkage. Whereas our results show that some TDT generalizations can have a substantially increased Type 1 error, these tests are often used in substantive research without caveats about the validity of their Type 1 error. For the association analysis of rare variants in sequencing studies, region-based extensions of the TDT generalizations, that rely on the postulated robustness of the single variant tests, have been proposed. We discuss the implications of our results for these region-based extensions.
ABSTRACTFor family‐based association studies, Horvath et al. proposed an algorithm for the association analysis between haplotypes and arbitrary phenotypes when the phase of the haplotypes is unknown, that is, genotype data is given. Their approach to haplotype analysis maintains the original features of the TDT/FBAT‐approach, that is, complete robustness against genetic confounding and misspecification of the phenotype. The algorithm has been implemented in the FBAT and PBAT software package and has been used in numerous substantive manuscripts. Here, we propose a simplification of the original algorithm that maintains the original approach but reduces the computational burden of the approach substantially and gives valuable insights regarding the conditional distribution. With the modified algorithm, the application to whole‐genome sequencing (WGS) studies becomes feasible; for example, in sliding window approaches or spatial‐clustering approaches. The reduction of the computational burden that our modification provides is especially dramatic when both parental genotypes are missing. For example, for eight variants and 441 nuclear families with mostly offspring‐only families, in a WGS study at the APOE locus, the running time decreased from approximately 21 hr for the original algorithm to 0.11 sec after our modification.
This chapter covers a mix of topics including splines, functional data, extreme values and density estimation. Being associated with continuous time monitoring processes, functional data are usually smooth curves or surfaces and are often treated as realizations of underlying random functions. The basic philosophy of functional data is to consider the observed curves as single entities, rather than only as a sequence of individual observations. Comprehensive surveys of statistical techniques for analyzing functional data can be found in Ramsay and Silverman and Ferraty and Vieu. Goldsmith study Diffusion Tensor Imaging (DTI) metrics of multiple sclerosis (MS) patients over multiple clinical visits. The data consist of 100 subjects, aged between 21 and 70 years at first visit. The number of visits per subject ranges from 2 to 8, and a total of 340 visits were recorded. Most statistical modeling is concerned with the mean …