The human leukocyte antigen (HLA) locus is associated with more complex diseases than any other locus in the human genome. In many diseases, HLA explains more heritability than all other known loci combined. In silico HLA imputation methods enable rapid and accurate estimation of HLA alleles in the millions of individuals that are already genotyped on microarrays. HLA imputation has been used to define causal variation in autoimmune diseases, such as type I diabetes, and in human immunodeficiency virus infection control. However, there are few guidelines on performing HLA imputation, association testing, and fine mapping. Here, we present a comprehensive tutorial to impute HLA alleles from genotype data. We provide detailed guidance on performing standard quality control measures for input genotyping data and describe options to impute HLA alleles and amino acids either locally or using the web-based Michigan Imputation Server, which hosts a multi-ancestry HLA imputation reference panel. We also offer best practice recommendations to conduct association tests to define the alleles, amino acids, and haplotypes that affect human traits. Along with the pipeline, we provide a step-by-step online guide with scripts and available software (https://github.com/immunogenomics/HLA_analyses_tutorial). This tutorial will be broadly applicable to large-scale genotyping data and will contribute to defining the role of HLA in human diseases across global populations.
Genome-wide association studies (GWAS) have laid the foundation for investigations into the biology of complex traits, drug development and clinical guidelines. However, the majority of discovery efforts are based on data from populations of European ancestry 1–3. In light of the differential genetic architecture that is known to exist between populations, bias in representation can exacerbate existing disease and healthcare disparities. Critical variants may be missed if they have a low frequency or are completely absent in European populations, especially as the field shifts its attention towards rare variants, which are more likely to be population-specific 4–10. Additionally, effect sizes and their derived risk prediction scores derived in one population may not accurately extrapolate to other populations 11, 12. Here we demonstrate the value of diverse, multi-ethnic participants in large-scale genomic studies. The Population Architecture using Genomics and Epidemiology (PAGE) study conducted a GWAS of 26 clinical and behavioural phenotypes in 49,839 non-European individuals. Using strategies tailored for analysis of multi-ethnic and admixed populations, we describe a framework for analysing diverse populations, identify 27 novel loci and 38 secondary signals at known loci, as well as replicate 1,444 GWAS catalogue associations across these traits. Our data show evidence of effect-size heterogeneity across ancestries for published GWAS associations, substantial benefits for fine-mapping using diverse cohorts and insights into clinical implications. In the United States—where minority populations have a disproportionately higher burden of …
In the version of this article originally published, one of the two authors with the name Wei Zhao was omitted from the author list and the affiliations for both authors were assigned to the single Wei Zhao in the author list. In addition, the ORCID for Wei Zhao (Department of Biostatistics and Epidemiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA) was incorrectly assigned to author Wei Zhou. The errors have been corrected in the HTML and PDF versions of the article.
ObjectiveTo investigate to what extent low-frequency genetic variants (with minor allele frequencies <5%) affect the risk of intracranial aneurysms (IAs).MethodsOne thousand fifty-six patients with IA and 2,097 population-based controls from the Netherlands were genotyped with the Illumina HumanExome BeadChip. After quality control (QC) of samples and single nucleotide variants (SNVs), we conducted a single variant analysis using the Fisher exact test. We also performed the variable threshold (VT) test and the sequence kernel association test (SKAT) at different minor allele count (MAC) thresholds of >5 and >0 to test the hypothesis that multiple variants within the same gene are associated with IA risk. Significant results were tested in a replication cohort of 425 patients with IA and 311 controls, and results of the 2 cohorts were combined in a meta-analysis.ResultsAfter QC, 995 patients with IA and 2,080 controls remained for further analysis. The single variant analysis comprising 46,534 SNVs did not identify significant loci at the genome-wide level. The gene-based tests showed a statistically significant association for fibulin 2 (FBLN2) (best p = 1 x 10(-6) for the VT test, MAC >5). Associations were not statistically significant in the independent but smaller replication cohort (p > 0.57) but became slightly stronger in a meta-analysis of the 2 cohorts (best p = 4.8 x 10(-7) for the SKAT, MAC >= 1).ConclusionGene-based tests indicated an association for FBLN2, a gene encoding an extracellular matrix protein implicated in vascular wall remodeling, but independent validation in larger cohorts is warranted. We did not identify any significant associations for single low-frequency genetic variants.
OBJECTIVE:To explore genetic and lifestyle risk factors of MRI-defined brain infarcts (BI) in large population-based cohorts. METHODS:We performed meta-analyses of genome-wide association studies (GWAS) and examined associations of vascular risk factors and their genetic risk scores (GRS) with MRI-defined BI and a subset of BI, namely, small subcortical BI (SSBI), in 18 population-based cohorts (n = 20,949) from 5 ethnicities (3,726 with BI, 2,021 with SSBI). Top loci were followed up in 7 population-based cohorts (n = 6,862; 1,483 with BI, 630 with SBBI), and we tested associations with related phenotypes including ischemic stroke and pathologically defined BI. RESULTS:The mean prevalence was 17.7% for BI and 10.5% for SSBI, steeply rising after age 65. Two loci showed genome-wide significant association with BI: FBN2, p = 1.77 × 10-8; and LINC00539/ZDHHC20, p = 5.82 × 10-9. Both have been associated with blood pressure (BP)-related phenotypes, but did not replicate in the smaller follow-up sample or show associations with related phenotypes. Age- and sex-adjusted associations with BI and SSBI were observed for BP traits (p value for BI, p [BI] = 9.38 × 10-25; p [SSBI] = 5.23 × 10-14 for hypertension), smoking (p [BI] = 4.4 × 10-10; p [SSBI] = 1.2 × 10-4), diabetes (p [BI] = 1.7 × 10-8; p [SSBI] = 2.8 × 10-3), previous cardiovascular disease (p [BI] = 1.0 × 10-18; p [SSBI] = 2.3 × 10-7), stroke (p [BI] = 3.9 × 10-69; p [SSBI] = 3.2 × 10-24), and MRI-defined white matter hyperintensity burden (p [BI] = 1.43 × 10-157; p [SSBI] = 3.16 × 10-106), but not with body mass index or cholesterol. GRS of BP traits were associated with BI and SSBI (p ≤ 0.0022), without indication of directional pleiotropy. CONCLUSION:In this multiethnic GWAS meta-analysis, including over 20,000 population-based participants, we identified genetic risk loci for BI requiring validation once additional large datasets become available. High BP, including genetically determined, was the most significant modifiable, causal risk factor for BI.
Body-fat distribution is a risk factor for adverse cardiovascular health consequences. We analyzed the association of body-fat distribution, assessed by waist-to-hip ratio adjusted for body mass index, with 228,985 predicted coding and splice site variants available on exome arrays in up to 344,369 individuals from five major ancestries (discovery) and 132,177 European-ancestry individuals (validation). We identified 15 common (minor allele frequency, MAF ≥5%) and nine low-frequency or rare (MAF <5%) coding novel variants. Pathway/gene set enrichment analyses identified lipid particle, adiponectin, abnormal white adipose tissue physiology and bone development and morphology as important contributors to fat distribution, while cross-trait associations highlight cardiometabolic traits. In functional follow-up analyses, specifically in Drosophila RNAi-knockdowns, we observed a significant increase in the total body triglyceride levels for two genes ( DNAH10 and PLXND1 ). We implicate novel genes in fat distribution, stressing the importance of interrogating low-frequency and protein-coding variants.
Holoprosencephaly is a pathology of forebrain development characterized by high phenotypic heterogeneity. The disease presents with various clinical manifestations at the cerebral or facial levels. Several genes have been implicated in holoprosencephaly but its genetic basis remains unclear: different transmission patterns have been described including autosomal dominant, recessive and digenic inheritance. Conventional molecular testing approaches result in a very low diagnostic yield and most cases remain unsolved. In our study, we address the possibility that genetically unsolved cases of holoprosencephaly present an oligogenic origin and result from combined inherited mutations in several genes. Twenty-six unrelated families, for whom no genetic cause of holoprosencephaly could be identified in clinical settings [whole exome sequencing and comparative genomic hybridization (CGH)-array analyses], were reanalysed under the hypothesis of oligogenic inheritance. Standard variant analysis was improved with a gene prioritization strategy based on clinical ontologies and gene co-expression networks. Clinical phenotyping and exploration of cross-species similarities were further performed on a family-by-family basis. Statistical validation was performed on 248 ancestrally similar control trios provided by the Genome of the Netherlands project and on 574 ancestrally matched controls provided by the French Exome Project. Variants of clinical interest were identified in 180 genes significantly associated with key pathways of forebrain development including sonic hedgehog (SHH) and primary cilia. Oligogenic events were observed in 10 families and involved both known and novel holoprosencephaly genes including recurrently mutated FAT1, NDST1, COL2A1 and SCUBE2. The incidence of oligogenic combinations was significantly higher in holoprosencephaly patients compared to two control populations (P < 10-9). We also show that depending on the affected genes, patients present with particular clinical features. This study reports novel disease genes and supports oligogenicity as clinically relevant model in holoprosencephaly. It also highlights key roles of SHH signalling and primary cilia in forebrain development. We hypothesize that distinction between different clinical manifestations of holoprosencephaly lies in the degree of overall functional impact on SHH signalling. Finally, we underline that integrating clinical phenotyping in genetic studies is a powerful tool to specify the clinical relevance of certain mutations.
BACKGROUND:Atherosclerosis is a chronic inflammatory disease in part caused by lipid uptake in the vascular wall, but the exact underlying mechanisms leading to acute myocardial infarction and stroke remain poorly understood. Large consortia identified genetic susceptibility loci that associate with large artery ischemic stroke and coronary artery disease. However, deciphering their underlying mechanisms are challenging. Histological studies identified destabilizing characteristics in human atherosclerotic plaques that associate with clinical outcome. To what extent established susceptibility loci for large artery ischemic stroke and coronary artery disease relate to plaque characteristics is thus far unknown but may point to novel mechanisms. METHODS:We studied the associations of 61 established cardiovascular risk loci with 7 histological plaque characteristics assessed in 1443 carotid plaque specimens from the Athero-Express Biobank Study. We also assessed if the genotyped cardiovascular risk loci impact the tissue-specific gene expression in 2 independent biobanks, Biobank of Karolinska Endarterectomy and Stockholm Atherosclerosis Gene Expression. RESULTS:A total of 21 established risk variants (out of 61) nominally associated to a plaque characteristic. One variant (rs12539895, risk allele A) at 7q22 associated to a reduction of intraplaque fat, P=5.09×10-6 after correction for multiple testing. We further characterized this 7q22 Locus and show tissue-specific effects of rs12539895 on HBP1 expression in plaques and COG5 expression in whole blood and provide data from public resources showing an association with decreased LDL (low-density lipoprotein) and increase HDL (high-density lipoprotein) in the blood. CONCLUSIONS:Our study supports the view that cardiovascular susceptibility loci may exert their effect by influencing the atherosclerotic plaque characteristics.
Motivation Genome-wide association studies have had great success in identifying human genetic variants associated with disease, disease risk factors, and other biomedical phenotypes. Many variants are associated with multiple traits, even after correction for trait-trait correlation. Discovering subsets of variants associated with a shared subset of phenotypes could help reveal disease mechanisms, suggest new therapeutic options, and increase the power to detect additional variants with similar pattern of associations. Here we introduce two methods based on a Bayesian framework, SNP And Pleiotropic PHenotype Organization (SAPPHO), one modeling independent phenotypes (SAPPHO-I) and the other incorporating a full phenotype covariance structure (SAPPHO-C). These two methods learn patterns of pleiotropy from genotype and phenotype data, using identified associations to discover additional associations with shared patterns. Results The SAPPHO methods, along with other recent approaches for pleiotropic association tests, were assessed using data from the Atherosclerotic Risk in Communities (ARIC) study of 8,000 individuals, whose gold-standard associations were provided by meta-analysis of 40,000 to 100,000 individuals from the CHARGE consortium. Using power to detect gold-standard associations at genome-wide significance (0.05 family-wise error rate) as a metric, SAPPHO performed best. The SAPPHO methods were also uniquely able to select the most significant variants in a parsimonious model, excluding other less likely variants within a linkage disequilibrium block. For meta-analysis, the SAPPHO methods implement summary modes that use sufficient statistics rather than full phenotype and genotype data. Meta-analysis applied to CHARGE detected 16 additional associations to the gold-standard loci, as well as 124 novel loci, at 0.05 false discovery rate. Reasons for the superior performance were explored by performing simulations over a range of scenarios describing different genetic architectures. With SAPPHO we were able to learn genetic structures that were hidden using the traditional univariate tests. Availability . SAPPHO software is available under the GNU General Public License, v2.
Electrocardiographic PR interval measures atrial and atrioventricular depolarization and conduction, and abnormal PR interval is a risk factor for atrial fibrillation and heart block. We performed a genome-wide association study in over 92,000 individuals of European descent and identified 44 loci associated with PR interval (34 novel). Examination of the 44 loci revealed known and novel biological processes involved in cardiac atrial electrical activity, and genes in these loci were highly over-represented in several cardiac disease processes. Nearly half of the 61 independent index variants in the 44 loci were associated with atrial or blood transcript expression levels, or were in high linkage disequilibrium with one or more missense variants. Cardiac regulatory regions of the genome as measured by cardiac DNA hypersensitivity sites were enriched for variants associated with PR interval, compared to non-cardiac regulatory regions. Joint analyses combining PR interval with heart rate, QRS interval, and atrial fibrillation identified additional new pleiotropic loci. The majority of associations discovered in European-descent populations were also present in African-American populations. Meta-analysis examining over 105,000 individuals of African and European descent identified additional novel PR loci. These additional analyses identified another 13 novel loci. Together, these findings underscore the power of GWAS to extend knowledge of the molecular underpinnings of clinical processes.
Objective We sought to assess whether genetic risk factors for atrial fibrillation (AF) can explain cardioembolic stroke risk. Methods We evaluated genetic correlations between a previous genetic study of AF and AF in the presence of cardioembolic stroke using genome-wide genotypes from the Stroke Genetics Network (N = 3,190 AF cases, 3,000 cardioembolic stroke cases, and 28,026 referents). We tested whether a previously validated AF polygenic risk score (PRS) associated with cardioembolic and other stroke subtypes after accounting for AF clinical risk factors. Results We observed a strong correlation between previously reported genetic risk for AF, AF in the presence of stroke, and cardioembolic stroke (Pearson r = 0.77 and 0.76, respectively, across SNPs with p < 4.4 × 10−4 in the previous AF meta-analysis). An AF PRS, adjusted for clinical AF risk factors, was associated with cardioembolic stroke (odds ratio [OR] per SD = 1.40, p = 1.45 × 10−48), explaining ∼20% of the heritable component of cardioembolic stroke risk. The AF PRS was also associated with stroke of undetermined cause (OR per SD = 1.07, p = 0.004), but no other primary stroke subtypes (all p > 0.1). Conclusions Genetic risk of AF is associated with cardioembolic stroke, independent of clinical risk factors. Studies are warranted to determine whether AF genetic risk can serve as a biomarker for strokes caused by AF.
55 Introduction 56 Risk factors for abdominal aortic aneurysm (AAA) are largely unknown and this 57 has hampered development of non-surgical treatments to alter the natural 58 history of disease. Here, we use Mendelian randomization (MR) analyses to 59 investigate the association between lipid-associated single nucleotide 60 polymorphisms and AAA. 61 Methods 62 Genetic risk scores, composed of lipid trait associated SNPs were constructed 63 and tested for their association with AAA using conventional (inverse-variance 64 weighted) MR in up to 4,914 cases and 47,912 controls across 5 studies. 65 Sensitivity analyses to account for potential genetic pleiotropy included MR66 Egger and median weighted MR, and to test the independent effect of lipids and 67 risk of AAA we used multivariable MR. Finally, we assessed the association 68 between AAA and SNPs in loci that can act as proxies for drug targets. 69 Results. 70 A 1 standard deviation (SD) genetic elevation of low-density lipoprotein 71 cholesterol (LDL-C) was associated with increased risk of AAA (odds ratio [OR] 72
To identify loci affecting the electrocardiographic QT interval, a measure of cardiac repolarisation associated with risk of ventricular arrhythmias and sudden cardiac death, we conducted a meta-analysis of three genome-wide association studies (GWAS) including 3,558 subjects from the TwinsUK and BRIGHT cohorts in the UK and the DCCT/EDIC cohort from North America. Five loci were significantly associated with QT interval at P,1610. To validate these findings we performed an in silico comparison with data from two QT consortia: QTSCD (n = 15,842) and QTGEN (n = 13,685). Analysis confirmed the association between common variants near NOS1AP (P = 1.4610) and the phospholamban (PLN) gene (P = 1.9610). The most associated SNP near NOS1AP (rs12143842) explains 0.82% variance; the SNP near PLN (rs11153730) explains 0.74% variance of QT interval duration. We found no evidence for interaction between these two SNPs (P = 0.99). PLN is a key regulator of cardiac diastolic function and is involved in regulating intracellular calcium cycling, it has only recently been identified as a susceptibility locus for QT interval. These data offer further mechanistic insights into genetic influence on the QT interval which may predispose to life threatening arrhythmias and sudden cardiac death. PLoS ONE | www.plosone.org 1 July 2009 | Volume 4 | Issue 7 | e6138 Citation: Nolte IM, Wallace C, Newhouse SJ, Waggott D, Fu J, et al. (2009) Common Genetic Variation Near the Phospholamban Gene Is Associated with Cardiac Repolarisation: Meta-Analysis of Three Genome-Wide Association Studies. PLoS ONE 4(7): e6138. doi:10.1371/journal.pone.0006138 Editor: Peter M. Visscher, Queensland Institute of Medical Research, Australia Received May 12, 2009; Accepted June 4, 2009; Published July 9, 2009 Copyright: 2009 Nolte et al. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Funding: For TwinsUK: This study was funded by the British Heart Foundation, Project grant No. 06/094. The study was also funded by the Wellcome Trust; European Community’s Seventh Framework Programme (FP7/2007-2013)/grant agreement HEALTH-F2-2008-201865-GEFOS and (FP7/2007-2013), ENGAGE project grant agreement HEALTH-F4-2007-201413 and the FP-5 GenomEUtwin Project (QLG2-CT-2002-01254). The study also receives support from the Dept of Health via the National Institute for Health Research (NIHR) comprehensive Biomedical Research Centre award to Guy’s & St Thomas’ NHS Foundation Trust in partnership with King’s College London. TDS is an NIHR senior Investigator. The project also received support from a Biotechnology and Biological Sciences Research Council (BBSRC) project grant(G20234). The authors acknowledge the funding and support of the National Eye Institute via an NIH/CIDR genotyping project (PI: Terri Young). CD is supported by a British Heart Foundation grant: SP/02/001. For the BRIGHT study: The BRIGHT study is supported by the Medical Research Council of Great Britain (grant number; G9521010D) and the British Heart Foundation (grant number PG02/128). CW was funded by the British Heart Foundation (grant number: FS/05/061/19501). SJN is funded by the Medical Research Council and The William Harvey Research Foundation. Profs Dominiczak and Samani are British Heart Foundation Chairholders. EZ is funded by the Wellcome Trust (WT088885/Z/09/Z). For the DCCT/EDIC study: The DCCT/EDIC Research Group is sponsored through research contracts from the National Institute of Diabetes, Endocrinology and Metabolic Diseases of the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) and the National Institutes of Health. The authors are grateful to the subjects in the DCCT/EDIC cohort for their longterm participation. A.D.P. holds a Canada Research Chair in the Genetics of Complex Diseases. This work has received support from National Institute of Diabetes and Digestive and Kidney Diseases Contract N01-DK-6-2204, National Institute of Diabetes and Digestive and Kidney Diseases Grant R01-DK-077510 and support from the Canadian Network of Centres of Excellence in Mathematics and Genome Canada through the Ontario Genomics Institute. For QTSCD: ARIC is carried out as a collaborative study supported by National Heart, Lung, and Blood Institute contracts N01-HC-55015, N01-HC-55016, N01-HC-55018, N01-HC-55019, N01-HC55020, N01-HC-55021, N01-HC-55022, R01HL087641, R01HL59367 and R01HL086694; National Human Genome Research Institute contract U01HG004402; and National Institutes of Health contract HHSN268200625226C. Infrastructure was partly supported by Grant Number UL1RR025005, a component of the National Institutes of Health and NIH Roadmap for Medical Research. In addition, we acknowledge support from NHLBI grants HL86694 and HL054512, and the Donald W. Reynolds Cardiovascular Clinical Research Center at Johns Hopkins University for genotyping and data analysis relevant to this study. AK is supported by a German Research Foundation Fellowship. The KORA study was funded by the State of Bavaria and by grants from the German Federal Ministry of Education and Research (BMBF) in the context of the German National Genome Research Network (NGFN), the German National Competence network on atrial fibrillation (AFNET) and the Bioinformatics for the Functional Analysis of Mammalian Genomes program (BFAM) by grants to Stefan Kaab (NGFN 01GS0499, 01GS0838 and AF-Net 01GI0204/N), Arne Pfeufer (NGFN 01GR0803, 01EZ0874), H. -Erich Wichmann (NGFN 01GI0204) and to Thomas Meitinger (NGFN 01GR0103). Stefan Kaab is also supported by a grant from the Fondation Leducq. The SardiNIA team was supported by Contract NO1-AG-1-2109 from the National Institute on Aging contract NO1-AG-1-2109 to the SardiNIA (‘‘ProgeNIA’’) team and in part by the Intramural Research Program of the US National Institute on Aging, NIH. The efforts of G.R.A. were supported in part by contract 263-MA-410953 from the National Institute on Aging to the University of Michigan and by research grants from the National Human Genome Research Institute and the National Heart, Lung, and Blood Institute (to G.R.A.). The GenNOVA study was supported by the Ministry of Health of the Autonomous Province of Bolzano and the South Tyrolean Sparkasse Foundation. The Heinz Nixdorf Recall Study was funded by a grant of the Heinz Nixdorf Foundation (Chairman: Dr. jur. G. Schmidt). For QTGEN: The Framingham Heart Study work was supported by the National Heart Lung and Blood Institute of the National Institutes of Health and Boston University School of Medicine (Contract No. N01-HC-25195), its contract with Affymetrix, Inc for genotyping services (Contract No. N02-HL-6-4278), and the Doris Duke Charitable Foundation (C.N.-C.) and Burroughs Wellcome Fund (C.N.-C.), based on analyses by Framingham Heart Study investigators participating in the SNP Health Association Resource (SHARe) project. The measurement of ECG intervals in Framingham Heart Study generation 1 and 2 samples was performed by eResearchTechnology and was supported by an unrestricted grant from Pfizer. The Rotterdam Study is funded by Erasmus Medical Center and Erasmus University, Rotterdam, Netherlands Organization for the Health Research and Development (ZonMw), the Research Institute for Diseases in the Elderly (RIDE), the Ministry of Education, Culture and Science, the Ministry for Health, Welfare and Sports, the European Commission (DG XII), and the Municipality of Rotterdam. The generation and management of GWAS genotype data for the Rotterdam Study is supported by the Netherlands Organisation of Scientific Research NWO Investments (#175.010.2005.011, 91103-012). This study is funded by the Research Institute for Diseases in the Elderly (014-93-015; RIDE2), and the Netherlands Genomics Initiative (NGI)/Netherlands Organisation for Scientific Research (NWO) project #050-060-810. The CHS research reported in this article was supported by contract numbers N01-HC-85079 through N01-HC-85086, N01-HC-35129, N01 HC-15103, N01 HC-55222, N01-HC-75150, N01-HC-45133, grant numbers U01 HL080295 and R01 HL087652 from the National Heart, Lung, and Blood Institute, with additional contribution from the National Institute of Neurological Disorders and Stroke. C.N.-C. is supported by NIH K23-HL-080025, a Doris Duke Charitable Foundation Clinical Scientist Development Award, and a Burroughs Wellcome Fund Career Award for Medical Scientists. M.E is funded by the Netherlands Heart Foundation 2007B221. J.I.R. is supported by the Cedars-Sinai Board of Governors’ Chair in Medical Genetics. The measurement of ECG intervals in the Framingham Heart Study generation 3 sample was completed by Alim Hirji and Sirisha Kovvali using AMPS software provided through an unrestricted academic license by AMPS, LLC (New York, NY,) USAA full list of principal CHS investigators and institutions can be found at http://www.chs-nhlbi.org/pi.htm. The authors acknowledge the essential role of the CHARGE (Cohorts for Heart and Aging Research in Genome Epidemiology) Consortium in development and support of this manuscript. CHARGE members include the Netherland’s Rotterdam Study, the NHLBI’s Atherosclerosis Risk in Communities (ARIC) Study, Cardiovascular Health Study (CHS) and Framingham Heart Study (FHS), and the NIA’s Iceland Age, Gene/Environment Susceptibility (AGES) Study. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing Interests: Aravinda Chakravarti is a paid member of the Scientific Advisory Board of Affymetrix, a role that is managed by the Committee on Conflict of Interest of the Johns Hopkins University School of Medicine. * E-mail: y.jamshidi@sgul.ac.uk . These authors contributed equally to this work. "For the full author list see Appendix S1
A correction to this article has been published and is linked from the HTML and PDF versions of this paper. The error has not been fixed in the paper.
Atrial fibrillation is a prevalent arrhythmia associated with a five-fold increased risk of ischemic stroke, and specifically the cardioembolic stroke subtype. Genome-wide association studies of these traits have yielded overlapping risk loci, but genome-wide investigation of genetic susceptibility shared between stroke and atrial fibrillation is lacking. Comparing the genetic architectures of the two diseases could inform whether cardioembolic strokes are driven by inherited atrial fibrillation susceptibility, and may help elucidate ischemic stroke mechanisms. Here, we analyze genome-wide genotyping data and estimate SNP-based heritability in atrial fibrillation and cardioembolic stroke to be nearly identical (20.0% and 19.5%, respectively). Further, we find that the traits are genetically correlated (r=0.77 for SNPs with p -4 in a previous atrial fibrillation meta-analysis). Clinical studies are warranted to assess whether genetic susceptibility to atrial fibrillation can be leveraged to improve the diagnosis and care of ischemic stroke patients.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Background: Genotype imputation is an important procedure in current genomic analysis such as genome-wide association studies, meta-analyses and fine mapping. Although high quality tools are available that perform the steps of this process, considerable effort and expertise is required to set up and run a best practice imputation pipeline, particularly for larger genotype datasets, where imputation has to scale out in parallel on computer clusters. Results: Here we present MOLGENIS-impute, an ‘imputation in a box’ solution that seamlessly and transparently automates the set up and running of all the steps of the imputation process. These steps include genome build liftover (liftovering), genotype phasing with SHAPEIT2, quality control, sample and chromosomal chunking/merging, and imputation with IMPUTE2. MOLGENIS-impute builds on MOLGENIS-compute, a simple pipeline management platform for submission and monitoring of bioinformatics tasks in High Performance Computing (HPC) environments like local/cloud servers, clusters and grids. All the required tools, data and scripts are downloaded and installed in a single step. Researchers with diverse backgrounds and expertise have tested MOLGENIS-impute on different locations and imputed over 30,000 samples so far using the 1000 Genomes Project and new Genome of the Netherlands data as the imputation reference. The tests have been performed on PBS/SGE clusters, cloud VMs and in a grid HPC environment. Conclusions: MOLGENIS-impute gives priority to the ease of setting up, configuring and running an imputation. It has minimal dependencies and wraps the pipeline in a simple command line interface, without sacrificing flexibility to adapt or limiting the options of underlying imputation tools. It does not require knowledge of a workflow system or programming, and is targeted at researchers who just want to apply best practices in imputation via simple commands. It is built on the MOLGENIS compute workflow framework to enable customization with additional computational steps or it can be included in other bioinformatics pipelines. It is available as open source from: (https://github.com/molgenis/molgenis-imputation) keywords: Imputation, genotyping, GWAS
Electrocardiographic PR interval measures atrio-ventricular depolarization and conduction, and abnormal PR interval is a risk factor for atrial fibrillation and heart block. Our genome-wide association study of over 92,000 European-descent individuals identifies 44 PR interval loci (34 novel). Examination of these loci reveals known and previously not-yet-reported biological processes involved in cardiac atrial electrical activity. Genes in these loci are over-represented in cardiac disease processes including heart block and atrial fibrillation. Variants in over half of the 44 loci were associated with atrial or blood transcript expression levels, or were in high linkage disequilibrium with missense variants. Six additional loci were identified either by meta-analysis of ~105,000 African and European-descent individuals and/or by pleiotropic analyses combining PR interval with heart rate, QRS interval, and atrial fibrillation. These findings implicate developmental pathways, and identify transcription factors, ion-channel genes, and cell-junction/cell-signaling proteins in atrio-ventricular conduction, identifying potential targets for drug development.
Background. Previous genetic association studies of human immunodeficiency virus-1 (HIV-1) progression have focused on common human genetic variation ascertained through genome-wide genotyping. Methods. We sought to systematically assess the full spectrum of functional variation in protein coding gene regions on HIV-1 progression through exome sequencing of 1327 individuals. Genetic variants were tested individually and in aggregate across genes and gene sets for an influence on HIV-1 viral load. Results. Multiple single variants within the major histocompatibility complex (MHC) region were observed to be strongly associated with HIV-1 outcome, consistent with the known impact of classical HLA alleles. However, no single variant or gene located outside of the MHC region was significantly associated with HIV progression. Set-based association testing focusing on genes identified as being essential for HIV replication in genome-wide small interfering RNA (siRNA) and clustered regularly interspaced short palindromic repeats (CRISPR) studies did not reveal any novel associations. Conclusions. These results suggest that exonic variants with large effect sizes are unlikely to have a major contribution to host control of HIV infection.