Multiple sclerosis (MS) is a complex immune-mediated disorder with polygenic and multicellular underpinnings, necessitating cell-type-specific molecular studies to delineate dysregulated pathways. Here, we profile 1,075 transcriptomes from 167 patients with MS and 42 healthy participants across six peripheral immune cell-type-states. MS-associated transcriptional differences are more pronounced in primary (unstimulated) immune cells than in in vitro-stimulated counterparts. We identify shared and cell-type-specific transcriptional alterations at the level of genes, pathways, and co-expressed gene modules, prioritizing regulators, such as ZBTB16, across T cells and monocytes, and replicating six MS-associated modules in independent datasets. The top T cell module is enriched for MS susceptibility genes and affects proliferation. The top monocyte module implicates dysregulated TNF-α/NF-κB signaling, for which an in silico drug screen and in vitro validation nominate alvespimycin as a candidate modulator. Together, these findings define stable peripheral immune dysregulation signatures in MS that may serve as diagnostic or prognostic biomarkers in at-risk individuals.
Apart from ancestry, personal or environmental covariates may contribute to differences in polygenic score (PGS) performance. We analyzed the effects of covariate stratification and interaction on body mass index (BMI) PGS (PGS BMI ) across four cohorts of European (N = 491,111) and African (N = 21,612) ancestry. Stratifying on binary covariates and quintiles for continuous covariates, 18/62 covariates had significant and replicable R 2 differences among strata. Covariates with the largest differences included age, sex, blood lipids, physical activity, and alcohol consumption, with R 2 being nearly double between best- and worst-performing quintiles for certain covariates. Twenty-eight covariates had significant PGS BMI –covariate interaction effects, modifying PGS BMI effects by nearly 20% per standard deviation change. We observed overlap between covariates that had significant R 2 differences among strata and interaction effects – across all covariates, their main effects on BMI were correlated with their maximum R 2 differences and interaction effects (0.56 and 0.58, respectively), suggesting high-PGS BMI individuals have highest R 2 and increase in PGS effect. Using quantile regression, we show the effect of PGS BMI increases as BMI itself increases, and that these differences in effects are directly related to differences in R 2 when stratifying by different covariates. Given significant and replicable evidence for context-specific PGS BMI performance and effects, we investigated ways to increase model performance taking into account nonlinear effects. Machine learning models (neural networks) increased relative model R 2 (mean 23%) across datasets. Finally, creating PGS BMI directly from GxAge genome-wide association studies effects increased relative R 2 by 7.8%. These results demonstrate that certain covariates, especially those most associated with BMI, significantly affect both PGS BMI performance and effects across diverse cohorts and ancestries, and we provide avenues to improve model performance that consider these effects.
An inverse correlation between stature and risk of coronary artery disease (CAD) has been observed in several epidemiologic studies, and recent Mendelian randomization (MR) experiments have suggested causal association. However, the extent to which the effect estimated by MR can be explained by cardiovascular, anthropometric, lung function, and lifestyle-related risk factors is unclear, with a recent report suggesting that lung function traits could fully explain the height-CAD effect. To clarify this relationship, we utilized a well-powered set of genetic instruments for human stature, comprising >1,800 genetic variants for height and CAD. In univariable analysis, we confirmed that a one standard deviation decrease in height (~6.5 cm) was associated with a 12.0% increase in the risk of CAD, consistent with previous reports. In multivariable analysis accounting for effects from up to 12 established risk factors, we observed a >3-fold attenuation in the causal effect of height on CAD susceptibility (3.7%, p = 0.02). However, multivariable analyses demonstrated independent effects of height on other cardiovascular traits beyond CAD, consistent with epidemiologic associations and univariable MR experiments. In contrast with published reports, we observed minimal effects of lung function traits on CAD risk in our analyses, indicating that these traits are unlikely to explain the residual association between height and CAD risk. In sum, these results suggest the impact of height on CAD risk beyond previously established cardiovascular risk factors is minimal and not explained by lung function measures.
Widespread availability of antiretroviral therapies (ART) for HIV-1 have generated considerable interest in understanding the pharmacogenomics of ART. In some individuals, ART has been associated with excessive weight gain, which disproportionately affects women of African ancestry. The underlying biology of ART-associated weight gain is poorly understood, but some genetic markers which modify weight gain risk have been suggested, with more genetic factors likely remaining undiscovered. To overcome limitations in available sample sizes for genome-wide association studies (GWAS) in people with HIV, we explored whether a multi-ancestry polygenic risk score (PRS) derived from large, publicly available non-HIV GWAS for body mass index (BMI) can achieve high cross-ancestry performance for predicting baseline BMI in diverse, prospective ART clinical trials datasets, and whether that PRSBMI is also associated with change in BMI over 48 weeks on ART. We show that PRSBMI explained ∼5-7% of variability in baseline (pre-ART) BMI, with high performance in both European and African genetic ancestry groups, but that PRSBMI was not associated with change in BMI on ART. This study argues against a shared genetic predisposition for baseline (pre-ART) BMI and ART-associated weight gain.
Multiple sclerosis is a leading cause of neurological disability in adults. Heterogeneity in multiple sclerosis clinical presentation has posed a major challenge for identifying genetic variants associated with disease outcomes. To overcome this challenge, we used prospectively ascertained clinical outcomes data from the largest international multiple sclerosis registry, MSBase. We assembled a cohort of deeply phenotyped individuals of European ancestry with relapse-onset multiple sclerosis. We used unbiased genome-wide association study and machine learning approaches to assess the genetic contribution to longitudinally defined multiple sclerosis severity phenotypes in 1813 individuals. Our primary analyses did not identify any genetic variants of moderate to large effect sizes that met genome-wide significance thresholds. The strongest signal was associated with rs7289446 (β = -0.4882, P = 2.73 × 10-7), intronic to SEZ6L on chromosome 22. However, we demonstrate that clinical outcomes in relapse-onset multiple sclerosis are associated with multiple genetic loci of small effect sizes. Using a machine learning approach incorporating over 62 000 variants together with clinical and demographic variables available at multiple sclerosis disease onset, we could predict severity with an area under the receiver operator curve of 0.84 (95% CI 0.79-0.88). Our machine learning algorithm achieved positive predictive value for outcome assignation of 80% and negative predictive value of 88%. This outperformed our machine learning algorithm that contained clinical and demographic variables alone (area under the receiver operator curve 0.54, 95% CI 0.48-0.60). Secondary, sex-stratified analyses identified two genetic loci that met genome-wide significance thresholds. One in females (rs10967273; βfemale = 0.8289, P = 3.52 × 10-8), the other in males (rs698805; βmale = -1.5395, P = 4.35 × 10-8), providing some evidence for sex dimorphism in multiple sclerosis severity. Tissue enrichment and pathway analyses identified an overrepresentation of genes expressed in CNS compartments generally, and specifically in the cerebellum (P = 0.023). These involved mitochondrial function, synaptic plasticity, oligodendroglial biology, cellular senescence, calcium and G-protein receptor signalling pathways. We further identified six variants with strong evidence for regulating clinical outcomes, the strongest signal again intronic to SEZ6L (adjusted hazard ratio 0.72, P = 4.85 × 10-4). Here we report a milestone in our progress towards understanding the clinical heterogeneity of multiple sclerosis outcomes, implicating functionally distinct mechanisms to multiple sclerosis risk. Importantly, we demonstrate that machine learning using common single nucleotide variant clusters, together with clinical variables readily available at diagnosis can improve prognostic capabilities at diagnosis, and with further validation has the potential to translate to meaningful clinical practice change.
Polygenic risk scores (PRS) have led to enthusiasm for precision medicine. However, it is well documented that PRS do not generalize across groups differing in ancestry or sample characteristics e.g., age. Quantifying performance of PRS across different groups of study participants, using genome-wide association study (GWAS) summary statistics from multiple ancestry groups and sample sizes, and using different linkage disequilibrium (LD) reference panels may clarify which factors are limiting PRS transferability. To evaluate these factors in the PRS generation process, we generated body mass index (BMI) PRS (PRSBMI) in the Electronic Medical Records and Genomics (eMERGE) network (N=75,661). Analyses were conducted in two ancestry groups (European and African) and three age ranges (adult, teenagers, and children). For PRSBMI calculations, we evaluated five LD reference panels and three sets of GWAS summary statistics of varying sample size and ancestry. PRSBMI performance increased for both African and European ancestry individuals using cross-ancestry GWAS summary statistics compared to European-only summary statistics (6.3% and 3.7% relative R-2 increase, respectively, p(African)=0.038, p(European)=6.26x10(-4)). The effects of LD reference panels were more pronounced in African ancestry study datasets. PRSBMI performance degraded in children; R-2 was less than half of teenagers or adults. The effect of GWAS summary statistics sample size was small when modeled with the other factors. Additionally, the potential of using a PRS generated for one trait to predict risk for comorbid diseases is not well understood especially in the context of cross-ancestry analyses - we explored clinical comorbidities from the electronic health record associated with PRSBMI and identified significant associations with type 2 diabetes and coronary atherosclerosis. In summary, this study quantifies the effects that ancestry, GWAS summary statistic sample size, and LD reference panel have on PRS performance, especially in cross-ancestry and age-specific analyses.
Loss or absence of hearing is common at both extremes of human lifespan, in the forms of congenital deafness and age-related hearing loss. While these are often studied separately, there is increasing evidence that their genetic basis is at least partially overlapping. In particular, both common and rare variants in genes associated with monogenic forms of hearing loss also contribute to the more polygenic basis of age-related hearing loss. Here, we directly test this model in the Penn Medicine BioBank–a healthcare system cohort of around 40,000 individuals with linked genetic and electronic health record data. We show that increased burden of predicted deleterious variants in Mendelian hearing loss genes is associated with increased risk and severity of adult-onset hearing loss. As a specific example, we identify one gene– TCOF1 , responsible for a syndromic form of congenital hearing loss–in which deleterious variants are also associated with adult-onset hearing loss. We also identify four additional novel candidate genes ( COL5A1 , HMMR , RAPGEF3 , and NNT ) in which rare variant burden may be associated with hearing loss. Our results confirm that rare variants in Mendelian hearing loss genes contribute to polygenic risk of hearing loss, and emphasize the utility of healthcare system cohorts to study common complex traits and diseases.
African ancestry populations are underrepresented in human genetic studies, which leaves a knowledge gap about genetic and environmental risk factors for metabolic disease which can bias healthcare treatment. To alleviate these shortcomings, we conducted a series of genome-wide association studies of cardiometabolic traits in a diverse sampling of ∼2,500 (max sample size) ethnically and geographically diverse Africans from populations practicing agriculturalist, hunter-gatherer, and pastoralist subsistence strategies. This study includes the Fulani pastoralists who have a relatively high incidence of adult-onset diabetes, despite having a low average body mass index (BMI) . All individuals in the present study are sampled from rural populations that have relatively homogeneous lifestyles and diet within a community. This unique aspect to the study cohort allows for within- and between- group comparisons to identify trait variation attributable to genetics vs. lifestyle variation. Subjects were genotyped on a new African-focused SNP array from the Human Heredity and Health in Africa (H3 Africa) consortium. Genome-wide data was imputed based on a panel of African whole-genome sequences and data from the 1000 genomes project, resulting in a total of million variants for trait association analysis. Genetic associations were tested for BMI, blood pressure, and blood biomarkers of cardiometabolic health. We checked for replication of genotype/phenotype associations by comparison to large non-African cohorts studied for the same traits. We find that most associations identified in non-African populations do not replicate in the Africans. However, we identified a number of novel loci associated with cardiometabolic traits in the African populations. This study has important implications for identifying genetic risk factors that may play a role in metabolic disease in individuals of African ancestry. Funded by ADA Pathway award 1-19-VSN-02. Disclosure D.Hui: None. T.B.Nyambo: None. S.Chanock: None. S.A.Tishkoff: None. D.Harris: None. M.Mcquillan: None. M.Hansen: None. A.Ranciaro: None. W.Beggs: None. S.W.Mpoloka: None. D.Woldemeskel: None. A.K.K.Njamnshi: None. Funding American Diabetes Association (1-19-VSN-02) ; National Institutes of Health (R35 GM134957-01, R01AR076241, 5T32DK007314-39, 1OT3HL142479-01, 1OT3HL142478-01, 1OT3HL142481-01, 1OT3HL142480-01, 1OT3HL147154-01)
Multiple sclerosis (MS) is a leading cause of neurological disability in adults. Heterogeneity in MS clinical presentation has posed a major challenge for identifying genetic variants associated with disease outcomes. To overcome this challenge, we used prospectively ascertained clinical outcomes data from the largest international MS Registry, MSBase. We assembled a cohort of deeply phenotyped individuals with relapse-onset MS. We used unbiased genome-wide association study and machine learning approaches to assess the genetic contribution to longitudinally defined MS severity phenotypes in 1,813 individuals. Our results did not identify any variants of moderate to large effect sizes that met genome-wide significance thresholds. However, we demonstrate that clinical outcomes in relapse-onset MS are associated with multiple genetic loci of small effect sizes. Using a machine learning approach incorporating over 62,000 variants and demographic variables available at MS disease onset, we could predict severity with an area under the receiver operator curve (AUROC) of 0.87 (95% CI 0.83 – 0.91). This approach, if externally validated, could quickly prove useful for clinical stratification at MS onset. Further, we find evidence to support central nervous system and mitochondrial involvement in determining MS severity.### Competing Interest StatementThe authors have declared no competing interest.### Funding StatementThis work was supported by a Research Fellowship awarded to Dr Vilija Jokubaitis from Multiple Sclerosis Australia (16-0206), and research grant support from the Royal Melbourne Hospital Home Lottery Grant (MH2013-055), Charity Works for MS (2012 Project grant), MSBase Foundation Project Grant, and Monash University.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:This study was approved by the Melbourne Health Human Research Ethics Committee, and by institutional review boards at all participating centres. All participants gave written informed consent for participation in the MSBase Registry, together with additional informed consent to participate in genetic research (HREC/13/MH/189 and per local approvals elsewhere).I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesClinical data from the MSBase Registry: To protect participant confidentiality, de-identified patient-level data sharing may be possible in principle, but will require permissions/consent from each contributing data controller. Genetic data will be deposited to an access controlled database shortly, whilst we continue to explore these data in further analyses. Access requests with scientifically sound proposals can be made in writing to Dr Vilija Jokubaitis (vilija.jokubaitis{at}monash.edu) or Prof Helmut Butzkueven (helmut.butzkueven{at}monash.edu).
Genome‐wide association studies (GWAS) are being conducted at an unprecedented rate in population‐based cohorts and have increased our understanding of the pathophysiology of many complex diseases. Regardless of the context, the practical utility of this information ultimately depends upon the quality of the data used for statistical analyses. Quality control (QC) procedures for GWAS are constantly evolving. Here, we enumerate some of the challenges in QC of genotyped GWAS data and describe the approaches involving genotype imputation of a sample dataset along with post‐imputation quality assurance, thereby minimizing potential bias and error in GWAS results. We discuss common issues associated with QC of the GWAS data (genotyped and imputed), including data file formats, software packages for data manipulation and analysis, sex chromosome anomalies, sample identity, sample relatedness, population substructure, batch effects, and marker quality. We provide detailed guidelines along with a sample dataset to suggest current best practices and discuss areas of ongoing and future research. © 2022 Wiley Periodicals LLC.
The polygenic and multi-cellular nature of multiple sclerosis (MS) immunopathology necessitates cell-type-specific molecular studies in order to improve our understanding of the diverse mechanisms underlying immune cell dysfunction in MS. Here, by generating a dataset of 1,075 transcriptomes from 209 participants (167 MS and 42 healthy), we assessed MS-associated transcriptional changes in six implicated cell-type-states: naïve and memory helper T cells and classical monocytes purified from peripheral blood, each in their primary ( ex vivo , unstimulated) and in vitro stimulated states. Our data suggest that primary profiles show larger MS-associated differences than the post-stimulation contexts. We further identified shared and distinct changes in individual genes, biological pathways, and co-expressed gene modules in MS T cells and monocytes, and prioritized genes such as ZBTB16 as MS-associated regulators in both cell types. Of six identified MS-associated co-expressed gene modules, three (two lymphoid and one myeloid) were replicated in independent data from peripheral blood mononuclear cells (PBMC) and monocyte-derived macrophages. A subsequent in silico drug screen prioritized small-molecule compounds for reversing the perturbation of the MS-associated modules. The effects of glucocorticoid receptor agonists as the top-identified therapeutic class for the replicated T cell modules were validated using targeted in silico analyses and in vitro experiments, suggesting the coordinated dysregulation of glucocorticoid-responsive genes in MS T cells. In summary, our study identifies and validates individual genes and co-expressed gene modules from T and myeloid cells that are perturbed in MS, offering new targets for therapeutic discovery and biomarker development to guide the management of MS.
One goal of genomic medicine is to uncover an individual's genetic risk for disease, which generally requires data connecting genotype to phenotype, as done in genome-wide association studies (GWAS). While there may be clinical promise to employing prediction tools such as polygenic risk scores (PRS), it currently stands that individuals of non-European ancestry may not reap the benefits of genomic medicine because of underrepresentation in large-scale genetics studies. Here, we discuss why this inequity poses a problem for genomic medicine and the reasons for the low transferability of PRS across populations. We also survey the ancestry representation of published GWAS and investigate how estimates of ancestry diversity in GWASparticipants might be biased. We highlight the importance of expanding genetic research in Africa, one of the most underrepresented regions in human genomics research, and discuss issues of ethics, resources, and technology for equitable advancement of genomic medicine.
OBJECTIVE:To investigate the importance of rare variants in adult-onset hearing loss. STUDY DESIGN:Genomic association study. SETTING:Large biobank from tertiary care center. METHODS:We investigated rare variants (minor allele frequency <5%) in 42 autosomal dominant (DFNA) postlingual hearing loss (HL) genes in 16,657 unselected individuals in the Penn Medicine Biobank. We determined the prevalence of known pathogenic and predicted deleterious variants in subjects with audiometric-proven sensorineural hearing loss. We scanned across known postlingual DFNA HL genes to determine those most significantly contributing to the phenotype. We replicated findings in an independent cohort (UK Biobank). RESULTS:While rare individually, when considering the accumulation of variants in all postlingual DFNA genes, more than 90% of participants carried at least 1 rare variant. Rare variants predicted to be deleterious were enriched in adults with audiometric-proven hearing loss (pure-tone average >25 dB; P = .015). Patients with a rare predicted deleterious variant had an odds ratio of 1.27 for HL compared with genotypic controls (P = .029). Gene burden in DIABLO, PTPRQ, TJP2, and POU4F3 were independently associated with sensorineural hearing loss. CONCLUSION:Although prior reports have focused on common variants, we find that rare predicted deleterious variants in DFNA postlingual HL genes are enriched in patients with adult-onset HL in a large health care system population. We show the value of investigating rare variants to uncover hearing loss phenotypes related to implicated genes.
Plasma lipids are known heritable risk factors for cardiovascular disease, but increasing evidence also supports shared genetics with diseases of other organ systems. We devised a comprehensive three-phase framework to identify new lipid-associated genes and study the relationships among lipids, genotypes, gene expression and hundreds of complex human diseases from the Electronic Medical Records and Genomics (347 traits) and the UK Biobank (549 traits). Aside from 67 new lipid-associated genes with strong replication, we found evidence for pleiotropic SNPs/genes between lipids and diseases across the phenome. These include discordant pleiotropy in the HLA region between lipids and multiple sclerosis and putative causal paths between triglycerides and gout, among several others. Our findings give insights into the genetic basis of the relationship between plasma lipids and diseases on a phenome-wide scale and can provide context for future prevention and treatment strategies. An analytical framework based on transcriptome-wide association and electronic medical records provides insights into the relationship between plasma lipids and complex diseases on a phenome-wide scale.
In complex trait genetics, the ability to predict phenotype from genotype is the ultimate measure of our understanding of genetic architecture underlying the heritability of a trait. A complete understanding of the genetic basis of a trait should allow for predictive methods with accuracies approaching the trait's heritability. The highly polygenic nature of quantitative traits and most common phenotypes has motivated the development of statistical strategies focused on combining myriad individually non-significant genetic effects. Now that predictive accuracies are improving, there is a growing interest in the practical utility of such methods for predicting risk of common diseases responsive to early therapeutic intervention. However, existing methods require individual-level genotypes or depend on accurately specifying the genetic architecture underlying each disease to be predicted. Here, we propose a polygenic risk prediction method that does not require explicitly modeling any underlying genetic architecture. We start with summary statistics in the form of SNP effect sizes from a large GWAS cohort. We then remove the correlation structure across summary statistics arising due to linkage disequilibrium and apply a piecewise linear interpolation on conditional mean effects. In both simulated and real datasets, this new non-parametric shrinkage (NPS) method can reliably allow for linkage disequilibrium in summary statistics of 5 million dense genome-wide markers and consistently improves prediction accuracy. We show that NPS improves the identification of groups at high risk for breast cancer, type 2 diabetes, inflammatory bowel disease, and coronary heart disease, all of which have available early intervention or prevention treatments.
Non-healing diabetic foot ulceration (DFU) is characterized by low grade chronic inflammation, both locally and systemically. We prospectively followed a group of DFU patients who either healed or developed non-healing chronic DFU. Serum and forearm skin analysis, both at protein expression and transcriptomic level, indicated that increased expression of factors such as IFNγ, VEGF and sVCAM-1 were associated with DFU healing. Furthermore, foot skin single-cell RNA-seq analysis showed multiple fibroblast cell clusters and increased inflammation in the dorsal skin of Diabetes Mellitus (DM) patients and DFU specimens when compared to controls. In addition, in myeloid cells of DM and DFU patients upstream regulator analysis, we observed inhibition of IL-13 and IFNγ and dysregulation of biological processes that included cell movement of monocytes, migration of dendritic cells and chemotaxis of antigen presenting cells pointing to an impaired migratory profile of immune cells in diabetic skin. The SLCO2A1 and CYP1A1 genes, which were up-regulated at the forearm of Non-Healers, were mainly expressed by the vascular endothelial cell cluster almost exclusively in DFU indicating a potential important role in wound healing. These results from integrated protein and transcriptome analyses identified individual genes and pathways that can potentially be targeted for enhancing DFU healing
Vitamin D is a nutrient and a hormone with multiple effects on immune regulation and respiratory viral infections, which can worsen asthma and lead to severe asthma exacerbations. We set up a complete experimental and analytical pipeline for ATAC-Seq and RNA-Seq to study genome-wide epigenetic changes in human bronchial epithelial cells of asthmatic subjects, following treatment of these cells with calcitriol (vitamin D3) and Poly (I:C)(a viral analogue). This approach led to the identification of biologically plausible candidate genes for viral infections and asthma, such as DUSP10 and SLC44A1.