Obesity is a major public health crisis associated with high mortality rates. Previous genome-wide association studies (GWAS) investigating body mass index (BMI) have largely relied on imputed data from European individuals. This study leveraged whole-genome sequencing (WGS) data from 88,873 participants from the Trans-Omics for Precision Medicine (TOPMed) Program, of which 51% were of non-European population groups. We discovered 18 BMI-associated signals (P < 5 × 10-9). Notably, we identified and replicated a novel low frequency single nucleotide polymorphism (SNP) in MTMR3 that was common in individuals of African descent. Using a diverse study population, we further identified two novel secondary signals in known BMI loci and pinpointed two likely causal variants in the POC5 and DMD loci. Our work demonstrates the benefits of combining WGS and diverse cohorts in expanding current catalog of variants and genes confer risk for obesity, bringing us one step closer to personalized medicine.
Background:Asthma exacerbations are episodes of symptom worsening requiring increased therapy, which affect patients across all asthma severities. Potential genetic associations with asthma exacerbations in an understudied population were investigated. Objective:We sought to perform a genome-wide association study on severe asthma exacerbations in an admixed adult population with varying asthma severities and to explore potential epigenetic roles. Methods:A genome-wide association study was conducted in 727 Brazilian patients (mean age, 43 years; 20% male; and 55% with exacerbations) from the Programa de Controle da Asma na Bahia study, analyzing 12 million variants. Severe exacerbation was defined as systemic corticosteroid use for 3 or more days, emergency room visits, or hospital admissions within the past year. Analyses were adjusted for age, sex, asthma severity, and genotype principal components. Replication was sought in the cohorts of the Unbiased Biomarkers for the Prediction of Respiratory Disease Outcomes (UBIOPRED) study, the Genes-environments and Admixture in Latino Americans II study, and the Study of African Americans, Asthma, Genes, and Environments using the same methodology. Epigenetic effects were assessed in silico via PhenoScanner v2. Results:Five intergenic variants (rs55670125, rs10854420, rs68160941, rs11910414, and rs35834033) in complete linkage disequilibrium reached genome-wide significance (odds ratio [OR], 2.5; P = 3.47 × 10-8), located between the CXADR and LOC105372741 genes on chromosome 21. Although not replicated, rs35834033 showed a nonsignificant trend (OR, 1.79; P = .17). Four variants were associated with H3K4me1 histone modification, linked to asthma pathogenesis. In addition, 88 suggestive variants were found; rs17697822 in FOXP1 was negatively associated with exacerbations (OR, 0.44; P = 4.03 × 10-6). Conclusions:The CXADR is highlighted as a potential novel susceptibility locus for asthma exacerbations, possibly tied to viral respiratory infections. Further replication and validation are needed.
Background:Genetic control of gene expression in asthma-related tissues is not well-characterized, particularly for African-ancestry populations, limiting advancement in our understanding of the increased prevalence and severity of asthma in those populations. Objective:To create novel transcriptome prediction models for asthma tissues (nasal epithelium and CD4+ T cells) and apply them in transcriptome-wide association study (TWAS) to discover candidate asthma genes. Methods:We developed and validated gene expression prediction databases for unstimulated CD4+ T cells (CD4+T) and nasal epithelium using an elastic net framework. Combining these with existing prediction databases (N=51), we performed TWAS of 9,284 individuals of African-ancestry to identify tissue-specific and cross-tissue candidate genes for asthma. For detailed Methods, please see the Supplemental Methods. Results:Novel databases for CD4+T and nasal epithelial gene expression prediction contain 8,351 and 10,296 genes, respectively, including four asthma loci (SCGB1A1, MUC5AC, ZNF366, LTC4S) not predictable with existing public databases. Prediction performance was comparable to existing databases and was most accurate for populations sharing ancestry with the training set (e.g. African ancestry). From TWAS, we identified 17 candidate causal asthma genes (adjusted P<0.1), including genes with tissue-specific (IL33 in nasal epithelium) and cross-tissue (CCNC and FBXW7) effects. Conclusions:Expression of IL33, CCNC, and FBXW7 may affect asthma risk in African ancestry populations by mediating inflammatory responses. The addition of CD4+T and nasal epithelium prediction databases to the public sphere will improve ancestry representation and power to detect novel gene-trait associations from TWAS.
In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole-genome sequencing (∼27X) of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls (∼37X). To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of the association of BD with single variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD genome-wide association study loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole-genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.
BACKGROUND:Genetic control of gene expression in asthma-related tissues is not well characterized, particularly for African-ancestry populations, limiting advancement in our understanding of the increased prevalence and severity of asthma in these populations. OBJECTIVE:We sought to create novel transcriptome prediction models for asthma tissues (nasal epithelium and CD4+ T cells) and apply them in a transcriptome-wide association study (TWAS) to discover candidate asthma genes. METHODS:We developed and validated gene expression prediction databases for unstimulated CD4+ T cells and nasal epithelium using an elastic net framework. Combining these with existing prediction databases (N = 51), we performed a TWAS of 9284 individuals of African ancestry to identify tissue-specific and cross-tissue candidate genes for asthma. RESULTS:Novel databases for CD4+ T cells and nasal epithelial gene expression prediction contain 8,351 and 10,296 genes, respectively, including 4 asthma loci (SCGB1A1, MUC5AC, ZNF366, and LTC4S) not predictable with existing public databases. Prediction performance was comparable to existing databases and was most accurate for populations sharing ancestry with the training set (eg, African ancestry). From the TWAS, we identified 17 candidate causal asthma genes (adjusted P < .1), including genes with tissue-specific (IL33 in nasal epithelium) and cross-tissue (CCNC and FBXW7) effects. CONCLUSIONS:Expression of IL33, CCNC, and FBXW7 may affect asthma risk in African- ancestry populations by mediating inflammatory responses. The addition of CD4+ T cell and nasal epithelium prediction databases to the public sphere will improve ancestry representation and power to detect novel gene-trait associations from TWAS.
Mosaic loss of Y (mLOY) is the most common somatic chromosomal alteration detected in human blood. The presence of mLOY is associated with altered blood cell counts and increased risk of Alzheimer disease, solid tumors, and other age-related diseases. We sought to gain a better understanding of genetic drivers and associated phenotypes of mLOY through analyses of whole-genome sequencing (WGS) of a large set of genetically diverse males from the Trans-Omics for Precision Medicine (TOPMed) program. We show that haplotype-based calling methods can be used with WGS data to successfully identify mLOY events. This approach enabled us to identify differences in mLOY frequencies across populations defined by genetic similarity, revealing a higher frequency of mLOY in the European (EUR) ancestry group compared to other ancestries. We identify multiple loci associated with mLOY susceptibility and show that subsets of human hematopoietic stem cells are enriched for the activity of mLOY susceptibility variants. Finally, we found that certain alleles on chromosome Y are more likely to be lost than others in detectable mLOY clones.
Precision medicine initiatives across the globe have led to a revolution of repositories linking large-scale genomic data with electronic health records, enabling genomic analyses across the entire phenome. Many of these initiatives focus solely on research insights, leading to limited direct benefit to patients. We describe the biobank at the Colorado Center for Personalized Medicine (CCPM Biobank) that was jointly developed by the University of Colorado Anschutz Medical Campus and UCHealth to serve as a unique, dual-purpose research and clinical resource accelerating personalized medicine. This living resource currently has more than 200,000 participants with ongoing recruitment. We highlight the clinical, laboratory, regulatory, and HIPAA-compliant informatics infrastructure along with our stakeholder engagement, consent, recontact, and participant engagement strategies. We characterize aspects of genetic and geographic diversity unique to the Rocky Mountain region, the primary catchment area for CCPM Biobank participants. We leverage linked health and demographic information of the CCPM Biobank participant population to demonstrate the utility of the CCPM Biobank to replicate complex trait associations in the first 33,674 genotyped individuals across multiple disease domains. Finally, we describe our current efforts toward return of clinical genetic test results, including high-impact pathogenic variants and pharmacogenetic information, and our broader goals as the CCPM Biobank continues to grow. Bringing clinical and research interests together fosters unique clinical and translational questions that can be addressed from the large EHR-linked CCPM Biobank resource within a HIPAA- and CLIA-certified environment.
Abstract Asthma has striking disparities across ancestral groups, but the molecular underpinning of these differences is poorly understood and minimally studied. A goal of the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to understand multi-omic signatures of asthma focusing on populations of African ancestry. RNASeq and DNA methylation data are generated from nasal epithelium including cases (current asthma, N = 253) and controls (never-asthma, N = 283) from 7 different geographic sites to identify differentially expressed genes (DEGs) and gene networks. We identify 389 DEGs; the top DEG, FN1, was downregulated in cases (q = 3.26 × 10−9) and encodes fibronectin which plays a role in wound healing. The top three gene expression modules implicate networks related to immune response (CEACAM5; p = 9.62 × 10−16 and CPA3; p = 2.39 × 10−14) and wound healing (FN1; p = 7.63 × 10−9). Multi-omic analysis identifies FKBP5, a co-chaperone of glucocorticoid receptor signaling known to be involved in drug response in asthma, where the association between nasal epithelium gene expression is likely regulated by methylation and is associated with increased use of inhaled corticosteroids. This work reveals molecular dysregulation on three axes – increased Th2 inflammation, decreased capacity for wound healing, and impaired drug response – that may play a critical role in asthma within the African Diaspora.
The following sections are included:OverviewDealing with the lack of diversity in current research datasetsDevelopment of fair machine learning algorithmsRace, genetic ancestry, and population structureConclusionAcknowledgments.
Background: Asthma is a chronic inflammatory disease of the airways that is heterogeneous and multifactorial, making its accurate characterization a complex process. Therefore, identifying the genetic variations associated with asthma and discovering the molecular interactions between the omics that confer risk of developing this disease will help us to unravel the biological pathways involved in its pathogenesis. Objective: We sought to develop a predictive genetic panel for asthma using machine learning methods. Methods: We tested 3 variable selection methods: Boruta's algorithm, the top 200 genome-wide association study markers according to their respective P values, and an elastic net regression. Ten different algorithms were chosen for the classification tests. A predictive panel was built on the basis of joint scores between the classification algorithms. Results: Two variable selection methods, Boruta and genomewide association studies, were statistically similar in terms of the average accuracies generated, whereas elastic net had the worst overall performance. The predictive genetic panel was completed with 155 single-nucleotide variants, with 91.18% accuracy, 92.75% sensitivity, and 89.55% specificity using the support vector machine algorithm. The markers used range from known single-nucleotide variants to those not previously described in the literature. Our study shows potential in creating genetic prediction panels with tailored penalties per marker, aiding in the identification of optimal machine learning methods for intricate results. Conclusions: This method is able to classify asthma and nonasthma effectively, proving its potential utility in clinical 2024;3:100282.)
AbstractThe role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N = 200,003), TOPMed (N = 87,652) and All of Us (N = 45,445). We performed rare ( < 0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P < $$6\times {10}^{-10}$$ 6 × 10 − 10 after conditioning on previously reported variants, with effect sizes ranging from −7cm to +4.7 cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5 cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.
Genome-wide association studies (GWAS) have become well-powered to detect loci associated with telomere length. However, no prior work has validated genes nominated by GWAS to examine their role in telomere length regulation. We conducted a multi-ancestry meta-analysis of 211,369 individuals and identified five novel association signals. Enrichment analyses of chromatin state and cell-type heritability suggested that blood/immune cells are the most relevant cell type to examine telomere length association signals. We validated specific GWAS associations by overexpressing KBTBD6 or POP5 and demonstrated that both lengthened telomeres. CRISPR/Cas9 deletion of the predicted causal regions in K562 blood cells reduced expression of these genes, demonstrating that these loci are related to transcriptional regulation of KBTBD6 and POP5. Our results demonstrate the utility of telomere length GWAS in the identification of telomere length regulation mechanisms and validate KBTBD6 and POP5 as genes affecting telomere length regulation.
Introduction: Obesity is recognized as a chronic condition with a multifactorial etiology, marked by persistent systemic low-level inflammation. Objectives: This study aims to explore and evaluate the role of genetic factors in predisposition to obesity within a diverse, mixed-race population. Methods: We conducted a Genome-Wide Association Study (GWAS) involving 1,036 individuals, comprising 333 eutrophic and 703 people with overweight, to pinpoint genetic variants linked to obesity. Genotyping was carried out using the MEGA chip by Illumina. Following this, imputation was performed using the CAAPA reference panel. We conducted in silico analyses using different platforms. Additionally, in a subset of 657 participants, we quantified levels of 11 cytokines (Eotaxin, IFNγ, IL-10, IL-6, IL-12, IL-13, IL-17A, IL-1β, IL-5, IL-8, and TNFα) in peripheral blood. We examined their relationship with the genotypes of the variants identified in the GWAS study. Results: We identified thirty-five variants that exhibited suggestive associations (5 x 10-8 < p-value < 1 x 10-5) with weight excess. Chromosome 4 harbored the main genes linked to this outcome (SPON2, RNF212, COL4A3, TMED11P and PCSK2) expressed in adipose tissue. Furthermore, we found a variant within the ZZEF1 gene on chromosome 17. Notably, variants such as rs10014526-T and rs77703123-T, in TMED11P displayed high linkage disequilibrium (LD) with variants in SPON2, rs75448245-G, rs11538062-T and rs75654334-T and in COL4A3, rs13419630-A, all negatively associated with the outcome. In contrast, the rs781851-G variant, in the ZZEF1 gene showed a positive association with the outcome. These polymorphic alleles were associated with variations in serum levels of the cytokines IL-6, IL-10, and IL-12. Conclusions: Our study implies that candidate genes linked to weight excess, notably SPON2, are connected to a perturbed immune pathway that underlies the characteristic inflammation seen in obesity. These novel uncovered associations in our study could potentially advance the field of precision medicine to treat obesity.
Most transcriptome-wide association studies (TWASs) so far focus on European ancestry and lack diversity. To overcome this limitation, we aggregated genome-wide association study (GWAS) summary statistics, whole-genome sequences and expression quantitative trait locus (eQTL) data from diverse ancestries. We developed a new approach, TESLA (multi-ancestry integrative study using an optimal linear combination of association statistics), to integrate an eQTL dataset with a multi-ancestry GWAS. By exploiting shared phenotypic effects between ancestries and accommodating potential effect heterogeneities, TESLA improves power over other TWAS methods. When applied to tobacco use phenotypes, TESLA identified 273 new genes, up to 55% more compared with alternative TWAS methods. These hits and subsequent fine mapping using TESLA point to target genes with biological relevance. In silico drug-repurposing analyses highlight several drugs with known efficacy, including dextromethorphan and galantamine, and new drugs such as muscle relaxants that may be repurposed for treating nicotine addiction.
Mutations in a diverse set of driver genes increase the fitness of haematopoietic stem cells (HSCs), leading to clonal haematopoiesis 1 . These lesions are precursors for blood cancers 2 – 6 , but the basis of their fitness advantage remains largely unknown, partly owing to a paucity of large cohorts in which the clonal expansion rate has been assessed by longitudinal sampling. Here, to circumvent this limitation, we developed a method to infer the expansion rate from data from a single time point. We applied this method to 5,071 people with clonal haematopoiesis. A genome-wide association study revealed that a common inherited polymorphism in the TCL1A promoter was associated with a slower expansion rate in clonal haematopoiesis overall, but the effect varied by driver gene. Those carrying this protective allele exhibited markedly reduced growth rates or prevalence of clones with driver mutations in TET2 , ASXL1 , SF3B1 and SRSF2 , but this effect was not seen in clones with driver mutations in DNMT3A . TCL1A was not expressed in normal or DNMT3A -mutated HSCs, but the introduction of mutations in TET2 or ASXL1 led to the expression of TCL1A protein and the expansion of HSCs in vitro. The protective allele restricted TCL1A expression and expansion of mutant HSCs, as did experimental knockdown of TCL1A expression. Forced expression of TCL1A promoted the expansion of human HSCs in vitro and mouse HSCs in vivo. Our results indicate that the fitness advantage of several commonly mutated driver genes in clonal haematopoiesis may be mediated by TCL1A activation.
Anthropometric traits, measuring body size and shape, are highly heritable and significant clinical risk factors for cardiometabolic disorders. These traits have been extensively studied in genome-wide association studies (GWASs), with hundreds of genome-wide significant loci identified. We performed a whole-exome sequence analysis of the genetics of height, body mass index (BMI) and waist/hip ratio (WHR). We meta-analyzed single-variant and gene-based associations of whole-exome sequence variation with height, BMI, and WHR in up to 22,004 individuals, and we assessed replication of our findings in up to 16,418 individuals from 10 independent cohorts from Trans-Omics for Precision Medicine (TOPMed). We identified four trait associations with single-nucleotide variants (SNVs; two for height and two for BMI) and replicated the LECT2 gene association with height. Our expression quantitative trait locus (eQTL) analysis within previously reported GWAS loci implicated CEP63 and RFT1 as potential functional genes for known height loci. We further assessed enrichment of SNVs, which were monogenic or syndromic variants within loci associated with our three traits. This led to the significant enrichment results for height, whereas we observed no Bonferroni-corrected significance for all SNVs. With a sample size of ∼20,000 whole-exome sequences in our discovery dataset, our findings demonstrate the importance of genomic sequencing in genetic association studies, yet they also illustrate the challenges in identifying effects of rare genetic variants.
BACKGROUND:Atopic dermatitis (AD) is characterized by TH2-dominated skin inflammation and systemic response to cutaneously encountered antigens. The TH2 cytokines IL-4 and IL-13 play a critical role in the pathogenesis of AD. The Q576->R576 polymorphism in the IL-4 receptor alpha (IL-4Rα) chain common to IL-4 and IL-13 receptors alters IL-4 signaling and is associated with asthma severity.OBJECTIVE:We sought to investigate whether the IL-4Rα R576 polymorphism is associated with AD severity and exaggerates allergic skin inflammation in mice.METHODS:Nighttime itching interfering with sleep, Rajka-Langeland, and Eczema Area and Severity Index scores were used to assess AD severity. Allergic skin inflammation following epicutaneous sensitization of mice 1 or 2 IL-4Rα R576 alleles (QR and RR) and IL-4Rα Q576 (QQ) controls was assessed by flow cytometric analysis of cells and quantitative RT-PCR analysis of cytokines in skin.RESULTS:The frequency of nighttime itching in 190 asthmatic inner-city children with AD, as well as Rajka-Langeland and Eczema Area and Severity Index scores in 1116 White patients with AD enrolled in the Atopic Dermatitis Research Network, was higher in subjects with the IL-4Rα R576 polymorphism compared with those without, with statistical significance for the Rajka-Langeland score. Following epicutaneous sensitization of mice with ovalbumin or house dust mite, skin infiltration by CD4+ cells and eosinophils, cutaneous expression of Il4 and Il13, transepidermal water loss, antigen-specific IgE antibody levels, and IL-13 secretion by antigen-stimulated splenocytes were significantly higher in RR and QR mice compared with QQ controls. Bone marrow radiation chimeras demonstrated that both hematopoietic cells and stromal cells contribute to the mutants' exaggerated allergic skin inflammation.CONCLUSIONS:The IL-4Rα R576 polymorphism predisposes to more severe AD and increases allergic skin inflammation in mice.
Background DNA methylation of cytosines at cytosine-phosphate-guanine (CpG) dinucleotides (CpGs) is a widespread epigenetic mark, but genome-wide variation has been relatively unexplored due to the limited representation of variable CpGs on commercial high-throughput arrays. Objectives To explore this hidden portion of the epigenome, this study combined whole-genome bisulfite sequencing with in silico evidence of gene regulatory regions to design a custom array of high-value CpGs. This study focused on airway epithelial cells from children with and without allergic asthma because these cells mediate the effects of inhaled microbes, pollution, and allergens on asthma and allergic disease risk. Methods This study identified differentially methylated regions from whole-genome bisulfite sequencing in nasal epithelial cell DNA from a total of 39 children with and without allergic asthma of both European and African ancestries. This study selected CpGs from differentially methylated regions, previous allergy or asthma epigenome-wide association studies (EWAS), or genome-wide association study loci, and overlapped them with functional annotations for inclusion on a custom Asthma&Allergy array. This study used both the custom and EPIC arrays to perform EWAS of allergic sensitization (AS) in nasal epithelial cell DNA from children in the URECA (Urban Environment and Childhood Asthma) birth cohort and using the custom array in the INSPIRE [Infant Susceptibility to Pulmonary Infections and Asthma Following RSV Exposure] birth cohort. Each CpG on the arrays was assigned to its nearest gene and its promotor capture Hi-C interacting gene and performed expression quantitative trait methylation (eQTM) studies for both sets of genes. Results Custom array CpGs were enriched for intermediate methylation levels compared to EPIC CpGs. Intermediate methylation CpGs were further enriched among those associated with AS and for eQTMs on both arrays. Conclusions This study revealed signature features of high-value CpGs and evidence for epigenetic regulation of genes at AS EWAS loci that are robust to race/ethnicity, ascertainment, age, and geography.
Megabase-scale mosaic chromosomal alterations (mCAs) in blood are prognostic markers for a host of human diseases. Here, to gain a better understanding of mCA rates in genetically diverse populations, we analyzed whole-genome sequencing data from 67,390 individuals from the National Heart, Lung, and Blood Institute Trans-Omics for Precision Medicine program. We observed higher sensitivity with whole-genome sequencing data, compared with array-based data, in uncovering mCAs at low mutant cell fractions and found that individuals of European ancestry have the highest rates of autosomal mCAs and the lowest rates of chromosome X mCAs, compared with individuals of African or Hispanic ancestry. Although further studies in diverse populations will be needed to replicate our findings, we report three loci associated with loss of chromosome X, associations between autosomal mCAs and rare variants in DCPS, ADM17, PPP1R16B and TET2 and ancestry-specific variants in ATM and MPL with mCAs in cis.