Most genetic variants associated with complex traits are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterized 14,324 RNA-sequencing samples from the Trans-Omics for Precision Medicine program and performed expression and splicing quantitative trait locus (e/sQTL) analyses in six tissues and cell types, including whole blood (n = 6454) and lung (n = 1291). We detected tens of thousands of secondary cis-e/sQTLs, showing that secondary cis-e/sQTL discovery remains unsaturated. We fine-mapped UK Biobank-derived genome-wide association study (GWAS) signals from 164 traits and identified e/sQTL colocalizations for 10,611 GWAS signals, including 7096 that colocalize with secondary e/sQTLs. Our results suggest that even larger e/sQTL analyses will uncover additional secondary e/sQTLs, further benefiting GWAS interpretation.
Thyroid diseases are common and highly heritable. We performed a meta-analysis of genome-wide association studies from 19 biobanks for five thyroid diseases: thyroid cancer (ThC), benign nodular goiter, Graves’ disease, lymphocytic thyroiditis and primary hypothyroidism. We analyzed genetic association data from ~2.9 million genomes and identified 313 known and 570 new independent loci linked to thyroid diseases. We discovered genetic correlations between ThC, benign nodular goiter and autoimmune thyroid diseases ( rg = 0.16–0.97). Telomere maintenance genes contributed to benign and malignant thyroid nodular disease risk, whereas cell cycle, DNA repair and damage response genes were associated with ThC. We propose a paradigm that explains genetic predisposition to benign and malignant thyroid nodules. We found polygenic risk score associations with ThC risk of structural disease recurrence, tumor size, multifocality, lymph node metastases and extranodal extension. Polygenic risk scores identified individuals with aggressive ThC in a biobank, creating an opportunity for genetically informed population screening.
Obesity is a major public health crisis associated with high mortality rates. Previous genome-wide association studies (GWAS) investigating body mass index (BMI) have largely relied on imputed data from European individuals. This study leveraged whole-genome sequencing (WGS) data from 88,873 participants from the Trans-Omics for Precision Medicine (TOPMed) Program, of which 51% were of non-European population groups. We discovered 18 BMI-associated signals (P < 5 × 10-9). Notably, we identified and replicated a novel low frequency single nucleotide polymorphism (SNP) in MTMR3 that was common in individuals of African descent. Using a diverse study population, we further identified two novel secondary signals in known BMI loci and pinpointed two likely causal variants in the POC5 and DMD loci. Our work demonstrates the benefits of combining WGS and diverse cohorts in expanding current catalog of variants and genes confer risk for obesity, bringing us one step closer to personalized medicine.
BACKGROUND:Existing asthma polygenic risk scores (PRSs) have minimal validation in African-ancestry populations, leaving gaps in our understanding of the wide applicability of PRSs. To widen our understanding of the applicability of asthma PRSs, we apply published PRSs in African-ancestry individuals and quantify the extent to which the PRS-asthma relationship is mediated by clinical biomarkers and gene-expression signatures of asthma. METHODS:We applied 22 PRSs from the PGS Catalog in 673 individuals from the Consortium on Asthma among African-Ancestry Populations in the Americas (CAAPA) and calculated the percent of the PRS-asthma relationship that is statistically mediated by clinical and nasal epithelium transcriptomic biomarkers of asthma. Asthma case/control status was defined as ever/never having a doctor's diagnosis of disease. For gene expression mediation analysis, we limited the cases to those with current disease. RESULTS:The PRS (PGS001782) created by the Global Biobank Meta-analysis Initiative (N = 32,658 individuals of African ancestry) performed the best (ΔAUC = 0.104, AUC = 0.657) adjusted for age, sex, study site, and the first two genetic principal components (PC1-2). The PRS's effect on asthma was mediated by total IgE (tIgE) (38.8%, p.adj < 0.0002), multi-allergen ImmunoCAP phadiatop specific IgE (sIgE) (38.7%, p.adj < 0.0002), and eosinophils (7.3%, p.adj = 0.004). Mediation was observed for gene expression modules related to T2 inflammation (21.9%, p.adj < 0.0024), wound healing (11.9%, p.adj = 0.008), and medication response (6.8%, p.adj = 0.049). CONCLUSION:We found the best PRS to be the one derived using the largest sample size and including African-ancestry individuals. Mediation supports the well-documented biology of T2 inflammation in asthma as well as pathophysiological components of asthma like wound healing and medication response.
Background:Genetic control of gene expression in asthma-related tissues is not well-characterized, particularly for African-ancestry populations, limiting advancement in our understanding of the increased prevalence and severity of asthma in those populations. Objective:To create novel transcriptome prediction models for asthma tissues (nasal epithelium and CD4+ T cells) and apply them in transcriptome-wide association study (TWAS) to discover candidate asthma genes. Methods:We developed and validated gene expression prediction databases for unstimulated CD4+ T cells (CD4+T) and nasal epithelium using an elastic net framework. Combining these with existing prediction databases (N=51), we performed TWAS of 9,284 individuals of African-ancestry to identify tissue-specific and cross-tissue candidate genes for asthma. For detailed Methods, please see the Supplemental Methods. Results:Novel databases for CD4+T and nasal epithelial gene expression prediction contain 8,351 and 10,296 genes, respectively, including four asthma loci (SCGB1A1, MUC5AC, ZNF366, LTC4S) not predictable with existing public databases. Prediction performance was comparable to existing databases and was most accurate for populations sharing ancestry with the training set (e.g. African ancestry). From TWAS, we identified 17 candidate causal asthma genes (adjusted P<0.1), including genes with tissue-specific (IL33 in nasal epithelium) and cross-tissue (CCNC and FBXW7) effects. Conclusions:Expression of IL33, CCNC, and FBXW7 may affect asthma risk in African ancestry populations by mediating inflammatory responses. The addition of CD4+T and nasal epithelium prediction databases to the public sphere will improve ancestry representation and power to detect novel gene-trait associations from TWAS.
Thyroid diseases are common and highly heritable. Under the Global Biobank Meta-analysis Initiative, we performed a meta-analysis of genome-wide association studies from 19 biobanks for five thyroid diseases: thyroid cancer, benign nodular goiter, Graves' disease, lymphocytic thyroiditis, and primary hypothyroidism. We analyzed genetic association data from ~2.9 million genomes and identified 235 known and 501 novel independent variants significantly linked to thyroid diseases. We discovered genetic correlations between thyroid cancer, benign nodular goiter, and autoimmune thyroid diseases (r 2 =0.21-0.97). Telomere maintenance genes contribute to benign and malignant thyroid nodular disease risk, whereas cell cycle, DNA repair, and DNA damage response genes are predominantly associated with thyroid cancer. We proposed a paradigm explaining genetic predisposition to benign and malignant thyroid nodules. We evaluated thyroid cancer polygenic risk scores (PRS) for clinical applications in thyroid cancer diagnosis. We found PRS associations with thyroid cancer risk features: multifocality, lymph node metastases, and extranodal extension.
BACKGROUND:Genetic control of gene expression in asthma-related tissues is not well characterized, particularly for African-ancestry populations, limiting advancement in our understanding of the increased prevalence and severity of asthma in these populations. OBJECTIVE:We sought to create novel transcriptome prediction models for asthma tissues (nasal epithelium and CD4+ T cells) and apply them in a transcriptome-wide association study (TWAS) to discover candidate asthma genes. METHODS:We developed and validated gene expression prediction databases for unstimulated CD4+ T cells and nasal epithelium using an elastic net framework. Combining these with existing prediction databases (N = 51), we performed a TWAS of 9284 individuals of African ancestry to identify tissue-specific and cross-tissue candidate genes for asthma. RESULTS:Novel databases for CD4+ T cells and nasal epithelial gene expression prediction contain 8,351 and 10,296 genes, respectively, including 4 asthma loci (SCGB1A1, MUC5AC, ZNF366, and LTC4S) not predictable with existing public databases. Prediction performance was comparable to existing databases and was most accurate for populations sharing ancestry with the training set (eg, African ancestry). From the TWAS, we identified 17 candidate causal asthma genes (adjusted P < .1), including genes with tissue-specific (IL33 in nasal epithelium) and cross-tissue (CCNC and FBXW7) effects. CONCLUSIONS:Expression of IL33, CCNC, and FBXW7 may affect asthma risk in African- ancestry populations by mediating inflammatory responses. The addition of CD4+ T cell and nasal epithelium prediction databases to the public sphere will improve ancestry representation and power to detect novel gene-trait associations from TWAS.
Precision medicine initiatives across the globe have led to a revolution of repositories linking large-scale genomic data with electronic health records, enabling genomic analyses across the entire phenome. Many of these initiatives focus solely on research insights, leading to limited direct benefit to patients. We describe the biobank at the Colorado Center for Personalized Medicine (CCPM Biobank) that was jointly developed by the University of Colorado Anschutz Medical Campus and UCHealth to serve as a unique, dual-purpose research and clinical resource accelerating personalized medicine. This living resource currently has more than 200,000 participants with ongoing recruitment. We highlight the clinical, laboratory, regulatory, and HIPAA-compliant informatics infrastructure along with our stakeholder engagement, consent, recontact, and participant engagement strategies. We characterize aspects of genetic and geographic diversity unique to the Rocky Mountain region, the primary catchment area for CCPM Biobank participants. We leverage linked health and demographic information of the CCPM Biobank participant population to demonstrate the utility of the CCPM Biobank to replicate complex trait associations in the first 33,674 genotyped individuals across multiple disease domains. Finally, we describe our current efforts toward return of clinical genetic test results, including high-impact pathogenic variants and pharmacogenetic information, and our broader goals as the CCPM Biobank continues to grow. Bringing clinical and research interests together fosters unique clinical and translational questions that can be addressed from the large EHR-linked CCPM Biobank resource within a HIPAA- and CLIA-certified environment.
Abstract Asthma has striking disparities across ancestral groups, but the molecular underpinning of these differences is poorly understood and minimally studied. A goal of the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to understand multi-omic signatures of asthma focusing on populations of African ancestry. RNASeq and DNA methylation data are generated from nasal epithelium including cases (current asthma, N = 253) and controls (never-asthma, N = 283) from 7 different geographic sites to identify differentially expressed genes (DEGs) and gene networks. We identify 389 DEGs; the top DEG, FN1, was downregulated in cases (q = 3.26 × 10−9) and encodes fibronectin which plays a role in wound healing. The top three gene expression modules implicate networks related to immune response (CEACAM5; p = 9.62 × 10−16 and CPA3; p = 2.39 × 10−14) and wound healing (FN1; p = 7.63 × 10−9). Multi-omic analysis identifies FKBP5, a co-chaperone of glucocorticoid receptor signaling known to be involved in drug response in asthma, where the association between nasal epithelium gene expression is likely regulated by methylation and is associated with increased use of inhaled corticosteroids. This work reveals molecular dysregulation on three axes – increased Th2 inflammation, decreased capacity for wound healing, and impaired drug response – that may play a critical role in asthma within the African Diaspora.
Context Thyroid nodule ultrasound-based risk stratification schemas rely on the presence of high-risk sonographic features. However, some malignant thyroid nodules have benign appearance on thyroid ultrasound. New methods for thyroid nodule risk assessment are needed. Objective We investigated polygenic risk score (PRS) accounting for inherited thyroid cancer risk combined with ultrasound-based analysis for improved thyroid nodule risk assessment. Methods The convolutional neural network classifier was trained on thyroid ultrasound still images and cine clips from 621 thyroid nodules. Phenome-wide association study (PheWAS) and PRS PheWAS were used to optimize PRS for distinguishing benign and malignant nodules. PRS was evaluated in 73 346 participants in the Colorado Center for Personalized Medicine Biobank. Results When the deep learning model output was combined with thyroid cancer PRS and genetic ancestry estimates, the area under the receiver operating characteristic curve (AUROC) of the benign vs malignant thyroid nodule classifier increased from 0.83 to 0.89 (DeLong, P value = .007). The combined deep learning and genetic classifier achieved a clinically relevant sensitivity of 0.95, 95% CI [0.88-0.99], specificity of 0.63 [0.55-0.70], and positive and negative predictive values of 0.47 [0.41-0.58] and 0.97 [0.92-0.99], respectively. AUROC improvement was consistent in European ancestry-stratified analysis (0.83 and 0.87 for deep learning and deep learning combined with PRS classifiers, respectively). Elevated PRS was associated with a greater risk of thyroid cancer structural disease recurrence (ordinal logistic regression, P value = .002). Conclusion Augmenting ultrasound-based risk assessment with PRS improves diagnostic accuracy.
Most transcriptome-wide association studies (TWASs) so far focus on European ancestry and lack diversity. To overcome this limitation, we aggregated genome-wide association study (GWAS) summary statistics, whole-genome sequences and expression quantitative trait locus (eQTL) data from diverse ancestries. We developed a new approach, TESLA (multi-ancestry integrative study using an optimal linear combination of association statistics), to integrate an eQTL dataset with a multi-ancestry GWAS. By exploiting shared phenotypic effects between ancestries and accommodating potential effect heterogeneities, TESLA improves power over other TWAS methods. When applied to tobacco use phenotypes, TESLA identified 273 new genes, up to 55% more compared with alternative TWAS methods. These hits and subsequent fine mapping using TESLA point to target genes with biological relevance. In silico drug-repurposing analyses highlight several drugs with known efficacy, including dextromethorphan and galantamine, and new drugs such as muscle relaxants that may be repurposed for treating nicotine addiction.
Anthropometric traits, measuring body size and shape, are highly heritable and significant clinical risk factors for cardiometabolic disorders. These traits have been extensively studied in genome-wide association studies (GWASs), with hundreds of genome-wide significant loci identified. We performed a whole-exome sequence analysis of the genetics of height, body mass index (BMI) and waist/hip ratio (WHR). We meta-analyzed single-variant and gene-based associations of whole-exome sequence variation with height, BMI, and WHR in up to 22,004 individuals, and we assessed replication of our findings in up to 16,418 individuals from 10 independent cohorts from Trans-Omics for Precision Medicine (TOPMed). We identified four trait associations with single-nucleotide variants (SNVs; two for height and two for BMI) and replicated the LECT2 gene association with height. Our expression quantitative trait locus (eQTL) analysis within previously reported GWAS loci implicated CEP63 and RFT1 as potential functional genes for known height loci. We further assessed enrichment of SNVs, which were monogenic or syndromic variants within loci associated with our three traits. This led to the significant enrichment results for height, whereas we observed no Bonferroni-corrected significance for all SNVs. With a sample size of ∼20,000 whole-exome sequences in our discovery dataset, our findings demonstrate the importance of genomic sequencing in genetic association studies, yet they also illustrate the challenges in identifying effects of rare genetic variants.
Abstract Disclosure: N. Pozdeyev: None. M. Dighe: None. M. Barrio: None. C. Raeburn: None. H.A. Smith: None. M. Fisher: None. S. Chavan: None. N. Rafaels: None. J. Shortt: None. M. Leu: None. T. Clark: None. C. Marshall: None. B.R. Haugen: None. D. Subramanian: None. K. Crooks: None. C. Gignoux: None. T.A. Cohen: None. Purpose. Evaluating thyroid nodules to rule out malignancy is a common clinical task. Image-based risk stratification schemas rely on the presence of high-risk thyroid nodule sonographic features and, therefore, are less suitable for the diagnosis of malignant thyroid nodules that have a benign appearance on the ultrasound. To mitigate the deficiency of thyroid nodule evaluation relying solely on the sonographic characteristics, we used thyroid cancer polygenic risk score (PRS) to complement deep learning analysis of ultrasound images. Methods. A supervised deep learning classifier of thyroid nodules was trained on 32,545 thyroid US images from 621 nodules and tested on an independent set of 232 nodules from patients genotyped on the Illumina's MEGAEX platform. The deep-learning thyroid nodule classifier was developed by fine-tuning a BiT-M ResNet-50x1 convolutional neural network (CNN) pre-trained on the ImageNet-21k dataset. A polygenic risk score (PRS) was calculated using thyroid cancer genome-wide association meta-analysis summary statistics from the Global Biobank Meta-analysis Initiative. The thyroid cancer PRS was defined as a weighted sum of five alleles with the strongest association with thyroid cancer. CNN predictions and PRS were combined into a meta-classifier using logistic regression with or without genetic ancestry and demographic covariates. Results. The CNN classifier achieved an area under the receiver operating characteristic curve (AUC) of 0.83 on the out-of-sample test set of 232 thyroid nodules. The CNN classifier incorrectly classified thyroid nodules without suspicious sonographic characteristics belonging to difficult-to-diagnose subtypes such as follicular thyroid cancer and follicular variant of papillary thyroid cancer. Combining predictions from the CNN classifier with PRS into a cross-validated classifier improved AUC to 0.868 (DeLong test, p = 0.05). Incorporating genetic ancestry in the form of five genetic principal components further improved AUC of the benign vs. malignant CNN + PRS thyroid nodule classifier to 0.885 (p = 0.007). Finally, when age, sex, and nodule dimensions were considered, the AUC of the meta-classifier increased to 0.915 (p = 2.3e-4). The meta-classifier including predictions from CNN, PRS, and genetic principal components showed a sensitivity of 0.95, specificity of 0.61, NPV of 0.97, and PPV of 0.5. This performance was superior to that of the clinical Thyroid Imaging Reporting and Data System (TI-RADS) as reported by radiologists in ultrasound reports. Conclusions. For the first time, we showed PRS provides an orthogonal thyroid cancer risk assessment complementary to the ultrasound image-based thyroid nodule risk evaluation. This proof-of-concept study opens an opportunity for developing next-generation schemas for thyroid nodule evaluation incorporating clinical, imaging, and genetic data. Presentation: Sunday, June 18, 2023
BACKGROUND:While numerous genetic loci associated with atopic dermatitis (AD) have been discovered, to date, work leveraging the combined burden of AD risk variants across the genome to predict disease risk has been limited. OBJECTIVES:This study aims to determine whether polygenic risk scores (PRSs) relying on genetic determinants for AD provide useful predictions for disease occurrence and severity. It also explicitly tests the value of including genome-wide association studies of related allergic phenotypes and known FLG loss-of-function (LOF) variants. METHODS:AD PRSs were constructed for 1619 European American individuals from the Atopic Dermatitis Research Network using an AD training dataset and an atopic training dataset including AD, childhood onset asthma, and general allergy. Additionally, whole genome sequencing data were used to explore genetic scoring specific to FLG LOF mutations. RESULTS:Genetic scores derived from the AD-only genome-wide association studies were predictive of AD cases (PRSAD: odds ratio [OR], 1.70; 95% CI, 1.49-1.93). Accuracy was first improved when PRSs were built off the larger atopy genome-wide association studies (PRSAD+: OR, 2.16; 95% CI, 1.89-2.47) and further improved when including FLG LOF mutations (PRSAD++: OR, 3.23; 95% CI, 2.57-4.07). Importantly, while all 3 PRSs correlated with AD severity, the best prediction was from PRSAD++, which distinguished individuals with severe AD from control subjects with OR of 3.86 (95% CI, 2.77-5.36). CONCLUSIONS:This study demonstrates how PRSs for AD that include genetic determinants across atopic phenotypes and FLG LOF variants may be a promising tool for identifying individuals at high risk for developing disease and specifically severe disease.
Human genetic studies support an inverse causal relationship between leukocyte telomere length (LTL) and coronary artery disease (CAD), but directionally mixed effects for LTL and diverse malignancies. Clonal hematopoiesis of indeterminate potential (CHIP), characterized by expansion of hematopoietic cells bearing leukemogenic mutations, predisposes both hematologic malignancy and CAD. TERT (which encodes telomerase reverse transcriptase) is the most significantly associated germline locus for CHIP in genome-wide association studies. Here, we investigated the relationship between CHIP, LTL, and CAD in the Trans-Omics for Precision Medicine (TOPMed) program (n = 63,302) and UK Biobank (n = 47,080). Bidirectional Mendelian randomization studies were consistent with longer genetically imputed LTL increasing propensity to develop CHIP, but CHIP then, in turn, hastens to shorten measured LTL (mLTL). We also demonstrated evidence of modest mediation between CHIP and CAD by mLTL. Our data promote an understanding of potential causal relationships across CHIP and LTL toward prevention of CAD.
Common genetic variants explain less variation in complex phenotypes than inferred from family-based studies, and there is a debate on the source of this ‘missing heritability’. We investigated the contribution of rare genetic variants to tobacco use with whole-genome sequences from up to 26,257 unrelated individuals of European ancestries and 11,743 individuals of African ancestries. Across four smoking traits, single-nucleotide-polymorphism-based heritability ( $$h^2_{\mathrm{SNP}}$$ ) was estimated from 0.13 to 0.28 (s.e., 0.10–0.13) in European ancestries, with 35–74% of it attributable to rare variants with minor allele frequencies between 0.01% and 1%. These heritability estimates are 1.5–4 times higher than past estimates based on common variants alone and accounted for 60% to 100% of our pedigree-based estimates of narrow-sense heritability ( $$h^2_{\mathrm{ped}}$$ , 0.18–0.34). In the African ancestry samples, $$h^2_{\mathrm{SNP}}$$ was estimated from 0.03 to 0.33 (s.e., 0.09–0.14) across the four smoking traits. These results suggest that rare variants are important contributors to the heritability of smoking. The team of authors led by Seon-Kyeong Jang use whole-genome sequencing data and show that rare genetic variants explain much of the ‘missing heritability’ in smoking behaviours. These results help address a long-standing mystery in behavioural genetics.