Most genetic variants associated with complex traits are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterized 14,324 RNA-sequencing samples from the Trans-Omics for Precision Medicine program and performed expression and splicing quantitative trait locus (e/sQTL) analyses in six tissues and cell types, including whole blood (n = 6454) and lung (n = 1291). We detected tens of thousands of secondary cis-e/sQTLs, showing that secondary cis-e/sQTL discovery remains unsaturated. We fine-mapped UK Biobank-derived genome-wide association study (GWAS) signals from 164 traits and identified e/sQTL colocalizations for 10,611 GWAS signals, including 7096 that colocalize with secondary e/sQTLs. Our results suggest that even larger e/sQTL analyses will uncover additional secondary e/sQTLs, further benefiting GWAS interpretation.
Whole genome sequencing (WGS) studies play a pivotal role in studying the genetic underpinnings of human diseases and traits. High quality and reproducible variant calling is the cornerstone for the success of downstream analyses, including WGS association studies and polygenic risk prediction. This paper compares the data quality, performance, and concordance of two widely used WGS variant callers, the Genome Analysis Toolkit (GATK) and Variant Tool set that discovers short variants (VT), using 60 532 multi-ancestry whole genomes sequenced by the Centers for Common Disease Genomics (CCDGs) of the NHGRI Genome Sequencing Program. Our findings show that both QCed GATK and VT pipelines yield highly consistent and reliable called Single Nucleotide Variants (SNVs) in large-scale WGS studies, supporting their agreements in joint variants calling. However, the two pipelines exhibit greater discrepancies in calling insertions and deletions (INDELs).
STUDY OBJECTIVES:Excessive daytime sleepiness (EDS), influenced by environmental and social-behavioral factors, is reported by a subset of patients with sleep apnea-a group that may be at elevated cardiovascular risk. However, it is unclear whether sleep apnea with and without EDS have distinct genetic underpinnings. In this study, we perform gene-by-EDS interaction analyses for apnea hypopnea index, a diagnostic marker of sleep apnea severity, to understand EDS's influence on its underlying genetic risk. METHODS:Discovery interaction analyses for common variants and gene-based rare variants were conducted respectively using multi-ethnic Trans-Omics for Precision Medicine (N = 11 619) data, followed by replication and subsequent meta-analysis in additional Trans-Omics for Precision Medicine-imputed data (N = 8904). The 1 degree-of-freedom (1df) G × E test and the 2df joint G,G × E tests were utilized. Sex-stratified analyses were additionally performed. RESULTS:Discovery analysis revealed two common intronic variants-rs13118183 (CCDC3) and rs281851 (MARCHF1)-and three rare variant gene sets mapped to SCUBE2, TMEM26, and CPS4FL-to exhibit interaction with EDS. Meta-analysis revealed EDS interaction with 11 rare variant gene sets mapped to UBLCP1, MED31, RAP1GAP, CPNE5, MYMX, YY1, ZNF773, YBEY, IQCB1, PI4K2B, and CORO1A. CONCLUSION:Genetic loci reveal connections to cardiovascular risk, insulin resistance, thiamine deficiency, and resveratrol mechanism. Discovered genetic signals may offer insight into pertinent biological pathways for sleep apnea patients with an excessively sleepy subtype. Statement of Significance Sleep apnea is a complex sleep disorder. Exemplifying this is the disparately varying estimates of presence of excessive daytime sleepiness (EDS) in patients, and persistent EDS that lingers despite treatment. Some data indicate that the excessively sleepy subtype of sleep apnea carries heightened cardiovascular risk. Whether EDS influences genetic risk factors underlying sleep apnea has not yet been investigated. This study addresses this gap, as the first genome-wide gene × EDS interaction study for apnea hypopnea index, the standard sleep apnea severity metric. Genetic loci that have been previously unconsidered for sleep apnea are revealed. Discovered interaction signals highlight pathways in metabolism, genes associated with cardiometabolic traits, and therapeutic agents influencing obesity, blood pressure, oxidative stress, and apnea hypopnea index.
Multiple germline and somatic genomic factors are associated with risk of coronary artery disease, but there is no single measure of risk that integrates all information from a DNA sample. To address this gap, we develop an integrated genomic model that includes six germline and somatic genetic drivers for coronary artery disease, including polygenic risk score, genetically-proxied proteomic/metabolomic risk scores, and clonal hematopoiesis of indeterminate potential. We evaluated its predictive power in the UK Biobank (N = 391,536), and validate it using data from the TOPMed program (N = 34,177). The 10-year coronary artery disease risk based on the integrated genomic model profile ranges from 1.1% to 15.5% in the UK Biobank and from 3.8% to 33.0% in TOPMed, with a more pronounced gradient in males than females. The integrated genomic model captures the cumulative effect of multiple genetic drivers, identifying individuals at high risk for coronary artery disease despite lacking any single high-risk genetic factor, as well as individuals at low risk despite carrying known high-risk factors. In middle age, the integrated genomic model augments the performance of the Pooled Cohort Equations, a clinical risk calculator for coronary artery disease. While the integrated genomic model yields only modest incremental predictive value over polygenic risk score at the population level, it identifies approximately 13% of high-risk individuals not detected by polygenic risk score alone.
The mapping of protein quantitative trait loci (pQTLs) can provide molecular links between genotype and phenotype. Most such studies focus on common variants, but the effects of low-frequency (LF) variants remain underexplored. Focusing on cis-pQTLs, we integrated serum measurements of 7,596 proteins with genomic data, including LF variants (minor allele frequency [MAF] 0.1-1%), in 5,291 Icelanders to identify independent cis-pQTLs for 2,166 SOMAmers. Incorporating LF variants increased the number of detected genetic signals per protein, demonstrating widespread allelic heterogeneity in cis-acting regulation of serum proteins. LF pQTLs were enriched for coding variants in the respective protein-encoding gene, but also among distal secondary signals, revealing additional regulatory layers not captured by common variants alone. Proteins affected by common variant cis-pQTLs were more often secreted and exhibited tissue-specific expression, whereas proteins exclusively affected by LF variants were primarily from more constrained and biologically essential pathways. Expanding both protein coverage and the allele-frequency spectrum reveals a more complex and heterogeneous cis-regulatory architecture of circulating proteins.
ABSTRACT Background Chronic obstructive pulmonary disease (COPD) is associated with musculoskeletal comorbidities, including cachexia. Weight loss (WL) is the major criterion for cachexia and increases risk for mortality in COPD. Risk factors for WL in COPD are incompletely understood. We performed this whole genome sequencing (WGS) analysis to identify genetic risk variants for WL in COPD. Methods We studied 16 972 participants from the Trans‐Omics for Precision Medicine (TOPMed) Initiative and All of Us Research Program. COPD was diagnosed using spirometry in TOPMed, while diagnosis codes were used in All of Us. WL was defined as WL ≥ 5% or a final body mass index (BMI) < 20 kg/m2. WGS data came from white blood cells in all cohorts. Single‐variant testing was conducted on both race‐ and study‐stratified cohorts and in a cosmopolitan, ancestry‐independent manner using GENESIS in TOPMed and SAIGE in All of Us. SAIGE‐GENE+ gene‐based analyses were performed on race‐stratified and cosmopolitan cohorts. Single variant meta‐analyses were conducted using METAL within (B/AA and NHW analyses of Black/African–American and non‐Hispanic white participants, respectively) and across racial groups (COSMO). Rare variant gene‐based results were combined using Fisher's method. Transcriptomic effects were predicted using MetaXcan. We used the GWAS Catalogue to analyse for colocalization with other related traits. Results Two single variants were associated with WL in COPD among All of Us participants: one intronic variant in HCN1 in Black/African–American participants (chr5:45271359:TACACAC:T, odds ratio with 95% confidence interval (OR (CI95)) = 2.43(1.78–3.31), p = 1.95 × 10−8) and one intergenic variant between PPP4R2 and PDZRN3 in the cosmopolitan and NHW cohorts (chr3:73345901:A:G, OR (CI95) = 0.21(0.12–0.35), p = 8.84 × 10−9 in cosmopolitan and OR (CI95) = 0.18(0.10–0.33), p = 9.44 × 10−9 in NHW). Single‐variant meta‐analysis identified two loci associated with WL in COPD: five variants within DRAIC in the B/AA meta‐analysis (lead variant chr15:69571341:A:G, OR (CI95) = 1.37(1.23–1.51), p = 1.29 × 10−9) and two intronic variants within RFX3 in the cosmopolitan meta‐analysis (lead variant chr9:3390983:T:C, OR (CI95) = 1.50(1.31–1.73), p = 1.06 × 10−8). Rare variants within RNU6‐565P (NHW analysis in All of Us; p = 2.83 × 10−7) and LOC339298 (B/AA analysis in All of Us, p = 1.85 × 10−6) were associated with WL in COPD. The RNU‐565P signal remained significant after combination with TOPMed results (pcombined = 1.23 × 10−6). MetaXcan predicted differential expression of LNC00959 in visceral adipose tissue (p = 1.16 × 10−6 in COSMO analysis). Colocalization analyses identified genomic associations between BMI and variants in or near DRAIC, RFX3, PDZRN3, and LINC00959. Conclusions In the first WGS analysis of WL in COPD, we have identified seven novel loci. Further characterization of these loci will validate our findings and improve our understanding of the molecular pathophysiology of this condition.
Rationale:Pulmonary artery (PA) enlargement is a non-invasive imaging biomarker associated with pulmonary hypertension and mortality in COPD; however, its genetic determinants remain incompletely understood. Objectives:To characterize the genetic architecture of PA size across COPD-enriched and population-based cohorts. Methods:We performed genome-wide association analyses of PA diameter using whole-genome sequencing in COPDGene (n=9,418) and ECLIPSE (n=1,859), and imputed-genotype data from the UK Biobank (n=37,073). We replicated lead variants in the Framingham Heart Study (FHS; n=3,289), incorporated all four studies into a joint meta-analysis, and identified independent signals through conditional analyses. Candidate effector genes were prioritized using coding variant annotation, colocalization, and integrative regulatory evidence. Measurements and Main Results:We identified 44 independent genome-wide significant PA diameter signals within 39 loci, including 8 variants replicated in FHS, novel associations near FRMD4B , SLC20A2 , BORCS7-ASMT , and KCNRG , and 5 signals in conditional analysis including multiple signals at ANO1 . Genetic effects were concordant across imaging modalities and cohorts of differing COPD burden. Effector-gene prioritization nominated ABCC8 , PDGFD , HMCN1 , CCNE1 , and TBX20 , implicating pathways in vascular remodeling, developmental regulation, smooth muscle and endothelial function, ion-channel signaling, and extracellular matrix organization. Colocalization with pulse pressure GWAS demonstrated substantial shared causal variation between pulmonary and systemic vascular biology. Conclusions:In this largest genetic study of pulmonary vascular imaging to date, PA diameter exhibits a polygenic architecture consistent across imaging modalities and cohorts of differing COPD burden. The prioritized effector genes bridge rare-variant pulmonary hypertension biology with common-variant systemic vascular biology.
Supplementary Table 2 shows additional results for non smokers and never smokers in the MESA and WHI cohorts.
Most genetic variants associated with complex traits and diseases occur in non-coding genomic regions and are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterize 14,324 ancestrally diverse RNA-sequencing samples from the NHLBI Trans-Omics for Precision Medicine (TOPMed) program and integrate whole genome sequencing data to perform cis and trans expression and splicing quantitative trait locus (cis-/trans-e/sQTL) analyses in six tissues and cell types, most notably whole blood (N=6,454) and lung (N=1,291). We show this dataset enables greater detection of secondary cis-e/sQTL signals than was achieved in previous studies, and that secondary cis-eQTL and primary trans-eQTL signal discovery is not saturated even though eGene discovery is. Most TOPMed trans-eQTL signals colocalize with cis-e/sQTL signals, suggesting many trans signals are mediated by cis signals. We fine-map European UK BioBank GWAS signals from 164 traits and colocalize the resulting 34,107 fine-mapped GWAS signals with TOPMed e/sQTL signals, finding that of 10,611 GWAS signals with a colocalization, 7,096 GWAS signals colocalize with at least one secondary e/sQTL signal. These results demonstrate that larger e/sQTL analyses will continue to uncover secondary e/sQTL signals, and that these new signals will benefit GWAS interpretation.
Polygenic scores (PGSs) for body mass index (BMI) may guide early prevention and targeted treatment of obesity. Using genetic data from up to 5.1 million people (4.6% African ancestry, 14.4% American ancestry, 8.4% East Asian ancestry, 71.1% European ancestry and 1.5% South Asian ancestry) from the GIANT consortium and 23andMe, Inc., we developed ancestry-specific and multi-ancestry PGSs. The multi-ancestry score explained 17.6% of BMI variation among UK Biobank participants of European ancestry. For other populations, this ranged from 16% in East Asian-Americans to 2.2% in rural Ugandans. In the ALSPAC study, children with higher PGSs showed accelerated BMI gain from age 2.5 years to adolescence, with earlier adiposity rebound. Adding the PGS to predictors available at birth nearly doubled explained variance for BMI from age 5 onward (for example, from 11% to 21% at age 8). Up to age 5, adding the PGS to early-life BMI improved prediction of BMI at age 18 (for example, from 22% to 35% at age 5). Higher PGSs were associated with greater adult weight gain. In intensive lifestyle intervention trials, individuals with higher PGSs lost modestly more weight in the first year (0.55 kg per s.d.) but were more likely to regain it. Overall, these data show that PGSs have the potential to improve obesity prediction, particularly when implemented early in life.
Circulating metabolite levels partly reflect the state of human health and diseases and can be impacted by genetic determinants. Hundreds of loci associated with circulating metabolites have been identified; however, most findings focus on predominantly European ancestry or single-study analyses. Leveraging the rich metabolomics resources generated by the National Heart, Lung, and Blood Institute (NHLBI) Trans-Omics for Precision Medicine (TOPMed) Program, we harmonized and accessibly cataloged 1,729 circulating metabolites among 25,058 ancestrally diverse samples. From our comparison of multiple methods, we provided a set of reasonable strategies for outlier and imputation handling to process metabolite data and show that inverse normalization by study and half-minimum imputation provide mostly similar results for pooled or meta-analysis. Following the practical analysis framework, we further performed a genome-wide association analysis on 1,135 selected metabolites using whole-genome sequencing data from 16,359 individuals passing the quality-control filters and discovered 1,775 independent loci associated with 667 metabolites. Among 160 unreported locus-metabolite pairs, we identified associations with loci locating within previously implicated metabolite-associated genes, as well as associations with loci locating in genes such as GAB3 and VSIG4 (located on the X chromosome) that may play a role in metabolic regulation. In the sex-stratified analysis, we revealed 85 independent locus-metabolite pairs with evidence of sexual dimorphism, which were located in well-known metabolic genes such as FADS2, D2HGDH, SUGP1, and UGT2B17, strongly supporting the importance of exploring sex difference in the human metabolome. Taken together, our study depicted the genetic contribution to circulating metabolite levels, providing additional insight into the understanding of human health.
To better characterize the potential biological mechanisms underlying insulin resistance (IR) and dementia, we derive cross-population and population specific polygenic scores [PSs] for fasting insulin and IR-related partitioned PSs [pPSs]. We conduct a cross-sectional study of the associations of these genetic scores with neurological outcomes in >17k participants (36% men, mean age 55 yrs) from the Trans-Omics for Precision Medicine (TOPMed) program (50% Non-Hispanic White, 23% Black/African American, 21% Hispanic/Latino American, and 4% Asian American). We report significant negative associations (P < 0.002) of the cross-population (P = 1.3 × 10-5) and European (PEA = 3.0 × 10-8) fasting insulin PSs with total cranial volume, and of a metabolic syndrome European PS with general cognitive function (BEA = -0.13, PEA = 0.0002) and lateral ventricular volume (BEA = 0.09, PEA = 0.002). We identify suggestive negative associations (P < 0.007) of metabolic syndrome and obesity pPSs with general cognitive function, and of lipodystrophy pPSs with total cranial volume. A higher genetic predisposition to IR is associated with lower brain size, and a genetic predisposition to specific IR-related type 2 diabetes subtypes, such as metabolic syndrome and mechanisms of IR mediated through obesity and lipodystrophy, is potentially involved in cognitive decline.
Clonal hematopoiesis (CH) is defined by the expansion of a lineage of genetically identical cells in blood. Genetic lesions that confer a fitness advantage, such as leukemogenic point mutations or mosaic chromosomal alterations (mCAs), are frequent mediators of CH. However, recent analyses of both single cell-derived colonies of hematopoietic cells and population sequencing cohorts have revealed CH frequently occurs in the absence of known driver genetic lesions. To characterize CH without known driver genetic lesions, we use 51,399 deeply sequenced whole genomes from the NHLBI TOPMed sequencing initiative to perform simultaneous germline and somatic mutation analyses among individuals without leukemogenic point mutations (LPM), which we term CH-LPMneg. We quantify CH by estimating the total mutation burden. Because estimating somatic mutation burden without a paired-tissue sample is challenging, we develop a novel statistical method, the Genomic and Epigenomic informed Mutation (GEM) rate, that uses external genomic and epigenomic data sources to distinguish artifactual signals from true somatic mutations. We perform a genome-wide association study of GEM to discover the germline determinants of CH-LPMneg. We identify seven genes associated with CH-LPMneg (TCL1A, TERT, SMC4, NRIP1, PRDM16, MSRA, SCARB1).Functional analyses of SMC4 and NRIP1 implicated altered hematopoietic stem cell self-renewal and proliferation as the primary mediator of mutation burden in blood. We then perform comprehensive multi-tissue transcriptomic analyses, finding that the expression levels of 404 genes are associated with GEM. Finally, we perform phenotypic association meta-analyses across four cohorts, finding that GEM is associated with increased white blood cell count, but is not significantly associated with incident stroke or coronary disease events. Overall, we develop GEM for quantifying mutation burden from WGS and use GEM to discover the genetic, genomic, and phenotypic correlates of CH-LPMneg.
Integrating multi-omics data may help researchers understand the genetic underpinnings of complex traits and diseases. However, the best ways to integrate multi-omics data and use them to address pressing scientific questions remain a challenge. One important and topical problem is how to assess the aggregate effect of multiple genomic data types (e.g. genotypes and gene expression levels) on a phenotype, particularly while accommodating routine issues, such as having related subjects' data in analyses. In this paper, we extend an existing composite kernel machine regression model to integrate two multi-omics data types, while accommodating for general correlation structures amongst outcomes. Due to the kernel machine regression framework, our methods allow for the integration of high-dimensional omics data with small, nonlinear, and interactive effects, and accommodation of general study designs. Here, we focus on scientific questions that aim to assess the association between a functional grouping (such as a gene or a pathway) and a quantitative trait of interest. We use a kernel machine regression to integrate the two multi-omics data types, as they may relate to the trait, and perform a global test of association. We demonstrate the advantage of this approach over single data type association tests via simulation. Finally, we apply this method to a large, multi-ethnic data set to investigate how predicted gene expression and rare genetic variation may be related to two platelet traits.
Serum lipid levels, which are influenced by both genetic and environmental factors, are key determinants of cardiometabolic health and are influenced by both genetic and environmental factors. Improving our understanding of their underlying biological mechanisms can have important public health and therapeutic implications. Although psychosocial factors, including depression, anxiety, and perceived social support, are associated with serum lipid levels, it is unknown if they modify the effect of genetic loci that influence lipids. We conducted a genome-wide gene-by-psychosocial factor interaction (G×Psy) study in up to 133,157 individuals to evaluate if G×Psy influences serum lipid levels. We conducted a two-stage meta-analysis of G×Psy using both a one-degree of freedom (1df) interaction test and a joint 2df test of the main and interaction effects. In Stage 1, we performed G×Psy analyses on up to 77,413 individuals and promising associations (P < 10−5) were evaluated in up to 55,744 independent samples in Stage 2. Significant findings (P < 5 × 10−8) were identified based on meta-analyses of the two stages. There were 10,230 variants from 120 loci significantly associated with serum lipids. We identified novel associations for variants in four loci using the 1df test of interaction, and five additional loci using the 2df joint test that were independent of known lipid loci. Of these 9 loci, 7 could not have been detected without modeling the interaction as there was no evidence of association in a standard GWAS model. The genetic diversity of included samples was key in identifying these novel loci: four of the lead variants displayed very low frequency in European ancestry populations. Functional annotation highlighted promising loci for further experimental follow-up, particularly rs73597733 (MACROD2), rs59808825 (GRAMD1B), and rs11702544 (RRP1B). Notably, one of the genes in identified loci (RRP1B) was found to be a target of the approved drug Atenolol suggesting potential for drug repurposing. Overall, our findings suggest that taking interaction between genetic variants and psychosocial factors into account and including genetically diverse populations can lead to novel discoveries for serum lipids.
Mosaic loss of Y (mLOY) is the most common somatic chromosomal alteration detected in human blood. The presence of mLOY is associated with altered blood cell counts and increased risk of Alzheimer disease, solid tumors, and other age-related diseases. We sought to gain a better understanding of genetic drivers and associated phenotypes of mLOY through analyses of whole-genome sequencing (WGS) of a large set of genetically diverse males from the Trans-Omics for Precision Medicine (TOPMed) program. We show that haplotype-based calling methods can be used with WGS data to successfully identify mLOY events. This approach enabled us to identify differences in mLOY frequencies across populations defined by genetic similarity, revealing a higher frequency of mLOY in the European (EUR) ancestry group compared to other ancestries. We identify multiple loci associated with mLOY susceptibility and show that subsets of human hematopoietic stem cells are enriched for the activity of mLOY susceptibility variants. Finally, we found that certain alleles on chromosome Y are more likely to be lost than others in detectable mLOY clones.
Brain structural volumes are highly heritable and are linked to multiple neuropsychological outcomes, including Alzheimer's disease (AD). Genome-wide association studies have successfully identified genetic variants associated with intracranial volume (ICV), total brain volume (TBV), hippocampal volume (HV), and lateral ventricular volume (LVV). However, these studies mostly focused on common genetic variants with minor allele frequencies (MAF) > 1%, and individuals included in most of these studies were of predominantly European ancestry. Here, we performed whole-genome sequence (WGS) association studies of MRI brain volumes in 7,674 individuals of diverse race and ethnicity from the Trans-Omics for Precision Medicine (TOPMed) program. We identified novel genetic loci on chromosomes 13 and 16 near LINC00598 and CACNG3 associated with HV and TBV, respectively (lead variants rs115674829, P-value = 1.7×10-9 in pooled analysis and rs150440001, P-value = 6.6×10-9 in black participants). Both lead variant minor A alleles are rarer in white participants (MAF = 0.14% and 0.03%) and in Hispanic participants (MAF = 1.5% and 0.17%) but more common in black participants (MAF = 13% and 1.5%). Rare variant aggregated analyses identified RIPK1, a gene encoding a kinase involved in neuroinflammation and promising target for AD treatment, suggestively associated with LVV (P-value=5×10-6). This study provides new insights into the genetic correlates of brain structural volumes and illustrates the importance of leveraging WGS data and cohorts of diverse race and ethnicity to better characterize the genetic architecture of complex polygenic traits.
Obstructive sleep apnea (OSA) is a multifactorial sleep disorder characterized by a strong genetic basis. Excessive daytime sleepiness (EDS) is a symptom that is reported by a subset of OSA patients, persisting even after treatment with continuous positive airway pressure (CPAP). It is recognized as a clinical subtype underlying OSA carrying alarming heightened cardiovascular risk. Thus, conceptualizing EDS as an exposure variable, we sought to investigate EDS’s influence on genetic variation linked to apnea-hypopnea index (AHI), a diagnostic measure of OSA severity. This study serves as the first large-scale genome-wide gene x environment interaction analysis for AHI, investigating the interplay between its genetic markers and EDS across and within specific sex. Our work pools together whole genome sequencing data from seven cohorts, enabling a diverse dataset (four population backgrounds) of over 11,500 samples. Among the total 16 discovered genetic targets with interaction evidence with EDS, eight are previously unreported for OSA, including CCDC3, MARCHF1, and MED31 identified in all sexes; TMEM26, CPSF4L, and PI4K2B identified in males; and RAP1GAP and YY1 identified in females. We discuss connections to insulin resistance, thiamine deficiency, and resveratrol use that may be worthy of therapeutic consideration for excessively sleepy OSA patients.
Genome-wide association studies (GWAS) have become well-powered to detect loci associated with telomere length. However, no prior work has validated genes nominated by GWAS to examine their role in telomere length regulation. We conducted a multi-ancestry meta-analysis of 211,369 individuals and identified five novel association signals. Enrichment analyses of chromatin state and cell-type heritability suggested that blood/immune cells are the most relevant cell type to examine telomere length association signals. We validated specific GWAS associations by overexpressing KBTBD6 or POP5 and demonstrated that both lengthened telomeres. CRISPR/Cas9 deletion of the predicted causal regions in K562 blood cells reduced expression of these genes, demonstrating that these loci are related to transcriptional regulation of KBTBD6 and POP5. Our results demonstrate the utility of telomere length GWAS in the identification of telomere length regulation mechanisms and validate KBTBD6 and POP5 as genes affecting telomere length regulation.