The high prevalence (>5%) of autoimmune hypothyroidism (AIHT) provides a unique opportunity to dissect genetic contributions to systemic and organ-specific autoimmunity. Here we performed a genome-wide association meta-analysis of 81,718 AIHT cases in FinnGen and the UK Biobank, identifying 418 independent signals (P < 5 × 10-8). At 48 of these loci, a protein-coding variant is, or is highly correlated (r2 > 0.95) with, the lead variant, including Finnish-enriched coding variants in LAG3, ZAP70 and TG. We demonstrated that ZAP70:T155M reduces T cell activation and broadly compare large-scale scans of nonthyroid autoimmunity and thyroid-stimulating hormone levels with a Bayesian classifier to assign loci into distinct groupings, estimating that 38% are involved in general autoimmunity whereas 20% are thyroid specific. We further identified substantial antagonistic pleiotropy, with 10% of AIHT loci showing a consistent protective effect against skin cancer. The AIHT results, including numerous genes encoding checkpoint proteins, support the causal role of natural immune variation influencing cancer outcomes.
Whole genome sequencing (WGS) studies play a pivotal role in studying the genetic underpinnings of human diseases and traits. High quality and reproducible variant calling is the cornerstone for the success of downstream analyses, including WGS association studies and polygenic risk prediction. This paper compares the data quality, performance, and concordance of two widely used WGS variant callers, the Genome Analysis Toolkit (GATK) and Variant Tool set that discovers short variants (VT), using 60 532 multi-ancestry whole genomes sequenced by the Centers for Common Disease Genomics (CCDGs) of the NHGRI Genome Sequencing Program. Our findings show that both QCed GATK and VT pipelines yield highly consistent and reliable called Single Nucleotide Variants (SNVs) in large-scale WGS studies, supporting their agreements in joint variants calling. However, the two pipelines exhibit greater discrepancies in calling insertions and deletions (INDELs).
The integration of genome data with electronic health records, driven by large biobank studies, has advanced human genetics by allowing systematic exploration of genotype-phenotype links. Regular donation enables large, longitudinal sample cohorts. Because blood donors are generally healthy, disease treatments or progression do not disturb interpretations in functional studies. We describe here a pipeline on how to collect blood donors' high quality plasma, serum, and living cell samples for multi-omics studies. Peripheral blood mononuclear cells (PBMC) were frozen and, after thawing, contained standard levels of immune cell subpopulations, responded to immune activation, and were of good quality starting material for single-nucleus multiome and cell imaging studies. We demonstrate that most genetic variants of interest to the major genomics study in Finland, FinnGen, could be found by random collection of samples during the standard blood donation without recall. Probing simple associations in the multi-omics data confirmed expected associations with e.g. age and sex, demonstrating good sample quality. As an example of interesting findings, we observed a significant association between frequent blood donation and lower levels of per- and polyfluoroalkyl substances (PFAS) (e.g., PFHxS β = -0.40 and p = 1.1 × 10-18). The study demonstrates that regular blood donors are a suitable target population for high-quality, cost-effective sample collections.
Polycystic ovary syndrome (PCOS) and its underlying features remain poorly understood. In this genetic study (n = 544,513), we expand the number of genetic loci from 16 to 29, and additionally identify 31 associated plasma proteins. Many risk-increasing loci were associated with later age at menopause, underscoring the reproductive longevity related to an increased oocyte number and/or availability across the lifespan. Hormonal regulation in the etiology of this condition, through metabolic and reproductive features, was emphasized. The proteomic analysis highlighted metabolic biology known to be related to PCOS. A polygenic risk score (PRS) was associated with adverse cardiometabolic outcomes, with differing relevance of testosterone and body mass index in women and men. Finally, while oligo-anovulation and anovulatory infertility are features of PCOS, we observed no impact of PCOS susceptibility on childlessness. We suggest that PCOS susceptibility confers balanced pleiotropic influences on fertility in women, and life-long adverse metabolic consequences in both sexes.
Large-scale genome-wide association studies (GWAS) and rare variant association studies (RVAS) from population biobanks provide valuable resources for gene discovery in complex human traits. We present an analysis of the All of Us Research Program v8 release, which includes whole genome sequencing data and harmonized phenotypic information of 392,030 participants after quality control, enabling a unified investigation of rare and common variants across a spectrum of human traits and diseases. We build an extensive phenome- and genome-wide ("All by All") computational framework to perform GWAS and RVAS on 3,602 phenotypes and identify 49,863 approximately independent, high-quality single-variant and gene-level associations. Meta-analyses of All of Us and UK Biobank, with sample sizes as large as 786,871 participants, further enhance statistical power and find 193 pLoF gene-phenotype associations that are not significant in either cohort alone, including 22 associations not highlighted by previous studies. We also present a public interactive browser that integrates association results for common and rare variants to facilitate interpretation and rapid querying of summary statistics, along with supporting documentation, and a Featured Workspace in the All of Us Researcher Workbench. Our framework will apply to iterative data releases as All of Us grows, empowering researchers worldwide to uncover insights into the functional effects of genetic components on complex traits and diseases.
Abstract Inflammatory bowel diseases (IBD), principally Crohn’s disease (CD) and ulcerative colitis (UC), are common chronic disorders involving inflammation and often progressive tissue damage. Genome-wide association studies have mapped many risk signals, but the causal variants, effector genes and relevant cellular contexts remain difficult to resolve, limiting mechanistic interpretation and therapeutic translation. Here we performed a multi-ancestry GWAS meta-analysis of 125,992 individuals with IBD and more than 1.2 million controls, identifying 619 independent association signals (374 novel) at 420 IBD regions that account for 77–80% of SNP-based heritability. Fine-mapping resolved 81 high-confidence variants, 41 not previously reported. Although most signals were shared between CD and UC, 39% showed IBD subtype specificity, with UC signals showing stronger enrichment in functional annotations from intestinal epithelial, secretory and enteroendocrine cells, and CD showing stronger genetic correlations with circulating inflammatory biomarkers, including C-reactive protein and glycoprotein acetylation. Latent causal modelling supported a causal effect of decreased high-density lipoprotein on CD risk. By integrating bulk and single-cell eQTL and pQTL resources using colocalisation and Mendelian randomisation, together with coding-variant evidence from exome sequencing, we prioritised 664 candidate effector genes across 341 signals, including 390 newly implicated IBD genes, revealing new biological mechanisms and candidate therapeutic targets supported by human genetics.
Objective:To define CSC genetic architecture and identify implicated ocular tissues, cell types, genes, and circulating proteins. Data Sources:Genome-wide data were assembled from FinnGen, All of Us, Mass General Brigham Biobank, Million Veteran Program, and a Dutch chronic CSC cohort. Serum protein quantitative trait loci, human single-cell ocular atlases, and UK Biobank macular optical coherence tomography (OCT) imaging were used for downstream analyses. Study Selection:Five European-ancestry cohorts with genome-wide data and cohort-specific CSC case-control definitions were included, comprising 2,584 cases and 1,044,455 controls. Variants present in at least 2 cohorts were meta-analyzed. Data Extraction and Synthesis:Cohort-level GWASs were adjusted for age, age squared, sex, genotyping array or batch, and 10 genetic principal components, then combined using fixed-effects inverse-variance meta-analysis. Post-GWAS analyses included gene prioritization, colocalization, Mendelian randomization, single-cell disease-relevance scoring, and testing of a CSC genetic risk score in UK Biobank OCT images. Main Outcomes and Measures:Genome-wide significant CSC loci, effector genes and proteins, tissue and cell-type enrichment, and CSC-relevant OCT abnormalities. Results:Across 11,068,938 variants, 10 loci reached genome-wide significance ( P < 5 × 10-8), including 3 novel loci near TGFB1, LINC00551 , and LOC105375630 and 7 replicated loci near CFH, CD46, NOTCH4, PREX1, PTPRB, GATA5 , and TNFRSF10A . Integrative analyses prioritized 10 candidate effector genes. Colocalization and Mendelian randomization implicated circulating TNFRSF10A, TGFB1, and CASP10 levels. Single-cell analyses localized genetic risk to sclera ( P = 2.0 × 10-4) and vascular endothelial cells ( P = 4.0 × 10-4), with fibroblast enrichment. In UK Biobank, OCT abnormalities were more frequent in the top vs bottom 1% of CSC genetic risk (18 of 109 [16.5%] vs 8 of 134 [6.0%]; odds ratio, 4.05; 95% CI, 1.65-10.87; P = .002). Conclusions and Relevance:In this GWAS meta-analysis, CSC susceptibility localized predominantly to scleral and vascular biology rather than primary retinal pigment epithelial dysfunction. These findings support CSC as a sclerovascular disorder and nominate complement regulation, endothelial signaling, and extracellular matrix pathways for future study. Key Points:Question: What genetic loci, genes, proteins, and ocular cell types underlie susceptibility to central serous chorioretinopathy (CSC)?Findings: In this GWAS meta-analysis of 2,584 CSC cases and 1,044,455 controls, 10 genome-wide significant loci were identified, including 3 novel loci. Integrative analyses implicated scleral fibroblasts and vascular endothelial cells, prioritized candidate effector genes and circulating proteins, and showed that high CSC genetic risk was associated with more frequent RPE abnormalities on macular OCT.Meaning: These findings support CSC as a primary sclerovascular disorder and nominate mechanism-linked pathways for future translational studies.Importance: The primary site of dysfunction in central serous chorioretinopathy (CSC) remains uncertain.
Intrahepatic cholestasis of pregnancy, which affects 0.2-2% of pregnancies, is characterized by pruritus, increased aminotransferase activity and elevated serum bile acids. Previous studies have implicated liver-enriched genes in intrahepatic cholestasis of pregnancy. We conducted a meta-analysis of intrahepatic cholestasis of pregnancy genome-wide association studies in the FinnGen study, deCODE, Estonian Biobank, the Danish Blood Donor Study and Copenhagen Hospital Biobank with 4,738 women with prior ICP and 436,834 female controls. The analysis found 26 genome-wide significant associations of which 10 were novel. Genes in the associated loci were prioritized using lead SNP expression quantitative trait loci associations and colocalization analysis to assess potential causality. The associated loci implicate bile acid synthesis, LDL cholesterol, and lipid metabolism. Additionally, comorbidity, genetic correlation and polygenic risk score analyses further indicated a link between intrahepatic cholestasis of pregnancy and pancreatitis, suggesting shared genetic underpinnings.
Background & Aims Genetic admixture of United States Hispanic individuals provides a unique opportunity to examine ancestral origins of inflammatory bowel disease (IBD) risk. In ∼7.3K Hispanic participants (1660 IBD cases; 5614 controls), we examined ancestral heterogeneity of IBD clinical phenotypes and sought to identify IBD risk loci that displayed heterogeneity of effect or were ancestry-specific. Methods Association of genetic ancestry with clinical phenotypes was evaluated. We conducted an ancestry-informed genome-wide (GW) association study for IBD, ulcerative colitis, and Crohn’s disease (CD) to obtain ancestry-specific effect size estimates for alleles from African (AFR), European (EUR), and Amerindian (AIAN) origin. Ancestry-specific replication was assessed in All of Us Hispanic participants and transferability was evaluated for populations with similar ancestral origin. Results Clinical phenotypes were associated with higher AFR (colonic, penetrating, or perianal CD; later age at diagnosis; IBD-related surgery) or AIAN ancestry (colonic CD). GW EUR-specific associations were observed within established loci for CD (NOD2, IL23R, HLA-DRA) and ulcerative colitis (HLA locus). For AFR or AIAN alleles, novel GW associations were observed in 14 loci. One AFR-specific IBD GW (PCGEM1) and 2 AFR-specific suggestive (TYROBP/LRFN3) associations replicated in All of Us. Several suggestive associations demonstrated transferability (AFR-specific TYROBP/LRFN3 and AIAN-specific GAD2). Several novel IBD risk variants also demonstrated association with clinical phenotypes. Conclusions Ancestry-informed regression enabled identification of novel AFR and AIAN-specific risk alleles, which may also inform observed phenotypic differences. We have shown that some previously identified IBD loci have associations that are EUR-specific. These findings highlight the importance of genetic ancestry for elucidating the biological underpinnings of IBD and may have important pharmacogenetic implications.
Autism spectrum disorder (ASD) is estimated to be up to four times as common in males as in females, yet the causes of this prevalence difference are not well established. One possible driver is genetic variation on the X chromosome, as it contains genes capable of contributing to ASD (e.g., PTCHD1, MECP2) and is known to play a role in genetic disorders with differential sex prevalence (e.g., color blindness). However, a lack of power compared to the autosomes combined with the complexities of modeling its biology have led to the X being largely overlooked in sequencing studies. Here, we develop quantitative X-linked TADA, a new model designed specifically for application to this chromosome, and use it to analyze rare variation from 50,663 individuals with ASD (and 136,670 individuals total). We find 9 genes on the X associated with ASD at a false discovery rate (FDR) < 0.05 and an additional 9 genes at FDR < 0.2, with many of these previously identified as involved in specific neurodevelopmental disorders. Point estimates of the liability conferred by de novo variants on the X are similar in females and males, with both sexes' estimates elevated >20% above the corresponding autosomal values. We also develop a general theory of how X-linked variation of any additive or non-additive effect influences liability and describe its implications for prevalence. Using this theory and our empirical results, we show how genetic variation on the X could contribute to the sex-differential prevalence of ASD.
Tourette syndrome (TS) is a neurodevelopmental disorder characterized by symptoms that emerge in childhood and often improve or even disappear in adulthood, providing a model for understanding how altered brain development shapes neural structure and function. We investigate brain structural alterations in TS and Chronic Tic Disorders (TS/CTD) across development, presenting the largest structural neuroimaging analysis for TS/CTD to date (1,803 individuals from the ENIGMA-TS Working Group), and integrating with large-scale genomewide association studies. Nonlinear age effects were observed in cortical thickness across development and in thalamic volume in children, indicating altered trajectories of brain maturation. Pediatric and adult TS/CTD showed distinct structural patterns, with widespread alterations in childhood and more focal changes in adulthood. Children also showed the most prominent effects highlighting the involvement of orbitofrontal cortex and putamen, alongside additional regions such as frontal and paralimbic areas. Genetic pleiotropy analyses identified overlap between TS/CTD-associated genetic effects on brain structure and neuroanatomical differences. Cross-disorder comparisons revealed correlations with ADHD and OCD and age-related patterns. These findings demonstrate altered neurodevelopmental trajectories in TS/CTD and implicate systems underlying inhibitory control and urge regulation.
Autism spectrum disorder is a heritable neurodevelopmental condition affecting approximately 3% of children that presents with core behavioral features and a range of possible comorbidities, including intellectual disability. While common variants contribute substantially to autism liability, the discovery of specific autism-associated genes has largely been driven by studies of rare and de novo variants. Many of these genes are also linked with broadly defined developmental disorders, but their involvement in other conditions has not been mapped at scale. Here, we analyze autosomal rare coding variation from 62,429 individuals with autism from research and clinical cohorts to identify 253 autism-associated genes at an estimated false discovery rate < 0.001. We cluster them based on association evidence from large-scale studies of developmental disorders, schizophrenia, bipolar disorder, and epilepsy, generating six clusters of genes with differing biological pathway enrichments and patterns of comorbidities. Investigating rare variant associations in the population using the UK Biobank and All of Us, we identify autism-associated genes displaying pleiotropy across physiological systems. In addition, we report 497 genes impacting development in a meta-analysis with 26,109 published developmental disorders samples. Collectively drawing upon data from over 1.5 million individuals, our study finds that rare variants across hundreds of genes contribute to autism with variable phenotypic outcomes.
The past decade has seen remarkable progress in identifying genes that, when impacted by deleterious coding variation, confer high likelihood for autism spectrum disorder (ASD), intellectual disability and other associated developmental disorders. However, most underlying gene discovery efforts have focused on individuals of European ancestry, limiting insights into genetic liability across diverse populations. To help address this, the Genomics of Autism in Latin American Ancestries (GALA) Consortium was formed, presenting here the largest sequencing study of autism in Latin American individuals (n > 15,000, including 4,717 participants with an ASD diagnosis). We identified 35 genome-wide significant (false discovery rate < 0.05) autism-associated genes, with substantial overlap with findings from European cohorts, and highly constrained genes showing consistent signal across populations. The results provide support for emerging (for example, MARK2, YWHAG, PACS1, RERE, SPEN, GSE1, GLS, TNPO3 and ANKRD17) and established autism genes and for the utility of genetic testing approaches for deleterious variants in individuals from diverse backgrounds; the results also demonstrate the ongoing need for more inclusive genetic research and testing. We conclude that the biology of autism is consistent across populations, with no detectable influence of ancestry.
Our understanding of the biological role of the Y chromosome remains limited. Here, we systematically profile germline Y haplogroups and somatic loss of the Y chromosome (LOY) in 122,683 East Asian males from BioBank Japan and 181,472 European males from the UK Biobank. A phenome-wide scan uncovers male-specific genetic regulation of complex traits, including pleiotropic effects of the Japanese-specific haplogroup D on height and type 2 diabetes (T2D). LOY increases T2D risk in East Asians but is associated with reduced T2D risk in Europeans. In East Asians, LOY contributes to T2D incidence particularly among males with lower polygenic risk scores, providing a compensatory explanation for disease risk beyond germline genetics. Incorporating sex-chromosome variation improves polygenic prediction of T2D risk in both sexes. Single-cell analyses reveal cell type-specific accumulation of LOY across tissues and disease contexts, with LOY in pancreatic β cells potentially impairing glucose metabolism. Our study demonstrates the clinical relevance of Y chromosome variation for diabetes risk prediction and management.
Circadian rhythms not only coordinate the timing of wake and sleep but also regulate homeostasis within the body, including glucose metabolism. The genetic variants that contribute to the temporal control of glucose levels have not been previously examined. Using genome-wide data from ~420,000 individuals from the UK Biobank and replication in ~100,000 individuals from the Estonian Biobank, ~500,000 from FinnGen, ~160,000 from the VA Million Veteran Program, and ~52,000 from the MGB Biobank, we show that glucose levels are under diurnal genetic control. We discover a robust temporal association of glucose levels at the Melatonin receptor 1B (MTNR1B, rs10830963, P = 1×10-22) and a canonical circadian pacemaker gene Cryptochrome 2 (CRY2) loci (rs12419690, P = 1×10-16). Furthermore, we show that sleep modulates glucose levels, and the genetic variants have an independent role in diurnal glucose control. Finally, we show that these variants independently modulate risk of type 2 diabetes and that sleep medications including melatonin associate with type 2 diabetes. Our findings, together with earlier genetic and epidemiological evidence, show a clear connection between sleep and metabolism and highlight genetic variation at MTNR1B and CRY2 in the control of diurnal glucose levels.
Inflammatory bowel disease (IBD) is a chronic immune-mediated disorder of the gastrointestinal tract whose genetic basis is only partly resolved because most risk variants identified by genome-wide association studies (GWAS) lie in non-coding regions, limiting direct gene assignment and biological interpretation1,2. Here we analysed whole-exome and whole-genome sequencing data from 86,213 IBD cases and 478,363 controls of European ancestry. We identified 68 IBD genes directly implicated by conditionally independent protein-coding associations across the allele frequency spectrum. Many newly implicated IBD genes are supported by orthogonal genomic or pleiotropic evidence, pointing to disease-related pathways and nominating targets with therapeutic relevance. We further identified allelic series and non-additive effects at key loci such as NOD2 and TYK2. These results show that large-scale sequencing can resolve disease genes and pathways that remain ambiguous from non-coding association alone, providing a more direct route from human genetics to biological insight and therapeutic hypotheses.
Abstract Many disease-associated variants are thought to act through gene regulation, yet conventional eQTL mapping explains only a fraction of GWAS loci, potentially because regulatory effects vary across cellular states and environments. We present CASTIE, a scalable Poisson mixed-model framework that directly models sparse single-cell read counts and enables genome-wide testing of genotype-by-context interactions without pre-screening for static effects. Applying CASTIE to 1.2 million peripheral blood mononuclear cells from 982 OneK1K donors identified 3,155 context-dependent eQTL associations, including 2,022 eGenes without detectable static effects. These associations yielded 374 colocalizations across 94 traits, representing 270 unique loci, of which 197 were not recovered using the corresponding static eQTLs. The colocalizations linked trait associations to specific cellular contexts and genes, including GCHFR , RNASET2 and ATP1A3 . In adipose-derived mesenchymal stem cells exposed to metabolic stimulations, CASTIE increased eGene discovery by 36 − 92% across cell populations and identified stimulation-dependent regulatory effects at metabolic trait loci. Thus, modeling cellular context reveals disease-relevant regulatory variation beyond static eQTL mapping.
Polygenic scores (PGSs) quantify individual genetic susceptibility to complex diseases and can identify high-risk individuals well before clinical onset. Their clinical translation, however, requires population-based reference resources, standardized benchmarking, and accessible tools for translating individual scores into disease likelihood. In this article, we systematically evaluate 3168 PGS models, primarily from the PGS Catalog, in 473,681 FinnGen participants, placing all models on a common performance scale to enable cross-model and cross-trait comparison. For each PGS, we create ancestry-adjusted reference distributions, providing a biobank-scale resource for interpreting individual scores. We perform phenome-wide association studies for each PGS, identifying 439,070 significant phenotypic associations, demonstratin g that integrating multiple scores improves predictive performance for most complex diseases, and providing public access to 11 top-performing interactive time-to-event models. All resources are accessible through the PGS Browser ( pgs.nchigm.org ), which offers a population-aware framework for score interpretation and lays groundwork for the clinical application of PGSs.
The Genome Aggregation Database (gnomAD) is a foundational resource for allele frequency data, widely used in genomic research and clinical interpretation. However, traditional estimates rely on individual-level genetic ancestry groupings that may obscure variation in recently admixed populations. To improve resolution, we applied local ancestry inference (LAI) to over 27 million variants in two admixed groups: Admixed American (n = 7,612) and African/African American (n = 20,250), deriving ancestry-specific allele frequencies. We show that 78.5% and 85.1% of variants in these groups, respectively, exhibit at least a twofold difference in ancestry-specific frequencies. Moreover, 81.49% of variants with LAI information would be assigned a higher gnomAD-wide maximum frequency after incorporating LAI, potentially altering clinical interpretations. This LAI-informed release reveals clinically relevant frequency differences that are masked in aggregate estimates and may support reclassifying some variants from Uncertain Significance to Benign or Likely Benign.