Most genetic variance in gene expression is due to trans-acting expression quantitative trait loci (eQTLs) spread across the genome. However, these loci are generally hard to map due to limited discovery power. Here, we simulate how local properties of expression regulation and global properties of regulatory networks alter the genome-wide proportions of cis- and trans-heritability. We find that network motifs and modular groups can reduce or enhance the effects of trans-eQTLs and that hub regulators shorten paths across the network and act as key sources of trans-acting variance. Critically, networks with all these features best recapitulate the observed distribution of cis- and trans-heritability. Taken together, our results suggest that the genome-wide genetic architecture of gene expression involves fewer regulators for each gene but implicates the same regulators more often across genes (i.e., is less polygenic and more pleiotropic) than previously anticipated.
The genome-wide burdens of deletions, loss-of-function mutations, and duplications correlate with many traits. Curiously, for most of these traits, variants that decrease expression have the same genome-wide average direction of effect as variants that increase expression. This seemingly contradicts the intuition that for individual genes reducing expression should have the opposite effect on a phenotype as increasing expression. To understand this paradox, we use the gene dosage response curve (GDRC), which relates changes in gene expression to expected changes in phenotype. We show that, for many traits, GDRCs are systematically biased in one trait direction relative to the other, and we develop a simple theoretical model that explains this bias in trait direction. Our results have broad implications for complex traits, drug discovery, and statistical genetics.
Gene families are groups of evolutionarily related genes. One large gene family that has experienced rapid evolution lies within the Major Histocompatibility Complex (MHC), whose proteins serve critical roles in innate and adaptive immunity. Across the ∼60 million year history of the primates, some MHC genes have turned over completely, some have changed function, some have converged in function, and others have remained essentially unchanged. Past work has typically focused on identifying MHC alleles within particular species or comparing gene content, but more work is needed to understand the overall evolution of the gene family across species. Thus, despite the immunologic importance of the MHC and its peculiar evolutionary history, we lack a complete picture of MHC evolution in the primates. We readdress this question using sequences from dozens of MHC genes and pseudogenes spanning the entire primate order, building a comprehensive set of gene and allele trees with modern methods. Overall, we find that the Class I gene subfamily is evolving much more quickly than the Class II gene subfamily, with the exception of the Class II MHC-DRB genes. We also pay special attention to the often-ignored pseudogenes, which we use to reconstruct different events in the evolution of the Class I region. We find that despite the shared function of the MHC across species, different species employ different genes, haplotypes, and patterns of variation to achieve a successful immune response. Our trees and extensive literature review represent the most comprehensive look into primate MHC evolution to date.
Natural selection on complex traits is difficult to study in part due to the ascertainment inherent to genome-wide association studies (GWAS). The power to detect a trait-associated variant in GWAS is a function of its frequency and effect size - but for traits under selection, the effect size of a variant determines the strength of selection against it, constraining its frequency. Recognizing the biases inherent to GWAS ascertainment, we propose studying the joint distribution of allele frequencies across populations, conditional on the frequencies in the GWAS cohort. Before considering these conditional frequency spectra, we first characterized the impact of selection and non-equilibrium demography on allele frequency dynamics forwards and backwards in time. We then used these results to understand conditional frequency spectra under realistic human demography. Finally, we investigated empirical conditional frequency spectra for GWAS variants associated with 106 complex traits, finding compelling evidence for either stabilizing or purifying selection. Our results provide insights into polygenic score portability and other properties of variants ascertained with GWAS, highlighting the utility of conditional frequency spectra.
Classical genes within the Major Histocompatibility Complex (MHC) are responsible for peptide presentation to T cells, thus playing a central role in immune defense against pathogens. These genes are subject to strong selective pressures including both balancing and directional selection, resulting in exceptional genetic diversity—thousands of alleles per gene in humans. Moreover, some allelic lineages appear to be shared between primate species, a phenomenon known as trans-species polymorphism (TSP) or incomplete lineage sorting, which is rare in the genome overall. However, despite the clinical and evolutionary importance of MHC diversity, we currently lack a full picture of primate MHC evolution. In particular, we do not know to what extent genes and allelic lineages are retained across speciation events. To start addressing this gap, we explore variation across genes and species in our companion paper (Fortier and Pritchard, 2025), and here we explore variation within individual genes. We used Bayesian phylogenetic methods to determine the extent of TSP at 17 MHC genes, including classical and non-classical Class I and Class II genes. We find strong support for ancient TSP in 7 of 10 classical genes, including—remarkably—between humans and old-world monkeys in MHC-DQB1. In addition to the long-term persistence of ancient lineages, we additionally observe rapid evolution at nucleotides encoding the proteins’ peptide-binding domains. The most rapidly-evolving amino acid positions are extremely enriched for autoimmune and infectious disease associations. Together, these results suggest complex selective forces—arising from differential peptide binding—that drive short-term allelic turnover within lineages while also maintaining deeply divergent lineages for at least 31 million years in some cases.
The ability of individual cells to maintain a distinct identity and respond to transient environmental signals requires tightly controlled regulation of gene networks. However, how discrete sets of regulators coordinate dynamic gene circuits remains poorly defined. The need for context-dependent regulation is prominent in human CD4+ T cells, where distinct cell lineages must respond to diverse signals to orchestrate effective adaptive immune responses and maintain homeostasis. We performed CRISPR screens in multiple primary human CD4+ T cell contexts to identify regulators that control expression of IL2RA, which is a canonical marker of T cell activation in pro-inflammatory effector T cells (Teffs) and constitutively expressed in anti-inflammatory regulatory T cells (Tregs) where it is required for fitness. Strikingly, the majority of identified regulators are required in discrete cell type and stimulation timepoints, and a subset even had opposite functional effects in different conditions. Using single-cell transcriptomics after pooled perturbation of context-specific screen hits, we characterized factors as regulators of overall rest or activation and constructed state-specific regulatory networks. Upstream of these networks, MED12 – a component of the Mediator complex – serves as a dynamic orchestrator of regulators across conditions, governing both cell type- and stimulation-specific gene expression. We determined that MED12 interacts with histone modifying proteins, affecting chromatin state and expression of genes encoding key state- and lineage-defining factors, including numerous regulators of IL2RA. Importantly, MED12 promotes the expression of several core genes including MYC, rest maintenance factor KLF2, and activation promoting gene GATA3. CRISPR ablation of MED12 blunted the transition from rest to activation and protected T cells from activation-induced cell death, resulting in increased durability of conventional effector T cells. Overall, CRISPR screens performed across primary cell conditions enabled the identification of regulatory circuits required to establish T cell rest and activation, which can be modulated to improve cellular persistence. Citation Format: Maya M Arce, Jennifer Umhoefer, Nadia Arang, Sivakanthan Kasinathan, Jacob W Freimer, Zachary Steinhart, Haolin Shen, Mineto Ota, Anika Wadhera, Minh T.N Pham, Rama Dajani, Dmytro Dorovsky, Yan Yi Chen, Qi Liu, Brian R Shy, Julia Carnevale, Ansuman T Satpathy, Nevan J Krogan, Jonathan K Pritchard, Alexander Marson. CD4+ T cell rest and activation is enabled by centralized control of state specific regulatory genes [abstract]. In: Proceedings of the AACR IO Conference: Discovery and Innovation in Cancer Immunology: Revolutionizing Treatment through Immunotherapy; 2025 Feb 23-26; Los Angeles, CA. Philadelphia (PA): AACR; Cancer Immunol Res 2025;13(2 Suppl):Abstract nr PR008.
Gene regulatory networks (GRNs) govern many core developmental and biological processes underlying human complex traits. Even with broad-scale efforts to characterize the effects of molecular perturbations and interpret gene coexpression, it remains challenging to infer the architecture of gene regulation in a precise and efficient manner. Key properties of GRNs, like hierarchical structure, modular organization, and sparsity, provide both challenges and opportunities for this objective. Here, we seek to better understand properties of GRNs using a new approach to simulate their structure and model their function. We produce realistic network structures with a novel generating algorithm based on insights from small-world network theory, and we model gene expression regulation using stochastic differential equations formulated to accommodate modeling molecular perturbations. With these tools, we systematically describe the effects of gene knockouts within and across GRNs, finding a subset of networks that recapitulate features of a recent genome-scale perturbation study. With deeper analysis of these exemplar networks, we consider future avenues to map the architecture of gene expression regulation using data from cells in perturbed and unperturbed states, finding that while perturbation data are critical to discover specific regulatory interactions, data from unperturbed cells may be sufficient to reveal regulatory programs.
The human leukocyte antigen (HLA) region plays an important role in human health through its involvement in immune cell recognition and maturation. While genetic variation in the HLA region is associated with many diseases, the pleiotropic patterns of these associations have not been systematically investigated. Here, we developed a haplotype approach to investigate disease associations phenome wide for 412,181 Finnish individuals and 2,459 diseases. Across the 1,035 diseases with a genome-wide association study association, we found a 17-fold average per-SNP enrichment of hits in the HLA region. Altogether, we identified 7,649 HLA associations across 647 diseases, including 1,750 associations uncovered by haplotype analysis. We found that some haplotypes show both risk-increasing and protective associations across different diseases, while others consistently increase risk across diseases, indicating a complex pleiotropic landscape involving a range of diseases. This study highlights the extensive impact of HLA variation on disease risk and underscores the importance of classical and non-classical genes as well as non-coding variation.
Genetic factors play an important role in prostate cancer (PCa) development with polygenic risk scores (PRS) predicting disease risk across genetic ancestries. However, there are few convincing modifiable factors for PCa and little is known about their potential interaction with genetic risk. Our study explores the role of neighborhood socioeconomic status (nSES)-and how it may interact with PRS-on PCa risk. We analyzed incident PCa cases and controls of European (cases = 5,960; controls = 93,990) and African (cases = 109; controls = 1,226) ancestry from the UK Biobank cohort. Using the English indices of deprivation, a set of validated metrics that quantify lack of resources within geographical areas, we performed logistic regression to investigate the main effects and interactions between nSES deprivation and genetic susceptibility to PCa, represented by a multi-ancestry PRS comprised of 269 genetic variants. The PRS was associated with PCa in the European (OR = 2.04; 95% confidence interval [CI], 2.00-2.09; p = 5.34 × 10-807) and African (OR = 1.35; 95% CI, 1.16-1.58; p = 1.05 × 10-4) ancestries. Additionally, nSES deprivation indices were inversely associated with PCa: employment, education, health, and income. From this, we suspect that PRS, through biological mechanisms, and nSES deprivation, likely through differences in screening, are associated with PCa, but act independently of each other. Our findings suggest that genetic factors and social determinants of health measured by neighborhood socioeconomic status do not synergistically increase risk of PCa.
Genetic association studies provide a unique tool for identifying causal links from genes to human traits and diseases. However, it is challenging to determine the biological mechanisms underlying most associations, and we lack genome-scale approaches for inferring causal mechanistic pathways from genes to cellular functions to traits. Here we propose new approaches to bridge this gap by combining quantitative estimates of gene-trait relationships from loss-of-function burden tests with gene-regulatory connections inferred from Perturb-seq experiments in relevant cell types. By combining these two forms of data, we aim to build causal graphs in which the directional associations of genes with a trait can be explained by their regulatory effects on biological programs or direct effects on the trait. As a proof-of-concept, we constructed a causal graph of the gene regulatory hierarchy that jointly controls three partially co-regulated blood traits. We propose that perturbation studies in trait-relevant cell types, coupled with gene-level effect sizes for traits, can bridge the gap between genetics and biology.
Genome-wide association studies have revealed that the genetic architectures of complex traits vary widely, including in terms of the numbers, effect sizes, and allele frequencies of significant hits. However, at present we lack a principled way of understanding the similarities and differences among traits. Here, we describe a probabilistic model that combines the effects of mutation, drift, and stabilizing selection at individual sites with a genome-scale model of phenotypic variation. In this model, the architecture of a trait arises from the distribution of selection coefficients of mutations and from two scaling parameters. We fit this model for 95 highly polygenic quantitative traits of different kinds from the UK Biobank. Notably, we infer that all these traits have fairly similar, though not identical, distributions of selection coefficients. This similarity suggests that differences in architectures of highly polygenic traits arise mainly from the two scaling parameters: the mutational target size and heritability per site, which vary by orders of magnitude among traits. When these two scale factors are accounted for, we find that the architectures of all 95 traits are very similar.
rRNA genes exhibit intra-individual hyper-variability and an outstanding question is their role in human health and disease. These include variants positioned at enigmatic regions of rRNA named Expansion-Segments (ESs) that protrude from the core of the ribosome, with poorly understood functions. In this study, we analyze rRNA variants in the UK Biobank population, revealing that common rRNA variations that give rise to ribosome subtypes, affect human physiology and rare rRNA-mutations affect diseases. We developed a Ribosome-Variation-Analysis (RiboVAn) method, identifying ribosome subtypes as heritable and a larger proportion of low-heritability rRNA-mutations. The heritable variants included ones in es15l associated with adiposity, es39l with body dimensions, and es27l with blood-related traits and diseases. Variant-chromosome specificity is observed where ribosome subtypes are linked to distinct acrocentric chromosomes. Burden analysis linked rRNA-mutations to diverse diseases including cancer and acute myocardial infarction. These findings causally link rRNA variation to human traits, disease, and establish that ESs have distinct and important functions in human physiology.
CRISPR screens are powerful tools to identify key genes that underlie biological processes. One important type of screen uses fluorescence activated cell sorting (FACS) to sort perturbed cells into bins based on the expression level of marker genes, followed by guide RNA (gRNA) sequencing. Analysis of these data presents several statistical challenges due to multiple factors including the discrete nature of the bins and typically small numbers of replicate experiments. To address these challenges, we developed a robust and powerful Bayesian random effects model and software package called Waterbear. Furthermore, we used Waterbear to explore how various experimental design parameters affect statistical power to establish principled guidelines for future screens. Finally, we experimentally validated our experimental design model findings that, when using Waterbear for analysis, high power is maintained even at low cell coverage and a high multiplicity of infection. We anticipate that Waterbear will be of broad utility for analyzing FACS-based CRISPR screens.
Circadian rhythms not only coordinate the timing of wake and sleep but also regulate homeostasis within the body, including glucose metabolism. However, the genetic variants that contribute to temporal control of glucose levels have not been previously examined. Using data from 420,000 individuals from the UK Biobank and replicating our findings in 100,000 individuals from the Estonian Biobank, we show that diurnal serum glucose is under genetic control. We discover a robust temporal association of glucose levels at the Melatonin receptor 1B (MTNR1B) (rs10830963, P = 1e-22) and a canonical circadian pacemaker gene Cryptochrome 2 (CRY2) loci (rs12419690, P = 1e-16). Furthermore, we show that sleep modulates serum glucose levels and the genetic variants have a separate mechanism of diurnal control. Finally, we show that these variants independently modulate risk of type 2 diabetes. Our findings, together with earlier genetic and epidemiological evidence, show a clear connection between sleep and metabolism and highlight variation at MTNR1B and CRY2 as temporal regulators for glucose levels.
Most human complex traits are enormously polygenic, with thousands of contributing variants with small effects, spread across much of the genome. These observations raise questions about why so many variants–and so many genes–impact any given phenotype. Here we consider a possible model in which variant effects are due to competition among genes for pools of shared intracellular resources such as RNA polymerases. To this end, we describe a simple theoretical model of resource competition for polymerases during transcription. We show that as long as a gene uses only a small fraction of the overall supply of polymerases, competition with other genes for this supply will only have a negligible effect on variation in the gene’s expression. In particular, although resource competition increases the proportion of heritability explained by trans-eQTLs, this effect is far too small to account for the roughly 70% of expression heritability thought to be due to trans-regulation. Similarly, we find that competition will only have an appreciable effect on complex traits under very limited conditions: that core genes collectively use a large fraction of the cellular pool of polymerases and their overall expression level is strongly correlated (or anti-correlated) with trait values. Our qualitative results should hold for a wide family of models relating to cellular resource limitations. We conclude that, for most traits, resource competition is not a major source of complex trait heritability.
Measures of selective constraint on genes have been used for many applications including clinical interpretation of rare coding variants, disease gene discovery, and studies of genome evolution. However, widely-used metrics are severely underpowered at detecting constraint for the shortest ~25% of genes, potentially causing important pathogenic mutations to be overlooked. We developed a framework combining a population genetics model with machine learning on gene features to enable accurate inference of an interpretable constraint metric, shet. Our estimates outperform existing metrics for prioritizing genes important for cell essentiality, human disease, and other phenotypes, especially for short genes. Our new estimates of selective constraint should have wide utility for characterizing genes relevant to human disease. Finally, our inference framework, GeneBayes, provides a flexible platform that can improve estimation of many gene-level properties, such as rare variant burden or gene expression differences.
Ancient DNA research in the past decade has revealed that European population structure changed dramatically in the prehistoric period (14,000-3,000 years before present, YBP), reflecting the widespread introduction of Neolithic farmer and Bronze Age Steppe ancestries. However, little is known about how population structure changed from the historical period onward (3,000 YBP - present). To address this, we collected whole genomes from 204 individuals from Europe and the Mediterranean, many of which are the first historical period genomes from their region (e.g. Armenia and France). We found that most regions show remarkable inter-individual heterogeneity. At least 7% of historical individuals carry ancestry uncommon in the region where they were sampled, some indicating cross-Mediterranean contacts. Despite this high level of mobility, overall population structure across western Eurasia is relatively stable through the historical period up to the present, mirroring geography. We show that, under standard population genetics models with local panmixia, the observed level of dispersal would lead to a collapse of population structure. Persistent population structure thus suggests a lower effective migration rate than indicated by the observed dispersal. We hypothesize that this phenomenon can be explained by extensive transient dispersal arising from drastically improved transportation networks and the Roman Empire’s mobilization of people for trade, labor, and military. This work highlights the utility of ancient DNA in elucidating finer scale human population dynamics in recent history.
Natural selection on complex traits is difficult to study in part due to the ascertainment inherent to genome-wide association studies (GWAS). The power to detect a trait-associated variant in GWAS is a function of frequency and effect size - but for traits under selection, the effect size of a variant determines the strength of selection against it, constraining its frequency. To account for GWAS ascertainment, we propose studying the joint distribution of allele frequencies across populations, conditional on the frequencies in the GWAS cohort. Before considering these conditional frequency spectra, we first characterized the impact of selection and non-equilibrium demography on allele frequency dynamics forwards and backwards in time. We then used these results to understand conditional frequency spectra under realistic human demography. Finally, we investigated empirical conditional frequency spectra for GWAS variants associated with 106 complex traits, finding compelling evidence for either stabilizing or purifying selection. Our results provide insight into polygenic score portability and other properties of variants ascertained with GWAS, highlighting the utility of conditional frequency spectra.