
Variant effect predictors (VEPs) are widely used to interpret the functional consequences of human genetic variation. Because most methods rely on sequence conservation, they implicitly treat conservation as evidence of functional constraint. However, substitution patterns across a phylogeny reflect not only selection but also differences in underlying mutation rates. Here, we show that this creates a systematic confounding: most VEPs capture mutation rate variation and misinterpret it as variation in functional importance. Widely used conservation metrics exhibit a related bias; in particular, phyloP scores correlate strongly with mutation rate even at putatively neutral sites. Consequently, variants at low-mutation-rate sites tend to be predicted as more damaging, and variants at highly mutable sites as more tolerated, than warranted by their true functional impact. We also identify a distinct biological signal in experimental measurements of mutational effects on protein stability: amino acid substitutions that are more likely to arise are, on average, less destabilizing than rarer substitutions. This provides empirical support for mutational robustness in the context of protein stability. However, this relationship is insufficient to explain the mutation-rate dependence observed in current VEP outputs. Together, our findings show that mutation rate heterogeneity systematically biases current variant effect prediction frameworks, highlight the need to model mutation probabilities explicitly in future VEPs, and reveal a genuine biological signal of mutational robustness.
ZNF536 encodes a C2H2 zinc-finger transcription factor that functions as a transcriptional repressor. While common noncoding variants at the ZNF536 locus have been reported to be associated with schizophrenia in a genome-wide association study (GWAS), the contribution of rare, protein-altering variants to human disease has not been systematically investigated. Through an international collaboration, we assembled a cohort of 21 affected individuals carrying 18 unique, rare, heterozygous, protein-altering ZNF536 variants. Most variants (15/18) were predicted loss-of-function (LoF) alleles, with the remainder being missense variants. Among families with available inheritance data (17/20), most variants arose de novo (12/17), while others were inherited from mosaic or mildly affected parents (5/17). Clinically, affected individuals presented with developmental delay along with high rates of autism spectrum disorder, intellectual disability, hyperactivity, aggressive behavior, anxiety, and hyperphagia; epilepsy and sleep disturbances were also frequently observed. To assess functional consequences of a proband-associated ZNF536 variant, we generated a Zfp536p.Gln169Ter knock-in mouse model. Homozygous mutants were non-viable, while heterozygotes survived but displayed autism-like behaviors, increased anxiety, and impaired recognition memory. Embryonic brain analysis revealed reduced cortical size, cortical thickness, and decreased deep-layer neuronal density. These features are consistent with phenotypes of a publicly available mouse knockout model and support our clinical cohort findings that rare monoallelic LoF variants in ZNF536 underlie a genetic neurodevelopmental disorder characterized by developmental delay, autism, and behavioral dysregulation. The pathogenicity of missense variants in disease remains to be determined. These results support a role for ZNF536 as a dosage-sensitive regulator of cortical development.
Sickle cell trait (SCT) is increasingly recognized as a risk factor for adverse health outcomes. We utilized three large US electronic health record-based biobanks (Vanderbilt University Medical Center’s BioVU, Penn Medicine Biobank, and All of Us) to conduct phenome-wide association (PheWAS) and clinical laboratory-wide association (LabWAS) meta-analyses of 4,813 individuals with SCT among 58,830 African genetic ancestry participants (∼60% female). Significant associations were replicated using a published PheWAS of SCT from the Million Veteran Program. Our PheWAS meta-analysis confirmed the association of SCT with increased risks of kidney disease, pulmonary embolism, and anemia, while also identifying associations with increased risk of splenomegaly, gout, acute pyelonephritis, and anemia of pregnancy. LabWAS confirmed prior associations of SCT with blood cell counts, red cell indices, kidney function, and urinary concentrating ability. We also identified associations with higher serum electrolytes, bilirubin, and reticulocyte count and lower blood urea nitrogen and platelet count. In sex-stratified analyses, the association of SCT with kidney-related disorders and kidney dysfunction was stronger in females, while the association with platelet and lymphocyte phenotypes was greater in males with SCT. Our results provide insights into the multi-system complications of SCT and have potential clinical implications both for general awareness of susceptibility and appropriate reference ranges for various clinical laboratory parameters in individuals with SCT.
The Evidence-based Network for the Interpretation of Germline Mutant Alleles (ENIGMA) research consortium conducted a comprehensive study to characterize spliceogenic variants in BRCA1 exon 18. The absence of systematic RNA-based assessment for these variants has led to inconsistent interpretation, limiting accurate classification and management of individuals and their families. The splicing profile of 166 variants was assessed using minigene assays; 32 were additionally analyzed in blood-derived RNA from 51 individuals and 18 in mouse embryonic stem cell (mESC)-based assays to evaluate homology-directed repair (HDR) capacity. mRNA assessment by RT-PCR in blood samples and minigene assays showed a significant positive correlation, with splicing analysis in mESCs displaying highly concordant results. The mESC-based HDR assay showed that the in-frame exon 18 skipping (Δ18) transcript encodes a non-functional protein lacking rescue activity. Linear regression analysis using mESC splicing and functional data indicated that ≥59% of full-length (FL) levels and <34% of Δ18 were associated with benign HDR activity. These thresholds differ from those recommended by the ClinGen ENIGMA BRCA1 and BRCA2 Variant Curation Expert Panel American College of Medical Genetics and Genomics (ACMG)/Association for Molecular Pathology (AMP) specifications for applying BP7_strong(RNA): >30% functional transcripts or <70% non-functional transcripts. Incorporation of RNA splicing evidence into variant interpretation increased pathogenic (28.6%-31.7%) and benign (3.7%-24.4%) classifications while reducing likely pathogenic (19.5%-17.7%), uncertain (18.9%-8.5%), and likely benign (29.3%-17.7%) categories. Experimental mRNA profiling impacted the interpretation of 34% of variants and resolved uncertainty in approximately 10% of cases. Exon 18 skipping was less tolerated, indicating that the degree of splice perturbation required to impair BRCA1 function may depend on the nature of the resulting non-functional transcript.
The question of how gene mutations of large effect and common variants of small effect relate to phenotypic variation dates from the origins of genetics. Mendelian diseases result from rare germline variants with major effects, while complex traits are associated with multiple, mostly common variants of small effect. High-dimensional phenotypes, such as facial shape, can shed new light on this age-old dichotomy, as their variation can be characterized in terms of directions in multivariate morphospace. Within such spaces, do Mendelian disease mutations move phenotypes along the same directions as common variants, or do they forge new directions that diverge from the common structure of background variation? Here, we analyze facial shape variation for 66 syndromes, quantify multivariate axes of facial shape variation for each syndrome, and test whether common genetic variants in cohorts of non-syndromic subjects are associated with phenotypic position along these same axes. We find that syndromic facial shape generally follows the background variance-covariance structure of facial shape in the general population. Furthermore, syndromic probands’ unaffected relatives have subtle facial morphology resembling the syndromes of their affected relatives. These results suggest that Mendelian disease variants act on facial shape in ways similar to common variants. Syndromic probands with higher “severity” likely occur on genetic backgrounds with higher cumulative severity of common variants for each syndromic axis. These findings position Mendelian diseases at extremes along phenotypic continua that exist in the background population rather than as qualitatively different phenotypes distinct from the overall structure of normal human phenotypic variation.
Nine-year-old brain tumor patient Gabriella Miller challenged members of Congress to “stop talking and start doing” when providing federal funding for research into cures for pediatric cancer and congenital anomalies. Though she ultimately lost her life to that cancer, her advocacy efforts resulted in the 2014 Gabriella Miller Kids First Research Act, launching the Gabriella Miller Kids First Pediatric Research Program at the National Institutes of Health (NIH). The overarching goal of the Gabriella Miller Kids First Pediatric Research Program is to help researchers uncover new insights into the biology of childhood cancer and congenital anomalies. Following the signing of the Gabriella Miller Kids First Research Act 2.0 in January 2025, the program has been extended at NIH through 2028 to advance the groundwork laid in the program’s first ten years. The Gabriella Miller Kids First Data Resource Center has since honored her legacy by building a comprehensive data resource for genomic research into pediatric conditions. Data from more than 30,000 participants annotated with demographic and clinical information related to their diagnoses have been released for secondary research and analysis using the center’s web-based platforms. This paper analyzes the outcomes of the initiative and highlights breakthroughs made by the larger research community resulting from the availability of this data resource. We explore the future expansion of the data resource to include new modalities and tools for supporting life-saving research for children like Gabriella Miller.
The sharing of data generated by clinical genetic and genomic testing without explicit consent is important for timely diagnosis and treatment. While many jurisdictions permit the sharing of identifiable data for direct clinical care, institutional policies vary in how clearly they specify key elements, including when sharing is permitted, what data are covered, and what safeguards apply. Greater clarity around these elements may support responsible data sharing while balancing timely care with transparency and appropriate protections. We conducted a mixed-methods content analysis of data-sharing and privacy policies from 33 clinical genomic institutions across 17 countries and regions. Using a predefined analytical framework, we assessed how policies document key governance elements relevant to sharing without explicit consent. Two independent reviewers extracted information about clinical contexts, data types, justifications, and protections. Although 70% of institutions described circumstances permitting data sharing without explicit consent, most policies did not clearly define the scope or governance of such sharing. Policies also rarely distinguished clinical from research or secondary use and inconsistently specified privacy and security safeguards. While sharing was commonly justified for clinical care (78.3%) or testing services (43.5%), data recipient roles and onward-sharing expectations were often left undefined. This uneven documentation could make it difficult for clinical teams and institutional decision-makers to identify and justify decisions about what is permitted and under what conditions. A guidance framework specifying core governance elements and corresponding protections could help institutions communicate their governance choices more clearly and support comparable baseline practices for responsible data sharing.
The genetics of complex traits in Africa has been historically understudied, which can contribute to healthcare inequalities. Here, we present observations of 27 anthropometric, cardiovascular, and blood biomarker measurements across 2,124 individuals from sub-Saharan Africa for whom we also have dense genotype data. First, we identified trait values that differ significantly across populations and subsistence lifestyles (e.g., hemoglobin levels and height). We then identified traits with high degrees of sexual dimorphism (e.g., weight and grip strength). ADMIXTURE analyses revealed substantial population structure in our dataset, and many of the phenotypes studied here are correlated with genetic ancestry components, particularly skin color and body size traits. A variance partitioning approach further revealed traits in which much of the SNP heritability is due to polymorphisms that also contribute to differences between ancestry components. Following genomic imputation, we performed genome-wide association studies (GWASs) for all 27 traits and identified >100 independent autosomal SNPs with genome-wide significant associations for at least one trait (p < 5 × 10-8). Many of these trait-associated variants are rare outside of Africa (minor-allele frequency [MAF] < 1%). We found that 100 kb windows surrounding the top GWAS hits from our African-ancestry cohort were enriched for trait associations in an identically sized European cohort and vice versa. We performed a more detailed analysis of height prediction from genetic data, finding that genome-wide admixture proportions predict height in Africans better than polygenic predictors based on large-scale European height GWASs.
RNA-binding proteins (RBPs) regulate gene expression, and a number of RBPs have been implicated in brain function and behavior. Here, we report 16 individuals with a neurodevelopmental disorder and de novo heterozygous variants in ELAVL2, encoding an RBP not previously linked to Mendelian disease. Thirteen individuals were identified through GeneMatcher. Their ELAVL2 variants include two structural, five nonsense, and six missense variants, supporting haploinsufficiency as the primary disease mechanism. The cohort presented with developmental delay, intellectual disability, autism spectrum disorder, seizures, sleep problems, sensory processing issues, emotional instability, and difficulty with socialization. Three additional variants (two missense and one terminal exon truncation), each previously reported in a different large cohort study, were also included for follow-up investigations. We provide multiple lines of evidence linking variants in ELAVL2 to the observed neurodevelopmental and behavioral phenotypes. First, we show that common genetic variants in ELAVL2 are significantly associated with intelligence, motor development, sleep-related traits, and sociability in the general population. Drosophila loss-of-function models provide further independent evidence for a conserved role in the regulation of seizure-like behavior, sensory processing, and sleep. Molecular studies confirm that some of the missense variants are deleterious, leading to decreased protein levels. Together, our integrative study combining Mendelian genetics, clinical and association studies, and animal and molecular modeling supports variants in ELAVL2 as a cause of a neurodevelopmental disorder, with haploinsufficiency as the disease mechanism, and identifies crucial roles of ELAVL2 in neuronal function, cognition, and behavior.
The capacity of cells to proliferate and survive is central to development and disease. Assays that measure cell fitness are therefore a cornerstone of biology, but traditional techniques lack donor diversity and have high technical variability that impedes scale and reproducibility. To overcome these barriers, we designed and validated a "cell village"-based fitness screening approach using pooled cultures of 12-39 genetically distinct human neural progenitor cell (NPC) lines. We also developed Townlet to establish a foundational statistical framework based on Dirichlet regression for analyzing proportional data from cell villages. Applying these systems, we identified hyperproliferation in NPCs harboring the autism risk factor chromosome 16p11.2 deletion, mapped common genetic variants near ZFHX3 associated with NPC proliferation rate, and discovered genetic modifiers of lead (Pb) sensitivity implicating ARNT2. Together, these experimental and analytical tools advance a scalable, genetically diverse in vitro platform for dissecting human variation in cell fitness and gene-environment interactions.
Pathogenic variants in TP53, the key tumor suppressor gene underlying Li-Fraumeni syndrome (LFS), are among the best-established causes of inherited cancer predisposition. However, large-scale sequencing has revealed that many apparently pathogenic TP53 variants detected in blood are the result of somatic clonal expansions, complicating risk interpretation. Using blood-derived whole-exome data from 469,391 UK Biobank participants, we combined the variant allele fraction (VAF) with haplotype-sharing analysis to distinguish germline and somatic TP53 variants. Germline variants were concentrated at sites linked to partial loss of p53 function and lower disease penetrance, whereas classic LFS alleles appeared to be predominantly somatically acquired. Classic LFS alleles at high VAF conferred markedly increased risk of hematological malignancy but not solid tumors, indicating an important contribution from large TP53-mutant clonal expansions. The prevalence of somatic clonal expansion also correlated with missense variant pathogenicity, suggesting that somatic activity provides an informative in vivo proxy for functional impact. These results provide new insights into TP53-associated cancer risk at the population level, demonstrate that somatic rather than germline risk predominates in middle-aged healthy adults, and provide a scalable framework for variant classification in large-scale population genomics.
DNA methylation (DNAm) episignatures are stable disorder-specific epigenetic patterns that serve as valuable biomarkers for assessing variant pathogenicity and phenotypic outcomes in neurodevelopmental disorders (NDDs). However, episignatures derived from whole blood are inherently tissue- and cell-type specific, limiting their applicability in prenatal diagnostics. To explore the feasibility of developing episignatures capable of informing variant pathogenicity across tissues and developmental stages, we conducted a proof-of-concept study using Down syndrome, a common NDD caused by trisomy 21 (T21). We generated a blood-derived T21 episignature using a large cohort of 266 samples. Next, we used that episignature and publicly available DNAm data for 850 T21 and control samples across six different pre- and postnatal tissues to train machine-learning models, thereby enabling accurate prediction of T21 status across all tested tissues. Notably, our results show that models trained on postnatal blood-derived signatures as well as other tissues can generate cell-type-agnostic disease-specific patterns. This method also supported integrating well-characterized postnatal episignatures with a limited set of prenatal samples, which generated an episignature capable of accurately classifying prenatal samples. This approach forges a path for cross-tissue DNAm biomarker development and lays the groundwork for a workflow to rapidly integrate episignatures into prenatal diagnostics.
Pathogenic rewiring of the three-dimensional (3D) genome architecture is increasingly being identified as the cause of genetic diseases, but recognizing the cis-regulatory effects of structural variation remains a challenge. The Xq27.1 region contains a quasi-palindrome identified as a pleiotropic hotspot for disease-causing interchromosomal insertions. In a large Danish family affected by X-linked recessive complex spastic paraplegia, we identified the segregation of a 149-kb interchromosomal insertion at Xq27.1 originating from 4q24. To understand the disease mechanism, we generated induced pluripotent stem cells (iPSCs) from affected individuals. Using CRISPR perturbation and neural differentiation experiments combined with high-throughput chromatin conformation capture (Hi-C) and transcriptomic analyses, we identify a 3D regulatory rewiring of SOX3 and transcriptional dysregulation of SOX3 targets in iPSC-derived neurons. Consistent with regulatory partitioning of the SOX3 topologically associating domain (TAD) in affected individuals, our experiments show that upstream cis-regulatory elements have a reduced ability to activate SOX3 expression and that the observed dysregulation depends on CTCF-binding sites within the insertion. This work provides mechanistic evidence that a position effect at the SOX3 locus can cause hereditary spastic paraplegia.
Understanding how selection shapes disease risk remains challenging. Variants influencing complex traits, including common diseases, can also impact fitness and thus be constrained by purifying selection. Consequently, genetic variance underlying disease susceptibility may be attributed to low-frequency, population-specific variants. We analyzed 509,817 genome-wide variants from 72,635 Han Taiwanese individuals to identify loci showing age-dependent allele frequency shifts that signal ongoing selection. After adjusting for potential age-related population structure, we detected 168 variants deviating from neutrality, with most showing declining frequencies in younger generations, consistent with purifying selection on deleterious alleles influencing disease risk. These variants were enriched for rare alleles (≤0.1%) and disease-associated variants. At BRCA1, we identified 16 rare pathogenic variants in strong linkage disequilibrium undergoing purifying selection that coexist with a positively selected haplotype, revealing temporally fluctuating selection; comparable patterns at BRCA2 and MLH1 suggest recurrent selective trade-offs in DNA repair genes. Phenome-wide association analysis across 30 hematologic and cardiometabolic traits linked a subset of candidates to increased erythrocyte volume and reduced hemoglobin concentration, suggesting subclinical physiological effects. These results demonstrate ongoing natural selection on disease-relevant variation, particularly affecting hematologic traits in the Han Taiwanese population, and highlight opportunities to refine precision-medicine risk models.
Whether polygenic risk, monogenic familial hypercholesterolemia (FH), and family history (FamHx) are additively informative for coronary heart disease (CHD) risk prediction across self-identified race/ethnicity (SIRE) groups has not been established. In two diverse cohorts-Electronic Medical Records and Genomics (eMERGE) phase IV (eIV; n = 19,348) and All of Us (AoU; n = 239,645)-we quantified the associations of a polygenic risk score (PRSCHD), pathogenic/likely pathogenic variants in genes associated with FH, and FamHx with CHD and evaluated their incremental value when added to the pooled cohort equations (PCEs). CHD was defined as myocardial infarction, unstable angina, or coronary revascularization. We modeled associations with multivariable logistic regression (prevalent CHD in eIV) and Cox proportional hazards (incident CHD in AoU) and characterized predictive performance with the c-statistic and reclassification and decision-curve net benefits across actionable 10-year risk thresholds. The effects of PRSCHD and FamHx were independent and additive in both cohorts and consistent across White, Black, and Latino SIRE groups. In eIV, adding PRSCHD and FamHx to the PCE increased the c-statistic for prevalent CHD from 0.719 to 0.753 (p-diff = 9.1 × 10-3) and reclassified 18.8% of participants at the 7.5% 10-year threshold, yielding approximately 4 additional true-positive CHD identifications per 1,000 screened. Net benefit gains were observed between the 7.5% and 10% thresholds across all three SIRE groups. In conclusion, PRSCHD and FamHx were independently and additively associated with CHD across major SIRE groups in two diverse cohorts in the United States (US), motivating the addition of these factors to clinical risk algorithms.
Early postzygotic mutations (PZMs) that arise after fertilization but prior to primordial germ cell specification may be present in both somatic and germ cells, causing mosaicism in a parent and constitutive inheritance in their offspring. In clinical family-trio whole-genome sequencing (WGS), such variants are systematically missed because their sub-heterozygous variant allele fraction (VAF) prevents heterozygous calling in the parent, while residual parental allele support disqualifies the variant as a candidate germline de novo mutation (DNM) in the child. Here, we developed a bioinformatic approach to ascertain parental PZMs from unfiltered DNM candidates in standard-depth (∼30×) trio WGS and applied it to 12,015 trios from the Genomics England 100,000 Genomes Project. We identified 1,015 high-confidence early autosomal parental PZMs, a large single-source catalog of this mutation class. These exhibited a monomodal VAF distribution centered around 5% in parental blood, consistent with empirically characterized ascertainment boundaries imposed by standard-depth sequencing and germline variant calling. PZMs showed no parental age or sex bias and displayed a mutational spectrum distinct from that of DNMs, with enrichment for C>A and T>A substitutions and depletion of T>C. Mutational signature analysis revealed that both mutation types are shaped by clock-like signatures SBS1 and SBS5 in similar proportions, suggesting that spectral differences reflect shifts within shared mutagenic processes. Exploratory genomic distribution analysis revealed a negative PZM association with GC content, in contrast to the positive association for DNMs. Among these, we found variants in DYNC1H1 and WT1 with potential clinical relevance that were missed by routine diagnostic pipelines.