Data within biobanks capture broad yet detailed indices of human variation, but biobank-wide insights can be difficult to extract due to complexity and scale. Here, using large-scale factor analysis, we distill hundreds of variables (diagnoses, assessments and survey items) into 35 latent constructs, using data from unrelated individuals with predominantly estimated European genetic ancestry in UK Biobank. These factors recapitulate known disease classifications, disentangle elements of socioeconomic status, highlight the relevance of psychiatric constructs to health and improve measurement of pro-health behaviours. We go on to demonstrate the power of this approach to clarify genetic signal, enhance discovery and identify associations between underlying phenotypic structure and health outcomes. In building a deeper understanding of ways in which constructs such as socioeconomic status, trauma, or physical activity are structured in the dataset, we emphasize the importance of considering the interwoven nature of the human phenome when evaluating public health patterns. Carey and colleagues reveal 35 major latent constructs (factors) in the phenotype data of unrelated individuals with predominantly estimated European genetic ancestry from UK Biobank.
Classical statistical genetics theory defines dominance as any deviation from a purely additive, or dosage, effect of a genotype on a trait, which is known as the dominance deviation. Dominance is well documented in plant and animal breeding. Outside of rare monogenic traits, however, evidence in humans is limited. We systematically examined common genetic variation across 1060 traits in a large population cohort (UK Biobank, N = 361,194 samples analyzed) for evidence of dominance effects. We then developed a computationally efficient method to rapidly assess the aggregate contribution of dominance deviations to heritability. Lastly, observing that dominance associations are inherently less correlated between sites at a genomic locus than their additive counterparts, we explored whether they may be leveraged to identify causal variants more confidently.
Broad yet detailed data collected in biobanks captures variation reflective of human health and behavior, but insights are hard to extract given their complexity and scale. In the largest factor analysis to date, we distill hundreds of medical record codes, physical assays, and survey items from UK Biobank into 35 understandable latent constructs. The identified factors recapitulate known disease classifications, highlight the relevance of psychiatric constructs, improve measurement of health-related behavior, and disentangle elements of socioeconomic status. We demonstrate the power of this principled data reduction approach to clarify genetic signal, enhance discovery, and identify associations between underlying phenotypic structure and health outcomes such as mortality. We emphasize the importance of considering the interwoven nature of the human phenome when evaluating large-scale patterns relevant to public health.
Both mild and severe epilepsies are influenced by variants in the same genes, yet an explanation for the resulting phenotypic variation is unknown. As part of the ongoing Epi25 Collaboration, we performed a whole-exome sequencing analysis of 13,487 epilepsy-affected individuals and 15,678 control individuals. While prior Epi25 studies focused on gene-based collapsing analyses, we asked how the pattern of variation within genes differs by epilepsy type. Specifically, we compared the genetic architectures of severe developmental and epileptic encephalopathies (DEEs) and two generally less severe epilepsies, genetic generalized epilepsy and non-acquired focal epilepsy (NAFE). Our gene-based rare variant collapsing analysis used geographic ancestry-based clustering that included broader ancestries than previously possible and revealed novel associations. Using the missense intolerance ratio (MTR), we found that variants in DEE-affected individuals are in significantly more intolerant genic sub-regions than those in NAFE-affected individuals. Only previously reported pathogenic variants absent in available genomic datasets showed a significant burden in epilepsy-affected individuals compared with control individuals, and the ultra-rare pathogenic variants associated with DEE were located in more intolerant genic sub-regions than variants associated with non-DEE epilepsies. MTR filtering improved the yield of ultra-rare pathogenic variants in affected individuals compared with control individuals. Finally, analysis of variants in genes without a disease association revealed a significant burden of loss-of-function variants in the genes most intolerant to such variation, indicating additional epilepsy-risk genes yet to be discovered. Taken together, our study suggests that genic and sub-genic intolerance are critical characteristics for interpreting the effects of variation in genes that influence epilepsy.
To date, the cellular and molecular mechanisms underlying sexual dimorphism in the incidence, prognosis, and treatment responses of cancer remain unclear. In a recent article published in Cancer Cell, Yuan et al. applied a pan-cancer analysis to identify sex-biased molecular signatures and revealed two sex-effect groups characterized by distinct incidence and mortality profiles.
To discover novel genes underlying amyotrophic lateral sclerosis (ALS), we aggregated exomes from 3,864 cases and 7,839 ancestry-matched controls. We observed a significant excess of rare protein-truncating variants among ALS cases, and these variants were concentrated in constrained genes. Through gene level analyses, we replicated known ALS genes including SOD1, NEK1 and FUS. We also observed multiple distinct protein-truncating variants in a highly constrained gene, DNAJC7. The signal in DNAJC7 exceeded genome-wide significance, and immunoblotting assays showed depletion of DNAJC7 protein in fibroblasts in a patient with ALS carrying the p.Arg156Ter variant. DNAJC7 encodes a member of the heat-shock protein family, HSP40, which, along with HSP70 proteins, facilitates protein homeostasis, including folding of newly synthesized polypeptides and clearance of degraded proteins. When these processes are not regulated, misfolding and accumulation of aberrant proteins can occur and lead to protein aggregation, which is a pathological hallmark of neurodegeneration. Our results highlight DNAJC7 as a novel gene for ALS.
ABSTRACT Bipolar disorder is a highly heritable psychiatric disorder that features episodes of mania and depression. We performed the largest genome-wide association study to date, including 20,352 cases and 31,358 controls of European descent, with follow-up analysis of 822 sentinel variants at loci with P<1×10 -4 in an independent sample of 9,412 cases and 137,760 controls. In the combined analysis, 30 loci reached genome-wide significant evidence for association, of which 20 were novel. These significant loci contain genes encoding ion channels and neurotransmitter transporters ( CACNA1C , GRIN2A , SCN2A , SLC4A1 ), synaptic components ( RIMS1 , ANK3 ), immune and energy metabolism components. Bipolar disorder type I (depressive and manic episodes; ~ 73% of our cases) is strongly genetically correlated with schizophrenia whereas bipolar disorder type II (depressive and hypomanic episodes; ~ 17% of our cases) is more strongly correlated with major depressive disorder. These findings address key clinical questions and provide potential new biological mechanisms for bipolar disorder.
Sequencing-based studies have identified novel risk genes for rare, severe epilepsies and revealed a role of rare deleterious variation in common epilepsies. To identify the shared and distinct ultra-rare genetic risk factors for rare and common epilepsies, we performed a whole-exome sequencing (WES) analysis of 9,170 epilepsy-affected individuals and 8,364 controls of European ancestry. We focused on three phenotypic groups; the rare but severe developmental and epileptic encephalopathies (DEE), and the commoner phenotypes of genetic generalized epilepsy (GGE) and non-acquired focal epilepsy (NAFE). We observed that compared to controls, individuals with any type of epilepsy carried an excess of ultra-rare, deleterious variants in constrained genes and in genes previously associated with epilepsy, with the strongest enrichment seen in DEE and the least in NAFE. Moreover, we found that inhibitory GABAA receptor genes were enriched for missense variants across all three classes of epilepsy, while no enrichment was seen in excitatory receptor genes. The larger gene groups for the GABAergic pathway or cation channels also showed a significant mutational burden in DEE and GGE. Although no single gene surpassed exome-wide significance among individuals with GGE or NAFE, highly constrained genes and genes encoding ion channels were among the top associations, including CACNA1G, EEF1A2 , and GABRG2 for GGE and LGI1, TRIM3 , and GABRG2 for NAFE. Our study confirms a convergence in the genetics of common and rare epilepsies associated with ultra-rare coding variation and highlights a ubiquitous role for GABAergic inhibition in epilepsy etiology in the largest epilepsy WES study to date.
To discover novel genetic risk factors underlying amyotrophic lateral sclerosis (ALS), we aggregated exomes from 3,864 cases and 7,839 ancestry matched controls. We observed a significant excess of ultra-rare and rare protein-truncating variants (PTV) among ALS cases, which was primarily concentrated in constrained genes; however, a significant enrichment in PTVs does persist in the remaining exome. Through gene level analyses, known ALS genes, SOD1, NEK1 , and FUS , were the most strongly associated with disease status. We also observed suggestive statistical evidence for multiple novel genes including DNAJC7 , which is a highly constrained gene and a member of the heat shock protein family (HSP40). HSP40 proteins, along with HSP70 proteins, facilitate protein homeostasis, such as folding of newly synthesized polypeptides, and clearance of degraded proteins. When these processes are not regulated, misfolding and accumulation of degraded proteins can occur leading to aberrant protein aggregation, one of the pathological hallmarks of neurodegeneration.