Externalizing spectrum disorders-spanning attention-deficit/hyperactivity disorder, conduct disorder, substance use disorders, and other disorders characterized by disinhibition-frequently co-occur within individuals due, in part, to shared genetic etiology. To advance understanding of this genetic architecture, we conducted a multi-ancestry, multivariate genome-wide association analysis of more than 4 million individuals, identifying 1,294 genomic regions linked to an externalizing factor. Fine-mapping and gene prioritization efforts identified 961 effector genes, with the putative causal variant associations showing robust replication in the All of Us Research Program sample. Bioinformatic analyses revealed a broadly distributed neural architecture with early and sustained involvement of GABAergic and glutamatergic neurons. Drug repurposing analyses further highlighted the role of GABAA receptors, as well as dopaminergic signaling, excitatory-inhibitory balance, and neurosteroid pathways. A genome-wide polygenic index predicted ~12% of the variance in externalizing in independent cohorts of individuals with European-like ancestry, compared to ~3% in individuals with African-like ancestry, and was associated with myriad health and life outcomes. Together, these findings map the shared genetic etiology of externalizing psychopathology and identify neurodevelopmental and synaptic mechanisms with translational relevance.
We conducted a genome-wide association study on income among individuals of European descent (N = 668,288) to investigate the relationship between socio-economic status and health disparities. We identified 162 genomic loci associated with a common genetic factor underlying various income measures, all with small effect sizes (the Income Factor). Our polygenic index captures 1-5% of income variance, with only one fourth due to direct genetic effects. A phenome-wide association study using this index showed reduced risks for diseases including hypertension, obesity, type 2 diabetes, depression, asthma and back pain. The Income Factor had a substantial genetic correlation (0.92, s.e. = 0.006) with educational attainment. Accounting for the genetic overlap of educational attainment with income revealed that the remaining genetic signal was linked to better mental health but reduced physical health and increased risky behaviours such as drinking and smoking. These findings highlight the complex genetic influences on income and health.
Behaviors and disorders related to self-regulation, such as substance use, antisocial behavior and attention-deficit/hyperactivity disorder, are collectively referred to as externalizing and have shared genetic liability. We applied a multivariate approach that leverages genetic correlations among externalizing traits for genome-wide association analyses. By pooling data from ~1.5 million people, our approach is statistically more powerful than single-trait analyses and identifies more than 500 genetic loci. The loci were enriched for genes expressed in the brain and related to nervous system development. A polygenic score constructed from our results predicts a range of behavioral and medical outcomes that were not part of genome-wide analyses, including traits that until now lacked well-performing polygenic scores, such as opioid use disorder, suicide, HIV infections, criminal convictions and unemployment. Our findings are consistent with the idea that persistent difficulties in self-regulation can be conceptualized as a neurodevelopmental trait with complex and far-reaching social and health correlates.
Failures of self-control can manifest as externalizing behaviors (e.g., aggression, rule-breaking) that have far-reaching negative consequences. Researchers have long been interested in measuring children's genetic risk for externalizing behaviors to inform efforts at early identification and intervention. Drawing on data from the Environmental Risk Longitudinal Twin Study (N = 862 twins) and the Millennium Cohort Study (N = 2,824 parent-child trios), two longitudinal cohorts from the UK, we leveraged molecular genetic data and within-family designs to test for genetic associations with externalizing behavior that are not affected by common sources of environmental influence. We found that a polygenic index (PGI) calculated from genetic variants discovered in previous studies of self-controlled behavior in adults captures direct genetic effects on externalizing problems in children and adolescents when evaluated with rigorous within-family designs (β's = 0.13-0.19 across development). The externalizing behavior PGI can usefully augment psychological studies of the development of self-control.
Behaviors and disorders characterized by difficulties with self-regulation, such as problematic substance use, antisocial behavior, and symptoms of attention-deficit/hyperactivity disorder (ADHD), incur high costs for individuals, families, and communities. These externalizing behaviors often appear early in the life course and can have far-reaching consequences. Researchers have long been interested in direct measurements of genetic risk for externalizing behaviors, which can be incorporated alongside other known risk factors to improve efforts at early identification and intervention. In a preregistered analysis drawing on data from the Environmental Risk (E-Risk) Longitudinal Twin Study (N=862 twins) and the Millennium Cohort Study (MCS; N=2,824 parent-child trios), two longitudinal cohorts from the UK, we leveraged molecular genetic data and within-family designs to test for genetic effects on externalizing behavior that are unbiased by the common sources of environmental confounding. Results are consistent with the conclusion that an externalizing polygenic index (PGI) captures causal effects of genetic variants on externalizing problems in children and adolescents, with an effect size that is comparable to those observed for other established risk factors in the research literature on externalizing behavior. Additionally, we find that polygenic associations vary across development (peaking from age 5-10 years), that parental genetics (assortment and parent-specific effects) and family-level covariates affect prediction little, and that sex differences in polygenic prediction are present but only detectable using within-family comparisons. Based on these findings, we believe that the PGI for externalizing behavior is a promising means for studying the development of disruptive behaviors across child development.
Mediation analysis is commonly used to identify mechanisms and intermediate factors between causes and outcomes. Studies drawing on polygenic scores (PGSs) can readily employ traditional regression-based procedures to assess whether traitMmediates the relationship between the genetic component of outcomeYand outcomeYitself. However, this approach suffers from attenuation bias, as PGSs capture only a (small) part of the genetic variance of a given trait. To overcome this limitation, we developed MA-GREML: a method for Mediation Analysis using Genome-based Restricted Maximum Likelihood (GREML) estimation.Using MA-GREML to assess mediation between genetic factors and traits comes with two main advantages. First, we circumvent the limited predictive accuracy of PGSs that regression-based mediation approaches suffer from. Second, compared to methods employing summary statistics from genome-wide association studies, the individual-level data approach of GREML allows to directly control for confounders of the association betweenMandY. In addition to typical GREML parameters (e.g., the genetic correlation), MA-GREML estimates (i) the effect ofMonY, (ii) thedirect effect(i.e., the genetic variance ofYthat is not mediated byM), and (iii) theindirect effect(i.e., the genetic variance ofYthat is mediated byM). MA-GREML also provides standard errors of these estimates and assesses the significance of the indirect effect.We use analytical derivations and simulations to show the validity of our approach under two main assumptions,viz., thatMprecedesYand that environmental confounders of the association betweenMandYare controlled for. We conclude that MA-GREML is an appropriate tool to assess the mediating role of traitMin the relationship between the genetic component ofYand outcomeY. Using data from the US Health and Retirement Study, we provide evidence that genetic effects on Body Mass Index (BMI), cognitive functioning and self-reported health in later life run partially through educational attainment. For mental health, we do not find significant evidence for an indirect effect through educational attainment. Further analyses show that the additive genetic factors of these four outcomes do partially (cognition and mental health) and fully (BMI and self-reported health) run through an earlier realization of these traits.
Measurement error in polygenic indices (PGIs) attenuates the estimation of their effects in regression models. We analyze and compare two approaches addressing this attenuation bias: Obviously Related Instrumental Variables (ORIV) and the PGI Repository Correction (PGI-RC). Through simulations, we show that the PGI-RC performs slightly better than ORIV, unless the prediction sample is very small ( N < 1000) or when there is considerable assortative mating. Within families, ORIV is the best choice since the PGI-RC correction factor is generally not available. We verify the empirical validity of the simulations by predicting educational attainment and height in a sample of siblings from the UK Biobank. We show that applying ORIV between families increases the standardized effect of the PGI by 12% (height) and by 22% (educational attainment) compared to a meta-analysis-based PGI, yet estimates remain slightly below the PGI-RC estimates. Furthermore, within-family ORIV regression provides the tightest lower bound for the direct genetic effect, increasing the lower bound for the standardized direct genetic effect on educational attainment from 0.14 to 0.18 (+29%), and for height from 0.54 to 0.61 (+13%) compared to a meta-analysis-based PGI.
Abstract Background Heritability and genetic correlation can be estimated from genome-wide single-nucleotide polymorphism (SNP) data using various methods. We recently developed multivariate genomic-relatedness-based restricted maximum likelihood (MGREML) for statistically and computationally efficient estimation of SNP-based heritability ( $$h^2_{\text{SNP}}$$ h SNP 2 ) and genetic correlation ( $$\rho _G$$ ρ G ) across many traits in large datasets. Here, we extend MGREML by allowing it to fit and perform tests on user-specified factor models, while preserving the low computational complexity. Results Using simulations, we show that MGREML yields consistent estimates and valid inferences for such factor models at low computational cost (e.g., for data on 50 traits and 20,000 individuals, a saturated model involving 50 $$h^2_{\text{SNP}}$$ h SNP 2 ’s, 1225 $$\rho _G$$ ρ G ’s, and 50 fixed effects is estimated and compared to a restricted model in less than one hour on a single notebook with two 2.7 GHz cores and 16 GB of RAM). Using repeated measures of height and body mass index from the US Health and Retirement Study, we illustrate the ability of MGREML to estimate a factor model and test whether it fits the data better than a nested model. The MGREML tool, the simulation code, and an extensive tutorial are freely available at https://github.com/devlaming/mgreml/ . Conclusion MGREML can now be used to estimate multivariate factor structures and perform inferences on such factor models at low computational cost. This new feature enables simple structural equation modeling using MGREML, allowing researchers to specify, estimate, and compare genetic factor models of their choosing using SNP data.
Understanding which biological pathways are specific versus general across diagnostic categories and levels of symptom severity is critical to improving nosology and treatment of psychopathology. Here, we combine transdiagnostic and dimensional approaches to genetic discovery for the first time, conducting a novel multivariate genome-wide association study of eight psychiatric symptoms and disorders broadly related to mood disturbance and psychosis. We identify two transdiagnostic genetic liabilities that distinguish between common forms of psychopathology versus rarer forms of serious mental illness. Biological annotation revealed divergent genetic architectures that differentially implicated prenatal neurodevelopment and neuronal function and regulation. These findings inform psychiatric nosology and biological models of psychopathology, as they suggest that the severity of mood and psychotic symptoms present in serious mental illness may reflect a difference in kind rather than merely in degree.
Human variation in brain morphology and behavior are related and highly heritable. Yet, it is largely unknown to what extent specific features of brain morphology and behavior are genetically related. Here, we introduce a computationally efficient approach for multivariate genomic-relatedness-based restricted maximum likelihood (MGREML) to estimate the genetic correlation between a large number of phenotypes simultaneously. Using individual-level data (N = 20,190) from the UK Biobank, we provide estimates of the heritability of gray-matter volume in 74 regions of interest (ROIs) in the brain and we map genetic correlations between these ROIs and health-relevant behavioral outcomes, including intelligence. We find four genetically distinct clusters in the brain that are aligned with standard anatomical subdivision in neuroscience. Behavioral traits have distinct genetic correlations with brain morphology which suggests trait-specific relevance of ROIs. These empirical results illustrate how MGREML can be used to estimate internally consistent and high-dimensional genetic correlation matrices in large datasets.
Joint diagonalization of a set of positive (semi)-definite matrices has a wide range of analytical applications, such as estimation of common principal components, estimation of multiple variance components, and blind signal separation. However, when the eigenvectors of involved matrices are not the same, joint diagonalization is a computationally challenging problem. To the best of our knowledge, currently existing methods require at least $O(KN^3)$ time per iteration, when $K$ different $N \times N$ matrices are considered. We reformulate this optimization problem by applying orthogonality constraints and dimensionality reduction techniques. In doing so, we reduce the computational complexity for joint diagonalization to $O(N^3)$ per quasi-Newton iteration. This approach we refer to as JADOC: Joint Approximate Diagonalization under Orthogonality Constraints. We compare our algorithm to two important existing methods and show JADOC has superior runtime while yielding a highly similar degree of diagonalization. The JADOC algorithm is implemented as open-source Python code, available at https://github.com/devlaming/jadoc.
We develop a polygenic index for individual income and examine random differences in this index with lifetime outcomes in a sample of ~35,000 biological siblings. We find that genetic fortune for higher income causes greater socio-economic status and better health, partly via intervenable environmental pathways such as education. The positive returns to schooling remain substantial even after controlling for now observable genetic confounds. Our findings illustrate that inequalities in education, income, and health are partly due the outcomes of a genetic lottery. However, the consequences of different genetic endowments are malleable, for example via policies that target education.
We conducted genome-wide association studies (GWAS) of relative intake from the macronutrients fat, protein, carbohydrates, and sugar in over 235,000 individuals of European ancestries. We identified 21 unique, approximately independent lead SNPs. Fourteen lead SNPs are uniquely associated with one macronutrient at genome-wide significance ( P < 5 × 10 −8 ), while five of the 21 lead SNPs reach suggestive significance ( P < 1 × 10 −5 ) for at least one other macronutrient. While the phenotypes are genetically correlated, each phenotype carries a partially unique genetic architecture. Relative protein intake exhibits the strongest relationships with poor health, including positive genetic associations with obesity, type 2 diabetes, and heart disease ( r g ≈ 0.15–0.5). In contrast, relative carbohydrate and sugar intake have negative genetic correlations with waist circumference, waist-hip ratio, and neighborhood deprivation (| r g | ≈ 0.1–0.3) and positive genetic correlations with physical activity ( r g ≈ 0.1 and 0.2). Relative fat intake has no consistent pattern of genetic correlations with poor health but has a negative genetic correlation with educational attainment ( r g ≈−0.1). Although our analyses do not allow us to draw causal conclusions, we find no evidence of negative health consequences associated with relative carbohydrate, sugar, or fat intake. However, our results are consistent with the hypothesis that relative protein intake plays a role in the etiology of metabolic dysfunction.
Meningiomas are common tumors in adults, which develop from the meningeal coverings of the brain and spinal cord. Loss-of-function mutations or deletion of the NF2 gene, resulting in loss of the encoded Merlin protein, lead to Neurofibromatosis type 2 (NF2), but also cause the formation of sporadic meningiomas. It was shown that inactivation of Nf2 in mice caused meningioma formation. Another meningioma tumor-suppressor candidate is the receptor-like density-enhanced phosphatase-1 (DEP-1), encoded by PTPRJ. Loss of DEP-1 enhances meningioma cell motility in vitro and invasive growth in an orthotopic xenograft model. Ptprj-deficient mice develop normally and do not show spontaneous tumorigenesis. Another genetic lesion may be required to interact with DEP-1 loss in meningioma genesis.In the present study we investigated in vitro and in vivo whether the losses of DEP-1 and Merlin/NF2 may have a combined effect.Human meningioma cells deficient for DEP-1, Merlin/NF2 or both showed no statistically significant changes in cell proliferation, while DEP-1 or DEP1/NF2 deficiency led to moderately increased colony size in clonogenicity assays. In addition, the loss of any of the two genes was sufficient to induce a significant reduction of cell size (p < .05) and profound morphological changes. Most important, in Ptprj knockout mice Cre/lox mediated meningeal Nf2 knockout elicited a four-fold increased rate of meningioma formation within one year compared with mice with Ptprj wild type alleles (25% vs 6% tumor incidence).Our data suggest that loss of DEP-1 and Merlin/NF2 synergize during meningioma genesis.
Genetic correlations estimated from genome-wide association studies (GWASs) reveal pervasive pleiotropy across a wide variety of phenotypes. We introduce genomic structural equation modelling (genomic SEM): a multivariate method for analysing the joint genetic architecture of complex traits. Genomic SEM synthesizes genetic correlations and single-nucleotide polymorphism heritabilities inferred from GWAS summary statistics of individual traits from samples with varying and unknown degrees of overlap. Genomic SEM can be used to model multivariate genetic associations among phenotypes, identify variants with effects on general dimensions of cross-trait liability, calculate more predictive polygenic scores and identify loci that cause divergence between traits. We demonstrate several applications of genomic SEM, including a joint analysis of summary statistics from five psychiatric traits. We identify 27 independent single-nucleotide polymorphisms not previously identified in the contributing univariate GWASs. Polygenic scores from genomic SEM consistently outperform those from univariate GWASs. Genomic SEM is flexible and open ended, and allows for continuous innovation in multivariate genetic analysis.