The major anxiety disorders (ANX; including generalized anxiety disorder, panic disorder and phobias) are highly prevalent, often onset early and cause substantial global disability. Although distinct in their clinical presentations, they probably represent differential expressions of a dysregulated threat-response system. Here, we present a genome-wide association meta-analysis comprising 122,341 European ancestry ANX cases and 729,881 controls. We identified 58 independent genome-wide significant risk variants and 66 genes with robust biological support. In an independent sample of 1,175,012 self-report ANX cases and 1,956,379 controls, 51 out of the 58 associations replicated. As predicted by twin studies, we found substantial genetic correlation between ANX and depression, neuroticism and other internalizing phenotypes. Follow-up analyses demonstrated enrichment in all major brain regions and highlighted GABAergic signaling as one potential mechanism implicated in ANX genetic risk. These results advance our understanding of the genetic architecture of ANX and prioritize genes for functional follow-up studies.
Genetics as a science has roots in studying phenotypes of relatives, but molecular approaches facilitate direct measurements of genomic variation between individuals. Agricultural and human biomedical research are both emphasizing genotype-based instruments, such as polygenic scores, but unlike in agriculture, there is an emerging consensus that family variables act nearly independently of genotypes in models of human disease. However, there is insufficient theoretical treatment of these scores, especially guiding our understanding of how and why scores derived from different sources of data may combine. To advance our understanding of this phenomenon, we use 2,066,057 family records of 99,645 genotyped probands from the Integrative Psychiatric Research (iPSYCH)2015 case-cohort study to show that state-of-the-field genotype- and phenotype-based genetic instruments explain largely independent components of liability to psychiatric disorders. We support these empirical results with theoretical analysis and simulations to describe, in a human biomedical context, parameters affecting current and future performance of the two approaches, their expected interrelationships, and consistency of observed results with expectations under simple additive, polygenic liability models of disease. We conclude, at least for psychiatric disorders, that the low correlation between current phenotype- and genotype-based genetic instruments is caused by both being noisy measures of additive genetic liability. We expect they should remain complementary over the near future and therefore expect approaches integrating both sources of information to achieve more power for genetic inference.
Genome-wide association studies (GWAS) of psychiatric disorders (PD) yield numerous loci with significant signals, but often do not implicate specific genes. Because GWAS risk loci are enriched in expression/protein/methylation quantitative loci (e/p/mQTL, hereafter xQTL), transcriptome/proteome/methylome-wide association studies (T/P/MWAS, hereafter XWAS) that integrate xQTL and GWAS information, can link GWAS signals to effects on specific genes. To further increase detection power, gene signals are aggregated within relevant gene sets (GS) by performing gene set enrichment (GSE) analyses. Often GSE methods test for enrichment of "signal" genes in curated GS while overlooking their linkage disequilibrium (LD) structure, allowing for the possibility of increased false positive rates. Moreover, no GSE tool uses xQTL information to perform mendelian randomization (MR) analysis. To make causal inference on association between PD and GS, we develop a novel MR GSE (MR-GSE) procedure. First, we generate a "synthetic" GWAS for each MSigDB GS by aggregating summary statistics for x-level (mRNA, protein or DNA methylation (DNAm) levels) from the largest xQTL studies available) of genes in a GS. Second, we use synthetic GS GWAS as exposure in a generalized summary-data-based-MR analysis of complex trait outcomes. We applied MR-GSE to GWAS of nine important PD. When applied to the underpowered opioid use disorder GWAS, none of the four analyses yielded any signals, which suggests a good control of false positive rates. For other PD, MR-GSE greatly increased the detection of GO terms signals (2,594) when compared to the commonly used (non-MR) GSE method (286). Some of the findings might be easier to adapt for treatment, e.g., our analyses suggest modest positive effects for supplementation with certain vitamins and/or omega-3 for schizophrenia, bipolar and major depression disorder patients. Similar to other MR methods, when applying MR-GSE researchers should be mindful of the confounding effects of horizontal pleiotropy on statistical inference.
MOTIVATION:As the availability of larger and more ethnically diverse reference panels grows, there is an increase in demand for ancestry-informed imputation of genome-wide association studies (GWAS), and other downstream analyses, e.g. fine-mapping. Performing such analyses at the genotype level is computationally challenging and necessitates, at best, a laborious process to access individual-level genotype and phenotype data. Summary-statistics-based tools, not requiring individual-level data, provide an efficient alternative that streamlines computational requirements and promotes open science by simplifying the re-analysis and downstream analysis of existing GWAS summary data. However, existing tools perform only disparate parts of needed analysis, have only command-line interfaces, and are difficult to extend/link by applied researchers. RESULTS:To address these challenges, we present Genome Analysis Using Summary Statistics (GAUSS)-a comprehensive and user-friendly R package designed to facilitate the re-analysis/downstream analysis of GWAS summary statistics. GAUSS offers an integrated toolkit for a range of functionalities, including (i) estimating ancestry proportion of study cohorts, (ii) calculating ancestry-informed linkage disequilibrium, (iii) imputing summary statistics of unobserved variants, (iv) conducting transcriptome-wide association studies, and (v) correcting for "Winner's Curse" biases. Notably, GAUSS utilizes an expansive, multi-ethnic reference panel consisting of 32 953 genomes from 29 ethnic groups. This panel enhances the range and accuracy of imputable variants, including the ability to impute summary statistics of rarer variants. As a result, GAUSS elevates the quality and applicability of existing GWAS analyses without requiring access to subject-level genotypic and phenotypic information. AVAILABILITY AND IMPLEMENTATION:The GAUSS R package, complete with its source code, is readily accessible to the public via our GitHub repository at https://github.com/statsleelab/gauss. To further assist users, we provided illustrative use-case scenarios that are conveniently found at https://statsleelab.github.io/gauss/, along with a comprehensive user guide detailed in Supplementary Text S1.
PTSD and AUD are frequently comorbid post-trauma outcomes. Much remains unknown about shared risk factors as PTSD and AUD work tends to be conducted in isolation. We examined how self-report measures of distress tolerance (DT), experiential avoidance (EA), and drinking motives (DM) differed across diagnostic groups in white, male combat-exposed veterans (n = 77). A MANOVA indicated a significant difference in constructs by group, F (5, 210) = 4.7, p = <.001. Follow-up ANOVAs indicated DM subscales (Coping: F (3,82) = 21.3; Social: F (3,82) = 13.1; Enhancement: F (3,82) = 10.4; ps = <.001) and EA (F (3,73) = 7.8, p < .001) differed by groups but not DT. Post hoc comparisons indicated that mean scores of the comorbid and AUD-only groups were significantly higher than controls for all DM subscales (all ps < .01). EA scores were significantly higher for the comorbid as compared to control (p < .001) and PTS-only (p = .007) groups. Findings support shared psychological factors in a comorbid PTSD-AUD population.
Post-traumatic stress disorder (PTSD) genetics are characterized by lower discoverability than most other psychiatric disorders. The contribution to biological understanding from previous genetic studies has thus been limited. We performed a multi-ancestry meta-analysis of genome-wide association studies across 1,222,882 individuals of European ancestry (137,136 cases) and 58,051 admixed individuals with African and Native American ancestry (13,624 cases). We identified 95 genome-wide significant loci (80 new). Convergent multi-omic approaches identified 43 potential causal genes, broadly classified as neurotransmitter and ion channel synaptic modulators (for example, GRIA1, GRM8 and CACNA1E), developmental, axon guidance and transcription factors (for example, FOXP2, EFNA5 and DCC), synaptic structure and function genes (for example, PCLO, NCAM1 and PDE4B) and endocrine or immune regulators (for example, ESR1, TRAF3 and TANK). Additional top genes influence stress, immune, fear and threat-related processes, previously hypothesized to underlie PTSD neurobiology. These findings strengthen our understanding of neurobiological systems relevant to PTSD pathophysiology, while also opening new areas for investigation. Multi-ancestry genome-wide analyses identify 95 loci associated with post-traumatic stress disorder and implicate candidate genes, pathways and neurobiological systems underlying its pathophysiology.
Large biobank samples provide an opportunity to integrate broad phenotyping, familial records, and molecular genetics data to study complex traits and diseases. We introduce Pearson-Aitken Family Genetic Risk Scores (PA-FGRS), a method for estimating disease liability from patterns of diagnoses in extended, age-censored genealogical records. We then apply the method to study a paradigmatic complex disorder, major depressive disorder (MDD), using the iPSYCH2015 case-cohort study of 30,949 MDD cases, 39,655 random population controls, and more than 2 million relatives. We show that combining PA-FGRS liabilities estimated from family records with molecular genotypes of probands improves three lines of inquiry. Incorporating PA-FGRS liabilities improves classification of MDD over and above polygenic scores, identifies robust genetic contributions to clinical heterogeneity in MDD associated with comorbidity, recurrence, and severity and can improve the power of genome-wide association studies. Our method is flexible and easy to use, and our study approaches are generalizable to other datasets and other complex traits and diseases.
Background Recently, gene expression (GE) imputation has become a viable alternative in studying the etiology of psychiatric phenotypes due to the substantial costs and ethical considerations associated with measuring GE in the human brain. Several popular methods for GE imputation were recently developed that capitalize on the existence of expression quantitative trait loci (eQTLs) data to predict GE in unrelated subjects using genome-wide association studies (GWAS) only. However, a primary challenge to the existing GE imputation methods is the relatively small number of accurately imputed genes, which limits their applicability in understanding the biological etiology of psychiatric disorders. To understand these limitations, we introduced several improvements to PrediXcan's elastic net methodology using eQTL data generated in the anterior cingulate cortex (ACC) of subjects with major depression (MDD), bipolar disorder (BP), and matched controls to build our own models and compare them against the off-the-shelf (OTS) PrediXcan models from predictdb.org. Our major improvement of these models is the incorporation of different neuronal cell fractions to impute with a greater level of accuracy a higher number of genes and evaluate their ability to generate biologically meaningful gene networks. Methods An ACC transcriptome dataset, which included 197 neurotypical controls, 226 MDD, and 123 BP subjects, was used to create 25,212 mRNA models with 10-fold cross-validation. To evaluate overfitting, which surprisingly most GE imputation methods do not assess, 30 % of the dataset was used for testing purposes only. We observed that incorporating estimates of neuronal cell fractions significantly increased the number of high-quality gene models. However, the accuracy of these models is heavily dependent on the robustness of those cell fraction estimates. Thus, we incorporated CATD software within our pipeline, which evaluated 28 different mRNA deconvolution methods for its ability to estimate cell fraction from pseudo-bulk ACC tissue. Weighted Gene Co-expression Network Analysis (WGCNA) was used to generate gene network modules that capture genes with similar expression profiles. The biological relevance of the gene networks to disease pathology was evaluated using Gene Set Enrichment Analysis (GSEA). Results Our large training dataset obtained 53 % more quality genotype-only gene models (N = 4,823) compared to the OTS gene models. Using the top-performing deconvolution method, 10 neuronal cell fractions were estimated and incorporated into our models, the number of quality models increased further by generating 17,430 quality gene models. The median Pearson correlation between the measured and imputed expression of the testing dataset was 0.465. Scale-free topology was achieved, and WGCNA generated 10 gene modules, with half being enriched for neurological and psychiatric-related pathways. Discussion Both our genotype-only and cell-fraction models produced significantly more quality gene models and higher imputation accuracy. With the incorporation of cell fraction estimates, enough gene models were created to conduct biologically relevant gene networks. However, given cell fraction estimates were generated using mRNA GE, its application is limited when only genotype data is available.
Background Antipsychotic medications are a mainstay of pharmacotherapy of psychiatric illnesses but have been shown to cause considerable metabolic adverse effects such as weight gain, diabetes, and hypercholesterolemia. Advancing methods to predict such effect could spur the development of Precision Psychiatry modalities. Methods We examined two cohorts of veterans receiving antipsychotics: the Corporate Data Warehouse (CDW, N=869,128) and the Million Veteran Program (MVP, N= 137,771 genotyped). We integrated multiple modalities of patient electronic Health Record (EHR) data including demographics, diagnoses, drug codes, and lab results, along with polygenic risk scores (PRSs) of psychiatric (ANX, BIP, MDD, SCZ) and metabolic (obesity, LDL, HDL, TGL and T2D) traits, to predict patient metabolic outcomes using multi-modal BERT architecture. Results In the CDW cohort, clinically significant weight gain (> 7%) during antipsychotic use was related to Asian ethnicity, pre-treatment elevation in triglycerides, and use of thiothixene, systemic contraceptives, and antipsoriatics, but inversely related to antimigraine agents, opioid antagonist analgesics, and immune suppressants. In the MVP cohort, BMI increase was related to Hispanic ancestry, first-generation antipsychotics, older age, higher T2D PRS, and inversely related to BP PRS. Discussion This is the largest study to date of genetic and environmental factors associated with antipsychotic-induced metabolic adverse effects. While our results require replication in independent samples, they suggest that multimodal AI could be useful in the identification of both risk and protective factors of psychotropic adverse effects and therefore, a potentially powerful tool in Precision Psychiatry. Disclosure Nothing to disclose.
Trauma exposure and drinking motives (e.g., social, enhancement, coping) are both associated with increased alcohol use and related problems. Studies have frequently investigated this relationship by examining drinking motives, such as drinking to cope with negative affect, in isolation, yet few studies have examined motives simultaneously in trauma-exposed populations. It is also unclear whether the relationship between drinking motives and alcohol use outcomes differs as a function of population characteristics (e.g., gender, trauma type). Using latent profile analysis, we aimed to (a) identify latent profiles characterized by drinking motives, assessed with the Drinking Motives Questionnaire (DMQ), in two samples: primarily male veterans with combat trauma (N = 174) and civilians with interpersonal trauma (N = 152), and (b) determine whether associations with alcohol use outcomes of consumption and binge drinking (BD) would differ by sample. A three-class solution was replicated across both samples: profiles characterized by moderate Social scores and low Enhancement and Coping scores (low ENH/COP), moderate scores across all domains (medium DMQ), and elevated scores across all domains (high DMQ). In both samples, profile membership was differentially associated with consumption and BD. Findings suggest patterns of drinking motives may be similar across different trauma-exposed populations, but associations with alcohol outcomes likely differ in meaningful ways. Results can help inform targeted interventions at different treatment settings, such as community health centers or VA hospitals.
Biobanks that collect deep phenotypic and genomic data across many individuals have emerged as a key resource in human genetics. However, phenotypes in biobanks are often missing across many individuals, limiting their utility. We propose AutoComplete, a deep learning-based imputation method to impute or ‘fill-in’ missing phenotypes in population-scale biobank datasets. When applied to collections of phenotypes measured across ~300,000 individuals from the UK Biobank, AutoComplete substantially improved imputation accuracy over existing methods. On three traits with notable amounts of missingness, we show that AutoComplete yields imputed phenotypes that are genetically similar to the originally observed phenotypes while increasing the effective sample size by about twofold on average. Further, genome-wide association analyses on the resulting imputed phenotypes led to a substantial increase in the number of associated loci. Our results demonstrate the utility of deep learning-based phenotype imputation to increase power for genetic discoveries in existing biobank datasets.
Antipsychotic drugs are widely used to treat psychiatric disorders such as schizophrenia and bipolar disorder. However, they are known to have important adverse effects related to metabolic syndrome (MS), which significantly impacts their effectiveness. Quantifying this effect is complicated because the temporal patterns of both antipsychotic medication usage and MS-related outcomes vary across patients. To minimize deleterious effects on patients, it is of the utmost importance to characterize and predict MS effects using genetic and non-genetic information available in electronic medical record (EmR) datasets. For our initial analysis, we aggregate patient's EMR data into three-month intervals, and perform model selection and linear regression to identify important covariates such as age, sex, comorbid medical conditions, oral hypoglycemic, statin, and metformin usage, and hospitalization. We further develop a longitudinal deep learning model that uses high dimensional EMR data to predict MSvar trajectories and utilize SHAP (Shapely additive explanations) values to identify both temporal and polypharmacy-related patterns, predictive of adverse metabolic effects. In addition, we utilize highly predictive phenotypes developed on the full cohort of VA patients to perform GWAS for the rate of change in MS outcomes as a function of SNP genotype in the MVP population (currently 658,582 patients with genomic profiles) and antipsychotic type, while adjusting for 20 ancestry principal components and other biologically relevant covariates. Subsequently, we use summary statistics from GWAS to perform transcriptome/proteome/methylome wide analyses, at both gene and gene set levels, employing state-of-the-art mendelian randomization tools. Our cohort was selected on the basis of antipsychotic medication usage and completeness of the characterization of our MS outcomes, from all 25 million patients with records in the VA EMR system, in use between 2000 to 2023. This results in a combined 12,303,200 patient-quarters of observation among 135,200 unique patients who have had at least one hundred measurements of weight and ten antipsychotic prescription fills. We find in a multivariate logistic regression that Quetiapine, Risperidone, Olanzapine, and Aripiprazole were associated with weight gain and Haloperidol, Ziprasidone, and Fluphenazine with weight loss in quarters (three month periods) when each medication was prescribed at least once, compared to quarters without the medication. This was when controlling for Anxiety and Bipolar diagnoses (associated with weight gain), as well as diabetes medications (metformin associated with weight loss, while insulin with weight gain) and statins, which were associated with weight gain. We further observed a strong impact of age on these associations, with diabetes medications and statins relatively more important in older patients and the antipsychotic medications with the young. GWAS results will be presented. Several second-generation antipsychotics are associated with weight gain, while two first-generation antipsychotics, as well as ziprasidone, are associated with weight loss. These are age-dependent and impacted by both psychiatric diagnoses as well as insulin, statin, and metformin use. GWAS results will be discussed.
Posttraumatic stress disorder (PTSD) genetics are characterized by lower discoverability than most other psychiatric disorders. The contribution to biological understanding from previous genetic studies has thus been limited. We performed a multi-ancestry meta-analysis of genome-wide association studies across 1,222,882 individuals of European ancestry (137,136 cases) and 58,051 admixed individuals with African and Native American ancestry (13,624 cases). We identified 95 genome-wide significant loci (80 novel). Convergent multi-omic approaches identified 43 potential causal genes, broadly classified as neurotransmitter and ion channel synaptic modulators (e.g., GRIA1, GRM8, CACNA1E ), developmental, axon guidance, and transcription factors (e.g., FOXP2, EFNA5, DCC ), synaptic structure and function genes (e.g., PCLO, NCAM1, PDE4B ), and endocrine or immune regulators (e.g., ESR1, TRAF3, TANK ). Additional top genes influence stress, immune, fear, and threat-related processes, previously hypothesized to underlie PTSD neurobiology. These findings strengthen our understanding of neurobiological systems relevant to PTSD pathophysiology, while also opening new areas for investigation.
Background Variation in genes involved in ethanol metabolism has been shown to influence risk for alcohol dependence (AD) including protective loss of function alleles in ethanol metabolizing genes. We therefore hypothesized that people with severe AD would exhibit different patterns of rare functional variation in genes with strong prior evidence for influencing ethanol metabolism and response when compared to genes not meeting these criteria. Objective Leverage a novel case only design and Whole Exome Sequencing (WES) of severe AD cases from the island of Ireland to quantify differences in functional variation between genes associated with ethanol metabolism and/or response and their matched control genes. Methods First, three sets of ethanol related genes were identified including those a) involved in alcohol metabolism in humans b) showing altered expression in mouse brain after alcohol exposure, and altering ethanol behavioral responses in invertebrate models. These genes of interest (GOI) sets were matched to control gene sets using multivariate hierarchical clustering of gene-level summary features from gnomAD. Using WES data from 190 individuals with severe AD, GOI were compared to matched control genes using logistic regression to detect aggregate differences in abundance of loss of function, missense, and synonymous variants, respectively. Results Three non-independent sets of 10, 117, and 359 genes were queried against control gene sets of 139, 1522, and 3360 matched genes, respectively. Significant differences were not detected in the number of functional variants in the primary set of ethanol-metabolizing genes. In both the mouse expression and invertebrate sets, we observed an increased number of synonymous variants in GOI over matched control genes. Post-hoc simulations showed the estimated effects sizes observed are unlikely to be under-estimated. Conclusion The proposed method demonstrates a computationally viable and statistically appropriate approach for genetic analysis of case-only data for hypothesized gene sets supported by empirical evidence.
Neuropsychiatric and substance use disorders (NPSUDs) have a complex etiology that includes environmental and polygenic risk factors with significant cross-trait genetic correlations. Genome-wide association studies (GWAS) of NPSUDs yield numerous association signals. However, for most of these regions, we do not yet have a firm understanding of either the specific risk variants or the effects of these variants. Post-GWAS methods allow researchers to use GWAS summary statistics and molecular mediators (transcript, protein, and methylation abundances) infer the effect of these mediators on risk for disorders. One group of post-GWAS approaches is commonly referred to as transcriptome/proteome/methylome-wide association studies, which are abbreviated as T/P/MWAS (or collectively as XWAS). Since these approaches use biological mediators, the multiple testing burden is reduced to the number of genes (∼20,000) instead of millions of GWAS SNPs, which leads to increased signal detection. In this work, our aim is to uncover likely risk genes for NPSUDs by performing XWAS analyses in two tissues-blood and brain. First, to identify putative causal risk genes, we performed an XWAS using the Summary-data-based Mendelian randomization, which uses GWAS summary statistics, reference xQTL data, and a reference LD panel. Second, given the large comorbidities among NPSUDs and the shared cis-xQTLs between blood and the brain, we improved XWAS signal detection for underpowered analyses by performing joint concordance analyses between XWAS results i) across the two tissues and ii) across NPSUDs. All XWAS signals i) were adjusted for heterogeneity in dependent instruments (HEIDI) (non-causality) p-values and ii) used to test for pathway enrichment. The results suggest that there were widely shared gene/protein signals within the major histocompatibility complex region on chromosome 6 (BTN3A2 and C4A) and elsewhere in the genome (FURIN, NEK4, RERE, and ZDHHC5). The identification of putative molecular genes and pathways underlying risk may offer new targets for therapeutic development. Our study revealed an enrichment of XWAS signals in vitamin D and omega-3 gene sets. So, including vitamin D and omega-3 in treatment plans may have a modest but beneficial effect on patients with bipolar disorder.
Background: The genome-wide association study (GWAS) is a common tool to identify genetic variants associated with complex traits, including psychiatric disorders (PDs). However, post-GWAS analyses are needed to extend the statistical inference to biologically relevant entities, e.g., genes, proteins, and pathways. To achieve this goal, researchers developed methods that incorporate biologically relevant intermediate molecular phenotypes, such as gene expression and protein abundance, which are posited to mediate the variant-trait association. Transcriptome-wide association study (TWAS) and proteome-wide association study (PWAS) are commonly used methods to test the association between these molecular mediators and the trait. Summary: In this review, we discuss the most recent developments in TWAS and PWAS. These methods integrate existing “omic” information with the GWAS summary statistics for trait(s) of interest. Specifically, they impute transcript/protein data and test the association between imputed gene expression/protein level with phenotype of interest by using (i) GWAS summary statistics and (ii) reference transcriptomic/proteomic/genomic datasets. TWAS and PWAS are suitable as analysis tools for (i) primary association scan and (ii) fine-mapping to identify potentially causal genes for PDs. Key Messages: As post-GWAS analyses, TWAS and PWAS have the potential to highlight causal genes for PDs. These prioritized genes could indicate targets for the development of novel drug therapies. For researchers attempting such analyses, we recommend Mendelian randomization tools that use GWAS statistics for both trait and reference datasets, e.g., summary Mendelian randomization (SMR). We base our recommendation on (i) being able to use the same tool for both TWAS and PWAS, (ii) not requiring the pre-computed weights (and thus easier to update for larger reference datasets), and (iii) most larger transcriptome reference datasets are publicly available and easy to transform into a compatible format for SMR analysis.
Alcohol use disorder (AUD) is moderately heritable with significant social and economic impact. Genome-wide association studies (GWAS) have identified common variants associated with AUD, however, rare variant investigations have yet to achieve well-powered sample sizes. In this study, we conducted an interval-based exome-wide analysis of the Alcohol Use Disorder Identification Test Problems subscale (AUDIT-P) using both machine learning (ML) predicted risk and empirical functional weights. This research has been conducted using the UK Biobank Resource (application number 30782.) Filtering the 200k exome release to unrelated individuals of European ancestry resulted in a sample of 147,386 individuals with 51,357 observed and 96,029 unmeasured but predicted AUDIT-P for exome analysis. Sequence Kernel Association Test (SKAT/SKAT-O) was used for rare variant (Minor Allele Frequency (MAF) < 0.01) interval analyses using default and empirical weights. Empirical weights were constructed using annotations found significant by stratified LD Score Regression analysis of predicted AUDIT-P GWAS, providing prior functional weights specific to AUDIT-P. Using only samples with observed AUDIT-P yielded no significantly associated intervals. In contrast, ADH1C and THRA gene intervals were significant (False discovery rate (FDR) <0.05) using default and empirical weights in the predicted AUDIT-P sample, with the most significant association found using predicted AUDIT-P and empirical weights in the ADH1C gene (SKAT-O P Default = 1.06 x 10 -9 and P Empirical weight = 6.25 x 10 -11 ). These findings provide evidence for rare variant association of the ADH1C gene with the AUDIT-P and highlight the successful leveraging of ML to increase effective sample size and prior empirical functional weights based on common variant GWAS data to refine and increase the statistical significance in underpowered phenotypes.
Large biobank samples provide an opportunity to integrate broad phenotyping, familial records, and molecular genetics data to study complex traits and diseases. We introduce Pearson-Aitken Family Genetic Risk Scores (PA-FGRS), a new method for estimating disease liability from patterns of diagnoses in extended, age-censored genealogical records. We then apply the method to study a paradigmatic complex disorder, Major Depressive Disorder (MDD), using the iPSYCH2015 case-cohort study of 30,949 MDD cases, 39,655 random population controls, and more than 2 million relatives. We show that combining PA-FGRS liabilities estimated from family records with molecular genotypes of probands improves the three lines of inquiry. Incorporating PA-FGRS liabilities improves classification of MDD over and above polygenic scores, identifies robust genetic contributions to clinical heterogeneity in MDD associated with comorbidity, recurrence, and severity, and can improve the power of genome-wide association studies (GWAS). Our method is flexible and easy to use and our study approaches are generalizable to other data sets and other complex traits and diseases.### Competing Interest StatementB.J.V. is a member of Allelica's scientific advisory board.### Funding StatementLundbeckfonden Fellowship R335-2019-2318; National Institute of Mental Health R01MH130581### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:The use of this data is according the guidelines provided by the Danish Scientific Ethics Committee, the Danish Health Data Authority, the Danish data protection agency and the Danish Neonatal Screening Biobank Steering Committee.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesAll data produced in the present study are available upon reasonable request to the authors and in accordance with Danish law.
Alcohol use disorder (AUD) is common and affects millions of people in the United States. Twin and family studies show that this disorder is heritable, and genome wide association studies (GWASs) report multiple loci associated with AUD. Post-GWAS analyses of GWAS signals yield enrichment in several gene-sets that show nominal enrichment in brain tissues. Unfortunately, only a handful of genes have been reported from rare-variant studies. However, recent studies strongly imply that for many complex disorders, common and rare variant findings, while not converging at gene level, converge at gene-set level. In addition, recent studies have showed that using phenotypic information predicted from biobanks can increase genetic discoveries. These suggest that using different approaches for rare variant analyses can improve detection power for genes associated with AUD. Here, we explore integrative approaches to prioritize genes associated with a proxy phenotype of AUD, the Alcohol Use Disorder Identification Test-Problems (AUDIT-P) phenotype. We analyzed the whole-exome-sequencing (WES) datasets of 500K people from the UK Biobank. First, we conducted WES analyses for individuals whose AUDIT scores are available. Second, we developed a pipeline to jointly model rare variants and gene-sets to improve statistical power. Finally, we used a machine learning approach developed by our group to predict AUDIT scores for all the 500K people, and re-analyzed their WES datasets. We first analyzed loss-of-function and missense variants from WES sequencing datasets of 133,914 people with AUDIT scores available. Three significant genes (ADH1C, FPR1, VPS29) were observed. The most significant signal was for ADH1C (adjusted p-value = 0.4 × 10-5). Next, to prioritize additional genes, we jointly analyzed gene-level statistics and 181 gene sets curated from previous studies. We prioritized several genes (max posterior probability > 0.8) including ADH1C. Finally, to see if the prediction of phenotypic information can help prioritize genes, we conducted WES analysis for all 414,508 samples of European descent with full predicted AUDIT scores. Statistical power for ADH1C was substantially improved (adjusted p-value = 3.3 × 10-21). Our results present top significant genes for AUD obtained by analyzing rare variants from a large-scale WES dataset. The results also include 1) the additional biological information into AUD, and 2) integrative approaches for incorporating functional genomics and health care record datasets to improve genetic discoveries. We are improving these integrate approaches to be able to increase statistical power for the prioritization of genes.