Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a chronic, multisystem disease characterized by post-exertional malaise and persistent fatigue. The cause of ME/CFS is not well understood, and there are no established biomarkers or FDA-approved pharmacotherapies. The clinical heterogeneity of ME/CFS presents challenges to diagnosis and treatment and necessitates collaborative efforts to generate robust findings. This study leveraged gene and protein expression data from the mapMECFS data repository and the DecodeME Genome-Wide Association Study (GWAS) to assess consistent gene signatures across studies. The mitochondrial genes MT-RNR1 and MT-RNR2 exhibited lower expression in ME/CFS cases in two studies. Combining this with increased expression of mitochondrial genes in platelets from another study, this supports mitochondrial dysregulation as having a role in ME/CFS. Furthermore, ME/CFS-associated genes were mapped to compounds in drug databases as possible treatments for further investigation. In muscle gene expression data, 107 approved compounds target 26 genes with functions relevant to mitochondrial support and immunomodulators. From the DecodeME GWAS, 83 approved compounds target 24 genes with functions related to energy metabolism and mitochondrial function. Though little consistency in specific genes was observed across studies, which highlights the need for larger studies, mitochondrial dysfunction in ME/CFS cases was evident across studies.
ABSTRACT Although short cervical length in the mid-trimester of pregnancy is a one of the strongest predictors of preterm birth ( i.e ., parturition before 37 completed weeks), there is limited understanding of how the dynamics of cervical remodeling ( i.e ., changes in cervical length) leading up to labor and delivery can inform obstetrical risk. In this study, latent growth curve analysis was applied to serial cervical length measurements across pregnancy (median of 6; IQR = 3-8) to quantify characteristics of cervical change in a cohort of 5,111 singleton pregnancies consisting predominantly of Black women. A conditional mediation model including nine common maternal risk factors for spontaneous preterm birth as exogenous predictors accounted for 26.5% of the variability in gestational age at delivery ( P < 0.001). This model provides insight into distinct mechanisms by which specific maternal risk factors influence preterm birth. For instance, effects of maternal parity and smoking status were fully mediated through cervical change parameters, whereas the influence of previous preterm birth was only partially explained, suggesting alternative pathways could be involved. This study provides the first account of the intermediary role of cervical dynamics in associations between known maternal risk factors and gestational age at delivery.
Genes influencing opioid use disorder (OUD) biology have been identified via genome-wide association studies (GWAS), gene expression, and network analyses. These discoveries provide opportunities to identifying existing compounds targeting these genes for drug repurposing studies. However, systematically integrating discovery results and identifying relevant available pharmacotherapies for OUD repurposing studies is challenging. To address this, we’ve constructed a framework that uses existing results and drug databases to identify candidate pharmacotherapies. For this study, two independent OUD related meta-analyses were used including a GWAS and a differential gene expression (DGE) study of post-mortem human brain. Protein-Protein Interaction (PPI) sub-networks enriched for GWAS risk loci were identified via network analyses. Drug databases Pharos, Open Targets, Therapeutic Target Database (TTD), and DrugBank were queried for clinical status and target selectivity. Cross-omic and drug query results were then integrated to identify candidate compounds. GWAS and DGE analyses revealed 3 and 335 target genes (FDR q < 0.05), respectively, while network analysis detected 70 genes in 22 enriched PPI networks. Four selection strategies were implemented, which yielded between 72 and 676 genes with statistically significant support and 110 to 683 drugs targeting these genes, respectively. After filtering out less specific compounds or those targeting well-established psychiatric-related receptors (OPRM1 and DRD2), between 2 and 329 approved drugs remained across the four strategies. By leveraging multiple lines of biological evidence and resources, we identified many FDA approved drugs that target genes associated with OUD. This approach a) allows high-throughput querying of OUD-related genes, b) detects OUD-related genes and compounds not identified using a single domain or resource, and c) produces a succinct summary of FDA approved compounds eligible for efficient expert review. Identifying larger pools of candidate pharmacotherapies and summarizing the supporting evidence bridges the gap between discovery and drug repurposing studies.
Most genome-wide association studies (GWASs) of depression focus on broad, heterogeneous outcomes, limiting the discovery of genomic risk loci specific to major depressive disorder (MDD). Previous UK Biobank (UKB) studies had limited ability to pinpoint MDD-associated loci due to a smaller sample with strictly defined MDD outcomes and further exclusion of many participants based on ancestry or relatedness, significantly underutilizing this resource's potential for elucidating the genetic architecture of MDD. Here, we present novel genomic insights into MDD by fully utilizing existing UKB data through (1) a trans-ancestry GWAS pipeline using two complementary approaches controlling for population structure and relatedness and (2) an increased sample with MDD symptom-level data across two mental health assessments. We identified strict MDD outcomes among 211,535 participants, representing a 38% increase in eligible participants from prior studies with only one assessment. Ancestrally inclusive analyses yielded 61 genomic risk loci across depression phenotypes, compared to 47 in the analyses restricted to participants genetically similar to European ancestry. Fourteen of these loci, including five novel, were associated with strict MDD phenotypes, whereas only one locus has been previously reported in UKB. MDD-associated genomic loci and predicted gene expression levels showed little overlap with broad depression, indicating higher specificity. Notably, polygenic scores based on these results were significantly associated with depression diagnoses across ancestry groups in the All of Us Research Program, highlighting the shared genetic architecture across populations. While the trans-ancestry analyses, which included non-European participants, increased the number of associated loci, the discovery of non-European ancestry-specific loci was limited, underscoring the need for larger, globally representative studies of MDD. Importantly, beyond these results, our GWAS pipeline will facilitate inclusive analyses of other traits and disorders, helping improve statistical power, representation, and generalizability in genomic studies.
Sonographic cervical length is a powerful predictor of maternal risk for spontaneous preterm birth (sPTB). Twin and family studies have established a maternal genetic heritability for sPTB ranging from 13 to 20
Background A disproportionate majority of individuals included in genome-wide association studies (GWAS) are of European ancestry (EUR), excluding significant proportions of potential study participants. Research indicates that incorporating participants from diverse ancestries can enhance the power of GWAS, improve the fine-mapping of loci, and increase the predictive performance of polygenic scores across populations. To realize these benefits, we developed a workflow for trans-ancestry GWAS that fully utilizes the genetic diversity present in large-scale electronic health record (EHR)-linked biobanks from the PsycheMERGE Network. Methods To date, published cross-ancestry psychiatric GWAS have been primarily conducted using an ancestrally stratified meta-analysis approach. This method involves running GWAS separately within ancestry groups, followed by meta-analysis across these groups, and typically requires the removal of ancestry outliers, individuals with admixed ancestry, and related individuals. The purpose of this presentation is to showcase our pipeline, which employs best practices to maximize data inclusion while addressing population stratification, thereby demonstrating enhanced power of genomic discovery and transferability of findings across populations. Our pipeline was implemented across PsycheMERGE Diversity Initiative Phase I sites, including the UK Biobank (conducted under application ID 30782), All of Us Research Program, Penn Medicine Biobank, and the Million Veteran Program. The proposed research aims to create the largest inclusive trans-ancestry GWAS effort across psychiatric outcomes defined through EHR data. Results Trans-ancestry GWAS was performed across Phase I sites, using two distinct but complementary approaches: 1) Stratified: ancestry-stratified mixed-effects GWAS followed by cross-ancestry meta-analysis and 2) Joint: joint trans-ancestry mixed-effects GWAS. In both methods, mixed-effects models allowed the inclusion of related samples and traditionally determined ancestry outliers. Samples from each site were statistically assigned to reference ancestry groups using our extended POP-MaD pipeline (Population grouping by Mahalanobis Distance), which includes an expanded reference population panel (including six non-EUR groups). Through this approach, we were able to incorporate an additional +118K individuals across the ancestry spectrum. Compared to European-ancestry GWAS alone, inclusion of multiple ancestries yielded more genome-wide significant loci, with the trans-ancestry mixed-effects GWAS having the highest power. For example, in UK Biobank, identified loci for lifetime major depression went from 1 in previous EUR reports to 61, without evidence of inflation from statistical artifacts from population stratification or familial relatedness, demonstrating the power of our approach. Discussion We view the stratified and joint approaches as complementary strategies to maximize inclusiveness, identify variants with broad global impact as well as ancestry-specific loci, and facilitate post-GWAS secondary analyses. The presented protocols provide GWAS summary statistics that can contribute to both ancestry-specific and cross-ancestry meta-analyses, thereby enhancing inclusiveness in the near term. Looking ahead, sustained recruitment efforts spanning the ancestry spectrum are vital for fostering globally representative genomic studies that offer equitable benefits across populations. Disclosure Nothing to disclose.
Background Strictly defined major depressive disorder (MDD) phenotypes in the UK Biobank (UKB) can yield more specific genome-wide association study (GWAS) results compared to broad depression. However, as MDD is deeply phenotyped in only a third of the UKB participants, prior GWASs in further subsets of European ancestry (EUR) or “White British” unrelated individuals have been underpowered for variant-level discovery. Here, we apply a pan-UKB, cross-ancestry GWAS pipeline that enhances power and provides deeper insights into the genetic architecture of MDD. Methods We assigned the UKB samples (application ID 30782) to reference ancestry groups using the POP-MaD pipeline (Population grouping by Mahalanobis Distance), identifying ∼458k EUR and 28k non-EUR individuals (across five non-EUR groups). We derived an expanded set of strictly defined MDD cases and controls by combining data from the 2016 and 2022 online mental health questionnaires. For comparison, we also obtained broad depression phenotypes based on prior studies. We then performed ancestry-stratified GWAS meta-analyses as well as joint trans-ancestry mixed-effects GWAS, while using REGENIE to include related samples in both approaches. Follow-up analyses involved conditional and joint association (COJO) analyses to identify independent genomic loci, fine-mapping and colocalization between different depression phenotypes, and gene-level post-GWAS analyses, including transcriptome-wide association studies (TWAS) and summary statistics-based Mendelian Randomization (SMR) using gene expression data from brain tissues. Results We identified over 62k cases and 118k controls of strictly defined MDD, with significant prevalence differences between ancestry groups. The EUR and cross-ancestry GWAS approaches revealed, in total, 61 independent genomic loci associated with various depression definitions. Notably, 14 of these loci were discovered in cross-ancestry analyses only, highlighting increased power. Sensitivity analyses and LD score regression in EUR showed no evidence of bias due to relatedness or more stringently defined ancestry outliers. Through additional trans-ancestry analyses with down sampled EUR, we also compared the impacts of increased sample size vs increased genetic diversity on genomic discovery for MDD.Analyses of strictly defined MDD phenotypes revealed 13 genomic risk loci, compared to only one locus previously reported in the unrelated, White-British subset. Of these, eleven loci were not significantly associated with broad depression despite the latter's much larger sample, suggesting higher specificity for MDD. Further, colocalization analyses indicated little overlap of causal SNPs underlying MDD and broad depression. Follow-up TWAS and SMR analyses revealed genes and brain tissues specific to MDD (e.g., LRRC37A4P in the hypothalamus, anterior cingulate cortex, and nucleus accumbens) vs those shared with broad depression (e.g., DFNA5 in the caudate nucleus), providing novel MDD-specific biological insights. Finally, for both MDD and broad depression, we compared the prediction accuracy of EUR and cross-ancestry polygenic scores in the All of Us Research Program. Discussion Inclusive, pan-UKB cross-ancestry analyses of depression maximize the genomic insights from extant data, augmenting the power for variant-level associations and post-GWAS analyses of deeply phenotyped MDD. These analyses help elucidate MDD-specific biological pathways and potential therapeutic targets. Disclosure Nothing to disclose.
Background Large biobanks such as the UK Biobank (UKB) offer great potential for discovering loci associated with complex disorders. However, data may not be collected uniformly with subsets of participants being assessed using clinically based instruments. We previously demonstrated that Missingness Adapted Group-wise Incremental Clustered (MAGIC)-LASSO can reliably predict CIDI-SF Major Depressive Disorder symptom score (MDDsx) in the UKB, with genetic correlation (rG) between measured (EUR n∼118k) and predicted (EUR n∼360k) MDDsx of 0.89. This machine learning (ML) method implements hierarchical clustering of variables according to missingness structure followed by sequential Group LASSO by cluster and therefore doesn't require any assumption of missingness at random. Final modeling is performed on features retained from within-cluster LASSO. However, alternative ML methods may be applied in the final step that may improve overall predictive performance. Methods The MAGIC framework was extended to use XGBoost, an ensemble ML technique allowing for higher order interactions, in the final model fitting and implemented to predict MDD in UKB. We also evaluated 3 variations of XGBoost including softprob, upsampling, and weighted classes. For MDDdx, the relative performance of assigned vs probability of case status was also examined. GWAS were performed on observed and predicted outcomes followed by rG analysis using LDSC. Results Using MAGIC-selected features (n=37), XGBoost accurately predicted MDDsx (MDDsxxgb) (rG=0.89). Although the LASSO and XGBoost rGs were equivalent, the distribution of MDDsxxgb more closely matched the observed. XGBoost also performed well for predicted lifetime MDDdx for case status (rG=0.91) and probability (rG=0.89). GWAS of predicted outcomes with increased effective sample sizes yielded additional genome-wide significant (GWS) loci. For MDDsxxgb, 28 independent GWS intervals were observed versus 2 for observed. For MDDdx, 20 and 39 GWS intervals were detected for assigned case and probability, respectively. Pilot assessment of softprob, upsampling, and weighted classes XGBoost predictions showed a small increase in overall RMSE but reduced MAE for high symptom classes (with better agreement between observed and predicted MDDsx distribution). Discussion We demonstrate that using XGBoost after MAGIC feature selection offers improved prediction of MDD, albeit at increased computational cost. Additionally, using predicted probabilities of MDDdx showed the most increase in detected loci. XGBoost and variants may further improve prediction for MDDsx where the observed distribution is zero-inflated, demonstrating the power of boosting frameworks for prediction of phenotypes with skewed and/or unbalanced distributions. Overall, these adaptive ML frameworks provide reliable predictions of missing or unobserved phenotypes which boost the power for genomic locus discovery. This research has been conducted using the UK Biobank (UKB) Resource (application number 30782)
Background Alcohol use disorder (AUD) has a profound public health impact. However, understanding of the molecular mechanisms underlying the development and progression of AUD remain limited. Here, we interrogate AUD-associated DNA methylation (DNAm) changes within and across addiction-relevant brain regions: the nucleus accumbens (NAc) and dorsolateral prefrontal cortex (DLPFC). Methods Illumina HumanMethylation EPIC array data from 119 decedents of European ancestry (61 cases, 58 controls) were analyzed using robust linear regression, with adjustment for technical and biological variables. Associations were characterized using integrative analyses of public gene regulatory data and published genetic and epigenetic studies. We additionally tested for brain region-shared and -specific associations using mixed effects modeling and assessed implications of these results using public gene expression data. Results At a false discovery rate ≤ 0.05, we identified 53 CpGs significantly associated with AUD status for NAc and 31 CpGs for DLPFC. In a meta-analysis across the regions, we identified an additional 21 CpGs associated with AUD, for a total of 105 unique AUD-associated CpGs (120 genes). AUD-associated CpGs were enriched in histone marks that tag active promoters and our strongest signals were specific to a single brain region. Of the 120 genes, 23 overlapped with previous genetic associations for substance use behaviors; all others represent novel associations. Conclusions Our findings identify AUD-associated methylation signals, the majority of which are specific within NAc or DLPFC. Some signals may constitute predisposing genetic and epigenetic variation, though more work is needed to further disentangle the neurobiological gene regulatory differences associated with AUD.
Excessive alcohol consumption is a leading cause of preventable death worldwide. To improve understanding of neurobiological mechanisms associated with alcohol use disorder (AUD) in humans, we compared gene expression data from deceased individuals with and without AUD across two addiction-relevant brain regions: the nucleus accumbens (NAc) and dorsolateral prefrontal cortex (DLPFC). Bulk RNA-seq data from NAc and DLPFC (N >= 50 with AUD, >= 46 non-AUD) were analyzed for differential gene expression using modified negative binomial regression adjusting for technical and biological covariates. The region-level results were meta-analyzed with those from an independent dataset (NNAc = 28 AUD, 29 non-AUD; NPFC = 66 AUD, 77 non-AUD). We further tested for heritability enrichment of AUD-related phenotypes, gene co-expression networks, gene ontology enrichment, and drug repurposing. We identified 176 differentially expressed genes (DEGs; 12 in both regions, 78 in NAc only, 86 in DLPFC only) for AUD in our new dataset. After meta-analyzing with published data, we identified 476 AUD DEGs (25 in both regions, 29 in NAc only, 422 in PFC only). Of these DEGs, 17 were significant when looked up in GWAS of problematic alcohol use or drinks per week. Gene co-expression analysis showed both concordant and unique gene networks across brain regions. We also identified 29 and 436 drug compounds that target DEGs from our meta-analysis in NAc and PFC, respectively. This study identified robust AUD-associated DEGs, contributing novel neurobiological insights into AUD and highlighting genes targeted by known drug compounds, generating opportunity for drug repurposing to treat AUD.
Most genome-wide association studies (GWAS) in the UK Biobank (UKB) have been restricted to ∼360,000 "White-British" unrelated individuals, excluding 27% of the participants. Recent studies show that combining relatively small cohorts of non-European ancestries with a larger European-ancestry sample can enhance the power of GWAS, the fine-mapping of loci, and the predictive performance of polygenic scores across populations. To realize these benefits, we present a workflow for trans-ancestry GWAS, leveraging the genetic diversity in the UKB.
Background Posttraumatic stress disorder (PTSD) is a common, clinically and economically impactful psychiatric disorder. Genetic discovery for PTSD has accelerated at a rapid pace with nearly 100 common variants identified from genome-wide association studies (GWAS). However, large-scale studies of rare variants associated with PTSD using next-generation sequencing data have not yet been conducted. Here, we analyzed the latest whole-exome-sequencing (WES) dataset from the UK Biobank (UKB) to detect genes and gene-sets containing rare variation associated with PTSD. Methods We conducted rare variant analyses using WES based genotypes and in 133,738 samples of European ancestry with mental health questionnaire data. PTSD phenotypes were defined from these questionnaire data. First, we used established quality-control steps for WES data. Next, synonymous, missense and loss-of-function variants were identified for analyses. We then ran SAIGE-GENE+ to generate variant and gene-level p-values for variants with a minor allele frequency (MAF) < 0.01. The GAUSS package was then used to conduct gene-set enrichment analysis for gene-level p-values. Gene-sets were selected from a wide variety of sources including the Psychiatric Genomics Consortium GWAS meta-analysis of PTSD in 2024 and curated gene-sets from studies of neuropsychiatric disorders. Results For variant-level analysis, there were no genome-wide significant variants detected with the most significant variant being chr1:152356894:A:G in the FLG2 gene (p = 4.98 × 10-7, MAF < 0.01). For gene-level analysis of 18,190 genes, only MAST4 was nominally significant (p = 8.2 × 10 6) for loss-of-function and missense variants. For gene-set enrichment analysis, 10 out of 189 tested gene-sets were nominally significant but not significant after multiple testing correction. Eight of these 10 gene-sets were from mouse mutants with central nervous system phenotypes with the most significant (p=0.003) being related to abnormal innervation. The remaining two gene-sets were from studies of other neuropsychiatric disorders. Finally, the gene-set from the latest PGC PTSD GWAS meta-analysis had a modest p-value of 0.089. Discussion For rare variant analysis, robustly statistically significant associations with a proxy phenotype of PTSD were not observed, even with a large sample (133k). However, nominally significant gene MAST4 has been previously associated with neurodevelopmental disorders. Similarly, some nominally enriched gene-sets were from studies of previous neuropsychiatric disorders. Notably, the gene-set from the latest common variant GWAS study had a modest p-value in this analysis. Future studies which focus on understanding the overlapping information of these types of variants can shed additional light on the genetic architecture of PTSD. Our future extensions of these analyses will include integrative analytic approaches and importantly expand to multiple ancestry samples across several biobanks in order to advance the understanding of PTSD.This study was conducted using the UK Biobank Resource application number 30782.
Little is known about how non-suicidal and suicidal self-injury are differentially genetically related to psychopathology and related measures. This research was conducted using the UK Biobank Resource, in participants of European ancestry ( N = 2320 non-suicidal self-injury [NSSI] only; N = 2648 suicide attempt; 69.18% female). We compared polygenic scores (PGS) for psychopathology and other relevant measures within self-injuring individuals. Logistic regressions and likelihood ratio tests (LRT) were used to identify PGS that were differentially associated with these outcomes. In a multivariable model, PGS for anorexia nervosa (odds ratio [OR] = 1.07; 95% confidence intervals [CI] 1.01; 1.15) and suicidal behavior (OR = 1.06; 95% CI 1.00; 1.12) both differentiated between NSSI and suicide attempt, while the PGS for other phenotypes did not. The LRT between the multivariable and base models was significant (Chi square = 11.38, df = 2, p = 0.003), and the multivariable model explained a larger proportion of variance (Nagelkerke's pseudo- R 2 = 0.028 vs. 0.025). While NSSI and suicidal behavior are similarly genetically related to a range of mental health and related outcomes, genetic liability to anorexia nervosa and suicidal behavior is higher among those reporting a suicide attempt than those reporting NSSI-only. Further elucidation of these distinctions is necessary, which will require a nuanced assessment of suicidal versus non-suicidal self-injury in large samples.
Background: Measuring and estimating alcohol consumption (AC) is important for individual health, public health, and Societal benefits. While self-report and diagnostic interviews are commonly used, incorporating biological-based indices can offer a complementary approach. Methods: We evaluate machine learning (ML) based predictions of AC using blood and urine-derived biomarkers. This research has been conducted using the UK Biobank (UKB) Resource. In addition to the prediction of the number of alcoholic Drinks Per Week (DPW), four other related phenotypes were predicted for performance comparison. Five ML models were assessed including LASSO, Ridge regression, Gradient Boosting Machines (GBM), Model Boosting (MBOOST), and Extreme Gradient Boosting (XGBOOST). Results: All five ML methods achieved moderate prediction of DPW (r2=0.304-0.356) with biomarkers significantly increasing prediction above using only known covariates and liver enzymes (r2=0.105). XGBOOST achieved the best prediction performance (r2=0.356, MAE=5.214) at the expense of increasing model complexity and training resources compared to other ML methods. All ML models were able to accurately predict if subjects were heavy drinkers (DPW>8 for women and DPW>15 for men) and produced explainable models that highlighted the role of biomarkers in predicting DPW. While phenotype correlations were similar across methods, XGBOOST produced similar heritability estimates for observed (h2=0.064) and predicted (h2=0.077) DPW. The estimated genetic correlation between observed and predicted DPW was 0.877. Conclusions: Predicting AC from ML-based biological measures provides an opportunity to identify individuals at increased risk of heavy AC, thereby offering complementary avenue for risk assessment beyond self-report, screening instruments, or structured interviews, which have some known biases. In addition, explainable AI tools identified a constellation of biomarkers associated with AC. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research has been conducted using the UK Biobank Resource application number 30782. MH, REP, BTW, and AEG were supported by NIH P50AA022537, REP, BTW, and AEG by NIH R01MH125938, REP by The Brain & Behavior Research Foundation NARSAD grant 28632 P&S Fund, AEG by NIH 1K01AA031748-01 and ECP by 1R01DA054313-01A1. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This research has been conducted using the UK Biobank Resource application number 30782. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present work are contained in the manuscript
Background: Non-suicidal self-injury and suicide attempt represent significant public health concerns. While these outcomes are related, there is prior evidence that their etiology does not entirely overlap. Efforts to directly differentiate risk across outcomes are uncommon, particularly among older, population-based cohorts.Methods: This research has been conducted using the UK Biobank. Data on individuals' self-reported history of non-suicidal self-injury only versus suicide attempt (maximum N = 6643) were analyzed. Applying LASSO and standard logistic regression, participants reporting one of these outcomes were assessed for differences across a range of sociodemographic, behavioral, and environmental features.Results: Sociodemographic features most strongly differentiated between the outcomes of non-suicidal self-injury only versus suicide attempt. Specifically, Black individuals were more likely to report a suicide attempt, as were those of mixed race, those endorsing higher levels of depressive symptoms or trauma history, and those who had experienced financial problems (odds ratios 1.02-3.92). Those more likely to engage in non-suicidal self-injury only were younger, female, had higher levels of education, those who resided with a partner, and those who had a recently injured relative.Limitations: Differences in timing across correlates and outcomes preclude the ability to establish causal pathways.Conclusions: The factors identified in the current study as differentially associated with non-suicidal self-injury only versus suicide attempt provide further evidence of at least partially distinct correlates, and warrant follow-up in independent samples to investigate causality.
Background Variation in genes involved in ethanol metabolism has been shown to influence risk for alcohol dependence (AD) including protective loss of function alleles in ethanol metabolizing genes. We therefore hypothesized that people with severe AD would exhibit different patterns of rare functional variation in genes with strong prior evidence for influencing ethanol metabolism and response when compared to genes not meeting these criteria. Objective Leverage a novel case only design and Whole Exome Sequencing (WES) of severe AD cases from the island of Ireland to quantify differences in functional variation between genes associated with ethanol metabolism and/or response and their matched control genes. Methods First, three sets of ethanol related genes were identified including those a) involved in alcohol metabolism in humans b) showing altered expression in mouse brain after alcohol exposure, and altering ethanol behavioral responses in invertebrate models. These genes of interest (GOI) sets were matched to control gene sets using multivariate hierarchical clustering of gene-level summary features from gnomAD. Using WES data from 190 individuals with severe AD, GOI were compared to matched control genes using logistic regression to detect aggregate differences in abundance of loss of function, missense, and synonymous variants, respectively. Results Three non-independent sets of 10, 117, and 359 genes were queried against control gene sets of 139, 1522, and 3360 matched genes, respectively. Significant differences were not detected in the number of functional variants in the primary set of ethanol-metabolizing genes. In both the mouse expression and invertebrate sets, we observed an increased number of synonymous variants in GOI over matched control genes. Post-hoc simulations showed the estimated effects sizes observed are unlikely to be under-estimated. Conclusion The proposed method demonstrates a computationally viable and statistically appropriate approach for genetic analysis of case-only data for hypothesized gene sets supported by empirical evidence.
Introduction: The availability of large-scale biobanks linking genetic data, rich phenotypes, and biological measures is a powerful opportunity for scientific discovery. However, real-world collections frequently have extensive missingness. While missing data prediction is possible, performance is significantly impaired by block-wise missingness inherent to many biobanks. Methods: To address this, we developed Missingness Adapted Group-wise Informed Clustered (MAGIC)-LASSO which performs hierarchical clustering of variables based on missingness followed by sequential Group LASSO within clusters. Variables are pre-filtered for missingness and balance between training and target sets with final models built using stepwise inclusion of features ranked by completeness. This research has been conducted using the UK Biobank (n > 500 k) to predict unmeasured Alcohol Use Disorders Identification Test (AUDIT) scores. Results: The phenotypic correlation between measured and predicted total score was 0.67 while genetic correlations between independent subjects was high >0.86. Discussion: Phenotypic and genetic correlations in real data application, as well as simulations, demonstrate the method has significant accuracy and utility for increasing power for genetic loci discovery.
Background Posttraumatic Stress Disorder (PTSD) tends to co-occur with greater alcohol consumption as well as alcohol use disorder (AUD). However, it is unknown whether the same etiologic factors that underlie PTSD-alcohol-related problems comorbidity also contribute to PTSD- alcohol consumption. Methods We used summary statistics from large-scale genome-wide association studies (GWAS) of European-ancestry (EA) and African-ancestry (AA) participants to estimate genetic correlations between PTSD and a range of alcohol consumption-related and alcohol-related problems phenotypes. Results In EAs, there were positive genetic correlations between PTSD phenotypes and alcohol-related problems phenotypes (e.g. Alcohol Use Disorders Identification Test (AUDIT) problem score) (rGs: 0.132-0.533, all FDR adjusted p < 0.05). However, the genetic correlations between PTSD phenotypes and alcohol consumption -related phenotypes (e.g. drinks per week) were negatively associated or non-significant (rGs: -0.417 to -0.042, FDR adjusted p: <0.05-NS). For AAs, the direction of correlations was sometimes consistent and sometimes inconsistent with that in EAs, and the ranges were larger (rGs for alcohol-related problems: -0.275 to 0.266, FDR adjusted p: NS, alcohol consumption-related: 0.145-0.699, FDR adjusted p: NS). Conclusions These findings illustrate that the genetic associations between consumption and problem alcohol phenotypes and PTSD differ in both strength and direction. Thus, the genetic factors that may lead someone to develop PTSD and high levels of alcohol consumption are not the same as those that lead someone to develop PTSD and alcohol-related problems. Discussion around needing improved methods to better estimate heritabilities and genetic correlations in diverse and admixed ancestry samples is provided.
Alcohol use disorder (AUD) is moderately heritable with significant social and economic impact. Genome-wide association studies (GWAS) have identified common variants associated with AUD, however, rare variant investigations have yet to achieve well-powered sample sizes. In this study, we conducted an interval-based exome-wide analysis of the Alcohol Use Disorder Identification Test Problems subscale (AUDIT-P) using both machine learning (ML) predicted risk and empirical functional weights. This research has been conducted using the UK Biobank Resource (application number 30782.) Filtering the 200k exome release to unrelated individuals of European ancestry resulted in a sample of 147,386 individuals with 51,357 observed and 96,029 unmeasured but predicted AUDIT-P for exome analysis. Sequence Kernel Association Test (SKAT/SKAT-O) was used for rare variant (Minor Allele Frequency (MAF) < 0.01) interval analyses using default and empirical weights. Empirical weights were constructed using annotations found significant by stratified LD Score Regression analysis of predicted AUDIT-P GWAS, providing prior functional weights specific to AUDIT-P. Using only samples with observed AUDIT-P yielded no significantly associated intervals. In contrast, ADH1C and THRA gene intervals were significant (False discovery rate (FDR) <0.05) using default and empirical weights in the predicted AUDIT-P sample, with the most significant association found using predicted AUDIT-P and empirical weights in the ADH1C gene (SKAT-O P Default = 1.06 x 10 -9 and P Empirical weight = 6.25 x 10 -11 ). These findings provide evidence for rare variant association of the ADH1C gene with the AUDIT-P and highlight the successful leveraging of ML to increase effective sample size and prior empirical functional weights based on common variant GWAS data to refine and increase the statistical significance in underpowered phenotypes.
Methods for predicting major depressive disorder (MDD) represent important tools for identifying individuals at increased risk for the disorder. In individuals who have not been directly screened for MDD using structured interviews, predictive approaches which utilize existing medical record data may be particularly useful for identifying patients who warrant such screenings. This research has been conducted using the UK Biobank Resource (UKB, ID30782). In this study, we investigate the utility of blood-based biomarkers, including 303 markers from blood chemistry and NMR assays, to predict various MDD-related phenotypes in the UKB. Using the CIDI-SF assessed in ∼⅓ of the UKB participants, we identified strictly-defined MDD cases based on self-reported symptoms and associated impairment: 45,341 probable (with ≥4 symptoms) and 39,524 definite cases (with ≥5 symptoms), along with 70,631 controls. Additionally, we defined three broad depression phenotypes: help-seeking (173,153 cases, 324,107 controls), cardinal depression symptoms (31,835 cases, 89,520 controls), and depression diagnosis (either self-reported or ICD-10 codes in electronic medical records: 60,517 cases, 441,963 controls). We apply 3 broad classes of machine learning (ML) algorithms, elastic nets, gradient boosting machines (GBM), and MBOOST, to predict these phenotypes, as well as CIDI-SF symptom count, solely using blood-based biomarkers. To assess prediction performance, we compare adjusted R-squared (r2 adj) values for covariate-only models using age and sex to predict the outcomes against the ML models utilizing blood-based measures in addition to covariates. We also estimate the genetic correlation with measured MDD phenotypes, as well as the Psychiatric Genomics Consortium (PGC) MDD. Furthermore, to demonstrate the broad utility of such a blood-based biomarker prediction approach, as well as generate data to serve as benchmark comparisons, we apply the same predictive approach to BMI, body fat percentage, height, and alcohol consumption, defined as drinks per week (DPW). Preliminary results show that predicted DPW were moderately well-predicted, with the best performing ML algorithm, GBM, producing adjusted R-squared of 0.354. However, MDD symptom count was not well predicted by blood-based biomarkers with the GBM prediction producing adjusted R-squared of 0.002. Given this performance, we expanded the framework to include all available UKB phenotypes as predictors in the MAGIC-LASSO approach and predicted MDD symptom counts with a phenotypic correlation of 0.69 between observed and measured symptom count. Genetic correlation between the predicted scores were 0.88 and 0.9 with measured scores and PGC MDD, respectively. These results support extending the MAGIC-LASSO approach to prediction via boosting algorithms and applying this approach to additional UKB MDD phenotypes. These results demonstrate blood-based biomarkers and ML based models can be useful in predicting some outcomes, such as DPW, but not MDD in the current implementation. Expanding the prediction space to include all available phenotypes results in improved prediction for MDD phenotypes. In persons where strictly-defined MDD is not directly assessed, such predictions may provide a useful screener to identify individuals for follow up clinical assessments and expand sample sizes for statistical genetic analyses.