Amidst the opioid crisis, understanding the genetic basis of opioid use disorder (OUD) is crucial for identifying biological mechanisms and intervention points. However, genome-wide association studies (GWASs) have been hampered by inadequate sample sizes and often the use of control populations not assessed for prior opioid exposure. Because opioid exposure is a prerequisite for the development of OUD, consideration of exposure history in controls is important. Electronic health record data (EHR) paired with genomic information allow a broader sampling of patients with OUD and exposed controls. We leveraged data across two healthcare systems to evaluate the impact of using controls not screened for opioid exposure ('generic') versus minimally opioid-exposed control ('exposed'). First, at the phenotypic level, we conducted phenome-wide association studies (PheWAS) to compare the medical comorbidity profiles of OUD cases when using generic versus exposed controls. While PheWAS results for OUD-related comorbidities were more pronounced when using the generic group, 83% of the disease associations were overlapping and of similar effect sizes. Second, at the genetic level, we conducted GWAS (cases vs. generic; cases vs. exposed) and assessed differences in genetic correlations and degrees of phenotypic misclassification. Genetic results were concordant across control groups based on heritability (generic: 0.16 ± 0.07 vs. 0.10 ± 0.07), associations with the coding OPRM1 variant rs1799971 (pgeneric = 8.83E-03 vs. pexposed = 1.83E-02) and genetic correlations with prior OUD GWAS (rg-generic = 0.83 ± 0.26 vs. rg-exposed = 0.78 ± 0.27). Although GWASs were limited by sample size (Ngeneric = 6269, Nexposed = 6365), compared to an independent OUD GWAS (N = 425 944), the dilution value for the two GWAS was not different from 1, suggesting no major impact of phenotypic misclassification. This study represents the first effort to enhance OUD genetic research through optimization of control definitions using EHR data. Generic controls ascertained within the US health systems, where exposure to prescription opioids is high, offer a practical alternative for genetic studies of OUD.
Background:Electronic health record (EHR)-based phenotyping underpins genome-wide association studies, yet current ICD-code phenotypes rely heavily on manually curated lists such as Phecodes. These definitions are labour-intensive to maintain, inherently subjective, and may omit clinically relevant diagnostic codes, reducing study power. Advances in text embedding models offer an opportunity to automate and standardize ICD-based phenotype construction. Methods:We developed Phecoder, an ensemble of pre-trained text embedding models that rank ICD codes by similarity from free-text phenotype descriptions. Nine embedding models and multiple unsupervised ensemble rank-fusion methods were evaluated against 1,125 PhecodeX phenotypes. Retrieval performance was assessed using recall and average precision at top-100 (R@100, AP@100). Expert clinical review of six neuropsychiatric phenotypes was undertaken to identify relevant ICD codes absent from PhecodeX. Cohort sizes under these new definitions were compared with PhecodeX across sex and ancestry strata in the Million Veteran Program (MVP). Findings:Among individual models, Qwen3-Embedding-4B achieved the highest median recall (R@100 = 0.86). Ensemble rank-fusion further improved R@100 by 3%, and median AP@100 by 8%. Expert review confirmed that Phecoder retrieved additional clinically relevant ICD codes beyond PhecodeX across all six neuropsychiatric case studies. Median potential case expansion increased by 200%, with 700% increases for bipolar disorder and 2000% increase for eating disorders. Interpretation:Manually defining ICD phenotypes has been critiqued as subjective, potentially yielding overly restrictive definitions that miss relevant codes. To address this issue, Phecoder algorithmically identifies relevant codes for ICD-based phenotyping. Phecoder extracts relevant ICD codes to expand the potential case pool across different demographic groups. Phecoder is easily applicable to future ICD-code releases and across different ICD coding versions that are used in different countries. Taken together, Phecoder has the potential to improve reproducibility in EHR data research. Funding:This research was supported by the Department of Veterans Affairs MVP (MVP-000, MVP-076 and MVP-096). The MVP is supported by the Office of Research and Development, Department of Veterans Affairs. The authors thank the MVP staff, researchers, and volunteers, who have contributed to MVP, and especially who previously served their country in the military and now generously agreed to enroll in the study (see mvp.va.gov for more information). The contents do not represent the views of the U.S. Department of Veterans Affairs or the United States Government. This study was supported by the Veterans Affairs Merit grants: BX006500 (to D.B.) and BX004189 (to P.R.). This work was supported by the National Institutes of Health (NIH): R01MH125246 (to P.R.), R01AG078657 (to G.V.), R01AG067025 (to P.R.), and U24AG087563 (to P.R.).
Schizophrenia and related psychoses occur in all human populations, with the highest rates of diagnosis among Black individuals and those of mainly African ancestry1. Decades of research have established a highly heritable and polygenic basis for schizophrenia, which is mostly shared across populations2-4. However, a recruitment bias towards European cohorts5 has led to discoveries that are poorly generalizable to African populations. This exclusion of the world's most genetically diverse populations narrows our understanding of disease biology and risks exacerbating health disparities. Here we show that electronic health records linked with genomic data from the Million Veteran Program (MVP)6-a national research programme that looks at the effects of genes, lifestyle, military experiences and exposures on the health and wellness of veterans-enable a comprehensive assessment of schizophrenia genetics in populations of African ancestry in the USA. We identify ancestry-independent associations in African populations and expand the catalogue of implicated regions by more than 100 loci. Through statistical fine-mapping and integrative transcriptomic analyses, we refine disease-associated signals to consensus genes with convergent neurobiological functions. These findings provide a much-needed view of schizophrenia's genetic architecture in populations of African ancestry, and offer biological insights that both extend previous work and broaden its global relevance.
Importance:Researchers commonly use counts of diagnostic codes from EHR-linked biobanks to infer phenotypic status. However, these approaches overlook temporal changes in EHR data, such as the discontinuation or "dropout" of diagnostic codes, which may exacerbate disparities in genomics research, as EHR data quality can be confounded with demographic attributes. Objective:To address this, we propose modeling diagnostic code dropout in EHR data to inform phenotyping for schizophrenia in genomic analyses. Design:We develop and test our diagnostic dropout model by analyzing EHR data from individuals with prior schizophrenia diagnoses. We further validate model performance on a subset of patients whose diagnoses were attained through chart review. Using PRS-CS and existing GWAS summary statistics, we first extrapolate polygenic weights. Then, we apply our dropout model's outputs to construct a data-driven filter defining our target cohort for measuring polygenic score performance. Setting:Our analysis utilizes EHR and genomic data from the Million Veteran Program. Participants:To model diagnostic dropout in schizophrenia, we leverage data from 12,739 patients with a history of schizophrenia, after excluding outliers. For polygenic score analyses, we incorporate data from a potential pool of 8,385 European ancestry and 6,806 African ancestry patients with a history of schizophrenia. Main outcomes and measures:We compare the performance of our diagnostic dropout model with alternative methodologies both in predicting diagnostic dropout on a holdout set, as well as on chart review labeled data. Using the top differential diagnosis predictors in our model, we select relevant cases by filtering out patients with a prior history of mood or anxiety disorders. We then test the impact of applying different filters for measuring polygenic score performance. Results:When evaluated on chart review-labeled data, our model improves the area under the precision-recall curve (AUPRC) by 9.6% compared to competing methods. By applying our data-driven filter for schizophrenia, we achieve a 62% increase in the association effect size when transferring a European polygenic score to an African ancestry target cohort. Conclusions and Relevance:These findings highlight the potential of modeling diagnostic code dropout to enhance the phenotypic quality of EHR-linked biobank data, advancing more equitable and accurate genomics research across diverse populations.
Background:Functional seizures (FS) are paroxysmal episodes that phenotypically resemble epileptic seizures but are not associated with brain epileptiform discharges; they are also known as psychogenic nonepileptic seizures. The exact etiology and pathophysiology of FS is unknown; however, trauma and stress-related disorders are known risk factors. Methods:We used a validated algorithm applied to electronic health records to identify individuals with FS in 6 international biobanks and hospital sites. We conducted a multisite FS genome-wide association study (GWAS) meta-analysis, including 10,910 FS cases (9040 predicted European ancestry [EA] and 1870 predicted African ancestry [AA]) and 664,500 (561,150 EA and 103,350 AA) control participants. Results:The EA meta-analysis identified significant single nucleotide polymorphism-based heritability on the liability scale of 2.21% (SE = 0.015%, p = 10-3; assumed FS prevalence = 0.14%), but no genome-wide significant loci. Nominal associations emerged from the EA GWAS within 16q23.3 (CDH13 intronic variant rs8056064, effect allele = A, z = -4.762, p = 1.92 × 10-6) and from the AA GWAS within 17q21.2 (rs34380994, effect allele = T, z = 5.28, p = 1.32 × 10-7). Significantly associated gene sets included magnesium ion transport, mitochondrial membrane complexes, and RNA polymerase II preinitiation complex assembly (Bonferroni-corrected p values = .0095, .012, and .03, respectively). MAGMA gene property analysis for tissue specificity showed significant enrichment of FS-associated genes within cerebellum-expressed genes (beta = 0.019, SE = 0.0059, p = 7.1 × 10-3). Conclusions:To our knowledge, this is the first GWAS of FS, and our results support a genetic basis of FS. Future large-scale genetic research studies are needed to corroborate these findings and identify genetic variants associated with FS.
Background Recent genome-wide association studies (GWAS) have identified genetic loci linked with Opioid Use Disorder (OUD), revealing promising therapeutic targets. However, variability in disease definitions and population structures among cohorts often dilute genetic association signals, a challenge known as effect size dilution. Additionally, existing OUD GWASs lack cell-type specificity and comprehensive cross-ancestry analyses. To address effect size dilution and enhance discovery power, our study employs an innovative strategy to mitigate phenotypic heterogeneity. Furthermore, by combining cell-type-specific gene expression data and performing comprehensive cross-ancestry analyses, we aim to pinpoint causal genetic variants associated with OUD. This integrative approach advances the understanding of gene dysregulation mechanisms underlying OUD in diverse populations. Methods To correct for effect size dilution, we employed PheMED, a cutting-edge statistical methodology designed to enhance the power of GWAS discoveries. Preliminary analyses were conducted on the largest reported OUD GWAS data from the Million Veterans Program for African and European genetic ancestries. We further incorporated state-of-the-art expression quantitative trait locus (eQTL) reference panels to map gene activity patterns to specific brain cell types. Mendelian Randomization, a causal inference statistical framework, was utilized to integrate cell-type-specific genetic variants that have direct targets of clinically available compounds. Results Our results demonstrate that correcting for effect size dilution using PheMED significantly enhances gene discovery power across ancestry. Leveraging eQTL data within a Mendelian Randomization framework revealed several novel and previously reported genes as potential treatment targets for OUD in multiple ancestral groups. For African genetic ancestry, RTN4 and EPHX2 were identified in astrocytes involved in lipid metabolism and neural signaling. For European genetic ancestry, NCAM1 was found in inhibitory neurons and GABRA2 in excitatory neurons previously linked to heroin dependence. Discussion By correcting for phenotypic heterogeneity and utilizing cell-type-specific gene expression data, we reveal novel mechanistic insights into OUD and identify actionable gene targets for therapeutic development. This research has the potential to expand personalized medicine approaches for OUD treatment and improve reproducibility in genetic studies involving diverse populations.
Neurodegenerative diseases and serious mental illnesses often exhibit overlapping characteristics, highlighting the potential for shared underlying mechanisms. To facilitate a deeper understanding of these diseases and pave the way for more effective treatments, we have generated a population-scale multi-omics dataset consisting of genotype and single-nucleus transcriptome data from the prefrontal cortex of frozen human brain specimens. Encompassing over 6.3 million nuclei from 1,494 donors, our dataset represents a diverse range of neurodegenerative and serious mental illnesses, including Alzheimer's and Parkinson's diseases, schizophrenia, bipolar disorder and diffuse Lewy body dementia, as well as neurotypical controls. Our dataset offers a unique opportunity to study disease interactions, as 21% of donors had comorbid diagnoses of two or more major brain disorders. Additionally, it includes detailed phenotypic information on neuropsychiatric symptoms, such as apathy and weight loss, which commonly accompany Alzheimer's disease and related dementias. We have performed stringent preprocessing and quality controls, ensuring the reliability and usability of the data. As a commitment to fostering collaborative research, we provide this valuable resource as an online repository, enabling widespread analyses across the scientific community.
Neurodegenerative and neuropsychiatric diseases impose a significant societal and public health burden. However, our understanding of the molecular mechanisms underlying these highly complex conditions remains limited. To gain deeper insights into the etiology of different brain diseases, we used specimens from 1,494 unique donors to generate a population-scale single-cell transcriptomic atlas of the human dorsolateral prefrontal cortex (DLPFC), comprising over 6.3 million individual nuclei. The cohort includes neurotypical controls as well as donors affected by eight common and complex brain disorders: Alzheimer's disease (AD), diffuse Lewy body disease (DLBD), vascular dementia (Vas), Parkinson's disease (PD), tauopathy, frontotemporal dementia, schizophrenia, and bipolar disorder. We show that inter-individual variation accounts for a substantial portion of gene expression variation in the DLPFC. By comparing transcriptomic variation across diseases, we reveal universal signatures enriched in basic cellular functions such as mRNA splicing and protein localization. After discounting these cross-disease signatures, we show strong genetic and transcriptomic concordance among AD, DLBD, Vas, and PD, largely driven by alteration of synaptic signaling functions in neurons. Furthermore, we characterize transcriptomic variation among different AD phenotypes that were distinct from healthy aging. We uncover mitigating effects of interneurons and aggravating effects of immune and vascular cells in AD dementia. Further exploring the effect of the neuropsychiatric symptoms frequently accompanying AD, we identify a link to deep layer excitatory neurons. By constructing transcriptome trajectories that capture AD progression, we show cell-type specific responses implicated in early and late stages of AD. Our atlas provides an unprecedented perspective of the transcriptomic landscape in neurodegenerative and neuropsychiatric diseases, shedding light on shared and distinct processes involving the neuro-immune-vascular systems, and identifying potential targets for therapeutic intervention.
Neuropsychiatric and neurodegenerative disorders exhibit cell–type–specific characteristics 1–8, yet most transcriptome–wide association studies have been constrained by the use of homogenate brain tissue9–11, limiting their resolution and power. Here, we present a single–nucleus transcriptome–wide association study (snTWAS) leveraging single–nucleus RNA sequencing of over 6 million nuclei from the dorsolateral prefrontal cortex of 1,494 donors across three ancestries–European, African, and Admixed American. We constructed ancestry–specific single–nucleus–derived transcriptomic imputation models (snTIMs) including up to 27 non–overlapping cellular populations, enhancing the resolution of genetically regulated gene expression (GReX) in the brain and uncovering novel gene–trait associations across 12 neuropsychiatric and neurodegenerative traits. Our snTWAS framework revealed cell–type–specific dysregulation of GReX, identifying over 4,000 novel gene–trait associations not detected by bulk tissue approaches. By applying these snTIMs to the Million Veteran Program, we validated major findings and explored the pleiotropy of cell–type–specific GReX, revealing cross–ancestry concordance and fine–mapping causal genes. This approach enhances the discovery of biologically relevant pathways and gene targets, highlighting the importance of cell–type resolution and ancestry–specific models in understanding the genetic architecture of complex brain disorders. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research is based on data from the Million Veteran Program, Office of Research and Development, Veterans Health Administration, and was supported by award I01BX004189. This publication does not represent the views of the Department of Veteran Affairs or the United States Government. We thank the participants of the Million Veteran Program, the scientists, clinicians and supportive staff involved in the construction of this biobank, and the scientific computing staff for the expertise that they provided. We thank the computational resources and staff expertise provided by the Scientific Computing at the Icahn School of Medicine at Mount Sinai. This study was also supported by the National Institutes of Health (NIH), Bethesda, MD under award numbers R01AG067025 (PR), R01AG082185 (PR), K08MH122911 (GV), R01AG078657 (GV), BX004189 (PR), R01AG065582 (PR), R01AG067025 (PR), R01MH125246 (PR) and T32MH087004 (KT). Human tissues were obtained from the NIH NeuroBioBank at the Mount Sinai Brain Bank (MSSM; supported by NIMH-75N95019C00049), the Rush Alzheimer's Disease Center (RADC; funding: P30AG10161, P30AG72975, R01AG15819, R01AG17917, R01AG22018, U01AG46152, and U01AG61356), and NIMH-IRP Human Brain Collection Core (HBCC, project # ZIC MH002903). This work was supported in part through the computational and data resources and staff expertise provided by Scientific Computing and Data at the Icahn School of Medicine at Mount Sinai and supported by the Clinical and Translational Science Award (CTSA) grant UL1TR004419 from the National Center for Advancing Translational Sciences. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved by the VA Central Institutional Review Board (IRB), and participating studies received approval from their respective IRBs. All participants provided written informed consent. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All results are included either in the main text or provided in supplementary tables or data.
Large-scale genome-wide association studies of schizophrenia have uncovered hundreds of associated loci but with extremely limited representation of African diaspora populations. We surveyed electronic health records of 200,000 individuals of African ancestry in the Million Veteran and All of Us Research Programs, and, coupled with genotype-level data from four case-control studies, realized a combined sample size of 13,012 affected and 54,266 unaffected persons. Three genome-wide significant signals - near PLXNA4, PMAIP1, and TRPA1 - are the first to be independently identified in populations of predominantly African ancestry. Joint analyses of African, European, and East Asian ancestries across 86,981 cases and 303,771 controls, yielded 376 distinct autosomal loci, which were refined to 708 putatively causal variants via multi-ancestry fine-mapping. Utilizing single-cell functional genomic data from human brain tissue and two complementary approaches, transcriptome-wide association studies and enhancer-promoter contact mapping, we identified a consensus set of 94 genes across ancestries and pinpointed the specific cell types in which they act. We identified reproducible associations of schizophrenia polygenic risk scores with schizophrenia diagnoses and a range of other mental and physical health problems. Our study addresses a longstanding gap in the generalizability of research findings for schizophrenia across ancestral populations, underlining shared biological underpinnings of schizophrenia across global populations in the presence of broadly divergent risk allele frequencies.
Objective: Treatment-resistant depression (TRD) occurs in roughly one-third of all individuals with major depressive disorder (MDD). Although research has suggested a significant common variant genetic component of liability to TRD, with heritability estimated at 8% when compared with nontreatment-resistant MDD, no replicated genetic loci have been identified, and the genetic architecture of TRD remains unclear. A key barrier to this work has been the paucity of adequately powered cohorts for investigation, largely because of the challenge in prospectively investigating this phenotype. The objective of this study was to perform a wellpowered genetic study of TRD. Methods: Using receipt of electroconvulsive therapy (ECT) as a surrogate for TRD, the authors applied standard machine learning methods to electronic health record data to derive predicted probabilities of receiving ECT. These probabilities were then applied as a quantitative trait in a genome-wide association study of 154,433 genotyped patients across four large biobanks. Results: Heritability estimates ranged from 2% to 4.2%, and significant genetic overlap was observed with cognition, attention deficit hyperactivity disorder, schizophrenia, alcohol and smoking traits, and body mass index. Two genome-wide significant loci were identified, both previously implicated in metabolic traits, suggesting shared biology and potential pharmacological implications. Conclusions: This work provides support for the utility of estimation of disease probability for genomic investigation and provides insights into the genetic architecture and biology of TRD.
Parkinson's Disease (PD) is a debilitating neurodegenerative disorder, characterized by motor and cognitive impairments, that affects >1% of the population over the age of 60. The pathogenesis of PD is complex and remains largely unknown. Due to the cellular heterogeneity of the human brain and changes in cell type composition with disease progression, this complexity cannot be fully captured with bulk tissue studies. To address this, we generated single-nucleus RNA sequencing and whole-genome sequencing data from 100 postmortem cases and controls, carefully selected to represent the entire spectrum of PD neuropathological severity and diverse clinical symptoms. The single nucleus data were generated from five brain regions, capturing the subcortical and cortical spread of PD pathology. Rigorous preprocessing and quality control were applied to ensure data reliability. Committed to collaborative research and open science, this dataset is available on the AMP PD Knowledge Platform, offering researchers a valuable tool to explore the molecular bases of PD and accelerate advances in understanding and treating the disease.
With the advent of healthcare-based genotyped biobanks, genome-wide association studies (GWAS) leverage larger sample sizes, incorporate patients with diverse ancestries and introduce noisier phenotypic definitions. Yet the extent and impact of phenotypic misclassification on large-scale datasets is not currently well understood due to a lack of statistical methods to estimate relevant parameters from empirical data. Here, we develop a statistical method and scalable software, PheMED, Phenotypic Measurement of Effective Dilution, to quantify phenotypic misclassification across GWAS using only summary statistics. We illustrate how the parameters estimated by PheMED relate to the negative and positive predictive value of the labeled phenotype, compared to ground truth, and how misclassification of the phenotype yields diluted effect-sizes of variant-phenotype associations. Furthermore, we apply our methodology to detect multiple instances of statistically significant dilution in real-world data. We demonstrate how effective dilution biases downstream GWAS replication and heritability analyses despite utilizing current best practices, and provide a dilution-aware meta-analysis approach that outperforms existing methods. Consequently, we anticipate that PheMED will be a valuable tool for researchers to address phenotypic data quality issues both within and across cohorts.
Schizophrenia (SCZ) and related psychoses occur in all human populations, but are diagnosed most frequently among Black and African ancestry individuals. Environmental exposures and adversities, disparities in access to care, and historical trends of over- and racialized diagnosis contribute to this discrepancy. Nonetheless, African diaspora populations remain underrepresented in large scale research initiatives, limiting the potential benefit of new biological insights to the communities most burdened by these illnesses. Building on our recent work in the Million Veteran Program (MVP), we undertook a collaborative effort to compile and harmonize extant genotypic and phenotypic data for admixed African ancestry individuals. These studies included the new All of Us (AOU) initiative, and the historic BiGS, COGS, GPC, PAARTNERS, and MGS studies; 15,101 SCZ, 11643 bipolar I (BIP), and 63,212 control participants were available for analysis. We applied a range of genomic methodologies including trans-ancestry meta-analysis and fine-mapping, polygenic risk score (PRS) profiling, TWAS, and genome-based restricted maximum likelihood (GREML). We combined African ancestry results with published findings based on European and East Asian populations to explore convergent (and divergent) effects in the most diverse genetic analysis of these disorders to date. In the discovery phase, we identified a secondary association in GRIN2A (P=5.05e-8); four loci attained genome-wide significance in the African ancestry meta-analysis. Across 270 PGC3 loci, 65% showed the same direction of allelic effect in African American veterans (P=1e-7), compared to 90% in European Americans (P=9e-48); importantly, when considering only European or African ancestry tracts (i.e., haplotypes) in African Americans, 69% (P=3.4e-10) and 58% (P=0.007) of index SNPs showed a consistent direction of allelic effect. We observed fewer total and novel associations when combining African results with published European findings than observed when combining European or East Asian datasets. More intriguingly, fine-mapping of PGC3 loci saw the total number of credible SNPs reduced by 13% following meta-analysis with African ancestry based results, and by 20% when only African tracts were analyzed. We also observed more drastic improvements in fine-mapping resolution, including narrowing a gene-dense signal at 12q24.3 down to a single locus, BCL7A, and honing of gene-spanning signal at CACNA1I to a singular intron. Notably, we did not observe cross-ancestry support for the MHC locus on chromosome 6p21; imputation of structural C4 alleles highlighted a remarkable "shift" with respect to copy number distributions, reflecting duplication event subsequent to human migrations out-of-Africa. Our expanded analyses of SCZ and BIP in African ancestry populations highlight challenges and opportunities of enhanced diversity in neuropsychiatric genetics research. We explore the implications of ancestry-based disparities in representation and generalizability of genetic effects. Leveraging two large-scale EHRs, we explore the broad pleiotropy of SCZ and BIP risk alleles, and benchmark the relevance of current instruments for risk stratification, and predicting hospitalization.
Binge eating disorder (BED) is the most common eating disorder, yet its genetic architecture remains largely unknown. Studying BED is challenging because it is often comorbid with obesity, a common and highly polygenic trait, and it is underdiagnosed in biobank data sets. To address this limitation, we apply a supervised machine-learning approach (using 822 cases of individuals diagnosed with BED) to estimate the probability of each individual having BED based on electronic medical records from the Million Veteran Program. We perform a genome-wide association study of individuals of African ( n = 77,574) and European ( n = 285,138) ancestry while controlling for body mass index to identify three independent loci near the HFE , MCHR2 and LRP11 genes and suggest APOE as a risk gene for BED. We identify shared heritability between BED and several neuropsychiatric traits, and implicate iron metabolism in the pathophysiology of BED. Overall, our findings provide insights into the genetics underlying BED and suggest directions for future translational research.
Background: Serious mental illnesses, including schizophrenia, bipolar disorder and depression are heritable, highly multifactorial disorders and major causes of disability worldwide. Polygenic risk scores (PRS) aggregate variants identified from genome-wide association studies (GWAS) into individual-level estimates of liability, and are a promising tool for clinical risk stratification. Methods: By leveraging the VA’s extensive electronic health record (EHR) and a cohort of 9 378 individuals with confirmed diagnoses of schizophrenia or bipolar I disorder, we validated automated case-control assignments based on ICD-9/10 codes, and benchmarked the performance of current PRS for schizophrenia, bipolar disorder, and major depression in 400 000 Million Veteran Program (MVP) participants. We explored broader relationships between PRS and 1 650 disease categories via phenome-wide association studies (PheWAS). Finally, we applied genomic structural equation modeling (gSEM) to derive novel PRS indexing common and disorder-specific latent genetic factors. Findings: Among 3 953 and 5 425 individuals with diagnoses of schizophrenia or bipolar disorder type I that were confirmed by structured clinical interviews, 95% were correctly identified using ICD-9/10 codes (2 or more). Current PRS were robustly associated with case status in European (p <10-254) and African (p<10-5) participants and were higher among more frequently hospitalized patients (p<10-4). PheWAS confirmed previous associations among higher neuropsychiatric PRS and elevated risk for psychiatric and physical health problems and extended these findings to African Americans. Interpretation: Using diagnoses confirmed by in-person structured clinical interviews and current neuropsychiatric PRS, we demonstrated the validity of an EHR-based phenotyping approach in US veterans, highlighting the potential of PRS for disentangling biological and mediated pleiotropy. Funding Information: Department of Veterans Affairs Cooperative Studies Program (CSP) #572; Million Veteran Program (MVP-000, MVP-006); Office of Research and Development, Department of Veterans Affairs. Declaration of Interests: Dr. Bigdeli is the recipient of a 2019 NARSAD Young Investigator Grant (#28276). Dr. Voloudakis is supported by the National Institutes of Health (NIH) under award number K08MH122911 and is the recipient of a 2020 NARSAD Young Investigator Grant (#29350). We are grateful to Drs. R. Karlsson Linnér and T.T. Mallard for sharing their script for plotting PheWAS results, which served as the basis of the figure presented herein.Dr. Harvey has served as a consultant to multiple pharmaceutical companies and device manufacturers on phase 2 or 3 treatment development; this consulting work has been determined to be unrelated to the content of the paper. No other authors report any relevant conflicts of interest. Ethics Approval Statement: This study was approved by the VA Central Institutional Review Board (IRB), and all patients provided written informed consent.
Binge-eating disorder (BED) is the most common eating disorder yet its genetic architecture remains largely unknown. Studying BED is challenging because it is often comorbid with obesity, a common and highly polygenic trait, and it is underdiagnosed in biobank datasets. To address this limitation, we apply a supervised machine learning approach to estimate the probability of each individual having BED based on electronic medical records from the Million Veteran Program. We perform a genome-wide association study on individuals of African (n = 77,574) and European (n = 285,138) ancestry while controlling for body mass index to identify three independent loci near the HFE, MCHR2 and LRP11 genes, which are reproducible across three independent cohorts. We identify genetic association between BED and several neuropsychiatric traits and implicate iron metabolism in the pathophysiology of BED. Overall, our findings provide insights into the genetics underlying BED and suggest directions for future translational research.
Accurate age of onset (AOO) measurement is vital to etiologic and preventive research. While AOO reports are known to be subject to recall error, few population-based studies have been used to investigate agreement in AOO reports over more than a decade. We examined AOO reports for depression, back/neck pain, and daily smoking, in a population-based cohort spanning 29 years. A stratified sample of participants from Zurich, Switzerland (n = 591) completed a psychiatric and physical health interview 7 times between 1979, at ages 20 (males) and 21 (females), and 2008. We used one-way ANOVA to estimate intraclass correlations (ICCs) and weighted mixed models to estimate mean change over time and test for interactions with sex and clinical characteristics. Stratum-specific ICCs among those with 2 + reports were 0.19 and 0.29 for depression, 0.46 and 0.35 for back pain, and 0.66 and 0.75 for smoking. The average yearly increases in AOO report from the wave of first 12-month diagnosis or reported smoking, estimated in mixed models, were 0.57 years (95% confidence interval: 0.35, 0.79) for depression, 0.44 (95%CI: 0.28, 0.59) years for back pain, and 0.08 (95%CI: 0.03, 0.14) years for smoking. Initial impairment and frequency of treatment were associated with differences in average yearly change for depression. There is substantial variability in AOO reports over time and systematic increase with age. The degree of increase may differ by outcome, and for some outcomes, by participant clinical characteristics. Future studies should identify predictors of AOO report stability to ultimately benefit etiologic and preventive research.