In an analysis of 173 multiplex families from the Portuguese Island Collection (PIC), we characterize the shared genetic architecture of serious mental illnesses (SMI) , including schizophrenia (SZ), bipolar disorder (BP), major depression (MDD), and autism (ASD). Within this cohort, co-segregation of psychotic and mood disorders occurred in 28% of families, while 7% demonstrated co-segregation of intellectual disability or ASD with SZ and mood disorder phenotypes. Whole-genome sequencing (WGS) was performed on a three-generation PIC family to identify rare, large-effect variants. We identified an extremely rare predicted loss-of-function (LoF) mutation in the Chromodomain Helicase DNA Binding Protein 2 (CHD2) gene. These findings highlight the utility of high-density multiplex families in founder populations for identifying rare, large-effect variants that span clinical diagnostic categories, with the identified CHD2 mutation suggesting that variation in a single neurodevelopmental gene may contribute to phenotypic heterogeneity across SMI. By combining population and family-based methodologies, this approach leverages shared genetic backgrounds and environments to provide a unique opportunity for cellular studies to explore the biological mechanisms underlying SMI, offering significant potential to inform future functional research and identify novel therapeutic targets.
Most genetic variants associated with complex heritability phenotypes lie in non-coding regions and are thought to influence disease risk by regulating gene expression. However, most transcriptome-wide association approaches primarily model local (cis) genetic effects, leaving much of gene regulation unexplained. Here, we show that incorporating distal (trans) regulatory effects improves the prediction of gene expression and the identification of disease-associated genes. Using RNA sequencing data from six human post-mortem brain regions, we developed INGENE and MODULE, two models capturing the combined influence of candidate trans-acting variants within gene coexpression networks. Integrating these models with conventional cis-based predictors improved gene expression imputation (maximum likelihood estimation, α = 0.05) for 18,744 genes across regions. Applying this framework to Psychiatric Genomics Consortium wave 3 genotypes identified 766 genes associated with schizophrenia (PFDR < 0.01), including 641 not previously reported by transcriptome-wide analyses. These findings highlight the contribution of distal regulatory mechanisms and gene network interactions to schizophrenia risk.
Here we developed and deployed the blended genome exome (BGE) method, a DNA library approach that generates low-pass whole-genome (1-4× mean depth) and deep whole-exome (30-40× mean depth) data in a single sequencing run. BGE is cost-effective, empowers most genomic discoveries possible with deep whole-genome sequencing and captures global common single-nucleotide polymorphism diversity. We applied BGE to sequence >53,000 samples from the PUMAS Project (Populations Underrepresented in Mental Illness Associations Studies), including African, African American and Latin American populations. Imputed genotypes showed high concordance with Illumina Global Screening Array calls (R2 ≥ 95% for minor allele frequency ≥1%; ≥90% for minor allele frequency <1%), with consistent performance across local ancestries in admixed cohorts. For protein-coding copy number variants, deletions and duplications spanning at least three exons had a positive predicted value of ~90% relative to deep whole-genome data. At ~28% of the cost of deep whole-genome sequencing, BGE provides a scalable, reliable platform to expand genomic discovery and equitable access to sequencing in underrepresented populations.
Non-coding genetic variants statistically associated with complex heritability phenotypes are thought to act primarily through transcriptome regulatory mechanisms. Predictions of gene expression in tissue like the human brain traditionally rely primarily on cis -eQTLs. Here, we introduce INGENE and MODULE, trans -eQTLs models designed to enhance the prediction of gene expression by capturing the collective impact of candidate trans -eQTLs acting within co-expression networks. Exploiting RNA-seq data in six post-mortem brain regions (amygdala, caudate nucleus, dorsal/subgenual anterior cingulate cortex, dorsolateral prefrontal cortex, and hippocampus), we validate our models on two testing datasets, demonstrating increased gene predictability compared to both an original cis -based model and to EpiXcan, the leading benchmark in cis -model performance. Integration of cis - and trans -predictions significantly improves gene-level expression imputation (MLE α= 0.05) for 18,744 genes across the six brain regions considered. Applying cis and trans models to PGC wave 3 genotypes identifies 766 SCZ-associated genes across brain regions (pFDR < .01), emphasizing the complementary nature of cis and trans predictions in trait association discovery. Of these genes, 641 represent novel transcriptome-wide associations with schizophrenia, highlighting the role of trans -heritability and genetic interactions underlying risk for this disorder, in addition to further supporting 125 previous candidates. ### Competing Interest Statement A. Bertolino received consulting fees from Biogen and lecture fees from Otsuka, Janssen, and Lundbeck. D. Weinberger serves on the scientific advisory boards of Sage Therapeutics and Pasithea Therapeutics. G. Pergola and G. C. Kikidis received lecture fees from Lundbeck. A. K. Malhotra is a consultant to Genomind, InformedDNA and Concert Pharmaceuticals. M. C. O Donovan, M. J. Owen, and J. T. R. Walters are supported by collaborative research grants from Takeda Pharmaceuticals. O. A. Andreassen is a consultant for HealthLytix and received speaker s honoraria from Lundbeck. C. Arango has been a consultant to or has received honoraria or grants from Acadia, Angelini, Gedeon Richter, Janssen Cilag, Lundbeck, Minerva, Otsuka, Roche, Sage, Servier, Shire, Schering Plough, Sumitomo Dainippon Pharma, Sunovion and Takeda. Research Projects of National Relevance 2020 (PRIN 2020; 2020WSCSLZ) Research Projects of National Relevance 2022 (PRIN 2022; 2022KXJYJA) Research Projects Of National Relevance PNRR 2022 (P2022HNBJX) The LIBD funded the collection and analysis of postmortem brain tissue
Schizophrenia and related psychoses occur in all human populations, with the highest rates of diagnosis among Black individuals and those of mainly African ancestry1. Decades of research have established a highly heritable and polygenic basis for schizophrenia, which is mostly shared across populations2-4. However, a recruitment bias towards European cohorts5 has led to discoveries that are poorly generalizable to African populations. This exclusion of the world's most genetically diverse populations narrows our understanding of disease biology and risks exacerbating health disparities. Here we show that electronic health records linked with genomic data from the Million Veteran Program (MVP)6-a national research programme that looks at the effects of genes, lifestyle, military experiences and exposures on the health and wellness of veterans-enable a comprehensive assessment of schizophrenia genetics in populations of African ancestry in the USA. We identify ancestry-independent associations in African populations and expand the catalogue of implicated regions by more than 100 loci. Through statistical fine-mapping and integrative transcriptomic analyses, we refine disease-associated signals to consensus genes with convergent neurobiological functions. These findings provide a much-needed view of schizophrenia's genetic architecture in populations of African ancestry, and offer biological insights that both extend previous work and broaden its global relevance.
Human induced pluripotent stem cells (hiPSC) and iPSC-differentiated neural cells, in combination with CRISPR editing, are commonly used for studying neurodevelopmental and other brain disorders. Female iPSCs undergo random X-chromosome inactivation (XCI) via epigenetic silencing by noncoding X inactive specific transcript (XIST). It is known that female iPSCs may lose XIST expression, leading to XCI erosion that affects both X-linked and autosomal gene expression. However, the effects of CRSIPR editing and neural differentiation on XCI erosion in iPSC-derived neurons and how this may confound a real-world transcriptomic analysis of differentially expressed genes (DEGs) are poorly understood. Here, leveraging bulk RNA-seq of hundreds of CRISPR-edited female iPSC lines from four donor lines for 66 genes and single-cell RNA-seq of iPSC-derived neurons of a subset of 42 edited genes, we investigated the effects of XCI erosion during CRISPR editing and in iPSC-derived neurons. We found that XCI erosion was variable in CRISPR-edited female iPSCs and largely preserved in iPSC-derived neurons. Like in iPSCs, XIST in neurons predominately influenced the expression of X-linked genes; however, its effect on autosomal genes was more pronounced in single neurons. Mechanistically, XIST epigenetically causes allelic imbalance of both X-linked and autosomal genes, with the former showing stronger allele-specific expression (ASE) bias. Notably, XIST-induced ASE bias exhibited a conserved positional pattern at loci affecting neurodevelopmental genes across different female lines and cell types. Finally, we demonstrated a confounding effect of XCI erosion on DEG analyses in iPSC-derived neurons. These results have significant implications in hiPSC modeling of neurodevelopmental and other brain disorders.
In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole-genome sequencing (∼27X) of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls (∼37X). To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of the association of BD with single variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD genome-wide association study loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole-genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.
Individuals with serious mental illness (SMI) contend with medical comorbidities like hypertension, lowering their life expectancies. Here we evaluated these comorbidities in the Genomic Psychiatry Cohort, a large, diverse sample of subjects with SMI and controls. We also assessed the relationship of these illnesses with factors such as ancestry. Participants were sorted into one of three groups: 1) schizophrenia and schizoaffective disorder-depressed type (SZ, n = 8188), 2) bipolar with and without psychosis and schizoaffective disorder-bipolar type (BD, n = 5577) and 3) controls (CON, n = 10,850). Logistic regression models provided the odds of subjects endorsing heart problems, hypertension, high blood sugar, high cholesterol, and cancer compared to CONs. Similar models were used to evaluate the effects of probable alcohol use disorder (pAUD), probable tobacco use disorder (pTUD), ancestry, and sex. Both SZ and BD had higher odds of all five comorbidities compared to CONs, with OR ranging between 1.61 and 2.38 for SZ and 2.08 to 3.14 for BD (p < 0.0033). Having both SMI and pAUD was associated with higher odds of heart problems, hypertension, and cancer, though for pTUD this association held true for heart problems and cancer in SZ only (p < 0.0033). When comparing African-ancestry (AA) and European-ancestry (EA), hypertension and high blood sugar impacted AA with SMI more, though high cholesterol and cancer impacted EA with SMI more. Our findings suggest that SZ and BD are associated with higher rates of medical comorbidities, and factors such as pAUD, pTUD, and ancestry may be implicated.
Bipolar disorder is a heritable mental illness with complex etiology. While the largest published genome-wide association study identified 64 bipolar disorder risk loci, the causal SNPs and genes within these loci remain unknown. We applied a suite of statistical and functional fine-mapping methods to these loci and prioritized 17 likely causal SNPs for bipolar disorder. We mapped these SNPs to genes and investigated their likely functional consequences by integrating variant annotations, brain cell-type epigenomic annotations, brain quantitative trait loci and results from rare variant exome sequencing in bipolar disorder. Convergent lines of evidence supported the roles of genes involved in neurotransmission and neurodevelopment, including SCN2A, TRANK1, DCLK3, INSYN2B, SYNE1, THSD7A, CACNA1B, TUBBP5, FKBP2, RASGRP1, FURIN, FES, MED24 and THRA among others in bipolar disorder. These represent promising candidates for functional experiments to understand biological mechanisms and therapeutic potential. Additionally, we demonstrated that fine-mapping effect sizes can improve performance of bipolar disorder polygenic risk scores across diverse populations and present a high-throughput fine-mapping pipeline.
Obsessive-compulsive disorder (OCD) affects ~1% of children and adults and is partly caused by genetic factors. We conducted a genome-wide association study (GWAS) meta-analysis combining 53,660 OCD cases and 2,044,417 controls and identified 30 independent genome-wide significant loci. Gene-based approaches identified 249 potential effector genes for OCD, with 25 of these classified as the most likely causal candidates, including WDR6, DALRD3 and CTNND1 and multiple genes in the major histocompatibility complex (MHC) region. We estimated that ~11,500 genetic variants explained 90% of OCD genetic heritability. OCD genetic risk was associated with excitatory neurons in the hippocampus and the cortex, along with D1 and D2 type dopamine receptor-containing medium spiny neurons. OCD genetic risk was shared with 65 of 112 additional phenotypes, including all the psychiatric disorders we examined. In particular, OCD shared genetic risk with anxiety, depression, anorexia nervosa and Tourette syndrome and was negatively associated with inflammatory bowel diseases, educational attainment and body mass index.
Anxiety disorders and substance use frequently co-occur, yet the moderating effects of sex and ancestry on these relationships remain understudied. This investigation examined associations between presumed panic disorder (pPD) and both presumed alcohol use disorder (pAUD) and tobacco use disorder (pTUD) in 10,953 individuals from the Genomic Psychiatry Cohort screened against severe mental illness. Our sample was demographically diverse (56% female; 55% European Ancestry, 45% African Ancestry), allowing robust comparison across these groups. Individuals with pPD ( n = 342) demonstrated significantly higher mean severity scores for both pAUD (1.26 vs. 0.33, p < 0.05) and pTUD (1.65 vs. 0.93, p < 0.05) compared with those without pPD. While female sex was associated with decreased risk for pAUD (B: −0.351, p < 0.05) compared with males, we observed no significant ancestry-based differences in substance use patterns among those with pPD. Two-way interaction analyses revealed that sex significantly moderated the relationship between pPD and pAUD (B: −0.97, p < 0.001), with the association being stronger among males than females. Additionally, comorbid presumed posttraumatic stress disorder was significantly associated with increased risk for both pAUD (B: 0.650, p < 0.05) and pTUD (B: 0.825, p < 0.05) but did not interact significantly with pPD. These findings advance our understanding of how biological sex influences the manifestation of comorbid panic and substance use disorders, offering clinical implications for assessment and treatment strategies that acknowledge sex-specific vulnerability patterns while highlighting the consistent relationship between these conditions across ancestral groups.
Background Accurate diagnosis of bipolar disorder (BPD) is difficult in clinical practice, with an average delay between symptom onset and diagnosis of about 7 years. A depressive episode often precedes the first manic episode, making it difficult to distinguish BPD from unipolar major depressive disorder (MDD). Aims We use genome-wide association analyses (GWAS) to identify differential genetic factors and to develop predictors based on polygenic risk scores (PRS) that may aid early differential diagnosis. Method Based on individual genotypes from case-control cohorts of BPD and MDD shared through the Psychiatric Genomics Consortium, we compile case-case-control cohorts, applying a careful quality control procedure. In a resulting cohort of 51 149 individuals (15 532 BPD patients, 12 920 MDD patients and 22 697 controls), we perform a variety of GWAS and PRS analyses. Results Although our GWAS is not well powered to identify genome-wide significant loci, we find significant chip heritability and demonstrate the ability of the resulting PRS to distinguish BPD from MDD, including BPD cases with depressive onset (BPD-D). We replicate our PRS findings in an independent Danish cohort (iPSYCH 2015, N = 25 966). We observe strong genetic correlation between our case-case GWAS and that of case-control BPD. Conclusions We find that MDD and BPD, including BPD-D are genetically distinct. Our findings support that controls, MDD and BPD patients primarily lie on a continuum of genetic risk. Future studies with larger and richer samples will likely yield a better understanding of these findings and enable the development of better genetic predictors distinguishing BPD and, importantly, BPD-D from MDD.
Bipolar disorder is a leading contributor to the global burden of disease1. Despite high heritability (60-80%), the majority of the underlying genetic determinants remain unknown2. We analysed data from participants of European, East Asian, African American and Latino ancestries (n = 158,036 cases with bipolar disorder, 2.8 million controls), combining clinical, community and self-reported samples. We identified 298 genome-wide significant loci in the multi-ancestry meta-analysis, a fourfold increase over previous findings3, and identified an ancestry-specific association in the East Asian cohort. Integrating results from fine-mapping and other variant-to-gene mapping approaches identified 36 credible genes in the aetiology of bipolar disorder. Genes prioritized through fine-mapping were enriched for ultra-rare damaging missense and protein-truncating variations in cases with bipolar disorder4, highlighting convergence of common and rare variant signals. We report differences in the genetic architecture of bipolar disorder depending on the source of patient ascertainment and on bipolar disorder subtype (type I or type II). Several analyses implicate specific cell types in the pathophysiology of bipolar disorder, including GABAergic interneurons and medium spiny neurons. Together, these analyses provide additional insights into the genetic architecture and biological underpinnings of bipolar disorder.
Importance:The clinical heterogeneity of bipolar disorder (BD) is a major obstacle to improving diagnosis, predicting patient outcomes, and developing personalized treatments. A genetic approach is needed to deconstruct the disorder and uncover its fundamental biology. Previous genetic studies focusing on broad diagnostic categories have been limited in their ability to parse this complexity. Objective:To test the hypothesis that clinically distinct subphenotypes of BD are associated with different underlying common variant genetic architectures. Design Setting and Participants:This multicenter study included a primary genome-wide association study (GWAS) of up to 23,819 bipolar disorder (BD) cases and 163,839 controls. These results were integrated via multi-trait analysis of GWAS (MTAG) with external summary statistics for BD (59,287 cases; 781,022 controls) and schizophrenia (SCZ; 53,386 cases; 77,258 controls). Sample overlap was statistically accounted for. Main Outcomes and Measures:The primary outcomes were the genetic dimensions underlying BD heterogeneity, differentiated by single nucleotide polymorphism (SNP)-heritability (h 2 SNP ), genetic correlations, genomic loci ( P ≤5×10 -8 ), and functional, cell-type, and gene-expression pathway analyses. Results:We identified four genetically-informed dimensions of BD: Severe Illness, Core Mania, Externalizing/Impulsive Comorbidity, and Internalizing/Affective Comorbidity. The analyses yielded up to 181 subphenotype-associated loci, 53 of which are novel. The Severe Illness Dimension was characterized by a unique neuro-immune signature (a protective association with HLA-DMB , P =2.50×10 -273 ) evident only when leveraging SCZ genetic data. The Internalizing/Affective dimension was associated with neurodevelopmental genes (e.g., DCC ). Notably, the rapid-cycling subphenotype showed a unique signature of strong negative selection, a finding not observed in other subphenotypes. Conclusions and Relevance:The clinical heterogeneity of bipolar disorder appears to be defined by a complex and multi-layered genetic architecture. The presented findings provide an empirical framework that may advance psychiatric nosology beyond its current diagnostic boundaries. These results may also inform future research to identify targets for personalized interventions. The delineation of these genetically-informed dimensions offers specific, biologically-grounded hypotheses for subsequent therapeutic discovery. Establishing such a framework is an essential step toward refining diagnostic criteria and developing more effective, personalized treatments. This work lays the foundation for a transition from a uniform treatment model to the paradigm of precision psychiatry. Key Points:Question: What are the distinct genetic architectures underlying the clinical heterogeneity of bipolar disorder?Findings: In this genetic study of 23,819 bipolar disorder (BD) cases and 163,839 controls, clinical heterogeneity mapped onto four genetically-informed dimensions. A severe illness dimension was defined by a neuro-immune signature ( HLA-DMB ) shared with schizophrenia. An affective comorbidity dimension was distinguished by neurodevelopmental pathways involving axonal guidance ( DCC ). Notably, the rapid-cycling phenotype showed evidence of purifying selection, suggesting influence by rare, highly penetrant alleles. Meaning: These findings provide a data-driven biological framework for bipolar disorder, guiding future research toward patient stratification and targeted therapeutics.
MOTIVATION:Many genetics studies report results tied to genomic coordinates of a legacy genome assembly. However, as assemblies are updated and improved, researchers are faced with either realigning raw sequence data using the updated coordinate system or converting legacy datasets to the updated coordinate system to be able to combine results with newer datasets. Currently available tools to perform the conversion of genetic variants have numerous shortcomings, including poor support for indels and multi-allelic variants, that lead to a higher rate of variants being dropped or incorrectly converted. As a result, many researchers continue to work with and publish using legacy genomic coordinates. RESULTS:Here we present BCFtools/liftover, a tool to convert genomic coordinates across genome assemblies for variants encoded in the variant call format with improved support for indels represented by different reference alleles across genome assemblies and full support for multi-allelic variants. It further supports variant annotation fields updates whenever the reference allele changes across genome assemblies. The tool has the lowest rate of variants being dropped with an order of magnitude less indels dropped or incorrectly converted and is an order of magnitude faster than other tools typically used for the same task. It is particularly suited for converting variant callsets from large cohorts to novel telomere-to-telomere assemblies as well as summary statistics from genome-wide association studies tied to legacy genome assemblies. AVAILABILITY AND IMPLEMENTATION:The tool is written in C and freely available under the MIT open source license as a BCFtools plugin available at http://github.com/freeseek/score.
Large-scale genome-wide association studies of schizophrenia have uncovered hundreds of associated loci but with extremely limited representation of African diaspora populations. We surveyed electronic health records of 200,000 individuals of African ancestry in the Million Veteran and All of Us Research Programs, and, coupled with genotype-level data from four case-control studies, realized a combined sample size of 13,012 affected and 54,266 unaffected persons. Three genome-wide significant signals - near PLXNA4, PMAIP1, and TRPA1 - are the first to be independently identified in populations of predominantly African ancestry. Joint analyses of African, European, and East Asian ancestries across 86,981 cases and 303,771 controls, yielded 376 distinct autosomal loci, which were refined to 708 putatively causal variants via multi-ancestry fine-mapping. Utilizing single-cell functional genomic data from human brain tissue and two complementary approaches, transcriptome-wide association studies and enhancer-promoter contact mapping, we identified a consensus set of 94 genes across ancestries and pinpointed the specific cell types in which they act. We identified reproducible associations of schizophrenia polygenic risk scores with schizophrenia diagnoses and a range of other mental and physical health problems. Our study addresses a longstanding gap in the generalizability of research findings for schizophrenia across ancestral populations, underlining shared biological underpinnings of schizophrenia across global populations in the presence of broadly divergent risk allele frequencies.