Background:Early initiation of alcohol, nicotine, cannabis, and other substances is a robust predictor of later substance use disorders and related psychopathology. We integrate time-varying environmental factors with polygenic risk scores (PRS) in a longitudinal framework to identify risk factors influencing substance initiation in adolescents. Methods:We analyzed data from the Adolescent Brain Cognitive Development (ABCD) Study® with repeated assessments from baseline through approximately four years of follow-up. For each substance (alcohol, nicotine, cannabis, and any substance), we defined a time-to-event outcome indexing age at first use. We assembled a high-dimensional panel of time-varying environmental covariates across multiple domains (family, neighborhood, school, mental health, cognition, and physical health etc.), along with time-invariant covariates and PRS for problematic alcohol use, cannabis use disorder, nicotine use disorder, and any substance use disorder. We constructed individual-level start-stop interval data using interview age in months and fitted time-varying Cox proportional hazards models with robust standard errors clustered by individual ID. First, we conducted univariate models for each predictor. Second, for each outcome, we fitted multivariable Cox models including core covariates (sex, age, site, ancestry principal components), all PRS, and selected predictors. In secondary analyses, we applied marginal structural models with inverse probability of treatment weighting to a subset of modifiable predictors to approximate causal effects under standard assumptions. Results:In univariate models, earlier initiation was broadly associated with multiple time-varying variables, including impulsivity and externalizing behaviors, sleep disturbance, parenting and monitoring, medication and caffeine use, school functioning and absenteeism, and cultural or value-based measures, alongside other mental health and behavioral factors. In multivariable Cox models, a smaller subset of environmental predictors remained robustly associated with the hazard of initiation across alcohol, nicotine, cannabis, and any substance, highlighting consistent signals in impulsivity traits, parental monitoring, and select health and lifestyle factors. PRS for alcohol use disorder (AUD), cannabis use disorder (CUD), nicotine dependence, and any SUD were positively associated with earlier initiation (hazard ratio [HR] > 1), with the strongest and most consistent signal observed for nicotine PRS (e.g., alcohol initiation HR ≈ 2.37; any-substance initiation HR ≈ 2.98). AUD PRS showed weaker associations for alcohol initiation but stronger associations for any-substance initiation. Causal analyses suggested that parental monitoring (PMQ mean), UPPS lack of planning, UPPS sensation seeking, and caffeine exposure may influence time to initiation: higher parental monitoring was protective (odds ratio [OR] ≈ 0.33-0.64), whereas higher impulsivity traits and caffeine exposure were associated with increased risk (OR ≈ 1.47-3.87) across outcomes, with conclusions robust across weighting specifications. Conclusions:Integrating time-varying environmental predictors with polygenic risk in a survival framework helps identify environmental factors most strongly associated with earlier substance use initiation beyond genetic liability. Follow-up causal analyses further highlight potentially actionable pathways, particularly parenting and monitoring and impulsivity-related traits, that may contribute to the developmental trajectory leading to adolescent substance use disorders.
Background:Substance use initiation in adolescence is influenced by both genetic and environmental factors; however, large-scale genetic studies often treat initiation as a binary outcome and underuse longitudinal timing information. Methods:We conducted time-to-event (survival) genome-wide association analyses (GWAS) of initiation for four outcomes-alcohol, nicotine, cannabis, and any substance use-using longitudinal follow-up data from the Adolescent Brain Cognitive Development (ABCD) Study. We performed ancestry-stratified GWAS within European (EUR), African (AFR), and Hispanic (HISP) groups, applying consistent quality control and covariate adjustment. Summary statistics were harmonized across ancestries and meta-analyzed using inverse-variance weighted fixed-effects and DerSimonian-Laird random-effects models. We evaluated genomic inflation and heterogeneity (Cochran's Q and I 2 ), identified independent lead variants at genome-wide and suggestive significance thresholds, and assessed cross-trait overlap of associated loci. Results:In the multi-ancestry meta-analysis, we observed suggestive association signals across traits (minimum p -values: alcohol ∼ 1 × 10 -7 , any ∼ 1 × 10 -7 , cannabis ∼ 5 × 10 -8 , nicotine ∼ 1 × 10 -8 ). Nicotine initiation showed one genome-wide significant variant in both fixed- and random-effects meta-analyses ( p < 5 × 10 -8 ). Across traits, suggestive loci demonstrated limited overlap, with the strongest concordance between alcohol and any substance use, consistent with shared liability. Heterogeneity statistics indicated that some loci exhibited cross-ancestry variation in effect estimates. Conclusions:Survival GWAS leveraging initiation timing can identify genetic signals that may be missed by binary designs and enables principled multi-ancestry synthesis. Our results highlight both shared and trait-specific genetic contributions to early substance initiation and provide a foundation for downstream functional annotation and integrative modeling with environmental risk factors. These findings demonstrate the value of incorporating developmental timing into genetic discovery and provide a framework for integrating longitudinal risk modeling with genomic analyses.
Background:Adolescent substance use initiation is shaped by multiple genetic and neurobiological factors. Externalizing liability, a transdiagnostic genetic dimension capturing shared predisposition to impulsivity, disinhibition, and related traits, is among the strongest polygenic predictors of early substance initiation. Yet how this genetic risk relates to brain structure and function, and whether baseline brain phenotypes statistically account for or instead act in parallel with genetic liability, remains unresolved. Methods:Using the ABCD Study, we analyzed an analytic cohort of 10,608 participants with genotype data, baseline multimodal neuroimaging-derived phenotypes (IDPs), and longitudinal substance initiation assessments. Outcome-specific models included up to 10,599 participants after complete-case filtering for survival variables and covariates. We implemented a multistage framework linking an externalizing polygenic risk score (extPRS) to baseline IDPs and longitudinal substance initiation outcomes, including alcohol, nicotine, cannabis, and any substance. Stage 1 screened extPRS-IDP associations using covariate-adjusted linear models with false discovery rate (FDR) control. Stage 2 estimated extPRS effects on time-to-initiation using Cox proportional hazards models. Stage 3 fit joint extPRS + IDP Cox models to identify IDPs that predicted initiation beyond extPRS. Stage 4 conducted bootstrap-based mediation analyses to quantify average causal mediation effects (ACME), average direct effects (ADE), and the proportion of the extPRS-initiation association statistically accounted for by individual IDPs. Results:Higher extPRS was robustly associated with earlier initiation across all substances: alcohol, hazard ratio (HR) = 1.13; nicotine, HR = 1.63; cannabis, HR = 1.67; and any substance, HR = 1.15. Thousands of extPRS-associated IDPs were identified at baseline, with highly concordant effect profiles across robustness specifications. In joint models, numerous IDPs independently predicted initiation timing above and beyond extPRS: 31 for alcohol, 32 for any substance, 137 for cannabis, and 459 for nicotine, with a replicated core set across specifications. Cannabis and nicotine initiation were jointly predicted by superficial white matter (SWM) microstructural integrity in sensorimotor cortex as a protective factor, and by irregular activity in a right-hemisphere region as a risk factor. Alcohol initiation was predicted by a largely distinct, strongly left-lateralized frontolimbic SWM intensity axis. Nicotine initiation additionally and uniquely involved restricted gray matter diffusion in the anterior cingulate cortex and subcallosal cortex. Despite these robust independent IDP associations, mediation analyses showed that indirect effects through individual baseline IDPs were very small in magnitude (ACME ≈ 10-4), accounting for less than 2% of the total extPRS effect, with FDR-significant mediation surviving only for alcohol and any-substance initiation. Conclusions:Within the scope of externalizing polygenic risk and baseline neuroimaging, the predominant pattern is one of largely parallel, additive contributions to adolescent substance initiation rather than a dominant genetic → brain → behavior pathway. Baseline brain features, particularly prefrontal functional variability and frontolimbic and sensorimotor white matter integrity, predict initiation risk beyond extPRS, indicating neurobiological vulnerabilities not captured by this genetic dimension. However, these baseline IDPs explain only a small fraction of the extPRS-initiation association, suggesting that externalizing genetic liability may operate through pathways not fully represented by cross-sectional baseline imaging. Whether other genetic risk dimensions, such as substance-specific PRS, or dynamic longitudinal brain measures show stronger mediation patterns remains an important open question.
Alcohol use disorder (AUD) is known to have a significant genetic component, yet there remains a gap between its heritability and findings from genome-wide association studies. One potential explanation for this could be genetic interactions, or epistasis, which remain largely unexplored in the context of AUD. We investigated the role of epistasis in AUD susceptibility among 742 American Indians. By analyzing 467 K variants in 3,736 genes and regulatory elements linked to AUD, we identified 97 interacting gene pairs significantly associated with AUD severity in an American Indian cohort. Five of these gene pairs: CNTNAP2-GRM8, CSMD1-DLGAP1, CSMD1-ERBB4, CSMD1-MAML2, and KCNQ5-ROBO2 - were replicated in All of Us research American Indian cohort (N = 5,037). These genes were enriched for immune system, cell adhesion, neuronal, and disease pathways. Their expressions were particularly enriched in midbrain GABAergic neurons. This large-scale epistasis study of AUD suggests that epistasis may contribute to the development of AUD.
Adolescent externalizing behavior is a major risk factor for later substance use and other psychiatric outcomes. Understanding its genetic architecture and its relationships with brain imaging phenotypes requires scalable genome-wide methods that can be applied to youth cohorts. Using data from the Adolescent Brain Cognitive Development (ABCD) Study®, we implemented a pipeline for conducting genome-wide association studies (GWAS) of longitudinal externalizing traits and multimodal imaging-derived phenotypes (IDPs). We performed quality-controlled genotype processing and constructed harmonized phenotype and covariate datasets. GWAS analyses were conducted using REGENIE in a two-step framework. In Step 1, ridge regression prediction models were trained using linkage disequilibrium (LD)-pruned variants. In Step 2, genome-wide association testing was performed for each phenotype. The analyses included three externalizing phenotypes-baseline, longitudinal mean, and longitudinal slope-and approximately 200 IDPs measured at baseline or summarized using their longitudinal means and slopes. We additionally constructed a custom LD reference panel using unrelated individuals and calculated LD scores using LD Score Regression software (LDSC). Genome-wide genetic correlations between externalizing traits and imaging phenotypes were subsequently estimated using cross-trait LD Score Regression. This exploratory study systematically evaluated genome-wide genetic correlations between regional cortical morphology and externalizing phenotypes during adolescence. Although several associations reached nominal statistical significance, none remained significant after correction for multiple comparisons. These results should not be interpreted as evidence for the absence of shared genetic architecture. Instead, the precision of the genetic-correlation estimates was limited by the available imaging GWAS sample size, uncertainty in SNP-based heritability estimates, and the large number of regional comparisons. Larger imaging-genetics samples and independent replication studies will be required to determine whether modest or regionally specific genetic correlations exist.
MOTIVATION:Epistasis, or genetic interaction, plays a crucial role in shaping complex traits and has been increasingly recognized for its widespread influence in genetic architectures. While epistasis detection has been extensively evaluated in case-control studies, its performance with quantitative phenotypes remains comparatively understudied. RESULTS:We identified and evaluated six epistasis detection methods applicable to quantitative trait analysis: EpiSNP, Matrix Epistasis, MIDESP, PLINK Epistasis, QMDR, and REMMA. Using the EpiGEN simulator, we generated synthetic datasets modeling four classes of pairwise SNP interactions-dominant, multiplicative, recessive, and XOR. We also assessed BOOST and MDR algorithms using discretized (case-control) versions of the same datasets. Performance varied notably by interaction type: REMMA achieved the highest overall detection rate (55%), particularly excelling with dominant interactions (100%). MDR excelled with multiplicative (57%) and XOR (69%) interactions. Meanwhile, EpiSNP attained the best performance for recessive interactions (67%). All methods except BOOST produced F1 scores below 0.05 for most interaction types. We further evaluated the methods using a real-world dataset. When applied to the Adolescent Brain Cognitive Development dataset to analyse the externalizing behavior phenotype, both PLINK Epistasis and PLINK BOOST identified SNPs within the DRD2 and DRD4 genes, consistent with previously reported genetic associations. Given the variability in tool performance across interaction types, no single method provides optimal detection across all scenarios. Leveraging multiple detection algorithms may therefore yield more comprehensive insights into epistatic effects in quantitative trait analyses. AVAILABILITY AND IMPLEMENTATION:All relevant code and simulated datasets can be found at github.com/staslist/Epistasis_Review repository.
Understanding the development of adolescent behavioral and mental health outcomes requires integrating genetic predisposition, environmental exposures, and neurobiological processes over time. Here, we present a unified quantitative framework that models the human body as a dynamic system, where genetic factors form the foundational state, environmental exposures act as time-varying inputs, the brain might serve as a mediation processor, and behavioral phenotypes emerge as system outputs. Using longitudinal data from the Adolescent Brain Cognitive Development (ABCD) Study, we construct harmonized multi-domain representations across six phenotypes: externalizing behavior, internalizing behavior, and four substance use initiation outcomes (alcohol, nicotine, cannabis, and any substance use). We integrate polygenic risk scores (PRS), multi-domain environmental features, and multimodal neuroimaging representations derived through stability selection and dimensionality reduction. Our framework supports both continuous longitudinal modeling and survival-based event modeling through a unified data structure. We further develop interpretable domain-level representations using principal components, weighted risk scores, and cluster-based summaries. These representations enable downstream modeling using survival analysis, state-space models, and machine learning approaches. This work establishes a scalable and interpretable framework for studying how genetic and environmental factors interact over time to shape behavioral outcomes, providing a foundation for identifying modifiable risk factors and informing early intervention strategies.
Psychiatric disorders are highly heritable and polygenic, influenced by environmental factors and often comorbid. Large-scale genome-wide association studies (GWASs) through consortium efforts have identified genetic risk loci and revealed the underlying biology of psychiatric disorders and traits. However, over 85% of psychiatric GWAS participants are of European ancestry, limiting the applicability of these findings to non-European populations. Latin America and the Caribbean, regions marked by diverse genetic admixture, distinct environments and healthcare disparities, remain critically understudied in psychiatric genomics. This threatens access to precision psychiatry, where diversity is crucial for innovation and equity. This Review evaluates the current state and advancements in psychiatric genomics within Latin America and the Caribbean, discusses the prevalence and burden of psychiatric disorders, explores contributions to psychiatric GWASs from these regions and highlights methods that account for genetic diversity. We also identify existing gaps and challenges and propose recommendations to promote equity in psychiatric genomics.
Background:Epistasis, or genetic interaction, has been increasingly recognized for its ubiquity and for its role in susceptibility to common human diseases, such as Alzheimer's. A wide variety of epistasis detection tools are currently available with several studies comparing the performance of methods suitable for case-control data. However, there is limited understanding of how well these tools perform with quantitative phenotypes. Methods:We identified six epistasis detection methods suitable for quantitative phenotype data: EpiSNP, Matrix Epistasis, MIDESP, PLINK Epistasis, QMDR, and REMMA. To evaluate these tools, we generated simulated datasets using EpiGEN. The datasets modeled various pairwise interactions between disease-associated SNPs, including dominant, multiplicative, recessive, and XOR interactions. Additionally, we assessed the BOOST and MDR algorithms on discretized (case-control) version of the datasets. These tools were then tested on the Adolescent Brain Cognitive Development (ABCD) dataset for the externalizing behavior phenotype. Results:Each tool exhibited strong performance for certain interaction types, but weaker performance for others. MDR achieved the highest overall detection rate of 60%, while EpiSNP had the lowest overall detection rate of 7%. MDR and MIDESP performed best at detecting multiplicative interactions with detection rates of 54% and 41% respectively. Both MDR and MIDESP were also effective at detecting XOR interactions with detection rates of 84% and 50% respectively. PLINK Epistasis, Matrix Epistasis, and REMMA excelled at detecting dominant interactions, all achieving a 100% detection rate. On the other hand, EpiSNP was particularly effective at detecting recessive interactions with a detection rate of 66%. When analyzing the ABCD dataset, Plink Epistasis and Plink BOOST identified SNPs within the DRD2 and DRD4 genes, which have been previously linked to externalizing behavior. Conclusion:Since no single method consistently outperforms others across all types of epistasis, and given that the specific types of epistasis present in a dataset are often unknown, it may be more effective to use multiple epistasis detection algorithms in combination to obtain comprehensive results.
Large disparities in the prevalence of cannabis use disorder (CUD) exist across ethnic groups in the U.S. Despite large GWAS meta-analyses identifying numerous genome-wide significant loci for CUD in European descents, little is known about other ethnic groups. While most GWAS and SNP-heritability studies focus on common genomic variants, rare and low-frequency variants, particularly those altering proteins, are known to be enriched for the heritability of complex traits and may contribute to disease in different ways across populations, either through converging or alternative pathways. In this study, we examined three populations including European Americans (EA) and two understudied populations: American Indians (AI) and Mexican Americans (MA). We focused on rare and low frequency functional variants in genes and pathways, and performed association analysis with CUD severity. We identified 10 significant loci in AI, the ARSA gene in MA, three significant pathways in MA, and one in EA associated with CUD severity. Notably, pathways related to arylsulfatases activation and heparan sulfate degradation were supported by both EA and MA, with additional evidence from AI. The integrin beta-1 cell surface interaction pathway, involved in cell adhesion, was uniquely significant in MA. Several immune-related pathways were also found, including an autoimmune condition significant in MA with evidence from EA as well, and a p38-gamma/delta mediated signaling pathway supported across all three cohorts. Although each population displayed distinct pathways linked to CUD, overlapping genes in top pathways suggested shared genetic factors, further highlighting the importance of considering diverse populations in genetic research on cannabis use disorder.
American Indians (AI) demonstrate the highest rates of both suicidal behaviors (SB) and alcohol use disorders (AUD) among all ethnic groups in the US. Rates of suicide and AUD vary substantially between tribal groups and across different geographical regions, underscoring a need to delineate more specific risk and resilience factors. Using data from over 740 AI living within eight contiguous reservations, we assessed genetic risk factors for SB by investigating: (1) possible genetic overlap with AUD, and (2) impacts of rare and low-frequency genomic variants. Suicidal behaviors included lifetime history of suicidal thoughts and acts, including verified suicide deaths, scored using a ranking variable for the SB phenotype (range 0-4). We identified five loci significantly associated with SB and AUD, two of which are intergenic and three intronic on genes AACSP1, ANK1, and FBXO11. Nonsynonymous rare and low-frequency mutations in four genes including SERPINF1 (PEDF), ZNF30, CD34, and SLC5A9, and non-intronic rare and low-frequency mutations in genes OPRD1, HSD17B3 and one lincRNA were significantly associated with SB. One identified pathway related to hypoxia-inducible factor (HIF) regulation, whose 83 nonsynonymous rare and low-frequency variants on 10 genes were significantly linked to SB as well. Four additional genes, and two pathways related to vasopressin-regulated water metabolism and cellular hexose transport, also were strongly associated with SB. This study represents the first investigation of genetic factors for SB in an American Indian population that has high risk for suicide. Our study suggests that bivariate association analysis between comorbid disorders can increase statistical power; and rare and low-frequency variant analysis in a high-risk population enabled by whole-genome sequencing has the potential to identify novel genetic factors. Although such findings may be population specific, rare functional mutations relating to PEDF and HIF regulation align with past reports and suggest a biological mechanism for suicide risk and a potential therapeutic target for intervention.
Background Although alcohol use disorder (AUD) is known to be significantly influenced by genetics, there is a notable gap between its heritability (estimated to be ∼50%) and the outcomes of genome-wide association studies (GWAS): only ∼10% of the AUD variation can be attributed to the additive effects of common genetic variants. One potential explanation for this disparity could be found in genetic interactions (GxG), or epistasis, a factor that has been largely unexplored in addiction research, primarily due to computational and statistical challenges. Our research sought to investigate how epistasis influences susceptibility to AUD in a population of American Indians, where 70% of individuals were diagnosed with AUD. American Indians as a whole exhibit the highest rates of AUD among all ethnic groups in the United States, although this demographic remains understudied. Methods We first identified a set of genes linked to AUD through disease and pathway databases, extensive GWAS studies, and previous analyses of our studied cohort. We further expanded this gene set by incorporating connections from protein-protein interaction (STRING) and regulatory interaction (GeneHancer) databases, yielding ∼800 genes and regulatory elements. Subsequently, we performed an epistasis analysis on an AUD severity phenotype utilizing ∼100K SNPs associated with the expanded gene set. A mixed model epistatic association analysis method was used to control for both admixed population structure and the relatedness in the American Indian cohort. The SNPxSNP interactions were further condensed into interactions between sets of neighboring SNPs. This was achieved by employing a bi-clustering algorithm to identify the sets of SNPs exhibiting the most significant interactions. The statistical significance was determined through hypergeometric test. Results Our analysis revealed nearly 400 significant GxG interactions, encompassing 45% AUD linked genes and 11% regulatory elements, to be associated with the AUD severity trait. When compared to the expanded gene set associated with AUD, the interacting genes were predominantly enriched in ion transport molecular functions, and microRNA targets involved in hypoxia inducible factor regulation and endothelial cell apoptosis. Further analysis at the cell type level unveiled that these interacting genes were most enriched in developing midbrain radial glia-like cells and neuronal cells such as GABAergic neurons, dopaminergic neurons, GABAergic neurons, and serotonergic neurons. Discussion We conducted the first large-scale epistasis study in addiction. Nearly half of genes potentially linked to AUD are involved in significant GxG interactions. Interacting genes are enriched in cell types highly involved in alcohol and substance use disorders. Our results suggest that epistasis may significantly contribute to AUD susceptibility. Disclosure Nothing to disclose.
BACKGROUND:To understand why some individuals who develop alcohol use disorders (AUD) first begin to drink heavily, a number of scales have been developed that index aspects of alcohol craving and restraint from drinking. We developed a new measure called the Alcohol Consumption Questionnaire (ACQ), based in part on items modified from scales used to index binge eating, because there are data to suggest that binge eating and binge drinking may share common antecedents. We present an initial validity study using data from a sample of Mexican Americans. METHODS:Data were from 699 Mexican American young adults in San Diego County, CA. A subsample (n = 60) had short-term test-retest data. Factor analysis and reliability assessment guided item reduction. Item response theory (IRT) analyses quantified item severity and identified questions with differential item functioning (DIF). Logistic regression assessed associations of mean scale scores with AUD, adjusting for key demographics, alcohol expectancies and subjective response to alcohol. We also examined associations with a protective genetic variant downstream from the alcohol dehydrogenase 7 (ADH7) gene. RESULTS:The scale was reduced from 20 to 14 questions, which can be summarized by a single overall score (Cronbach's alpha = 0.896) or by two sub-scores (Consumption: 12 items, Cronbach's alpha = 0.896; Enjoyment: 2 items, Cronbach's alpha = 0.780). Test-retest reliability was very high (0.80-0.98) in this sample. The overall ACQ score and each subdomain score were strongly associated with AUD (ORs = 5.95 mild; 11.41 moderate; 48.56 severe) and family history of AUD. Respondents with the protective genetic variant had significantly lower overall ACQ scores (p < 0.001). CONCLUSION:The ACQ is a novel measure of alcohol consumption with strong relationships with both the AUD phenotype and ADH7 gene variants in a sample of Mexican American young adults.
Cannabis use disorder (CUD) is common and has in part a genetic basis. The risk factors underlying its devel-opment likely involve multiple genes that are polygenetic and interact with each other and the environment to ultimately lead to the disorder. Co-morbidity and genetic correlations have been identified between CUD and other disorders and traits in select populations primarily of European descent. If two or more traits, such as CUD and another disorder, are affected by the same genetic locus, they are said to be pleiotropic. The present study aimed to identify specific pleiotropic loci for the severity level of CUD in three high-risk population cohorts: American Indians (AI), Mexican Americans (MA), and European Americans (EA). Using a previously developed computational method based on a machine learning technique, we leveraged the entire GWAS catalog and identified 114, 119, and 165 potentially pleiotropic variants for CUD severity in AI, MA, and EA respectively. Ten pleiotropic loci were shared between the cohorts although the exact variants from each cohort differed. While majority of the pleiotropic genes were distinct in each cohort, they converged on numerous enriched biological pathways. The gene ontology terms associated with the pleiotropic genes were predominately related to synaptic functions and neurodevelopment. Notable pathways included Wnt/beta-catenin signaling, lipoprotein assembly, response to UV radiation, and components of the complement system. The pleiotropic genes were the most significantly differentially expressed in frontal cortex and coronary artery, up-regulated in adipose tissue, and down-regulated in testis, prostate, and ovary. They were significantly up-regulated in most brain tissues but were down-regulated in the cerebellum and hypothalamus. Our study is the first to attempt a large-scale pleiotropy detection scan for CUD severity. Our findings suggest that the different population cohorts may have distinct genetic factors for CUD, however they share pleiotropic genes from underlying pathways related to Alzheimer's disease, neuroplasticity, immune response, and reproductive endocrine systems.
Runs of homozygosity (ROH) arise when an individual inherits two copies of the same haplotype segment. While ROH are ubiquitous across human populations, Native populations-with shared parental ancestry arising from isolation and endogamy-can carry a substantial enrichment for ROH. We have been investigating genetic and environmental risk factors for alcohol use disorders (AUD) in a group of American Indians (AI) who have higher rates of AUD than the general U. S. population. Here we explore whether ROH might be associated with incidence and severity of AUD in this admixed AI population (n = 742) that live on geographically contiguous reservations, using low-coverage whole genome sequences. We have found that the genomic regions in the ROH that were identified in this population had significantly elevated American Indian heritage compared with the rest of the genome. Increased ROH abundance and ROH burden are likely risk factors for AUD severity in this AI population, especially in those diagnosed with severe and moderate AUD. The association between ROH and AUD was mostly driven by ROH of moderate lengths between 1 and 2 Mb. An ROH island on chromosome 1p32.3 and a rare ROH pool on chromosome 3p12.3 were found to be significantly associated with AUD severity. They contain genes involved in lipid metabolism, oxidative stress and inflammatory responses; and OSBPL9 was found to reside on the consensus part of the ROH island. These data demonstrate that ROH are associated with risk for AUD severity in this AI population.
Alcohol and other substance use disorders (AUD and SUD) are complex diseases that are postulated to have a polygenic inheritance and are often comorbid with other disorders. The comorbidities may arise partially through genetic pleiotropy. Identification of specific gene variants accounting for large parts of the variance in these disorders has yet to be accomplished. We describe a flexible strategy that takes a variant-trait association database and determines if a subset of disease/straits are potentially pleiotropic with the disorder under study. We demonstrate its usage in a study of use disorders in two independent cohorts: alcohol, stimulants, cannabis (CUD), and multi-substance use disorders (MSUD) in American Indians (AI) and AUD and CUD in Mexican Americans (MA). Using a machine learning method with variants in GWAS catalog, we identified 229 to 246 pleiotropic variants for AI and 153 to 160 for MA for each SUD. Inflammation was the most enriched for MSUD and AUD in AIs. Neurological disorder was the most significantly enriched for CUD in both cohorts, and for AUD and stimulants in AIs. Of the select pleiotropic genes shared among substances-cohorts, multiple biological pathways implicated in SUD and other psychiatric disorders were enriched, including neurotrophic factors, immune responses, extracellular matrix, and circadian regulation. Shared pleiotropic genes were significantly up-regulated in brain regions playing important roles in SUD, down-regulated in esophagus mucosa, and differentially regulated in adrenal gland. This study fills a gap for pleiotropy detection in understudied admixed populations and identifies pleiotropic variants that may be potential targets of interest for SUD.