Abstract Cocaine use disorder (CUD) is a major public health crisis. The specific genes mediating CUD remain largely unknown. We conducted a genome-wide association study (GWAS) using outbred N/NIH Heterogeneous Stock (HS; n = 836, female = 415, male = 421) rats. We examined CUD-related phenotypes including acquisition of self-administration, escalation of intake, and compulsive-like responding. These traits were phenotypically correlated and exhibited modest SNP heritability (h2 = 0.07 – 0.16). We identified six genome-wide significant associations (>-log10(p)=5.58; α = 0.05 by permutation). One locus on chromosome 19 was associated with variable time between cocaine infusions (post infusion interval) and contains several carboxylesterase genes that are orthologous to the human CES1 gene. Notably, carboxylesterases metabolize cocaine. Three non-synonymous coding variants in Ces1c and Ces1d were in perfect linkage disequilibrium with this locus. The other five loci contained promising coding and expression variants, including Trak2, a gene previously associated with CUD in human GWAS and Slc10a7, Plcl1, and Satb2 which have been associated with alcohol and tobacco use disorder. This is the largest genetic study of cocaine self-administration ever conducted in rats. Our results replicate previous loci associated with CUD in humans and provide several novel biological insights including the potential of pharmacological strategies targeting carboxylesterases.
Studies have shown that substance use liability is associated with novelty seeking, anxiety-like behavior, and pain sensitivity. We examined whether common genetic variation in outbred Sprague-Dawley rats explained variation in behavioral measures from three assays with established links to substance use: locomotor response to a novel environment, elevated plus maze, and tail flick. We estimated single-nucleotide polymorphism heritability and performed genome-wide association analyses using permutation-derived significance thresholds (N=534-654 rats across traits). Heritability estimates ranged from 0.14-0.38 across eleven traits. Three independent loci were identified: chromosome 1 for elevated plus maze open-arm behavior (α=0.05), chromosome 14 for elevated plus maze immobility (α=0.10), and chromosome 17 for tail flick latency (α=0.05). Candidate genes included Slc18a2, Gfra1, and Pdzd8 (chromosome 1); Rel and Bcl11a (chromosome 14); and Eci2 and Eci3 (chromosome 17). We compared these loci with our genome wide association study of a F2 intercross of selectively bred high- and low-responder rats, originally derived from Sprague-Dawleys, that model individual differences in externalizing and internalizing behavior. The current loci are distinct from the ones identified in the bred lines. This difference likely reflects selection history in the high- and low-responder F2s, which focused on facets of exploratory locomotion, while loci for anxiety and pain sensitivity traits were identified in the outbreds. This highlights the benefit of using both outbred and selectively bred rats to probe causal variants contributing to individual differences in substance use liability. The current outbred findings implicate monoaminergic signaling, transcriptional control, and lipid metabolism as testable mechanisms for addiction-relevant behaviors.
Addiction is a complex and heritable trait which progresses through several developmental stages, each of which is presumably influenced by multiple partially overlapping genetic factors. Cocaine initially produces rewarding effects, followed by aversive effects including anxiety, craving, anhedonia, and withdrawal. These aversive effects have been suggested to contribute to the etiology of cocaine use disorders (CUD), as repeated exposure is thought to desensitize the rewarding effects and sensitize the aversive effects through a process involving both aberrant reward-based learning and aberrant avoidance-based learning. We examined the genetic basis of aversion learning using both food-based and cocaine-based behavioral assays in outbred Heterogenous Stock (HS) rats. A total of 1,074 HS rats (35.3% male) underwent runway operant cocaine-seeking, food-based progressive ratio and punishment testing, and locomotion testing. These phenotypes were significantly heritable (with h2 estimates as high as 0.307) and identified significant (p < 0.05) genetic loci related to avoidance-based learning including from the punishment task on Chromosomes 2, 3, 5, and 6, and the cocaine-operant runway latency task on Chromosome X. 172 positional candidate genes were identified from significant and suggestive loci, including Cdh10, Cdh12, Cdh18 which have previously been associated with smoking initiation from human GWAS, Adcy3, Cfap206, and Drc1 which are associated with primary neuronal cilia, as well as SNPs associated with novelty-related and social interaction phenotypes in independent samples of HS rats. Our results suggest that these aversion learning phenotypes are themselves complex heritable traits influenced by multiple genetic loci, which may pleiotropically affect other aspects of addiction biology.
Addiction vulnerability is associated with the tendency to attribute incentive salience to reward predictive cues. Both addiction and the attribution of incentive salience are influenced by environmental and genetic factors. To characterize the genetic contributions to incentive salience attribution, we performed a genome-wide association study (GWAS) in a cohort of 1596 heterogeneous stock (HS) rats. Rats underwent a Pavlovian conditioned approach task that characterized the responses to food-associated stimuli ("cues"). Responses ranged from cue-directed "sign-tracking" behavior to food-cup directed "goal-tracking" behavior (12 measures, SNP heritability: 0.051-0.215). Next, rats performed novel operant responses for unrewarded presentations of the cue using the conditioned reinforcement procedure. GWAS identified 14 quantitative trait loci (QTLs) for 11 of the 12 traits across both tasks. Interval sizes of these QTLs varied widely. Seven traits shared a QTL on chromosome 1 that contained a few genes (e.g., Tenm4, Mir708) that have been associated with substance use disorders and other psychiatric disorders in humans. Other candidate genes (e.g., Wnt11, Pak1) in this region had coding variants and expression-QTLs in mesocorticolimbic regions of the brain. We also conducted a Phenome-Wide Association Study (PheWAS) on addiction-related behaviors in HS rats and found that the QTL on chromosome 1 was also associated with nicotine self-administration in a separate cohort of HS rats. These results provide a starting point for the molecular genetic dissection of incentive motivational processes and provide further support for a relationship between the attribution of incentive salience and drug abuse-related traits.
The intestinal microbiome influences health and disease. Its composition is affected by host genetics and environmental exposures. Understanding host genetic effects is critical but challenging in humans, due to the difficulty of detecting, mapping and interpreting them. To address this, we analyse host genetic effects in four cohorts of outbred laboratory rats exposed to distinct but controlled environments. We show that polygenic host genetic effects are consistent across cohort environments. We identify three replicated microbiome-associated loci, one of which involves the sialyltransferase gene St6galnac1 and Paraprevotella. We find a similar association in a human cohort, between ST6GAL1 and Paraprevotella, both of which have been linked with immune and infectious diseases. Moreover, we find indirect (i.e. social) genetic effects on microbiome phenotypes, which substantially increase the total genetic variance. Finally, we identify a novel mechanism whereby indirect genetic effects can contribute to “missing heritability”. The intestinal microbiome is shaped by genetics and environment. Here, the authors show in rats that host genetic effects, including indirect social effects, influence microbiome composition, identify replicated loci, and reveal mechanisms contributing to microbiome heritability.
Age-related hearing loss (ARHL) is one of the most prevalent conditions affecting the elderly. ARHL is influenced by a combination of environmental and genetic factors; the identification of the genes that confer risk will aid in the prevention and treatment of ARHL. The mouse and human inner ears are functionally and genetically homologous. We used Carworth Farms White (CFW) mice to study the genetic basis of ARHL because they are genetically diverse and exhibit variability in the age of onset and severity of ARHL. Hearing at a range of frequencies was measured using auditory brainstem response (ABR) thresholds in 946 male and female CFW mice at the age of 1, 6, and 10 months. We genotyped the mice using low-coverage (mean coverage 0.27 ×) whole-genome sequencing (lcWGS) followed by imputation using STITCH. To determine the accuracy of the genotypes, we sequenced 8 samples at > 30 × coverage and used those data to estimate the accuracy of lcWGS genotyping, which was > 99.5
Age-related hearing impairment is the most common cause of hearing loss and is one of the most prevalent conditions affecting the elderly globally. It is influenced by a combination of environmental and genetic factors. The mouse and human inner ears are functionally and genetically homologous. Investigating the genetic basis of age-related hearing loss (ARHL) in an outbred mouse model may lead to a better understanding of the molecular mechanisms of this condition. We used Carworth Farms White (CFW) outbred mice, because they are genetically diverse and exhibit variation in the onset and severity of ARHL. The goal of this study was to identify genetic loci involved in regulating ARHL. Hearing at a range of frequencies was measured using Auditory Brainstem Response (ABR) thresholds in 946 male and female CFW mice at the age of 1, 6, and 10 months. We obtained genotypes at 4.18 million single nucleotide polymorphisms (SNP) using low-coverage (mean coverage 0.27x) whole-genome sequencing followed by imputation using STITCH. To determine the accuracy of the genotypes we sequenced 8 samples at >30x coverage and used calls from those samples to estimate the discordance rate, which was 0.45%. We performed genetic analysis for the ABR thresholds for each frequency at each age, and for the time of onset of deafness for each frequency. The SNP heritability ranged from 0 to 42% for different traits. Genome-wide association analysis identified several regions associated with ARHL that contained potential candidate genes, including Dnah11, Rapgef5, Cpne4, Prkag2, and Nek11. We confirmed, using functional study, that Prkag2 deficiency causes age-related hearing loss at high frequency in mice; this makes Prkag2 a candidate gene for further studies. This work helps to identify genetic risk factors for ARHL and to define novel therapeutic targets for the treatment and prevention of ARHL.
Affordable sequencing and genotyping methods are essential for large-scale genome-wide association studies. While genotyping microarrays and reference panels for imputation are available for human subjects, nonhuman model systems often lack such options. Our lab previously demonstrated an efficient and cost-effective method to genotype heterogeneous stock rats using double-digest genotyping by sequencing. However, low-coverage whole-genome sequencing offers an alternative method that has several advantages. Here, we describe a cost-effective, high-throughput, high-accuracy genotyping method for N/NIH heterogeneous stock rats that can use a combination of sequencing data previously generated by double-digest genotyping by sequencing and more recently generated by low-coverage whole-genome sequencing data. Using double-digest genotyping-by-sequencing data from 5,745 heterogeneous stock rats (mean 0.21x coverage) and low-coverage whole-genome sequencing data from 8,760 heterogeneous stock rats (mean 0.27x coverage), we can impute 7.32 million biallelic single-nucleotide polymorphisms with a concordance rate > 99.76% compared to high-coverage (mean 33.26x coverage) whole-genome sequencing data for a subset of the same individuals. Our results demonstrate the feasibility of using sequencing data from double-digest genotyping by sequencing or low-coverage whole-genome sequencing for accurate genotyping and demonstrate techniques that may also be useful for other genetic studies in nonhuman subjects.
Common genetic factors likely contribute to multiple psychiatric diseases including mood and substance use disorders. Certain stable, heritable traits reflecting temperament, termed externalizing or internalizing, play a large role in modulating vulnerability to these disorders. To model these heritable tendencies, we selectively bred rats for high and low exploration in a novel environment [bred High Responders (bHR) vs. Low Responders (bLR)]. To identify genes underlying the response to selection, we phenotyped and genotyped 538 rats from an F2 cross between bHR and bLR. Several behavioral traits show high heritability, including the selection trait: exploratory locomotion (EL) in a novel environment. There were significant phenotypic and genetic correlations between tests that capture facets of EL and anxiety. There were also correlations with Pavlovian conditioned approach (PavCA) behavior despite the lower heritability of that trait. Ten significant and conditionally independent loci for six behavioral traits were identified. Five of the six traits reflect different facets of EL that were captured by three behavioral tests. Distance traveled measures from the open field and the elevated plus maze map onto different loci, thus may represent different aspects of novelty-induced locomotor activity. The sixth behavioral trait, number of fecal boli, is the only anxiety-related trait mapping to a significant locus on chromosome 18 within which the Pik3c3 gene is located. There were no significant loci for PavCA. We identified a missense variant in the Plekhf1 gene on the chromosome 1:95 Mb QTL and Fancf and Gas2 as potential candidate genes that may drive the chromosome 1:107 Mb QTL for EL traits. The identification of a locomotor activity-related QTL on chromosome 7 encompassing the Pkhd1l1 and Trhr genes is consistent with our previous finding of these genes being differentially expressed in the hippocampus of bHR vs. bLR rats. The strong heritability coupled with identification of several loci associated with exploratory locomotion and emotionality provide compelling support for this selectively bred rat model in discovering relatively large effect causal variants tied to elements of internalizing and externalizing behaviors inherent to psychiatric and substance use disorders.
Behavioral diversity is critical for population fitness. Individual differences in risk-taking are observed across species, but underlying genetic mechanisms and conservation are largely unknown. We examined dark avoidance in larval zebrafish, a motivated behavior reflecting an approach-avoidance conflict. Brain-wide calcium imaging revealed significant neural activity differences between approach-inclined versus avoidance-inclined individuals. We used a population of ∼6,000 to perform the first genome-wide association study (GWAS) in zebrafish, which identified 34 genomic regions harboring many genes that are involved in synaptic transmission and human psychiatric diseases. We used CRISPR to study several causal genes: serotonin receptor-1b ( htr1b ), nitric oxide synthase-1 ( nos1 ), and stress-induced phosphoprotein-1 ( stip1 ). We further identified 52 conserved elements containing 66 GWAS significant variants. One encoded an exonic regulatory element that influenced tissue-specific nos1 expression. Together, these findings reveal new genetic loci and establish a powerful, scalable animal system to probe mechanisms underlying motivation, a critical dimension of psychiatric diseases.
Power analyses are often used to determine the number of animals required for a genome wide association analysis (GWAS). These analyses are typically intended to estimate the sample size needed for at least one locus to exceed a genome-wide significance threshold. A related question that is less commonly considered is the number of significant loci that will be discovered with a given sample size. We used simulations based on a real dataset that consisted of 3,173 male and female adult N/NIH heterogeneous stock (HS) rats to explore the relationship between sample size and the number of significant loci discovered. Our simulations examined the number of loci identified in sub-samples of the full dataset. The sub-sampling analysis was conducted for four traits with low (0.15 ± 0.03), medium (0.31 ± 0.03 and 0.36 ± 0.03) and high (0.46 ± 0.03) SNP-based heritabilities. For each trait, we sub-sampled the data 100 times at different sample sizes (500, 1,000, 1,500, 2,000, and 2,500). We observed an exponential increase in the number of significant loci with larger sample sizes. Our results are consistent with similar observations in human GWAS and imply that future rodent GWAS should use sample sizes that are significantly larger than those needed to obtain a single significant result.
1 Abstract Human unilateral renal agenesis is a congenital urinary tract malformation. Affected individuals have only one kidney, which is often an asymptomatic developmental defect. A total of 5,585 male and female HS rats were assessed for unilateral renal agenesis and genotyped for 3’513,321 markers. The R package SAIGEgds was used for the association analysis. The adjusted p-value threshold for the association analysis determined by permutation was equal to 5.6 (-log10). Two additional datasets were used as validation tests. Population two included 1,577 rats genotyped for 7,425,889 markers and a case-control imbalance equal to 1:174; population three included 1,407 rats, genotyped for 254,932 markers and case-control ratio equal to 1:38. The python package GxTheta was used to perform a polygenic epistasis analysis for the analyzed HS rat population. A founder haplotype mosaic determination was performed using the R package QTL2. Associated regions were selected for further analysis, including long-read PacBio sequencing for founder individuals and a founder haplotype prediction test. A similarity analysis at a genomic level and for loci encoding transcription factors predicted to interact with selected sequences inside the associated loci were accomplished. A total of 1,181 polymorphisms were associated with URA. All associated polymorphisms were located on chromosome 14 between 32.9 and 36.6 Mb. The most significant polymorphism was chr14:36,411,266, a G/T transversion. The same associated region was identified in population three. Polygenic epistasis was determined as not predominant for the presentation of URA. Based on the haplotype mosaic probability estimation, cases display a higher probability of inheriting the ACI allele. The long-read sequencing analysis showed the presence of an Erv insertion inside the intron one of the KIT gene located inside the associated region. The Erv insertion comprises one Erv sequence and two Ltr sequences located downstream and upstream of the former. No Erv insertion was identified for the founder strain BN. For ACI and HSRA, only one Ltr sequence was identified. One hundred and seven genes encoding TFs that recognize binding sites on the Erv insertion were analyzed for sequence similarity against the reference HSRA. The TF similarity score analysis for the interaction genotype and phenotype showed significance after FDR correction for 20 TFs, including AHR, HNF1B, JUNB, RARG, and RXRA. A mechanism identifying URA as a threshold phenotype is suggested in HS rats. It implies the existence of a minimum threshold for the final number of nephrons and kidney associated structures required for stalling the apoptotic process of the metanephric rudiments. Animals exhibiting a quantitative cumulative defect would express URA, being this malformation identified as a phenotype with decreased penetrance in the assessed population of HS rats. All these processes are described as mediated by KIT and TFs able to interact with sequences of the Erv insertion.
A vexing observation in genome-wide association studies (GWASs) is that parallel analyses in different species may not identify orthologous genes. Here, we demonstrate that cross-species translation of GWASs can be greatly improved by an analysis of co-localization within molecular networks. Using body mass index (BMI) as an example, we show that the genes associated with BMI in humans lack significant agreement with those identified in rats. However, the networks interconnecting these genes show substantial overlap, highlighting common mechanisms including synaptic signaling, epigenetic modification, and hormonal regulation. Genetic perturbations within these networks cause abnormal BMI phenotypes in mice, too, supporting their broad conservation across mammals. Other mechanisms appear species specific, including carbohydrate biosynthesis (humans) and glycerolipid metabolism (rodents). Finally, network co-localization also identifies cross-species convergence for height/body length. This study advances a general paradigm for determining whether and how phenotypes measured in model species recapitulate human biology.
Many personality traits are influenced by genetic factors. Rodents models provide an efficient system for analyzing genetic contribution to these traits. Using 1,246 adolescent heterogeneous stock (HS) male and female rats, we conducted a genome-wide association study (GWAS) of behaviors measured in an open field, including locomotion, novel object interaction, and social interaction. We identified 30 genome-wide significant quantitative trait loci (QTL). Using multiple criteria, including the presence of high impact genomic variants and co-localization of cis-eQTL, we identified 17 candidate genes (Adarb2, Ankrd26, Cacna1c, Cacng4, Clock, Ctu2, Cyp26b1, Dnah9, Gda, Grxcr1, Eva1a, Fam114a1, Kcnj9, Mlf2, Rab27b, Sec11a, and Ube2h) for these traits. Many of these genes have been implicated by human GWAS of various psychiatric or drug abuse related traits. In addition, there are other candidate genes that likely represent novel findings that can be the catalyst for future molecular and genetic insights into human psychiatric diseases. Together, these findings provide strong support for the use of the HS population to study psychiatric disorders.
Vanderbilt Genetics Institute, Vanderbilt University Medical Center, Nashville, TN, United States, Maize Research Institute, Sichuan Agricultural University, Chengdu, China, Department of Biostatistics, St. Jude Children’s Research Hospital, Memphis, TN, United States, Department of Statistics, Purdue University, West Lafayette, IN, United States, Department of Psychiatry, University of California San Diego, San Diego, CA, United States
There has been extensive discussion of the "Replication Crisis" in many fields, including genome-wide association studies (GWAS). We explored replication in a mouse model using an advanced intercross line (AIL), which is a multigenerational intercross between two inbred strains. We re-genotyped a previously published cohort of LG/J x SM/J AIL mice (F-34; n = 428) using a denser marker set and genotyped a new cohort of AIL mice (F39-43; n = 600) for the first time. We identified 36 novel genome-wide significant loci in the F-34 and 25 novel loci in the F39-43 cohort. The subset of traits that were measured in both cohorts (locomotor activity, body weight, and coat color) showed high genetic correlations, although the SNP heritabilities were slightly lower in the F39-43 cohort. For this subset of traits, we attempted to replicate loci identified in either F-34 or F39-43 in the other cohort. Coat color was robustly replicated; locomotor activity and body weight were only partially replicated, which was inconsistent with our power simulations. We used a random effects model to show that the partial replications could not be explained by Winner's Curse but could be explained by study-specific heterogeneity. Despite this heterogeneity, we performed a mega-analysis by combining F-34 and F39-43 cohorts (n = 1,028), which identified four novel loci associated with locomotor activity and body weight. These results illustrate that even with the high degree of genetic and environmental control possible in our experimental system, replication was hindered by study-specific heterogeneity, which has broad implications for ongoing concerns about reproducibility.
Objective Obesity is influenced by genetic and environmental factors. Despite the success of human genome-wide association studies, the specific genes that confer obesity remain largely unknown. The objective of this study was to use outbred rats to identify the genetic loci underlying obesity and related morphometric and metabolic traits. Methods This study measured obesity-relevant traits, including body weight, body length, BMI, fasting glucose, and retroperitoneal, epididymal, and parametrial fat pad weight in 3,173 male and female adult N/NIH heterogeneous stock (HS) rats across three institutions, providing data for the largest rat genome-wide association study to date. Genetic loci were identified using a linear mixed model to account for the complex family relationships of the HS and using covariates to account for differences among the three phenotyping centers. Results This study identified 32 independent loci, several of which contained only a single gene (e.g.,Epha5, Nrg1, Klhl14) or obvious candidate genes (e.g.,Adcy3, Prlhr). There were strong phenotypic and genetic correlations among obesity-related traits, and there was extensive pleiotropy at individual loci. Conclusions This study demonstrates the utility of HS rats for investigating the genetics of obesity-related traits across institutions and identify several candidate genes for future functional testing.
Field-grown plants have variable exposure to sunlight as a result of shifting cloud-cover, seasonal changes, canopy shading, and other environmental factors. As a result, they need to have developed a method for dissipating excess energy obtained from periodic excessive sunlight exposure. Non-photochemical quenching (NPQ) dissipates excess energy as heat, however, the physical and molecular genetic mechanics of NPQ variation are not understood. In this study, we investigated the genetic loci involved in NPQ by first growing different Arabidopsis thaliana accessions in local and seasonal climate conditions, then measured their NPQ kinetics through development by chlorophyll fluorescence. We used genome-wide association studies (GWAS) to identify 15 significant quantitative trait loci (QTL) for a range of photosynthetic traits, including a QTL co-located with known NPQ gene PSBS (AT1G44575). We found there were large alternative regulatory segments between the PSBS promoter regions of the functional haplotypes and a significant difference in PsbS protein concentration. These findings parallel studies in rice showing recurrent regulatory evolution of this gene. The variation in the PSBS promoter and the changes underlying other QTLs could give insight to allow manipulations of NPQ in crops to improve their photosynthetic efficiency and yield.
Muscle bulk in adult healthy humans is highly variable even after accounting for height, age and sex. Low muscle mass, due to fewer and/or smaller constituent muscle fibers, would exacerbate the impact of muscle loss occurring in aging or disease. Genetic variability substantially influences muscle mass differences, but causative genes remain largely unknown. In a genome-wide association study (GWAS) on appendicular lean mass (ALM) in a population of 85,750 middle-age (38-49 years) individuals from the UK Biobank (UKB) we found 182 loci associated with ALM ( P <5×10 −8 ). We replicated associations for 78% of these loci ( P <5×10 −8 ) with ALM in a population of 181,862 elderly (60-74 years) individuals from UKB. We also conducted a GWAS on hindlimb skeletal muscle mass of 1,867 mice from an advanced intercross between two inbred strains (LG/J and SM/J) which identified 23 quantitative trait loci. 38 positional candidates distributed across 5 loci overlapped between the two species. In vitro studies of positional candidates confirmed CPNE1 and STC2 as modifiers of myogenesis. Collectively, these findings shed light on the genetics of muscle mass variability in humans and identify targets for the development of interventions for treatment of muscle loss. The overlapping results between humans and the mouse model GWAS point to shared genetic mechanisms across species.