Identity-by-descent (IBD) segments are a useful tool for applications ranging from demographic inference to relationship classification, but most detection methods rely on phasing information and therefore require substantial computation time. As genetic datasets grow, methods for inferring IBD segments that scale well will be critical. We developed IBIS, an IBD detector that locates long regions of allele sharing between unphased individuals, and benchmarked it with Refined IBD, GERMLINE, and TRUFFLE on 3,000 simulated individuals. Phasing these with Beagle 5 takes 4.3 CPU days, followed by either Refined IBD or GERMLINE segment detection in 2.9 or 1.1 h, respectively. By comparison, IBIS finishes in 6.8 min or 7.8 min with IBD2 functionality enabled: speedups of 805-946× including phasing time. TRUFFLE takes 2.6 h, corresponding to IBIS speedups of 20.2-23.3×. IBIS is also accurate, inferring ≥7 cM IBD segments at quality comparable to Refined IBD and GERMLINE. With these segments, IBIS classifies first through third degree relatives in real Mexican American samples at rates meeting or exceeding other methods tested and identifies fourth through sixth degree pairs at rates within 0.0%-2.0% of the top method. While allele frequency-based approaches that do not detect segments can infer relationship degrees faster than IBIS, the fastest are biased in admixed samples, with KING inferring 30.8% fewer fifth degree Mexican American relatives correctly compared with IBIS. Finally, we ran IBIS on chromosome 2 of the UK Biobank dataset and estimate its runtime on the autosomes to be 3.3 days parallelized across 128 cores.
The de novo ceramide synthesis pathway is essential to human biology and health, but genetic influences remain unexplored. The core function of this pathway is the generation of biologically active ceramide from its precursor, dihydroceramide. Dihydroceramides have diverse, often protective, biological roles; conversely, increased ceramide levels are biomarkers of complex disease. To explore the genetics of the ceramide synthesis pathway, we searched for deleterious nonsynonymous variants in the genomes of 1,020 Mexican Americans from extended pedigrees. We identified a Hispanic ancestry-specific rare functional variant, L175Q, in delta 4-desaturase, sphingolipid 1 (DEGS1), a key enzyme in the pathway that converts dihydroceramide to ceramide. This amino acid change was significantly associated with large increases in plasma dihydroceramides. Indexes of DEGS1 enzymatic activity were dramatically reduced in heterozygotes. CRISPR/Cas9 genome editing of HepG2 cells confirmed that the L175Q variant results in a partial loss of function for the DEGS1 enzyme. Understanding the biological role of DEGS1 variants, such as L175Q, in ceramide synthesis may improve the understanding of metabolic-related disorders and spur ongoing research of drug targets along this pathway.
Simulations of close relatives and identical by descent (IBD) segments are common in genetic studies, yet most past efforts have utilized sex averaged genetic maps and ignored crossover interference, thus omitting features known to affect the breakpoints of IBD segments. We developed Ped-sim, a method for simulating relatives that can utilize either sex-specific or sex averaged genetic maps and also either a model of crossover interference or the traditional Poisson model for inter-crossover distances. To characterize the impact of previously ignored mechanisms, we simulated data for all four combinations of these factors. We found that modeling crossover interference decreases the standard deviation of pairwise IBD proportions by 10.4% on average in full siblings through second cousins. By contrast, sex-specific maps increase this standard deviation by 4.2% on average, and also impact the number of segments relatives share. Most notably, using sex-specific maps, the number of segments half-siblings share is bimodal; and when combined with interference modeling, the probability that sixth cousins have non-zero IBD sharing ranges from 9.0 to 13.1%, depending on the sexes of the individuals through which they are related. We present new analytical results for the distributions of IBD segments under these models and show they match results from simulations. Finally, we compared IBD sharing rates between simulated and real relatives and find that the combination of sex-specific maps and interference modeling most accurately captures IBD rates in real data. Ped-sim is open source and available from https://github.com/williamslab/ped-sim.
As genetic datasets increase in size, the fraction of samples with one or more close relatives grows rapidly, resulting in sets of mutually related individuals. We present DRUID-deep relatedness utilizing identity by descent-a method that works by inferring the identical-by-descent (IBD) sharing profile of an ungenotyped ancestor of a set of close relatives. Using this IBD profile, DRUID infers relatedness between unobserved ancestors and more distant relatives, thereby combining information from multiple samples to remove one or more generations between the deep relationships to be identified. DRUID constructs sets of close relatives by detecting full siblings and also uses an approach to identify the aunts/uncles of two or more siblings, recovering 92.2% of real aunts/uncles with zero false positives. In real and simulated data, DRUID correctly infers up to 10.5% more relatives than PADRE when using data from two sets of distantly related siblings, and 10.7%-31.3% more relatives given two sets of siblings and their aunts/uncles. DRUID frequently infers relationships either correctly or within one degree of the truth, with PADRE classifying 43.3%-58.3% of tenth degree relatives in this way compared to 79.6%-96.7% using DRUID.
Background: The Caribbean vervet monkey (Chlorocebus aethiops sabaeus) is a potentially valuable animal model of neurodegenerative disease. However, the trajectory of aging in vervets and its relationship to human disease is incompletely understood. Methods: To characterize biomarkers associated with neurodegeneration, we measured cerebrospinal fluid (CSF) concentrations of A beta(1-40), A beta(1-42), total tau, and p-tau(181) in 329 members of a multigenerational pedigree. Linkage and genome-wide association were used to elucidate a genetic contribution to these traits. Results: A beta(1-40) concentrations were significantly correlated with age, brain total surface area, and gray matter thickness. Levels of p-tau(181) were associated with cerebral volume and brain total surface area. Among the measured analytes, only CSF A beta(1-40) was heritable. No significant linkage (LOD>3.3) was found, though suggestive linkage was highlighted on chromosomes 4 and 12. Genome-wide association identified a suggestive locus near the chromosome 4 linkage peak. Conclusions: Overall, these results support the vervet as a non-human primate model of amyloid-related neurodegeneration, such as Alzheimer's disease and cerebral amyloid angiopathy, and highlight A beta(1-40) and p-tau(181) as potentially valuable biomarkers of these processes.
Processing speed is a psychological construct that refers to the speed with which an individual can perform any cognitive operation. Processing speed correlates strongly with general cognitive ability, declines sharply with age and is impaired across a number of neurological and psychiatric disorders. Thus, identifying genes that influence processing speed will likely improve understanding of the genetics of intelligence, biological aging and the etiologies of numerous disorders. Previous genetics studies of processing speed have relied on simple phenotypes (eg, mean reaction time) derived from single tasks. This strategy assumes, erroneously, that processing speed is a unitary construct. In the present study, we aimed to characterize the genetic architecture of processing speed by using a multidimensional model applied to a battery of cognitive tasks. Linkage and QTL‐specific association analyses were performed on the factors from this model. The randomly ascertained sample comprised 1291 Mexican‐American individuals from extended pedigrees. We found that performance on all three distinct processing‐speed factors (Psychomotor Speed; Sequencing and Shifting and Verbal Fluency) were moderately and significantly heritable. We identified a genome‐wide significant quantitative trait locus (QTL) on chromosome 3q23 for Psychomotor Speed (LOD = 4.83). Within this locus, we identified a plausible and interesting candidate gene for Psychomotor Speed (Z = 2.90, P = 1.86 × 10−03).
Background:The Caribbean vervet monkey (Chlorocebus aethiops sabaeus) is a potentially valuable animal model of neurodegenerative disease. However, the trajectory of aging in vervets and its relationship to human disease is incompletely understood. Methods:To characterize biomarkers associated with neurodegeneration, we measured cerebrospinal fluid (CSF) concentrations of Aβ1-40, Aβ1-42, total tau, and p-tau181 in 329 members of a multigenerational pedigree. Linkage and genome-wide association were used to elucidate a genetic contribution to these traits. Results:Aβ1-40 concentrations were significantly correlated with age, brain total surface area, and gray matter thickness. Levels of p-tau181 were associated with cerebral volume and brain total surface area. Among the measured analytes, only CSF Aβ1-40 was heritable. No significant linkage (LOD > 3.3) was found, though suggestive linkage was highlighted on chromosomes 4 and 12. Genome-wide association identified a suggestive locus near the chromosome 4 linkage peak. Conclusions:Overall, these results support the vervet as a non-human primate model of amyloid-related neurodegeneration, such as Alzheimer's disease and cerebral amyloid angiopathy, and highlight Aβ1-40 and p-tau181 as potentially valuable biomarkers of these processes.
Aims:The recent failures of HDL-raising therapies have underscored our incomplete understanding of HDL biology. Therefore there is an urgent need to comprehensively investigate HDL metabolism to enable the development of effective HDL-centric therapies. To identify novel regulators of HDL metabolism, we performed a joint analysis of human genetic, transcriptomic, and plasma HDL-cholesterol (HDL-C) concentration data and identified a novel association between trafficking protein, kinesin binding 2 (TRAK2) and HDL-C concentration. Here we characterize the molecular basis of the novel association between TRAK2 and HDL-cholesterol concentration.Methods and results:Analysis of lymphocyte transcriptomic data together with plasma HDL from the San Antonio Family Heart Study (n = 1240) revealed a significant negative correlation between TRAK2 mRNA levels and HDL-C concentration, HDL particle diameter and HDL subspecies heterogeneity. TRAK2 siRNA-mediated knockdown significantly increased cholesterol efflux to apolipoprotein A-I and isolated HDL from human macrophage (THP-1) and liver (HepG2) cells by increasing the mRNA and protein expression of the cholesterol transporter ATP-binding cassette, sub-family A member 1 (ABCA1). The effect of TRAK2 knockdown on cholesterol efflux was abolished in the absence of ABCA1, indicating that TRAK2 functions in an ABCA1-dependent efflux pathway. TRAK2 knockdown significantly increased liver X receptor (LXR) binding at the ABCA1 promoter, establishing TRAK2 as a regulator of LXR-mediated transcription of ABCA1.Conclusion:We show, for the first time, that TRAK2 is a novel regulator of LXR-mediated ABCA1 expression, cholesterol efflux, and HDL biogenesis. TRAK2 may therefore be an important target in the development of anti-atherosclerotic therapies.
Inferring relatedness from genomic data is an essential component of genetic association studies, population genetics, forensics, and genealogy. While numerous methods exist for inferring relatedness, thorough evaluation of these approaches in real data has been lacking. Here, we report an assessment of 12 state-of-the-art pairwise relatedness inference methods using a data set with 2485 individuals contained in several large pedigrees that span up to six generations. We find that all methods have high accuracy (92-99%) when detecting first-and second-degree relationships, but their accuracy dwindles to <43% for seventh-degree relationships. However, most identical by descent (IBD) segment-based methods inferred seventh-degree relatives correct to within one relatedness degree for >76% of relative pairs. Overall, the most accurate methods are Estimation of Recent Shared Ancestry (ERSA) and approaches that compute total IBD sharing using the output from GERMLINE and Refined IBD to infer relatedness. Combining information from the most accurate methods provides little accuracy improvement, indicating that novel approaches, such as new methods that leverage relatedness signals from multiple samples, are needed to achieve a sizeable jump in performance.
By analyzing multitissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalog of expression quantitative trait loci (eQTLs) in a nonhuman primate model. This catalog contains more genome-wide significant eQTLs per sample than comparable human resources and identifies sex-and age-related expression patterns. Findings include a master regulatory locus that likely has a role in immune function and a locus regulating hippocampal long noncoding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.
Environmental correlation of plasma lipid species with T2D-related traits. (XLSX 54Â kb)
Inferring relatedness from genomic data is an essential component of genetic association studies, population genetics, forensics, and genealogy. While numerous methods exist for inferring relatedness, thorough evaluation of these approaches in real data has been lacking. Here, we report an assessment of 12 state-of-the-art pairwise relatedness inference methods using a dataset with 2,485 individuals contained in several large pedigrees that span up to six generations. We find that all methods have high accuracy (~92% – 99%) when detecting first and second degree relationships, but their accuracy dwindles to less than 43% for seventh degree relationships. However, most IBD segment-based methods inferred seventh degree relatives correct to within one relatedness degree for more than 76% of relative pairs. Overall, the most accurate methods are ERSA and approaches that compute total IBD sharing using the output from GERMLINE and Refined IBD to infer relatedness. Combining information from the most accurate methods provides little accuracy improvement, indicating that novel approaches—such as new methods that leverage relatedness signals from multiple samples—are needed to achieve a sizeable jump in performance.
BACKGROUND:Differential plasma concentrations of circulating lipid species are associated with pathogenesis of type 2 diabetes (T2D). Whether the wide inter-individual variability in the plasma lipidome contributes to the genetic basis of T2D is unknown. Here, we investigated the potential overlap in the genetic basis of the plasma lipidome and T2D-related traits.RESULTS:We used plasma lipidomic data (1202 pedigreed individuals, 319 lipid species representing 23 lipid classes) from San Antonio Family Heart Study in Mexican Americans. Bivariate trait analyses were used to estimate the genetic and environmental correlation of all lipid species with three T2D-related traits: risk of T2D, presence of prediabetes and homeostatic model of assessment - insulin resistance. We found that 44 lipid species were significantly genetically correlated with one or more of the three T2D-related traits. Majority of these lipid species belonged to the diacylglycerol (DAG, 17 species) and triacylglycerol (TAG, 17 species) classes. Six lipid species (all belonging to the triacylglycerol class and containing palmitate at the first position) were significantly genetically correlated with all the T2D-related traits.CONCLUSIONS:Our results imply that: a) not all plasma lipid species are genetically informative for T2D pathogenesis; b) the DAG and TAG lipid classes partially share genetic basis of T2D; and c) 1-palmitate containing TAGs may provide additional insights into the genetic basis of T2D.
Significance Contributions of rare variants to common and complex traits such as type 2 diabetes (T2D) are difficult to measure. This paper describes our results from deep whole-genome analysis of large Mexican-American pedigrees to understand the role of rare-sequence variations in T2D and related traits through enriched allele counts in pedigrees. Our study design was well-powered to detect association of rare variants if rare variants with large effects collectively accounted for large portions of risk variability, but our results did not identify such variants in this sample. We further quantified the contributions of common and rare variants in gene expression profiles and concluded that rare expression quantitative trait loci explain a substantive, but minor, portion of expression heritability.
The hippocampal formation is a brain structure integrally involved in episodic memory, spatial navigation, cognition and stress responsiveness. Structural abnormalities in hippocampal volume and shape are found in several common neuropsychiatric disorders. To identify the genetic underpinnings of hippocampal structure here we perform a genome-wide association study (GWAS) of 33,536 individuals and discover six independent loci significantly associated with hippocampal volume, four of them novel. Of the novel loci, three lie within genes ( ASTN2 , DPP4 and MAST4 ) and one is found 200 kb upstream of SHH . A hippocampal subfield analysis shows that a locus within the MSRB3 gene shows evidence of a localized effect along the dentate gyrus, subiculum, CA1 and fissure. Further, we show that genetic variants associated with decreased hippocampal volume are also associated with increased risk for Alzheimer’s disease ( r g =−0.155). Our findings suggest novel biological pathways through which human genetic variation influences hippocampal volume and risk for neuropsychiatric illness.
Progranulin (GRN) loss-of-function mutations leading to progranulin protein (PGRN) haploinsufficiency are prevalent genetic causes of frontotemporal dementia. Reports also indicated PGRN-mediated neuroprotection in models of Alzheimer's and Parkinson's disease; thus, increasing PGRN levels is a promising therapeutic for multiple disorders. To uncover novel PGRN regulators, we linked whole-genome sequence data from 920 individuals with plasma PGRN levels and identified the prosaposin (PSAP) locus as a new locus significantly associated with plasma PGRN levels. Here we show that both PSAP reduction and overexpression lead to significantly elevated extracellular PGRN levels. Intriguingly, PSAP knockdown increases PGRN monomers, whereas PSAP overexpression increases PGRN oligomers, partly through a protein-protein interaction. PSAP-induced changes in PGRN levels and oligomerization replicate in human-derived fibroblasts obtained from a GRN mutation carrier, further supporting PSAP as a potential PGRN-related therapeutic target. Future studies should focus on addressing the relevance and cellular mechanism by which PGRN oligomeric species provide neuroprotection.
BACKGROUND AND AIMS:While the prevalence of major depression is elevated among cannabis users, the role of genetics in this pattern of comorbidity is not clear. This study aimed to estimate the heritability of cannabis use and major depression, quantify the genetic overlap between these two traits and localize regions of the genome that segregate in families with cannabis use and major depression.DESIGN:Family-based univariate and bivariate genetic analysis.SETTING:San Antonio, Texas, USA.PARTICIPANTS:Genetics of Brain Structure and Function study (GOBS) participants: 1284 Mexican Americans from 75 large multi-generation families and an additional 57 genetically unrelated spouses.MEASUREMENTS:Phenotypes of life-time history of cannabis use and major depression, measured using the semistructured MINI-Plus interview. Genotypes measured using ~1 M single nucleotide polymorphisms (SNPs) on Illumina BeadChips. A subselection of these SNPs were used to build multi-point identity-by-descent matrices for linkage analysis.FINDINGS:Both cannabis use [h2 = 0.614, P = 1.00 × 10-6 , standard error (SE) = 0.151] and major depression (h2 = 0.349, P = 1.06 × 10-5 , SE = 0.100) are heritable traits, and there is significant genetic correlation between the two (ρg = 0.424, P = 0.0364, SE = 0.195). Genome-wide linkage scans identify a significant univariate linkage peak for major depression on chromosome 22 [logarithm of the odds (LOD) = 3.144 at 2 centimorgans (cM)], with a suggestive peak for cannabis use on chromosome 21 (LOD = 2.123 at 37 cM). A significant pleiotropic linkage peak influencing both cannabis use and major depression was identified on chromosome 11 using a bivariate model (LOD = 3.229 at 112 cM). Follow-up of this pleiotropic signal identified a SNP 20 kb upstream of NCAM1 (rs7932341) that shows significant bivariate association (P = 3.10 × 10-5 ). However, this SNP is rare (seven minor allele carriers) and does not drive the linkage signal observed.CONCLUSIONS:There appears to be a significant genetic overlap between cannabis use and major depression among Mexican Americans, a pleiotropy that appears to be localized to a region on chromosome 11q23 that has been linked previously to these phenotypes.
SLC30A8 encodes zinc transporter 8 which is involved in packaging and release of insulin. Evidence for the association of SLC30A8 variants with type 2 diabetes (T2D) is inconclusive. We interrogated single nucleotide polymorphisms (SNPs) around SLC30A8 for association with T2D in high-risk, pedigreed individuals from extended Mexican American families. This study of 118 SNPs within 50 kb of the SLC30A8 locus tested the association with eight T2D-related traits at four levels: (i) each SNP using measured genotype approach (MGA); (ii) interaction of SNPs with age and sex; (iii) combinations of SNPs using Bayesian Quantitative Trait Nucleotide (BQTN) analyses; and (iv) entire gene locus using the gene burden test. Only one SNP (rs7817754) was significantly associated with incident T2D but a summary statistic based on all T2D-related traits identified 11 novel SNPs. Three SNPs and one SNP were weakly but interactively associated with age and sex, respectively. BQTN analyses could not demonstrate any informative combination of SNPs over MGA. Lastly, gene burden test results showed that at best the SLC30A8 locus could account for only 1-2% of the variability in T2D-related traits. Our results indicate a lack of association of the SLC30A8 SNPs with T2D in Mexican American families.
The Genetic Analysis Workshops (GAW) are a forum for development, testing, and comparison of statistical genetic methods and software. Each contribution to the workshop includes an application to a specified data set. Here we describe the data distributed for GAW19, which focused on analysis of human genomic and transcriptomic data.