Summary CRISPR-Cas9 editing is a scalable technology for mapping of biological pathways, but it has been reported to cause a variety of undesired large-scale structural changes to the genome. We performed an arrayed CRISPR-Cas9 scan of the genome in primary human cells, targeting 17,065 genes for knockout with 101,029 guides. High-dimensional phenomics reveals a “proximity bias” in which CRISPR knockouts bear unexpected phenotypic similarity to knockouts of biologically-unrelated genes on the same chromosome arm, recapitulating both canonical genome structure and structural variants. Transcriptomics connects proximity bias to chromosome-arm truncations. Analysis of published large-scale knockout and knockdown experiments confirms that this effect is general across cell types, labs, Cas9 delivery mechanisms, and assay modalities, and suggests proximity bias is caused by DNA double-strand-breaks with cell cycle control in a mediating role. Finally, we demonstrate a simple correction for large-scale CRISPR screens to mitigate this pervasive bias while preserving biological relationships.
BACKGROUND:Genome-wide association studies have identified hundreds of loci associated with lipid levels. However, the genetic mechanisms underlying most of these loci are not well-understood. Recent work indicates that changes in the abundance of alternatively spliced transcripts contribute to complex trait variation. Consequently, identifying genetic loci that associate with alternative splicing in disease-relevant cell types and determining the degree to which these loci are informative for lipid biology is of broad interest. METHODS:We analyze gene splicing in 83 sample-matched induced pluripotent stem cell (iPSC) and hepatocyte-like cell lines (n=166), as well as in an independent collection of primary liver tissues (n=96) to perform discovery of splicing quantitative trait loci (sQTLs). RESULTS:We observe that transcript splicing is highly cell type specific, and the genes that are differentially spliced between iPSCs and hepatocyte-like cells are enriched for metabolism pathway annotations. We identify 1384 hepatocyte-like cell sQTLs and 1455 iPSC sQTLs at a false discovery rate of <5% and find that sQTLs are often shared across cell types. To evaluate the contribution of sQTLs to variation in lipid levels, we conduct colocalization analysis using lipid genome-wide association data. We identify 19 lipid-associated loci that colocalize either with an hepatocyte-like cell expression quantitative trait locus or sQTL. Only 2 loci colocalize with both a sQTL and expression quantitative trait locus, indicating that sQTLs contribute information about genome-wide association studies loci that cannot be obtained by analysis of steady-state gene expression alone. CONCLUSIONS:These results provide an important foundation for future efforts that use iPSC and iPSC-derived cells to evaluate genetic mechanisms influencing both cardiovascular disease risk and complex traits in general.
BACKGROUND:Multi-phenotype analysis of genetically correlated phenotypes can increase the statistical power to detect loci associated with multiple traits, leading to the discovery of novel loci. This is the first study to date to comprehensively analyze the shared genetic effects within different hemostatic traits, and between these and their associated disease outcomes. OBJECTIVES:To discover novel genetic associations by combining summary data of correlated hemostatic traits and disease events. METHODS:Summary statistics from genome wide-association studies (GWAS) from seven hemostatic traits (factor VII [FVII], factor VIII [FVIII], von Willebrand factor [VWF] factor XI [FXI], fibrinogen, tissue plasminogen activator [tPA], plasminogen activator inhibitor 1 [PAI-1]) and three major cardiovascular (CV) events (venous thromboembolism [VTE], coronary artery disease [CAD], ischemic stroke [IS]), were combined in 27 multi-trait combinations using metaUSAT. Genetic correlations between phenotypes were calculated using Linkage Disequilibrium Score Regression (LDSC). Newly associated loci were investigated for colocalization. We considered a significance threshold of 1.85 × 10-9 obtained after applying Bonferroni correction for the number of multi-trait combinations performed (n = 27). RESULTS:Across the 27 multi-trait analyses, we found 4 novel pleiotropic loci (XXYLT1, KNG1, SUGP1/MAU2, TBL2/MLXIPL) that were not significant in the original individual datasets, were not described in previous GWAS for the individual traits, and that presented a common associated variant between the studied phenotypes. CONCLUSIONS:The discovery of four novel loci contributes to the understanding of the relationship between hemostasis and CV events and elucidate common genetic factors between these traits.
Although affecting different arterial territories, the related atherosclerotic vascular diseases coronary artery disease (CAD) and peripheral artery disease (PAD) share similar risk factors and have shared pathobiology. To identify novel pleiotropic loci associated with atherosclerosis, we performed a joint analysis of their shared genetic architecture, along with that of common risk factors. Using summary statistics from genome-wide association studies of nine known atherosclerotic (CAD, PAD) and atherosclerosis risk factors (body mass index, smoking initiation, type 2 diabetes, low density lipoprotein, high density lipoprotein, total cholesterol, and triglycerides), we perform 15 separate multi-trait genetic association scans which resulted in 25 novel pleiotropic loci not yet reported as genome-wide significant for their respective traits. Colocalization with single-tissue eQTLs identified candidate causal genes at 14 of the detected signals. Notably, the signal between PAD and LDL-C at the PCSK6 locus affects PCSK6 splicing in human liver tissue and induced pluripotent derived hepatocyte-like cells. These results show that joint analysis of related atherosclerotic disease traits and their risk factors allowed identification of unified biology that may offer the opportunity for therapeutic manipulation. The signal at PCSK6 represent possible shared causal biology where existing inhibitors may be able to be leveraged for novel therapies.
ABSTRACT Background High-dimensional electronic health records (EHR) data can be used to phenotype complex diseases. The aim of this study is to apply unsupervised clustering to EHR-based traits derived in a cohort of patients with heart failure (HF) from a large integrated health system. Methods Using the institutional EHR, we identified 8569 patients with HF and extracted 1263 EHR-based input features, including clinical, echocardiographic, and comorbidity data, prior to the time of HF diagnosis. Principal component analysis, Uniform Manifold Approximation and Projection, and spectral clustering were applied to the input features after sex stratification of the cohort. The optimal number of clusters for each sex-stratified group was selected by highest Silhouette score and by within-cluster and between-cluster sums of squares. Determinants of cluster assignment were evaluated. Results We identified four clusters in each of the female-only (44%) and male-only (56%) cohorts. Sex-specific cohorts differed significantly by age of HF diagnosis, left ventricular chamber size, markers of renal and hepatic function, and comorbidity burden (all p<0.001). Left ventricular ejection fraction was not a strong driver of cluster assignment. Conclusion Readily available EHR data collected in the course of routine care can be leveraged to accurately classify patients into major phenotypic HF subtypes using data driven approaches.
Heart failure is a leading cause of cardiovascular morbidity and mortality. However, the contribution of common genetic variation to heart failure risk has not been fully elucidated, particularly in comparison to other common cardiometabolic traits. We report a multi-ancestry genome-wide association study meta-analysis of all-cause heart failure including up to 115,150 cases and 1,550,331 controls of diverse genetic ancestry, identifying 47 risk loci. We also perform multivariate genome-wide association studies that integrate heart failure with related cardiac magnetic resonance imaging endophenotypes, identifying 61 risk loci. Gene-prioritization analyses including colocalization and transcriptome-wide association studies identify known and previously unreported candidate cardiomyopathy genes and cellular processes, which we validate in gene-expression profiling of failing and healthy human hearts. Colocalization, gene expression profiling, and Mendelian randomization provide convergent evidence for the roles of BCKDHA and circulating branch-chain amino acids in heart failure and cardiac structure. Finally, proteome-wide Mendelian randomization identifies 9 circulating proteins associated with heart failure or quantitative imaging traits. These analyses highlight similarities and differences among heart failure and associated cardiovascular imaging endophenotypes, implicate common genetic variation in the pathogenesis of heart failure, and identify circulating proteins that may represent cardiomyopathy treatment targets.
SUMMARY:Identifying genomic features responsible for genome-wide association study (GWAS) signals has proven to be a difficult challenge; many researchers have turned to colocalization analysis of GWAS signals with expression quantitative trait loci (eQTL) and splicing quantitative trait loci (sQTL) to connect GWAS signals to candidate causal genes. The ColocQuiaL pipeline provides a framework to perform these colocalization analyses at scale across the genome and returns summary files and locus visualization plots to allow for detailed review of the results. As an example, we used ColocQuiaL to perform colocalization between a recent type 2 diabetes GWAS and Genotype-Tissue Expression (GTEx) v8 single-tissue eQTL and sQTL data. AVAILABILITY AND IMPLEMENTATION:ColocQuiaL is primarily written in R and is freely available on GitHub: https://github.com/bvoightlab/ColocQuiaL.
Clinical and epidemiological studies have shown that circulatory system diseases and nervous system disorders often co-occur in patients. However, genetic susceptibility factors shared between these disease categories remain largely unknown. Here, we characterized pleiotropy across 107 circulatory system and 40 nervous system traits using an ensemble of methods in the eMERGE Network and UK Biobank. Using a formal test of pleiotropy, five genomic loci demonstrated statistically significant evidence of pleiotropy. We observed region-specific patterns of direction of genetic effects for the two disease categories, suggesting potential antagonistic and synergistic pleiotropy. Our findings provide insights into the relationship between circulatory system diseases and nervous system disorders which can provide context for future prevention and treatment strategies.
Introduction: Unsupervised machine learning (UML) applied to high dimensional data has been used to discover cardiovascular disease subtypes; however, the reproducibility of subtypes identified by different algorithms has not been explored. We compared the ability of several promising UML and clustering algorithms to identify heart failure (HF) subtypes using high dimensional electronic health record (EHR) data. Methods: Using the Penn Medicine EHR, we identified all patients who had > 2 instances of ICD-10-CM HF diagnosis. We extracted 1272 EHR-based features (vital signs, demographics, echocardiographic measurements, laboratories, comorbidities) from time of HF diagnosis and limited the cohort based on data completeness (n=8569). We selected the following methods based on prior success in simulation studies and used them to identify HF subtypes: Similarity Network Fusion (SNF), Locally Linear Embedding (LLE), Modified LLE, Uniform Manifold Approximation and Projection (UMAP), and Principal Component Analysis (PCA) followed by several clustering algorithms including K-means, Density-based spatial clustering of applications with noise (DBSCAN), and Spectral Clustering. K groups 2-12 were evaluated. Clustering performance was assessed by silhouette score and visual separation. Results: Model visualizations are shown in the Figure. Highest silhouette score achieved for each model varied widely from 0.02-0.62; optimal cluster number ranged from 2-4 across models. Normalization and standardization of continuous data did not significantly alter silhouette scores or optimal cluster number. Conclusions: HF subtypes identified through UML applied to EHR data may vary substantially depending on the algorithms used. Benchmarking strategies to evaluate reproducibility of UML in the EHR are needed to ensure valid HF patient stratification and phenotypic refinement.
ABSTRACT IMPACT: Measuring and analyzing qualitative and quantitative traits using phenomics approaches will yield previously unrecognized heart failure subphenotypes and has the potential to improve our knowledge of heart failure pathophysiology, identify novel biomarkers of disease, and guide the development of targeted therapeutics for heart failure. OBJECTIVES/GOALS: Current classification schemes fail to capture the broader pathophysiologic heterogeneity in heart failure. Phenomics offers a newer unbiased approach to identify subtypes of complex disease syndromes, like heart failure. The goal of this research is to use data-driven associations to redefine the classification of the heart failure syndrome. METHODS/STUDY POPULATION: We will identify < 10 subphenotypes of patients with heart failure using unsupervised machine learning approaches for dense multidimensional quantitative (i.e. demographics, comorbid conditions, physiologic measurements, clinical laboratory, imaging, and medication variables; disease diagnosis, procedure, and billing codes) and qualitative data extracted from an integrated health system electronic health record. The heart failure subphenotypes we identify from the integrated health system electronic health record will be replicated in other heart failure population datasets using unsupervised learning approaches. We will explore the potential to establish associations between identified subphenotypes and clinical outcomes (e.g. all-cause mortality, cardiovascular mortality). RESULTS/ANTICIPATED RESULTS: We expect to identify < 10 mutually exclusive phenogroups of patients with heart failure that have differential risk profiles and clinical trajectories. DISCUSSION/SIGNIFICANCE OF FINDINGS: We will attempt to derive and validate a data-driven unbiased approach to the categorization of novel phenogroups in heart failure. This has the potential to improve our knowledge of heart failure pathophysiology, identify novel biomarkers of disease, and guide the development of targeted therapeutics for heart failure.
Atherosclerosis, which is the narrowing of the arterial walls via accumulation of cholesterol-rich arterial plaques, is the leading cause of vascular disease worldwide, including myocardial infarction and ischemic stroke. Although atherosclerosis affects arteries throughout the body, previous genome-wide association studies (GWAS) have been performed on specific atherosclerotic phenotypes such as coronary artery disease (CAD) and peripheral artery disease (PAD). There is substantial evidence to suggest that these more specific atherosclerosis phenotypes share a common genetic etiology. We performed a series of multi-trait GWAS using combinations of two atherosclerosis traits and seven atherosclerosis risk factor traits and detected 31 novel pleiotropic loci. We performed these multi-trait GWAS using the N-GWAMA multi-trait GWAS method and summary statistics for CAD (van der Harst et al. 2018), PAD (Klarin et al. 2019), body mass index (Pulit et al. 2019), type II diabetes (Vujkovic et al. 2020), smoking initiation (Wootton et al. 2020), and lipid traits (Klarin et al. 2019). We identified candidate causal genes for 14 of these loci through colocalization analysis with GTEx expression quantitative trait locus (eQTL) data. VDAC2 and PCSK6 are two candidate causal genes that our results and previous literature suggest are potential therapeutic targets. VDAC2 eQTLs in aorta and tibial artery colocalized with a multi-trait GWAS signal detected in the CAD PAD multi-trait GWAS. Previous work has shown that VDAC2 regulates apoptosis, and our results suggest increased VDAC2 expression in smooth muscle cells could increase smooth muscle cell accumulation in atherosclerotic plaques. A sQTL (splicing QTL) for PCSK6 in liver colocalized with a multi-trait GWAS signal between PAD and LDL. Further analysis of the sQTL signal suggested that the effect allele correlates with a more active isoform of PCSK6 , which could increase lipid fractions and risk of atherosclerosis. These results show that joint analysis of atherosclerotic disease traits and their risk factors allows for identification of unified biology that may offer the opportunity for therapeutic manipulation.
Introduction: The contribution of common genetic variation to heart failure (HF) risk has not been fully elucidated. Here, we applied multi-ancestry and multivariate methods to summary genetic data from >1 million individuals to improve power to identify common genetic variants, genes, cells, tissues, and circulating proteins/metabolites associated with HF and related cardiac imaging traits. Methods: Trans-ancestry meta-analysis of HF (56,722 cases and 1,133,054 controls) was performed using METAL. Multivariate analysis including GWAS of HF and imaging traits (MRI and echocardiogram) was performed using N-GWAMA. Downstream transcriptome-wide association studies (TWAS; S-PrediXcan), tissue/cell enrichment (LDSC-SEG using RNAseq and snRNAseq of human left ventricle samples from the MAGnet consortium), and Mendelian randomization (MR) analyses were performed. Results: The multi-ancestry HF GWAS identified 15 loci associated with all-cause HF (p < 5 x 10 -8 ). Multivariate analysis identified 48 (16 novel) loci (p < 5 x 10 -8 ), with enrichment for loci near Mendelian cardiomyopathy genes (p < 1 x 10 -4 ). Genetic associations were enriched (FDR < 0.05) for cardiac and musculoskeletal gene expression and chromatin marks. Gene expression (p = 0.007) and splicing events (p = 0.01) were enriched for established cardiomyopathy genes. Branch chain amino acid dehydrogenase ( BCKDHA ) expression was prioritized in TWAS, and MR identified causal associations between circulating branch chain amino acids and cardiac imaging traits: LVEDV MRI (leucine β = -0.137, 95% CI -0.25 to -0.022, p = 0.02; isoleucine β = -0.276, 95% CI -0.38 to -0.17, p = 3 x 10 -7 ) and LVSEV MRI (leucine β = -.131, 95% CI -0.24 to -0.026, p = 0.01; isoleucine β = -0.217, 95% CI -0.33 to -0.11, p = 1 x 10 -4 ). Unbiased proteome-wide MR of 725 circulating proteins identified 18 significant (FDR < 0.05) causal protein-trait associations, including between lipoprotein(a) and HF (OR 1.09 per 1-SD increase in lipoprotein(a), 95% CI 1.06 to 1.11, p = 1.6 x 10 -11 ). Conclusion: These analyses implicate novel common genetic variation in the pathogenesis of HF, highlight Mendelian cardiomyopathy genes in common HF, and identify circulating metabolites and proteins that may represent new treatment targets.
Background Identification of genetic risk factors that are shared between Alzheimer's disease (AD) and other traits, i.e., pleiotropy, can help improve our understanding of the etiology of AD and potentially detect new therapeutic targets. Previous epidemiological correlations observed between cardiometabolic traits and AD led us to assess the pleiotropy between these traits. Methods We performed a set of bivariate genome-wide association studies coupled with colocalization analysis to identify loci that are shared between AD and eleven cardiometabolic traits. For each of these loci, we performed colocalization with Genotype-Tissue Expression (GTEx) project expression quantitative trait loci (eQTL) to identify candidate causal genes. Results We identified three previously unreported pleiotropic trait associations at known AD loci as well as four novel pleiotropic loci. One associated locus was tagged by a low-frequency coding variant in the gene DOCK4 and is potentially implicated in its alternative splicing. Colocalization with GTEx eQTL data identified additional candidate genes for the loci we detected, including ACE, the target of the hypertensive drug class of ACE inhibitors. We found that the allele associated with decreased ACE expression in brain tissue was also associated with increased risk of AD, providing human genetic evidence of a potential increase in AD risk from use of an established anti-hypertensive therapeutic. Conclusion Our results support a complex genetic relationship between AD and these cardiometabolic traits, and the candidate causal genes identified suggest that blood pressure and immune response play a role in the pleiotropy between these traits.
We hereby provide the initial portrait of lincNORS , a spliced lincRNA generated by the MIR193BHG locus, entirely distinct from the previously described miR-193b-365a tandem. While inducible by low O 2 in a variety of cells and associated with hypoxia in vivo, our studies show that lincNORS is subject to multiple regulatory inputs, including estrogen signals. Biochemically, this lincRNA fine-tunes cellular sterol/steroid biosynthesis by repressing the expression of multiple pathway components. Mechanistically, the function of lincNORS requires the presence of RALY, an RNA-binding protein recently found to be implicated in cholesterol homeostasis. We also noticed the proximity between this locus and naturally occurring genetic variations highly significant for sterol/steroid-related phenotypes, in particular the age of sexual maturation. An integrative analysis of these variants provided a more formal link between these phenotypes and lincNORS , further strengthening the case for its biological relevance.
Significance Generation of heat by brown and beige fat cells is a potential avenue to increased energy expenditure, and thus management of obesity and metabolic syndrome. PM20D1 plays a role in thermogenesis based on mouse studies, but its expression had not been investigated in human adipocytes. Here we show that human PM20D1 expression is genetically variable at 2 levels. Genotype at certain distant variants correlates with overall PM20D1 expression levels across all human tissues (an “on/off switch”), while a different variant near the gene determines its regulation specifically in adipocytes by the PPARγ receptor and the antidiabetic drugs that target it (a “rheostat”). Human regulatory genetic variation in PM20D1 expression is associated with obesity and may ultimately inform individualized medicine approaches.
The link between cardiovascular diseases and neurological disorders has been widely observed in the aging population. Disease prevention and treatment rely on understanding the potential genetic nexus of multiple diseases in these categories. In this study, we were interested in detecting pleiotropy, or the phenomenon in which a genetic variant influences more than one phenotype. Marker-phenotype association approaches can be grouped into univariate, bivariate, and multivariate categories based on the number of phenotypes considered at one time. Here we applied one statistical method per category followed by an eQTL colocalization analysis to identify potential pleiotropic variants that contribute to the link between cardiovascular and neurological diseases. We performed our analyses on ~530,000 common SNPs coupled with 65 electronic health record (EHR)-based phenotypes in 43,870 unrelated European adults from the Electronic Medical Records and Genomics (eMERGE) network. There were 31 variants identified by all three methods that showed significant associations across late onset cardiac- and neurologic- diseases. We further investigated functional implications of gene expression on the detected "lead SNPs" via colocalization analysis, providing a deeper understanding of the discovered associations. In summary, we present the framework and landscape for detecting potential pleiotropy using univariate, bivariate, multivariate, and colocalization methods. Further exploration of these potentially pleiotropic genetic variants will work toward understanding disease causing mechanisms across cardiovascular and neurological diseases and may assist in considering disease prevention as well as drug repositioning in future research.
Traditionally, the use of genomic information for personalized medical decisions relies on prior discovery and validation of genotype-phenotype associations. This approach constrains care for patients presenting with undescribed problems. The National Institutes of Health (NIH) Undiagnosed Diseases Program (UDP) hypothesized that defining disease as maladaptation to an ecological niche allows delineation of a logical framework to diagnose and evaluate such patients. Herein, we present the philosophical bases, methodologies, and processes implemented by the NIH UDP. The NIH UDP incorporated use of the Human Phenotype Ontology, developed a genomic alignment strategy cognizant of parental genotypes, pursued agnostic biochemical analyses, implemented functional validation, and established virtual villages of global experts. This systematic approach provided a foundation for the diagnostic or non-diagnostic answers provided to patients and serves as a paradigm for scalable translational research.
The National Institutes of Health Undiagnosed Diseases Program (NIH UDP) applies translational research systematically to diagnose patients with undiagnosed diseases. The challenge is to implement an information system enabling scalable translational research. The authors hypothesized that similar complex problems are resolvable through process management and the distributed cognition of communities. The team, therefore, built the NIH UDP integrated collaboration system (UDPICS) to form virtual collaborative multidisciplinary research networks or communities. UDPICS supports these communities through integrated process management, ontology-based phenotyping, biospecimen management, cloud-based genomic analysis, and an electronic laboratory notebook. UDPICS provided a mechanism for efficient, transparent, and scalable translational research and thereby addressed many of the complex and diverse research and logistical problems of the NIH UDP. Full definition of the strengths and deficiencies of UDPICS will require formal qualitative and quantitative usability and process improvement measurement.