Abstract At rare human genomic regions, DNA methylation states are established in the early embryo and maintained during cellular differentiation, yielding systemic (i.e. not tissue-specific) interindividual epigenetic variation. Previous screens for such correlated regions of systemic interindividual variation (CoRSIVs) were limited to White Americans. Here, we describe the first human CoRSIV screen including self-identified Black and White Americans. We integrate deep whole-genome bisulfite sequencing data for three tissues from each of ten Black and ten White donors in the NIH Genotype-Tissue Expression program. This approach identifies twice as many CoRSIVs among Black than White Americans. Establishment of CoRSIV methylation is sensitive to periconceptional environmental exposures including assisted reproduction, seasonal variation, and famine. CoRSIV-associated genes are enriched for GWAS variants linked to cancer and neurodevelopment. Although only 15% of Black CoRSIVs overlap with those among White individuals, both sets are associated with the same subfamilies of transposable elements. Within multiple cell lines, ranked enrichments of transcription factor binding to Black and White CoRSIVs are exquisitely coordinated and related to genome organization, indicating that CoRSIV methylation states established in the early embryo play an important role in guiding subsequent cellular differentiation.
BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58 706 SVs in a study sample of 11 556 CAD cases and 42 907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.
Atrial fibrillation (AF) is a prevalent and morbid abnormality of the heart rhythm with a strong genetic component. Here, we meta-analyzed genome and exome sequencing data from 36 studies that included 52,416 AF cases and 277,762 controls. In burden tests of rare coding variation, we identified novel associations between AF and the genes MYBPC3, LMNA, PKP2, FAM189A2 and KDM5B. We further identified associations between AF and rare structural variants owing to deletions in CTNNA3 and duplications of GATA4. We broadly replicated our findings in independent samples from MyCode, deCODE and UK Biobank. Finally, we found that CRISPR knockout of KDM5B in stem-cell-derived atrial cardiomyocytes led to a shortening of the action potential duration and widespread transcriptomic dysregulation of genes relevant to atrial homeostasis and conduction. Our results highlight the contribution of rare coding and structural variants to AF, including genetic links between AF and cardiomyopathies, and expand our understanding of the rare variant architecture for this common arrhythmia.
ABSTRACT Prediabetes is one of the main health concerns in public health, and various etiological factors contribute to its onset. This study aimed to evaluate genetic associations and gene‐macronutrient interaction with prediabetes‐related metabolites to understand how genetic variation and dietary intake contribute to dysglycemia. We analyzed a total of 482 self‐identified Mexican American participants recruited from Starr County, Texas in 2018‐2019. Untargeted metabolomic profiling was performed using LC‐MS. Nutrient densities of six macronutrients were derived from a 106‐item food frequency questionnaire. Genetic associations for each metabolite were tested using Generalized linear Mixed Model Association Tests (GMMAT). Gene‐macronutrient interactions on prediabetes‐associated metabolites were assessed with the Mixed Model Association Test for GEne‐Environment Interaction (MAGEE). Age, gender, and BMI were included as covariates in all association tests. Among 308 named and 2,471 unnamed metabolites, 17 novel variant‐metabolite pairs were discovered, including rs10947898 in LRFN2 associated with diacylglycerol DG32:1(p‐value: 8.95E‐09). Among 145 named and 687 unnamed metabolites after filtering, gene‐macronutrient interaction analyses identified seven named metabolites, including a variant(rs111251222) in MXD3 that interacted with monounsaturated fat to influence eicosadienoic acid levels (Interaction p‐value: 9.88E‐09). Prediabetes and nutrient‐related metabolites in Mexican Americans showed significant genetic associations and gene‐nutrient interactions.
Omics approaches have emerged as indispensable tools in unravelling the intricate molecular landscape of cardiovascular disease (CVD) by providing comprehensive insights into the underlying mechanisms driving CVD pathogenesis, progression, and response to therapy. Integrative omics approaches further enhance our understanding by integrating multi-omics data to explain complex molecular networks and identify novel disease pathways. In this collection of BMC Cardiovascular Disorders, we invited submissions on omics approaches to investigate and understand CVD. Successfully achieving the goals of this collection, we were able to publish some interesting and impactful research articles after a rigorous peer review process.
Any failure in frontonasal development can lead to malformations at the middle facial region, such as frontonasal dysplasia, midfacial clefts, and hyper/hypotelorism. Various environmental factors influence morphogenesis through epigenetic regulations, including the action of noncoding microRNAs (miRNAs). However, it remains unclear how miRNAs are involved in the frontonasal development. In our analysis of publicly available miRNA-seq and RNA-seq datasets, we found that miR-28a-5p, miR-302a-3p, miR-302b-3p, and miR-302d-3p were differentially expressed in the frontonasal process during embryonic days 10.5 to 13.5 (E10.5-E13.5) in mice. Overexpression of these miRNAs led to a suppression of cell proliferation in cultured mouse embryonic frontonasal mesenchymal (MEFM) cells as well as in O9-1 cells, a cranial neural crest cell line. Through advanced bioinformatic analyses and miRNA-gene regulation assays, we identified that miR-28a-5p regulated a total of 25 genes, miR-302a-3p regulated 23 genes, miR-302b-3p regulated 22 genes, and miR-302d-3p regulated 20 genes. Notably, the expression of miR-302a/b/d-3p-unlike miR-28a-5p-was significantly upregulated by excessive exposure to all-trans retinoic acid (atRA) that induces craniofacial malformations. Inhibition of these miRNAs restored the reduced cell proliferation caused by atRA by normalizing the expression of target genes associated with frontonasal anomalies. Therefore, our findings suggest that miR-302a/b/d-3p plays a crucial role in the development of frontonasal malformations.
Epigenome-wide association studies (EWAS) profile DNA methylation across the human genome to identify associations with diseases and exposures. Most employ Illumina methylation arrays; this platform, however, under-samples interindividual epigenetic variation. The systemic and stable nature of epigenetic variation at correlated regions of systemic interindividual variation (CoRSIVs) should be advantageous to EWAS. Here, we analyze 2,203 published EWAS to determine whether Illumina probes within CoRSIVs are over-represented in the literature. Enrichment of CoRSIV-overlapping probes was observed for most classes of disease, particularly for neurodevelopmental disorders and type 2 diabetes, indicating an opportunity to improve the power of EWAS by over 200- and over 100-fold, respectively. EWAS targeting all known CoRSIVs should accelerate discovery of associations between individual epigenetic variation and risk of disease.
Spatial transcriptomics (ST) enables systematic profiling of whole-transcriptome gene expression in tissues while preserving spatial context. Recent advances in sequencing- and imaging-based ST technologies have ushered in the era of microscopic-resolution ST (μST), allowing transcriptome mapping at cellular and even subcellular scales with unprecedented precision. Despite these advances, μST faces substantial challenges, including sparse transcript discovery per submicron (or micron)-sized spatial units and data fragmentation across platforms, hindering integration and analysis. There is also a growing demand for scalable, segmentation-free, and universally applicable analysis methods, as well as strategies for 3D mapping, multi-omics integration, and artificial intelligence (AI)-driven spatial analysis. In this review, we highlight recent breakthroughs, outline key challenges, and discuss emerging experimental and computational solutions shaping the future of μST.
Skeletal muscle is essential for both movement and metabolic processes, characterized by a complex and ordered structure. Despite its importance, a detailed spatial map of gene expression within muscle tissue has been challenging to achieve due to the limitations of existing technologies, which struggle to provide high-resolution views. In this study, we leverage the Seq-Scope technique, an innovative method that allows for the observation of the entire transcriptome at an unprecedented submicron spatial resolution. By applying this technique to the mouse soleus muscle, we analyze and compare the gene expression profiles in both healthy conditions and following denervation, a process that mimics aspects of muscle aging. Our approach reveals detailed characteristics of muscle fibers, other cell types present within the muscle, and specific subcellular structures such as the postsynaptic nuclei at neuromuscular junctions, hybrid muscle fibers, and areas of localized expression of genes responsive to muscle injury, along with their histological context. The findings of this research significantly enhance our understanding of the diversity within the muscle cell transcriptome and its variation in response to denervation, a key factor in the decline of muscle function with age. This breakthrough in spatial transcriptomics not only deepens our knowledge of muscle biology but also sets the stage for the development of new therapeutic strategies aimed at mitigating the effects of aging on muscle health, thereby offering a more comprehensive insight into the mechanisms of muscle maintenance and degeneration in the context of aging and disease.
Background & Aims:Cirrhosis is a leading cause of liver-related mortality and a multifactorial disease. To date, the complex genetic architecture of non-viral cirrhosis has not been fully explored. Cross-trait genetic correlations can elucidate the common genetic etiology of genetically correlated phenotypes. This study aims to identify polygenic and pleiotropic traits associated with cirrhosis using the linkage disequilibrium score regression analysis. Methods:We conducted genome-wide association analysis of 9,622,842 imputed SNPs on 3,368 non-viral cirrhosis cases and 258,258 controls, and cross-trait analysis between non-viral cirrhosis and various polygenic and pleiotropic traits using the UK Biobank cohort study. We further performed sensitivity analyses by removing genomic regions of alcohol intake, smoking behaviors, and obesity. We observed multiple traits showing robust genetic correlations (rg) with non-viral cirrhosis. Results:We found strong genetic correlations between the genetic architectures of non-viral cirrhosis and clinical/physiologic factors, including BMI (rg=0.82), alanine aminotransferase (0.71), diabetes (0.70), number of cigarettes currently smoked daily (0.67), amount of alcohol drunk on a typical drinking day (0.60), insomnia (0.59), gout (0.57), depression (0.50), apoliprotein-A (-0.33), HDL cholesterol (-0.49). Exclusion of genomic regions associated with alcohol intake, smoking behaviors, and obesity demonstrated consistent directions and persistent associations in genetic patterns. The inheritability of cirrhosis on the observed scale showed 0.56%. Conclusions:This study provides a comprehensive assessment of the shared genetic architecture of non-viral cirrhosis predisposition and numerous polygenic and pleiotropic traits, most notably BMI, alanine aminotransferase, and diabetes. These findings provide new information on underlying comorbid conditions that can increase the non-viral cirrhosis risk.
The male predominance in sporadic thoracic aortic aneurysm and dissection (TAD) suggests that the X chromosome contributes to TAD, but this has not been tested. We investigated whether X-linked variation-common (minor allele frequency [MAF] ≥0.01) and rare (MAF <0.01)-was associated with sporadic TAD in three cohorts of European descent (Discovery: 364 cases, 874 controls; Replication: 516 cases, 440,131 controls, and ARIC [Atherosclerosis Risk in Communities study]: 753 cases, 2247 controls). For analysis of common variants, we applied a sex-stratified logistic regression model followed by a meta-analysis of sex-specific odds ratios. Furthermore, we conducted a meta-analysis of overlapping common variants between the Discovery and Replication cohorts. For analysis of rare variants, we used a sex-stratified optimized sequence kernel association test model. Common variants results showed no statistically significant findings in the Discovery cohort. An intergenic common variant near SPANXN1 was statistically significant in the Replication cohort (p = 1.81 × 10-8). The highest signal from the meta-analysis of the Discovery and Replication cohorts was a ZNF182 intronic common variant (p = 3.5 × 10-6). In rare variants results, RTL9 reached statistical significance (p = 5.15 × 10-5). Although most of our results were statistically insignificant, our analysis is the most comprehensive X-chromosome association analysis of sporadic TAD to date.
Background: Studies have found higher hemoglobin A1c (HbA1c) at similar levels of glucose for adults who self-identify as Black vs. White, but the mechanism for this difference is poorly understood. Genetic variants that are more common in persons of African descent that affect hemoglobin (Hb) or red cell turnover may explain some of this difference. A recently discovered deletion in a non-coding region of the HBA2 gene is common in persons of African descent and may contribute to differences in Hb and HbA1c. Methods: We evaluated 5178 participants (mean age 56 years, 18% Black, 55% female) who attended Visit 2 (1990-1992) of the Atherosclerosis Risk in Communities Study without diagnosed diabetes and BMI <30 kg/m 2 . We compared the distributions of Hb and HbA1c in Black and White participants by the number of deletions in the HBA2 region (zero, one, or two). We used linear regression to evaluate the extent to which this genetic variant explained glucose-independent Black-White differences in HbA1c and Hb. Results: Deletions in the HBA2 region were common in Black participants (30% one, 4% two) but rare in White participants (1%: one, 0%: two). Mean Hb was lower in Black participants with one (-0.38 g/dL, p<0.001) and two deletions (-0.76 g/dL, p=0.002) and HbA1c was higher (one: +0.11%, p<0.001; two deletions: +0.13%, p=0.11) ( Figure ). Results were similar in White participants with one deletion (N=38), but non-significant. In a regression model adjusted for age, sex, fasting glucose, and BMI, the differences in HbA1c and Hb comparing Blacks vs White adults were +0.2% and -0.6 g/dL, respectively. Additional adjustment for HBA2 deletions explained 16% of the difference in HbA1c and 19% of the difference in Hb. Conclusions: A variation in the non-coding region of the HBA2 gene common in those with African ancestry results in significantly lower Hb and significantly higher HbA1c, partially explaining the small observed racial differences in HbA1c. Other undiscovered variants could more fully explain observed racial differences in HbA1c.
Spatial transcriptomics technologies aim to advance gene expression studies by profiling the entire transcriptome with intact spatial information from a single histological slide. However, the application of spatial transcriptomics is limited by low resolution, limited transcript coverage, complex procedures, poor scalability and high costs of initial setup and/or individual experiments. Seq-Scope repurposes the Illumina sequencing platform for high-resolution, high-content spatial transcriptome analysis, overcoming these limitations. It offers submicrometer resolution, high capture efficiency, rapid turnaround time and precise annotation of histopathology at a much lower cost than commercial alternatives. This protocol details the implementation of Seq-Scope with an Illumina NovaSeq 6000 sequencing flow cell, allowing the profiling of multiple tissue sections in an area of 7 mm × 7 mm or larger. We describe the preparation of a fresh-frozen tissue section for both histological imaging and sequencing library preparation and provide a streamlined computational pipeline with comprehensive instructions to integrate histological and transcriptomic data for high-resolution spatial analysis. This includes the use of conventional software tools for single-cell and spatial analysis, as well as our recently developed segmentation-free method for analyzing spatial data at submicrometer resolution. Aside from array production and sequencing, which can be done in batches, tissue processing, library preparation and running the computational pipeline can be completed within 3 days by researchers with experience in molecular biology, histology and basic Unix skills. Given its adaptability across various biological tissues, Seq-Scope establishes itself as an invaluable tool for researchers in molecular biology and histology.
Spatial transcriptomics (ST) technologies have advanced to enable transcriptome-wide gene expression analysis at submicron resolution over large areas. Analysis of high-resolution ST data relies heavily on image-based cell segmentation or gridding, which often fails in complex tissues due to diversity and irregularity of cell size and shape. Existing segmentation-free analysis methods scale only to small regions and a small number of genes, limiting their utility in high-throughput studies. Here we present FICTURE, a segmentation-free spatial factorization method that can handle transcriptome-wide data labeled with billions of submicron resolution spatial coordinates. FICTURE is orders of magnitude more efficient than existing methods and it is compatible with both sequencing- and imaging-based ST data. FICTURE reveals the microscopic ST architecture for challenging tissues, such as vascular, fibrotic, muscular, and lipid-laden areas in real data where previous methods failed. FICTURE's cross-platform generality, scalability, and precision make it a powerful tool for exploring high-resolution ST.
Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency ≤ 0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.
ABSTRACTSpatial transcriptomics (ST) technologies represent a significant advance in gene expression studies, aiming to profile the entire transcriptome from a single histological slide. These techniques are designed to overcome the constraints faced by traditional methods such as immunostaining and RNAin situhybridization, which are capable of analyzing only a few target genes simultaneously. However, the application of ST in histopathological analysis is also limited by several factors, including low resolution, a limited range of genes, scalability issues, high cost, and the need for sophisticated equipment and complex methodologies. Seq-Scope—a recently developed novel technology—repurposes the Illumina sequencing platform for high-resolution, high-content spatial transcriptome analysis, thereby overcoming these limitations. Here we provide a detailed step-by-step protocol to implement Seq-Scope with an Illumina NovaSeq 6000 sequencing flow cell that allows for the profiling of multiple tissue sections in an area of 7 mm × 7 mm or larger. In addition to detailing how to prepare a frozen tissue section for both histological imaging and sequencing library preparation, we provide comprehensive instructions and a streamlined computational pipeline to integrate histological and transcriptomic data for high-resolution spatial analysis. This includes the use of conventional software tools for single cell and spatial analysis, as well as our recently developed segmentation-free method for analyzing spatial data at submicrometer resolution. Given its adaptability across various biological tissues, Seq-Scope establishes itself as an invaluable tool for researchers in molecular biology and histology.KEY POINTSThe protocol outlines a method for repurposing an Illumina NovaSeq 6000 flow cell as a spatial transcriptomics array, enabling the generation of high-resolution spatial datasets.The protocol introduces a streamlined data analysis pipeline that produces a spatial digital gene expression matrix suitable for various single-cell and spatial transcriptome analysis methods.The protocol allows for the capture of histology images from the same tissue section subjected to spatial transcriptomics analysis and allows users to precisely align the transcriptome dataset with the histological image using fiducial marks engraved on the flow cell surface.Leveraging commonly available Illumina equipment, the protocol offers researchers ultra-high submicrometer resolution in spatial transcriptomics analysis with a comprehensive pipeline, rapid turnaround, cost efficiency, and versatility.
Ever larger Structural Variant (SV) catalogs highlighting the diversity within and between populations help researchers better understand the links between SVs and disease. The identification of SVs from DNA sequence data is non-trivial and requires a balance between comprehensiveness and precision. Here we present a catalog of 355,667 SVs (59.34% novel) across autosomes and the X chromosome (50bp+) from 138,134 individuals in the diverse TOPMed consortium. We describe our methodologies for SV inference resulting in high variant quality and >90% allele concordance compared to long-read de-novo assemblies of well-characterized control samples. We demonstrate utility through significant associations between SVs and important various cardio-metabolic and hemotologic traits. We have identified 690 SV hotspots and deserts and those that potentially impact the regulation of medically relevant genes. This catalog characterizes SVs across multiple populations and will serve as a valuable tool to understand the impact of SV on disease development and progression.
Coronary artery calcification (CAC) is a measure of atherosclerosis and a well-established predictor of coronary artery disease (CAD) events. Here we describe a genome-wide association study of CAC in 22,400 participants from multiple ancestral groups. We confirmed associations with four known loci and identified two additional loci associated with CAC (ARSE and MMP16), with evidence of significant associations in replication analyses for both novel loci. Functional assays of ARSE and MMP16 in human vascular smooth muscle cells (VSMCs) demonstrate that ARSE is a promoter of VSMC calcification and VSMC phenotype switching from a contractile to a calcifying or osteogenic phenotype. Furthermore, we show that the association of variants near ARSE with reduced CAC is likely explained by reduced ARSE expression with the G allele of enhancer variant rs5982944. Our study highlights ARSE as an important contributor to atherosclerotic vascular calcification and a potential drug target for vascular calcific disease. de Vries, Conomos, Singh and Nicholson et al. identify two additional loci associated with coronary artery calcification (ARSE and MMP16) via a genome-wide association study in 22,400 participants from multiple ancestral groups and prove that ARSE is a mediator of vascular smooth muscle cell calcification and phenotype switching.