Genome-wide association studies (GWAS) have identified >1,200 signals associated with type 2 diabetes (T2D), yet identifying functional variants remains challenging because the majority of them lie in noncoding regions of the genome and are in areas of high linkage disequilibrium (LD). While chromatin accessibility QTL (caQTL) and expression QTL (eQTL) analyses are useful for nominating regulatory mechanisms underlying GWAS signals, limitations still exist in pinpointing functional variants within regions of high LD. A complementary approach that has been less frequently applied is to focus on the allele-specific effect on chromatin accessibility at heterozygous single-nucleotide polymorphisms (SNPs), hereafter referred to as "allelic imbalance". We analyzed the allelic imbalance of reads generated from an assay for transposase-accessible chromatin with sequencing (ATAC-seq) across genotyped samples from 490 donors in T2D-relevant tissues: skeletal muscle, liver, pancreatic islets, adipose tissue, and relevant cell types. We identified 119,949 allelically imbalanced SNPs (FDR<0.05) across the genome. The allelic imbalance was often most prominent in one tissue and showed an enrichment overlapping with tissue-specific transcription factor (TF) binding footprints. Focusing on the 8,581 SNPs in previously published 99% credible sets from 338 T2D GWAS signals, we identified 256 imbalanced SNPs across 123 (36.4% of) signals, each showing allelic imbalance in at least one tissue or cell type. Of these, 71 signals contained only a single imbalanced SNP, representing excellent candidate causative variants. As a proof-of-concept, we showed that 23 of the 256 imbalanced SNPs were supported by allelic assays from previous studies. Further, we experimentally validated two imbalanced SNPs as likely functional variants: rs34584161 among a seven-SNP T2D credible set at the RNF6 signal in islets and rs849134 among a 13-SNP credible set at the JAZF1 signal in liver. This study demonstrates the power of integrating ATAC-seq allelic imbalance (ASAI) with GWAS statistical fine-mapping to identify candidate functional regulatory variants from among tightly linked GWAS variants in disease-relevant tissues. While applied here in T2D, this approach represents a widely applicable high-throughput framework for refining the genetic architecture of complex traits.
The hypothalamus, composed of multiple nuclei, is essential for maintaining the body’s homeostasis. Within the mediobasal hypothalamus, the arcuate nucleus (ARC) contains key neuronal populations, including appetite-suppressing pro-opiomelanocortin (POMC) neurons that regulate energy and glucose balance. Here, we present a chemically defined, scalable method for differentiating human pluripotent stem cells (hPSCs) into hypothalamic neurons enriched for POMC cells, compatible with robotic cell culture platforms for high-throughput use. Neuronal identity was validated by MERFISH single-cell transcriptomics, RNA-Seq, ATAC-Seq, and comparison to human hypothalamus. The method is robust across multiple hPSC lines, showing consistent induction of ventral diencephalon and hypothalamic markers. Derived neurons display metabolic disease-relevant features, including body mass index (BMI)-associated gene enrichment, and ATAC-Seq identifies potential candidate regulatory regions linked to hypothalamic development and metabolic traits. Functional assays reveal neuronal responses to insulin and the GLP-1 receptor agonist Exendin-4, and transcriptional responses to altered glucose conditions. This platform delivers a physiologically relevant model of human hypothalamic neurons that enables deeper mechanistic and therapeutic studies of metabolic disease.
Induced pluripotent stem cells (iPSCs) enabled the generation of diverse cell types; however, certain fundamental biological properties, such as the genetic and epigenetic determinants of proliferation, remain poorly characterized. We quantified proliferation across 602 unique donors with a time-lapse imaging-based growth area under the curve (gAUC) phenotype and correlated gAUC with cell line gene expression and genotype. We identified 3,091 differentially expressed genes and found that rare deleterious variants in WDR54, TMEM250, and C2orf81 were associated with reduced iPSC growth. Notably, WDR54 was differentially expressed with respect to gAUC. Although no common variants were associated, common genetic variation explained 71%-75% of the variance. These results indicate a complex genetic architecture of iPSC growth rates, where rare, large-effect variants in important growth regulators are layered onto a highly polygenic background. These findings can impact the design of pooled iPSC-based studies and disease models, which may be confounded by intrinsic growth differences.
Fibro-adipogenic progenitors (FAPs) in skeletal muscle have been implicated in type 2 diabetes (T2D) risk, yet their heterogeneity and context-dependent regulation remain poorly understood. Here, we establish induced pluripotent stem cell (iPSC)-derived FAPs as a faithful model of primary FAPs by leveraging a unique resource: iPSC lines and skeletal muscle biopsies obtained from the same 30 individuals. Donor-matched comparisons reveal that iPSC-FAPs recapitulate the transcriptome, epigenome, and subtype composition of muscle tissue FAPs. Using single-nucleus multiomics, we show that high-insulin exposure drives iPSC-FAPs toward an adipogenic fate - and that this adipogenic subtype is enriched for T2D GWAS signals, an enrichment undetectable under baseline conditions. We map the T2D-associated rs3814707 non-coding signal to LTBP3, a gene that influences FAP adipogenic differentiation. These findings reveal how disease-relevant regulatory mechanisms can be masked in unstimulated cells and establish iPSC-FAPs as a powerful platform for dissecting the state-dependent biology of complex metabolic disease.
Introduction and Objective: Intra-islet responses to proinflammatory stressors are central to the pathogenesis of type 1 diabetes (T1D), yet the gene-regulatory dynamics triggered by inflammatory and viral insults in islet cell populations are not fully understood. Methods: Here, we leveraged 10x Genomics multiome single nucleus sequencing to jointly profile transcriptomes and chromatin accessibility in human islets from 55 donors under control conditions and following 48-hour exposure to either a pro-inflammatory cytokine cocktail (TNFα, IFN-γ, and IL-1β) or coxsackievirus. To capture transcript diversity at higher resolution, we performed long-read sequencing (Oxford Nanopore) on selected multiome cDNA libraries. Results: By multiplexing samples with a randomized block design and implementing robust quality control and data integration pipelines, we identified 135,565 high quality nuclei and resolved exposure-specific response pathways in nine pancreatic cell types - acinar (41,260 nuclei), ductal (35,662), alpha (19,646), beta (29,419), delta (2,555), gamma (1,401), immune (709), stellate (2,204), and endothelial (3,709) - including dynamic shifts in isoform expression in response to treatments. For example, in ductal cells, cytokine treatment upregulated the transcript isoform ENST00000297350.9 (log2 fold change = 2.0, p = 1.9 × 10−5), a splice variant of TNFRSF11B (osteoprotegerin), a member of the TNF receptor superfamily, consistent with previous reports of cytokine-mediated induction in islets. Conclusion: Through ongoing long-read sequencing and integration with joint chromatin accessibility profiling, we are modeling cell-type-specific gene regulatory dynamics that orchestrate islet response to proinflammatory stimuli. This resource highlights transcripts and regulatory nodes with potential functional relevance for immune-β-cell crosstalk, offering mechanistic insights to inform pathogenic immune-islet interactions in T1D. Disclosure C.C. Robertson: None. Z. Meng: None. S. Lin: None. A.K. Huber: None. X. Wang: None. P. Orchard: None. M.R. Erdos: None. F. Collins: Board Member; Current; ImmunoBrain Checkpoint, Inc. S. Chen: Stock/Shareholder; Current; iOrganBio Inc. Stock/Shareholder; Ended; Oncobeat. S. Parker: Research Support; Current; Pfizer Inc. Consultant; Ended; Novo Nordisk. Funding National Institutes of Health (5U01DK127777)
Hutchinson-Gilford progeria syndrome (HGPS) is a premature aging disorder affecting tissues of mesenchymal origin. Most patients harbor a c.1824C>T/p.G608= variant, commonly described as G608G, in exon 11 of LMNA that leads to aberrant splicing and production of the toxic progerin protein. In addition to cardiovascular, dermal, and adipose tissue deterioration, HGPS mouse models also develop progressive bone dysplasia that occurs in patients. Here we characterize the efficacy of in vivo mutation correction with an adenine base editor (ABE) to rescue structural and functional defects in HGPS transgenic murine bone tissue. Treatment of double-copy transgenic osteoblast cultures with a lentiviral-delivered CRISPR-Cas9 ABE achieved nearly 40% gene correction in vitro, resulting in significant reduction of progerin transcripts and protein, in the absence of selective agents. Furthermore, gene correction improved progeroid osteoblasts' capacity to deposit and mineralize extracellular matrix compared to untreated cultures. In vivo, a single intravenous dose of AAV9-delivered ABE corrected the mutation, achieving ~14%, ~22%, ~10% and < 1% correction in bone by six months of age when administered at P3, P14, 1 and 4 months of age, respectively. Partially rescued bone structural and physical parameters were observed in P14-treated mice with concomitant normalization of gene transcriptional programs and intracellular signaling pathways involved in bone remodeling. This work demonstrates in vivo delivery of a locus-specific DNA base editor to bone tissue, delineates the timing of treatment required for maximum efficacy, and suggests that this system might be tailored for application to other monogenic bone disorders.
Abstract Skeletal muscle, a primary site of insulin-mediated glucose uptake, plays a central role in the pathogenesis of type 2 diabetes. It is therefore critical to understand the disease-associated alterations in skeletal muscle and identify the underlying drivers of this dysregulation. Here, we characterize type 2 diabetes associated transcriptional dysregulation using 301 skeletal muscle biopsies from living donors with and without diabetes. Using weighted gene co-expression network analysis, we identify 56 distinct gene modules, which we further characterize using single-nucleus RNA-seq-derived cell type signatures and pathway enrichment analysis. We identify numerous cell type-associated dysregulated pathways in skeletal muscle tissue from individuals with diabetes, including muscle fiber-associated mitochondrial function and mRNA splicing and processing; endothelial vascularization and phospholipase D signaling; and macrophage- and T-cell-associated inflammation. Through analysis of module hub genes and transcription factor regulatory network analysis, we further identify candidate driver genes of this dysregulation including ATP5L , ATF2 , SIRT1, and THRAP3 in muscle fibers; JAM2 and CLEC14A in endothelial cells; and F13A1 and IRF8 in immune cells. Finally, we integrate our co-expression networks with single-nucleus ATAC-seq data to identify proximal and distal genomic regulatory elements and identify context-specific enrichment for type 2 diabetes and related trait GWAS signals in muscle fiber and endothelial modules. Together, our results reveal dysregulation in pathways in muscle tissue from individuals with diabetes, identify candidate drivers, and connect the genomic drivers of this dysregulation across type 2 diabetes and related metabolic traits.
The identification of sex-differential gene regulatory elements is essential for understanding sex-differential patterns of health and disease. We leveraged bulk and single-nucleus RNA sequencing (RNA-seq) and single-nucleus ATAC-seq data from 281 skeletal muscle biopsies to characterize sex differences in gene expression and regulation at the cell-type and whole-tissue levels. We found highly concordant sex-biased expression of over 2,100 genes across the three muscle fiber types and bulk tissue. Gene pathways related to mitochondrial activity and energy metabolism were enriched for male-biased expression, whereas those related to signal transduction and cell differentiation were enriched for female-biased expression. We found widespread sex-biased chromatin accessibility enriched in proximal and distal gene regulatory states; in gene promoters, sex-biased chromatin accessibility was positively associated with sex-biased expression. Long noncoding RNAs (lncRNAs) and microRNAs (miRNAs) also showed extensive sex-biased expression in the fiber-type and bulk data, respectively. Together, these results highlight nuclear and cytoplasmic mechanisms for sex-differential gene regulation in skeletal muscle.
To investigate the interplay between physical activity and cardiometabolic traits in human skeletal muscle, we characterized gene expression and chromatin accessibility across skeletal muscle cell types in 263 Finnish individuals from the FUSION Tissue Biopsy Study. We analyzed skeletal muscle single-nucleus RNA-seq data (168,309 nuclei, 23,849 genes), ATAC-seq data (242,069 nuclei, 927,588 peaks), and bulk RNA-seq data (22,309 genes). Lower insulin resistance (HOMA-IR) and higher total physical activity were both associated with higher proportions of Type 1 nuclei and lower proportions of Type 2x nuclei. We identified cell-type-level and tissue-level gene expression-trait and gene set-trait associations for cardiometabolic and physical activity traits, and a smaller proportion of cell-type-level chromatin accessibility-trait associations. Traits typically associated with better health-lower trait values of cardiometabolic traits (BMI, HOMA-IR, normal glucose tolerance vs. type 2 diabetes, 2-hour plasma glucose) and higher physical activity levels (total and vigorous)-were associated with higher expression of energy metabolism genes and lower expression of signaling pathway genes across muscle fiber types, total pseudobulk, and to some extent in bulk tissue. For HOMA-IR and physical activity, these directions of association remained when adjusting for both traits in the same model, indicating apparently independent associations in the same pathways.
RNA modifications are critical regulators of gene expression and cellular processes; however, the epitranscriptome is less well studied than the epigenome. Here, we studied transcriptome-wide changes in RNA modifications and expression levels in two human pancreatic beta-cell lines, EndoC-BH1 and EndoC-BH3, after one hour of glucose stimulation. Using direct RNA nanopore sequencing (dRNA-seq), we measured N6-methyladenosine (m6A), 5-methylcytosine (m5C), inosine, and pseudouridine concurrently across the transcriptome. We developed a differential RNA modification method and identified 1,697 differentially modified sites (DMSs) across all modifications. These DMSs were largely independent of changes in gene expression levels and enriched in transcripts for type 2 diabetes (T2D) genes. Our study demonstrates how dRNA-seq can be used to detect and quantify RNA modification changes in response to cellular stimuli at the single-nucleotide level and provides new insights into RNA-mediated mechanisms that may contribute to normal beta-cell response and potential dysfunction in T2D.
Polygenic scores (PGSs) for body mass index (BMI) may guide early prevention and targeted treatment of obesity. Using genetic data from up to 5.1 million people (4.6% African ancestry, 14.4% American ancestry, 8.4% East Asian ancestry, 71.1% European ancestry and 1.5% South Asian ancestry) from the GIANT consortium and 23andMe, Inc., we developed ancestry-specific and multi-ancestry PGSs. The multi-ancestry score explained 17.6% of BMI variation among UK Biobank participants of European ancestry. For other populations, this ranged from 16% in East Asian-Americans to 2.2% in rural Ugandans. In the ALSPAC study, children with higher PGSs showed accelerated BMI gain from age 2.5 years to adolescence, with earlier adiposity rebound. Adding the PGS to predictors available at birth nearly doubled explained variance for BMI from age 5 onward (for example, from 11% to 21% at age 8). Up to age 5, adding the PGS to early-life BMI improved prediction of BMI at age 18 (for example, from 22% to 35% at age 5). Higher PGSs were associated with greater adult weight gain. In intensive lifestyle intervention trials, individuals with higher PGSs lost modestly more weight in the first year (0.55 kg per s.d.) but were more likely to regain it. Overall, these data show that PGSs have the potential to improve obesity prediction, particularly when implemented early in life.
Identifying genetic variants that regulate gene expression can help uncover mechanisms underlying complex traits. We performed a meta-analysis of skeletal muscle expression quantitative trait locus (eQTL) using data from 1,002 individuals from two studies. A stepwise analysis identified 18,818 conditionally distinct signals for 12,283 genes, and 35% of these genes contained two or more signals. Colocalization of these eQTL signals with 26 muscular and cardiometabolic trait genome-wide association studies (GWASs) identified 2,252 GWAS-eQTL colocalizations that nominated 1,342 candidate genes. Notably, 22% of the GWAS-eQTL colocalizations involved non-primary eQTL signals. Additionally, 37% of the colocalized GWAS-eQTL signals corresponded to the closest protein-coding gene, while 44% were located >50 kb from the transcription start site of the nominated gene. To assess tissue specificity for a heterogeneous trait, we compared colocalizations with type 2 diabetes (T2D) signals across muscle, adipose, liver, and islet eQTLs; we identified 551 candidate genes for 309 T2D signals representing 36% of T2D signals tested and over 100 more than were detected with any one tissue alone. We then functionally validated the allelic regulatory effect of an eQTL variant for INHBB linked to T2D in both muscle and adipose tissue. Together, these results further demonstrate the value of skeletal muscle eQTLs in elucidating mechanisms underlying complex traits.
Understanding the spatial distribution of gene expression in the pancreas is essential for establishing the molecular basis of pancreatic function in healthy and disease contexts. Recent platforms offer a robust method for quantifying gene expression within a spatial context. Here, we report spatial transcriptomic profiling from pancreas samples obtained from three donors with type 2 diabetes (T2D) and three donors with normal glucose tolerance (NGT). Our analysis identified a major technical challenge: substantial transcript bleed of highly abundant genes (e.g., INS and GCG) into adjacent tissue regions. We demonstrate that this bleed can be computationally corrected using probabilistic models. Our analysis highlights the importance of incorporating bleed-correction techniques in the preprocessing of spatial transcriptomic profiling data. In summary, this study provides a dataset, methods, and resources to investigate the spatial regulation of gene expression in normal and T2D-affected human pancreas.
Characterization of DNA binding sites for specific proteins is of fundamental importance in molecular biology. It is commonly addressed experimentally by chromatin immunoprecipitation and sequencing (ChIP-seq) of bulk samples (103-107 cells). We have developed an alternative method that uses a Chromatin Antibody-mediated Methylating Protein (ChAMP) composed of a GpC methyltransferase fused to protein G. By tethering ChAMP to a primary antibody directed against the DNA-binding protein of interest, and selectively switching on its enzymatic activity in situ, we generated distinct and identifiable methylation patterns adjacent to the protein binding sites. This method is compatible with methods of single-cell methylation-detection and single molecule methylation identification. Indeed, as every binding event generates multiple nearby methylations, we were able to confidently detect protein binding in long single molecules.
Genome-wide association studies (GWASs) have identified over 100 signals associated with type 1 diabetes (T1D). However, it has been challenging to translate any given T1D GWAS signal into mechanistic insights, such as causal variants, their target genes, and the specific cell types involved. Here, we present a comprehensive multi-omic integrative analysis of single-cell/nucleus resolution profiles of gene expression and chromatin accessibility in human pancreatic islets under baseline and T1D-stimulating conditions. We nominate effector cell types for all T1D GWAS signals and the regulatory elements and genes for three independent T1D signals acting through β cells at the DLK1/MEG3, RASGRP1, and TOX loci. Subsequently, we validated the functional impact of these genes and regulatory regions using isogenic human embryonic stem cells (hESCs). We found that loss of RASGRP1 or DLK1, as well as disruption of their corresponding regulatory regions, led to increased β cell apoptosis. Furthermore, β cells derived from isogenic hESCs carrying the T1D risk allele of rs3783355 associated with DLK1 showed elevated β cell death. Through additional RNA sequencing (RNA-seq) and assay for transposase-accessible chromatin using sequencing (ATAC-seq) analyses, we identified five genes upregulated in both RASGRP1-/- and DLK1-/- β-like cells, four of which are near T1D GWAS signals. This integrative approach combining single-cell multi-omics, GWASs, and isogenic human pluripotent stem cell (hPSC)-derived β-like cells illuminates cell type context, genes, single nucleotide polymorphisms (SNPs), and regulatory elements underlying T1D-associated signals, providing insights into the biological functions and molecular mechanisms involved.
Complete characterization of the genetic effects on gene expression is needed to elucidate tissue biology and the etiology of complex traits. Here, we analyzed 2,344 subcutaneous adipose tissue samples and identified 34K conditionally distinct expression quantitative trait locus (eQTL) signals in 18K genes. Over half of eQTL genes exhibited at least two eQTL signals. Compared to primary signals, non-primary signals had lower effect sizes, lower minor allele frequencies, and less promoter enrichment; they corresponded to genes with higher heritability and higher tolerance for loss of function. Colocalization of eQTL with conditionally distinct genome-wide association study signals for 28 cardiometabolic traits identified 3,605 eQTL signals for 1,861 genes. Inclusion of non-primary eQTL signals increased colocalized signals by 46%. Among 30 genes with ≥2 pairs of colocalized signals, 21 showed a mediating gene dosage effect on the trait. Thus, expanded eQTL identification reveals more mechanisms underlying complex traits and improves understanding of the complexity of gene expression regulation.
Hutchison-Gilford progeria syndrome (HGPS) is a rare genetic disease caused by a mutation in LMNA, the gene encoding A-type lamins, leading to premature aging with severely reduced life span. HGPS is characterized by growth deficiency, subcutaneous fat and muscle issues, wrinkled skin, alopecia, and atherosclerosis. Patients also develop a bone phenotype with reduced bone mineral density, osteolysis and striking demineralization of long bones. To further clarify the tissue modifications in HGPS, we characterized bone mineralization in the LmnaG609G/G609G progeria mouse model. Femurs from 8-week-old mice and humeri from 15-week-old mice were analyzed using quantitative backscattered electron imaging to assess bone mineralization density distribution, osteocyte lacunae sections and structural bone histomorphometry. Tissue sections were stained with Giemsa and Goldner trichrome for histologic evaluation. Bone tissue from Lmna+/+ and LmnaG609G/G609G mice had similar mineral content at 3 different bone sites with specific tissue ages. The osteocyte lacunae features were not statistically different, but more empty lacunae were found in LmnaG609G/G609G at both animal ages. Bone histomorphometry and histology demonstrated decreased bone volume per tissue volume in primary (8W: -23%, p=0.001; 15W: -38%, p=0.002) and secondary spongiosa (8W: -36%, p=0.001; 15W: -49 %, ns), as well as growth plate dysplasia with thinner unmineralized resting and proliferative zones in the LmnaG609G/G609G mice versus controls (8W: -18%, p=0.006; 15W: -25%, p=0.001). Overall, the LmnaG609G/G609G mouse develops chondrodysplasia with reduced trabecular bone volume. Mineral content findings at several tissue sites and ages suggest that bone dysplasia results from impaired bone formation with normal bone turnover.
Skeletal muscle, the largest human organ by weight, is relevant to several polygenic metabolic traits and diseases including type 2 diabetes (T2D). Identifying genetic mechanisms underlying these traits requires pinpointing the relevant cell types, regulatory elements, target genes, and causal variants. Here, we used genetic multiplexing to generate population-scale single nucleus (sn) chromatin accessibility (snATAC-seq) and transcriptome (snRNA-seq) maps across 287 frozen human skeletal muscle biopsies representing 456,880 nuclei. We identified 13 cell types that collectively represented 983,155 ATAC summits, cataloging caQTL peaks that atlas-level snATAC maps often miss. We integrated genetic variation to discover 6,866 expression quantitative trait loci (eQTL) and 100,928 chromatin accessibility QTL (caQTL) (5% FDR) across the five cell types. We identified 1,973 eGenes colocalized with caQTL and used mediation analyses to construct causal directional maps for chromatin accessibility and gene expression. 3,378 genome-wide association study (GWAS) signals across 43 relevant traits colocalized with sn-e/caQTL, 52% in a cell-specific manner. 77% of GWAS signals colocalized with caQTL and not eQTL, highlighting the critical importance of population-scale chromatin profiling for GWAS functional studies. A C2CD4A/B T2D GWAS signal colocalized with caQTL in muscle fibers and multiple chromatin loop models nominated VPS13C, a glucose uptake gene. Sequence of the caQTL peak overlapping caSNP rs7163757 showed allelic regulatory activity differences in a human myocyte cell line massively parallel reporter assay. These results illuminate the genetic regulatory architecture of human skeletal muscle at high-resolution epigenomic, transcriptomic, and cell state scales and serve as a template for population-scale multi-omic mapping in complex tissues and traits. Disclosure A. Varshney: None. N. Manickam: Employee; 10x Genomics. P. Orchard: None. A. Tovar: None. Z. Zhang: None. F. Feng: None. T.A. Lakka: None. M. Laakso: None. J. Tuomilehto: Stock/Shareholder; Orion Pharma, Aktivolabs, Digostics. K.L. Mohlke: None. J. Kitzman: Advisory Panel; Myome Inc. H.A. Koistinen: Other Relationship; AstraZeneca, Novo Nordisk. J. Liu: None. M. Boehnke: None. F.S. Collins: None. L. Scott: None. S.C. Parker: Research Support; Pfizer Inc. Funding R01DK072193UM1DK126185R01DK117960P30AR069620