Knee-osteoarthritis (knee OA) is a prevalent joint disorder lacking Food and Drug Administration-approved cell therapies to halt progression. This study uses single-cell RNA sequencing to analyze bone marrow aspirate concentrate (BMAC) and stromal vascular fraction (SVF) samples in a clinical trial of autologous cell therapies. Trial site-specific variability was significant in BMAC, necessitating tailored normalization, whereas SVF was less affected, likely due to uniform subcutaneous fat sampling. Variance partitioning and tensor decomposition identified site effects in BMAC but revealed shared pathways across cell types in both tissues. Differential gene expression (DEG) analysis between responders and non-responders yielded no significant findings, although likelihood ratio test (LRT) revealed enrichment for DEG patterns linked to disease severity, potentially masked by patient heterogeneity. Key BMAC pathways included oxidative phosphorylation, unfolded protein response, and tumor necrosis factor alpha (TNF-α) signaling. Cell-cell communication analysis suggested enhanced human leukocyte antigen (HLA) signaling in non-responder MSCs (mesenchymal stromal cells), consistgenesent with inflammation, while responders showed more coordinated immune interactions. BMAC-MSCs promoted chondrocyte proliferation, whereas SVF-MSCs emphasized immune regulation. This study suggests that variability in therapy outcomes reflects patient heterogeneity beyond genomic factors, complicating the immediate use of genomic profiling to guide treatment. Nonetheless, as molecular pathways become better understood, integrating genomic insights into personalized strategies may become feasible.
The long-standing notion that genotypes map to phenotypes through simple one gene-one trait relationships continues to shape both research in the life sciences and public understanding, with implications for policy and funding priorities. Yet this paradigm is increasingly recognized as inadequate for explaining continuous phenotypic variation and the complex genetic architectures of the genotype-phenotype map. Modern genetics emerged from the early 20th-century synthesis of Mendelian and biometric schools of heredity, with R.A. Fisher demonstrating early on how multiple discrete loci could collectively produce continuous variation. Despite this fundamental insight, Mendelism-with its focus on single genes and standardized genetic backgrounds-became the dominant framework, shaping current genetics research and molecular biology as well as science education. The advent of large-scale genomic data has revealed yet again the limitations of this reductionist approach. Evidence from quantitative genetics now shows that most phenotypes arise from complex networks of many interdependent genes and their dynamic responses to environmental perturbations. Here we trace the historical roots of how Mendelian classical genetics departed from the biometric school to create the current predominant paradigm in genetics, despite fundamentally unresolved issues. Moving on from this one-sided paradigm will require systematic development of integrative, evolutionarily grounded experimental approaches that better capture the multigenic and context-dependent nature of inheritance. Achieving such an extended perspective will require methodological innovation, including advances in large-scale (e.g. automated) phenotyping. Dedicated research programs will be necessary to advance a new era of genetic research into the complex mechanisms underlying phenotypic variation.
The generalizability of polygenic scores (PGS) remains a major hurdle in the pursuit of equitable genomic medicine. Differences in disease prevalence across groups, potentially including social strata, influence the relationship between PGS and risk. Here we quantified the magnitude of PGS-by-context (PGS×C) interactions for seven human diseases and pairs of 75 contexts in the UK Biobank (n = 408,801). Across 24,198 PGS×C models, 746 (3.1%) had significant interactions by two criteria, up to fourfold more than expected by chance, improving predictive accuracy. The predominant mechanism for PGS×C is the amplification of genetic effects in adverse contexts, such as low polyunsaturated fatty acids or social determinants of ill health. We introduce the notion of the proportion needed to benefit as a metric that quantifies the expected effectiveness of interventions as a function of polygenic risk. Our results highlight the need for more comprehensive sampling across groups experiencing adverse exposures.
Accurate estimates of gene penetrance are critical for evaluating disease risk, guiding clinical decision-making, supporting genetic counseling, and advancing research. Multiple methodologies have been developed to estimate penetrance, aiding in the classification of human genetic disorders. Penetrance estimating approaches for rare and low-frequency/common variants may vary. Building on these approaches, we suggest the adoption of quantitative thresholds to categorize variant penetrance as high (≥70%), moderate (30%–70%), or low (<30%). This penetrance stratification may approximately reflect the continuous influence of genetic variants on disease. Traditionally, human diseases have been classified dichotomously as monogenic or polygenic. Based on these models, high-penetrance variants are typically associated with Mendelian disorders, while low-penetrance variants contribute modestly to disease risk. Since the completion of the human genome project, there have been hundreds of thousands of genetic variants identified in the gray zone between monogenic and polygenic variants. To bridge the conceptual gap between monogenic and polygenic disorders, we previously proposed a novel framework termed genetically transitional disease. This concept highlights disorders influenced by genetic variants of low- to moderate-penetrance along with environment and aligns with the ClinGen Working Group's framework on risk alleles and low-penetrance variants. By overcoming the inadequacy and obstacle of a strictly binary classification of genetic disease and moving beyond, the genetically transitional disease model offers a refined and more advantageous approach to capturing genetic nuances and represents a potential paradigm shift in the interpretation of genetic variation, with more important and broader implications for both patient care and research.
Vaso-occlusive episodes (VOEs) or acute pain events, involving complex interactions between sickle erythrocytes and other blood cells, are a hallmark of sickle cell disease (SCD). In this study, we analyzed changes in peripheral blood transcriptomes between steady state and VOEs in individuals with SCD. We followed a cohort of 174 individuals with SCD with or without chronic pain and collected peripheral blood at clinic visits (steady state) and during hospitalizations (VOEs). We performed RNA-Seq profiling of CD45 + leukocytes and CD71 + erythroid cells. Pathways linked to complement activation, coagulation, and IL-6/JAK/STAT3 signaling were enriched during VOEs in the CD45 + cells. Contrastingly, the CD71 + cells showed an enrichment of pathways related to the cell cycle, such as mTORC1 signaling and the G 2 M checkpoint during VOEs. We then analyzed the expression changes of genes in patients with longitudinal data to determine potential biomarkers for VOEs. Expression of 4 genes — FAM20A , IL1B , MS4A4A , and SERPINB2 — was elevated during VOEs compared with steady state in the majority of patients. Furthermore, our results indicate that patients experiencing chronic pain exhibited 44% increased enrichment of significant pathways during VOEs when compared with patients without chronic pain.
Many non-coding variants influence complex traits and diseases through gene regulation, yet the mechanisms linking these variants to downstream biology remain poorly understood. Here, we present eQTLGen Phase 2, a comprehensive genome-wide analysis of gene expression quantitative trait loci (eQTLs) in 43,301 blood samples from 52 datasets. Beyond local ciseffects, this sample size enabled the first systematic mapping of trans-eQTLs at scale. We identify cis-eQTLs for nearly all expressed genes (94.7%) and trans-eQTLs for over half (56.2%). Second, by colocalizing cis-eQTLs with trans-eQTLs, we infer a directed gene regulatory network comprising 47,554 directed gene regulatory relationships. These networks reveal how genetic perturbations in upstream regulators produce dose-dependent downstream effects, supported by Perturb-seq and ChIP-seq data. Third, integrating this network with 87 genome-wide association studies allows us to systematically prioritize trait-relevant pathways and candidate genes. Variants exerting both cis- and trans-effects are markedly more likely to colocalize with trait associations than cis-only variants, delineating a subset of functionally active cis-eQTLs from a large group with limited downstream impact. This distinction provides a conceptual framework for identifying regulatory variants that truly mediate complex trait biology. Together, these results provide a publicly available resource of cis- and trans-eQTLs and an in vivo scaffold for human gene-regulatory networks, elucidating how propagation of cis-effects modulates complex disease.
BACKGROUND:Crohn's disease (CD) is characterized by chronic intestinal inflammation. Previous single-cell transcriptomic studies have mostly focused on established disease, leaving a knowledge gap in relation to treatment-naive profiles across multiple regions of the gut. METHODS:To study disease onset, a treatment-naive pediatric CD cohort was recruited, and single-cell transcriptomics was performed on ileum, colon, and rectum biopsies collected at initial endoscopy. A clustering stability assessment workflow was developed to ensure clustering and downstream results were robust. RESULTS:Inflammation did not strongly influence cellular proportion due to heterogeneity across donor and tissue. Tensor decomposition revealed distinct mesenchymal and myeloid cell-mediated sources of disease pathology, corresponding to previously identified fibrotic and pro-inflammatory disease progression. Integrating transcriptomics and genome-wide association summary statistics for CD suggested myeloid and T cells drive disease, highlighting potential cellular therapeutic targets. CONCLUSION:Tensor decomposition stratified donors into clinically meaningful groups based on their transcriptomic profile, suggesting these signatures can be utilized for personalized medicine.
Background:Single cell multi-omic investigation opens-up new opportunities to understand mechanisms of gene regulation. Existing methods for inferring transcript abundance from chromatin accessibility fail to prioritize the most relevant peaks and tend to assume positive associations between ATAC peaks and RNA counts. We hypothesize that gene regulation can be modeled as a function of combined positive and negative interactions among peaks and that causal regulatory variants are enriched in the vicinity of the most critical peaks. Results:A machine learning pipeline leveraging single nuclear multiomic transcriptome and chromatin accessibility data is developed to model gene expression as a function of ATAC peak intensity. Multiome data was available for 18 immune cell types from 29 donors, 19 with Crohn's disease. The pipeline aggregates results from three machine learning approaches (random forest regression, XGBoost, and Light GBM) as well as linear regression to identify which ATAC peaks contribute to explaining variation among donors and cell types in pseudobulk gene expression. The coefficient of determination with cross-validation was used to identify robust models which typically explain between 5% and 40% of transcript abundance, utilizing on average 47% of the ATAC peaks, representing a significant gain in predictive accuracy. The most important peaks are enriched in GWAS variants for inflammatory bowel disease and the autoimmune disease systemic lupus erythematosus, but not for rheumatoid arthritis. Conclusion:Atlanta Plots visualize the proportion of ATAC peaks contributing to a predictive model of gene expression as well as the proportion of variance explained by the model. Software implementing our pipeline, "snATAC-Express", is freely available on GitHub.
Long-lived plasma cells (LLPCs) sustain lifelong antibody secretion, forming the foundation of durable immune memory. However, the molecular mechanisms underpinning LLPC survival within the bone marrow (BM) remain incompletely understood. Using single-nucleus RNA-seq and ATAC-seq (10X Multiome) from matched blood and BM antibody-secreting cells (ASCs) of healthy adults, we profiled 8,059 nuclei, generating an integrated transcriptional and epigenetic map of ASCs across their developmental continuum with unparalleled resolution. In the BM LLPC cluster, Gene Set Enrichment Analysis (GSEA) validated the enrichment of the TNF/NF-κB signaling pathway and TLR cascades, emphasizing their anti-apoptotic roles. These clusters also demonstrate increased chromatin accessibility at key loci such as RELA, TLR1, TLR6, TLR10, NFKBIA, and TAB1, underscoring the central roles of TLR and TNF-α signaling pathways in activating the NF-κB pathway. Upregulated genes such as BCL2, NFKBIZ, TNFAIP3, and BIRC3 highlighted key survival mechanisms, while motif enrichment analysis revealed significant enrichment for REL and NF-κB binding motifs, supporting chromatin-level regulation of LLPC longevity. Together, these findings advance our understanding of LLPC survival and resilience and provide a framework for improving vaccine strategies and targeted immunotherapies. Computational and Systems Immunology (COMP)
BACKGROUND & AIMS:Despite widespread biologic use, more than 70% of patients with Crohn's disease require resectional surgery, most commonly of the terminal ileum. Gene expression and genetics of the neoterminal ileum at postoperative, surveillance colonoscopies highlight pathways of disease recurrence. Postoperative colonoscopy transcriptomes were interrogated to evaluate the hypothesis that specific molecular mechanisms contribute to recurrent Crohn's pathophysiology. METHODS:Ribo-depleted, paired-end sequencing was run on 339 neoterminal ileal pinch biopsies from 267 patients with (Rutgeerts i2b+) and without (Rutgeerts i2a or lower) recurrent disease. Differential gene and transcript usage were assessed. Expression quantitative trait loci link genetic variation with gene expression. Serial sampling was performed on 70 patients. RESULTS:At colonoscopy, 4171 genes increased and 3579 genes decreased in recurring vs nonrecurring patients. Although gene expression was highly correlated (r = 0.71), we observed and replicated higher dynamic ranges of gene expression in male compared with female patients. Activation of both pro- (tumor necrosis factor, interferon gamma) and anti-inflammatory (transforming growth factor beta) pathways was observed; importantly, multiple nuclear hormone receptor pathways demonstrated activation, including both estrogen and dihydrotestosterone pathways, whereas progesterone was inhibited. We observed sex-specific expression quantitative trait loci driven by recurrent samples and differential transcript usage related to lipid metabolism, membrane trafficking, and the extracellular matrix. CONCLUSIONS:In recurrent disease of the postoperative neoterminal ileum, markedly greater dynamic ranges of gene expression occur in male compared with female patients. Pathway analyses implicate numerous nuclear hormone pathways, highlighting new mechanisms for therapeutic targeting beyond pro-inflammatory cytokine blockade. This study identifies key covariates and pathways of disease recurrence, many of which are distinct from drivers of initial disease susceptibility.
Volumetric muscle loss (VML) injuries result in chronic fibrosis, inflammation, and persistent functional deficits. Fibro-adipogenic progenitor (FAP) cells are a heterogeneous, muscle-resident stromal cell population that play a crucial role in muscle regeneration, but also contribute to fibrosis in muscle disease. The role of FAPs in VML is not well established and may be critical target to ensure functional muscle regeneration after VML. We utilized a VML model in the mouse quadriceps to study the location, secretome, surface marker distribution, gene expression, and single-cell transcriptional profile of FAPs after VML. After VML, a subpopulation of FAPs highly expressed β1-integrin and were elevated in the post-VML muscle tissue; these FAPs had increased fibrotic gene expression and increased myofibroblast differentiation potential. Transforming growth factor-β1 (TGF-β1) and tissue inhibitor of matrix metalloproteinase 1 (TIMP1) were identified as secreted proteins from VML derived FAPs that produced both pro-fibrotic and anti-myogenic signaling. These data establish an aberrant FAP sub-population that are elevated in VML injury and provides novel targets for future scarless muscle regeneration in VML.
As the building blocks of proteins and precursors of many other important compounds, amino acids play a vital role in the biochemical processes needed to sustain life. The branched-chain amino acids (BCAAs) are unique in their structure and function, as they are metabolized in muscle tissue and play important roles in protein synthesis and energy production. However, despite their physiological importance, relatively little integrative research has been conducted into the direct relationships between this class of metabolites and their effect on risk for metabolic diseases. Utilizing an integrative PheWAS approach using UK Biobank data, we were able to identify strong, high confidence, metabolite-disease correlations for the three BCAAs: leucine, isoleucine, and valine. Relationships were established through comparison of metabolite level-disease prevalence associations with polygenic scores for BCAAs, followed by Mendelian randomization analysis. All BCAAs studied demonstrated especially strong relationships with type II diabetes, and robust relationships with obesity, hypertension, sleep apnea, and chronic kidney disease. We illustrate this with a set of metabolite prevalence-disease risk plots that suggest differing potential for disease based on varying levels of branched-chain amino acid metabolites. Similar results are observed with polygenic scores for plasma BCAAs. Mendelian randomization shows positive effects of leucine and isoleucine on hypertension, and either reverse causality or no clear directional relationship for other associations, notably effects of obesity and type II diabetes on all three BCAAs, with limited or borderline evidence for other outcomes. Overall, the results of our study highlight a relatively unexplored area of metabolite-disease associations and provide a blueprint for uncovering additional relationships using readily available biobank data.
Genome-wide association studies typically identify hundreds to thousands of loci, many of which harbor multiple independent peaks, each parsimoniously assumed to be due to the activity of a single causal variant. Fine-mapping of such variants has become a priority and since most associations are located within regulatory regions, it is also assumed that they colocalize with regulatory variants that influence the expression of nearby genes. Here we examine these assumptions by using a moderate throughput expression CROPseq protocol in which Cas9 nuclease is used to induce small insertions and deletions across the credible set of SNPs that may account for expression quantitative trait loci (eQTL) for genes associated with inflammatory bowel disease (IBD). Of the 4,384 SNPs targeted in 88 loci (an average of 50 per locus), 439 were significant and further examined for validation. From these, 98 significantly altered target gene expression in HL-60 myeloid cell line, 74 in induced macrophages from these HL-60 cells, and 78 in induced neutrophils for a total of 201 validated effects (46%), 43 of which were observed in at least two of the cell types. Considering the observed sensitivity and specificity of the controls, we estimate that there are at least 150 true positives per cell type, an average of almost 2.4 for each of the 64 eQTL for which putative causal variants have been fine-mapped. This implies that haplotype effects are likely to explain many of the associations. We also demonstrate that the same approach can be used to investigate the activity of very rare variants in regulatory regions for 89 genes, providing a rapid strategy for establishing clinical relevance of non-coding mutations.
The past 2 decades have witnessed extraordinary advances in our understanding of the genetic factors influencing inflammatory bowel disease (IBD), providing a foundation for the approaching era of genomic medicine. On behalf of the NIDDK IBD Genetics Consortium, we herein survey 11 grand challenges for the field as it embarks on the next 2 decades of research utilizing integrative genomic and systems biology approaches. These involve elucidation of the genetic architecture of IBD (how it compares across populations, the role of rare variants, and prospects of polygenic risk scores), in-depth cellular and molecular characterization (fine-mapping causal variants, cellular contributions to pathology, molecular pathways, interactions with environmental exposures, and advanced organoid models), and applications in personalized medicine (unmet medical needs, working toward molecular nosology, and precision therapeutics). We review recent advances in each of the 11 areas and pose challenges for the genetics and genomics communities of IBD researchers.
Genetically transitional disease (GTD) is emerging as a new concept in genomic medicine to straddle between the traditional binary classification of monogenic and polygenic disease. Genetic testing result reports in molecular laboratories have been predicated on the monogenic disease model, which focuses on pathogenic and likely pathogenic variants. While variants of uncertain significance (VUS) are reported by laboratories, there are challenges with regard to their clinical application so that these variants are often dismissed by ordering physicians. Unlike Mendelian disorders, where genetic variants are of high penetrance and highly probabilistic, the GTD concept is employed to highlight the impact of low-to-moderate effect gene variants whose influence on disease is modified by the genetic background. The GTD concept may explain health conditions associated with variants that are necessary but not sufficient for pathogenesis, lying in the mid gray zone between Mendelian and polygenic diseases. Although VUSs may not reach the level of pathogenicity based on American College of Medical Genetics and Genomics guidelines, they could be provisionally classified as GTD-associated variants to annotate and interpret the relationship between VUS and human genetic disease. The appropriate implementation of the GTD concept could impact patient care and research by focusing attention on the individual variability of responses in various diseases.
The transferability of polygenic scores across population groups is a major concern with respect to the equitable clinical implementation of genomic medicine. Since genetic associations are identified relative to the population mean, inevitably differences in disease or trait prevalence among social strata influence the relationship between PGS and risk. Here we quantify the magnitude of PGS-by-Exposure (PGSxE) interactions for seven human diseases (coronary artery disease, type 2 diabetes, obesity thresholded to body mass index and to waist-to-hip ratio, inflammatory bowel disease, chronic kidney disease, and asthma) and pairs of 75 exposures in the White-British subset of the UK Biobank study (n=408,801). Across 24,198 PGSxE models, 746 (3.1%) were significant by two criteria, at least three-fold more than expected by chance under each criterion. Predictive accuracy is significantly improved in the high-risk exposures and by including interaction terms with effects as large as those documented for low transferability of PGS across ancestries. The predominant mechanism for PGS×E interactions is shown to be amplification of genetic effects in the presence of adverse exposures such as low polyunsaturated fatty acids, mediators of obesity, and social determinants of ill health. We introduce the notion of the proportion needed to benefit (PNB) which is the cumulative number needed to treat across the range of the PGS and show that typically this is halved in the 70th to 80th percentile. These findings emphasize how individuals experiencing adverse exposures stand to preferentially benefit from interventions that may reduce risk, and highlight the need for more comprehensive sampling across socioeconomic groups in the performance of genome-wide association studies.
At the beginning of the COVID-19 pandemic, the Georgia Institute of Technology made the decision to keep the university doors open for on-campus attendance. To manage COVID-19 infection rates, internal resources were applied to develop and implement a mass asymptomatic surveillance program. The objective was to identify infections early for proper follow-on verification testing, contact tracing, and quarantine/isolation as needed. Program success depended on frequent and voluntary sample collection from over 40,000 students, faculty, and staff personnel. At that time, the nasopharyngeal (NP) swab, not saliva, was the main accepted sample type for COVID-19 testing. However, due to collection discomfort and the inability to be self-collected, the NP swab was not feasible for voluntary and frequent self-collection. Therefore, saliva was selected as the clinical sample type and validated. A saliva collection kit and a sample processing and analysis workflow were developed. The results of a clinical sample-type comparison study between co-collected and matched NP swabs and saliva samples showed 96.7% positive agreement and 100% negative agreement. During the Fall 2020 and Spring 2021 semesters, 319,988 samples were collected and tested. The program resulted in maintaining a low overall mean positivity rate of 0.78% and 0.54% for the Fall 2020 and Spring 2021 semesters, respectively. For this high-throughput asymptomatic COVID-19 screening application, saliva was an exceptionally good sample type.