Epigenetic clocks based on DNA methylation patterns are among the most accurate molecular correlates of chronological age, yet widely used clocks are predominantly empirical models with limited explicit characterization of the underlying methylation variability, lacking a direct connection to the physical mechanisms of aging. In this work, we bridge this gap by introducing an information-theoretic framework for DNA methylation dynamics combined with nonlinear machine learning to develop a competitive and interpretable age predictor. We model the population distribution of methylation β-values at each CpG site using a reparameterized three-parameter Generalized Gamma Distribution (GGD) and derive a closed-form expression for its differential Shannon entropy. The resulting CpG-level entropy is used to characterize methylation variability and as a criterion for locus filtering. We introduce the Stacy Gradient Boosting Clock (Stacy-GB), which combines this GGD-based representation with a LightGBM regressor. The model was evaluated across independent cohorts using the ComputAgeBench epigenetic clock benchmark. Stacy-GB achieved a mean absolute error (MAE) of 3.74 years and a median error (bias) of 2.41 years, significantly outperforming state-of-the-art epigenetic clock baselines. Furthermore, age acceleration estimated by Stacy-GB was associated with several clinical pathologies, including ischemic heart disease, HIV infection, multiple sclerosis, and Werner syndrome, supporting its potential as an accurate and biophysically grounded tool for clinical aging research.
Background: Hepatic stellate cells (HSCs) drive cirrhosis and hepatocellular carcinoma (HCC). Statin use decreases cirrhosis and HCC through unclear mechanisms. We aimed to uncover statin-responsive pathways preventing activation of HSC subpopulations in human MASLD-HCC. Methods: We used single-nucleus RNA-sequencing and spatial imaging on matched human MASLD-HCC samples in cirrhotic and non-cirrhotic livers, along with RNA-sequencing of human primary HSCs to identify Yes-Associated Protein (YAP) as a statin-responsive pathway, which was assessed for cellular localization and function, that is, downstream gene expression after statin exposure. We used pharmacologic inhibitors, gene silencing, pharmacological repletion, and direct quantification by UPLC-MS/MS to interrogate the mevalonate pathway as a statin-responsive YAP regulator. Results: We identified an HCC-HSC subcluster enriched in activated and deactivated marker genes. Spatial resolution of each cell type revealed that myofibroblast-like HSC subpopulations and YAP-effector genes colocalized in the peritumoral pseudocapsule. In fact, although Rho GTPases were differentially expressed across most cell types, YAP-effector genes were largely upregulated in HCC-HSCs by single-nucleus RNA sequencing. Bulk RNA-sequencing of statin-treated activated human primary HSCs showed decreased expression of YAP-related genes and downregulation of Rho GTPase pathways in response to statins. Statins reduced GGPP levels in LX2 HSCs, determined by direct measurement of intracellular GGPP, which was associated with cytoskeletal restructuring, YAP cytosolic retention, and reduced YAP-effector gene expression. Exogenous GGPP repletion reversed these effects. GGPP synthase 1 knockdown and direct Rho inhibition recapitulated the cellular responses of YAP following statin exposure. Conclusions: Upregulation of YAP-effector genes was largely restricted to HCC-HSCs. Our mechanistic studies support that statins lower GGPP, reduce Rho GTPases prenylation, cause YAP cytosolic retention, and decrease YAP nuclear activity. This study expands on the chemoprotective mechanisms of statins in HCC.
Pseudoxanthoma Elasticum (PXE) is a rare disease caused by loss of function of the ATP-binding cassette C (ABC) member 6 (Abcc6) gene and characterized by ectopic calcification of multiple tissues, but the physiological reasons underlying ectopic calcification in PXE remain unclear. In a murine model of Abcc6-deficient PXE in which animals developed robust cardiac calcification after heart injury, we show the critical importance of the liver in mediating ectopic cardiac calcification. Tissue-specific deletion of Abcc6 in the liver, but not in the heart, was sufficient to cause post-injury cardiac calcification. Metabolomics and gene expression analysis demonstrated deficiencies in nucleotide metabolism, cellular energetics, and defects in cellular respiration underlying ectopic calcification in PXE. Functional abnormalities in cellular respiration in the injured heart were similar in animals with global or liver-specific Abcc6 deficiency, showing that hepatic Abcc6 expression regulated cellular respiration in the injured heart. We show that ectopic calcification in PXE was primarily dystrophic and that treatment with clodronate or etidronate, which prevent the growth of calcium hydroxyapatite mineralization, was sufficient to rescue the phenotype of ectopic cardiac calcification in Abcc6-deficient states. Taken together, these observations highlight the role of the liver in regulating target tissue metabolic and mitochondrial function in causing ectopic calcification in Abcc6-deficient states.
Background:While immunologic aging impacts immune responses to vaccination, consistent biomarkers associated with aging of the immune system and suboptimal serologic response to influenza vaccination have not been well-studied. Identification of readily measurable biomarkers of immunosenescence may have predictive clinical utility and inform targeted influenza vaccination strategies and future research into aging of the immune system. Methods:We quantified multiple serum/plasma and cell-based parameters related to immune aging (CMV serostatus, plasma cytokines/chemokines, TREC, TERT, NK cell functionality, and DNA methylation clock) at baseline in an adult (age range 18-85) cohort of 2019-2020 influenza vaccine recipients (n=337) and evaluated their associations with vaccine-induced HAI response to influenza A/H1N1, A/H3N2 and B/Victoria strains. Results:CMV IgG titers were significantly positively correlated with vaccine-induced increases in HAI antibody titers to influenza A/H1N1 ( p= 0.02) and A/H3N2 ( p =0.014). CMV IgG titers ( p= 0.00096) and CMV seropositivity ( p= 0.003) were also associated with Day 28 HAI seropositivity against influenza A/H3N2 in subjects seronegative at baseline. Conversely, plasma MCP-1 levels were negatively associated with HAI responses to the A/H3N2 ( p= 0.04) strain. These findings were significant independent of age, sex or vaccine type received (high vs standard-dose seasonal influenza vaccine). Conclusions:Our identification of significant relationships between easily quantifiable immune markers and HAI responses to influenza A vaccine strains across sex and age enhances our knowledge of specific links between immune aging and influenza vaccine-induced immunity. These markers could be leveraged for predicting response to influenza immunization.
Cellulosomes are large, surface-displayed enzyme complexes that enable anaerobic bacteria to degrade recalcitrant plant polysaccharides, yet cellulosome-expressing bacteria are thought to be rare in the human gut. Here, we show that extensive sequence divergence obscures the detection of many ruminococcal cellulosomes by conventional sequence homology-based methods. Using proteome-scale AlphaFold2 structural predictions, we uncovered a substantially expanded set of putative cellulosome-producing Ruminococcus species, including six previously unrecognized human symbionts. Structure-based clustering identifies several novel cohesin families that retain conserved folds despite extreme sequence divergence and define distinct, phylogenetically conserved cellulosome architectures. The analysis reveals R. callidus and related human symbionts encode elaborate cellulosomes that are invisible to sequence-based annotation. Similarly, R. difficilis, a human gut symbiont, has been found to possess genes for an atypical cohesin-based assembly enriched in amylases and related starch-binding proteins, which may enable this microbe to degrade resistant starches that evade digestion in the upper gastrointestinal tract. Together, these findings reveal that ruminococcal cellulosomes are far more prevalent and diverse than previously appreciated and demonstrate the power of structural proteomics to uncover deeply divergent functional systems in the gut microbiome.IMPORTANCEPlant cell wall polysaccharides are a major dietary carbon source, yet their degradation relies on rare, highly specialized microbial enzyme assemblies known as cellulosomes, which have long been considered uncommon in the human gut. Using proteome-scale structure prediction combined with experimental validation, we show that cellulosomes are far more widespread and structurally diverse in human-associated Ruminococcus species than previously appreciated. We identify multiple new cohesin families and reveal distinct cellulosome architectures likely adapted to degrade different dietary substrates. Together, these findings redefine the distribution and evolution of cellulosomes in gut microbes and demonstrate the power of structural proteomics to uncover deeply diverged biological systems.
Humanized mouse models are essential for evaluating the engraftment capacity and genetic integrity of gene-modified hematopoietic stem and progenitor cells (HSPCs). Here, we compared two widely used xenotransplantation platforms, NSG and NBSGW mice, in the context of lentiviral vector (LVV) transduction and CRISPR/Cas9-mediated gene correction. HSPCs harboring high LVV copy numbers exhibited engraftment deficits in NSG mice that were not observed in NBSGW mice. This discrepancy highlights the potential for the NBSGW model to mask safety liabilities of LVV-modified products due to its higher overall levels of human chimerism. In contrast, CRISPR/Cas9 editing with a single-stranded oligodeoxynucleotide donor yielded comparable correction rates in both models, even across decreasing input cell doses, demonstrating that long-term repopulating hematopoietic stem cells (HSCs) retain equivalent engraftment capacity in each strain. Single-cell RNA-sequencing revealed distinct progenitor populations that were markedly under-represented in the NSG model but preserved in NBSGW recipients, emphasizing the greater capacity of NBSGW mice to better support multilineage human hematopoiesis. Together, these findings establish that both NSG and NBSGW mice are suitable for assessing long-term engraftment and gene modification outcomes in human HSPCs. However, the significantly higher percentage of human cell chimerism in the NBSGW model may obscure cell populations with engraftment deficits. Careful selection of in vivo models is therefore critical for rigorous preclinical evaluation of gene therapy products prior to clinical translation.
We report the genome sequences of two syntrophic bacteria, Syntrophomonas wolfei subsp. saponavida SD2T and Syntrophomonas wolfei subsp. methylbutyratica 4J5T, that degrade a variety of short, branched, or long-chain saturated fatty acids to obtain energy when grown in co-culture with a suitable hydrogen-consuming microbial partner strain.
Abstract The measurement of inbreeding has gained significance across diverse fields, including population and conservation genetics, agricultural genetics, breeding programs for animals and plants, and wildlife management. This is due to the fact that inbreeding leads to increased homozygosity and results in lower genetic diversity, rendering populations more vulnerable to environmental changes, diseases, and other stressors. High or mid-coverage whole genome sequencing (WGS) has been widely used for inbreeding estimation, but it is resource-intensive. We aimed to investigate the use of ultra low-coverage whole genome sequencing (ulcWGS) as a cost-effective alternative for inbreeding analysis. Domestic dogs were used for our study as their extensive breeding histories lead to populations with a wide range of inbreeding levels. We constructed a multi-breed reference panel from high-coverage WGS samples. Inbreeding in independent ulcWGS samples was then estimated using runs of homozygosity (RoH) and inbreeding coefficients ( F ). We modeled the relationship between these measures and sequencing depth using nonlinear regression, to generate inbreeding estimates relative to sequencing depth. Resulting relative RoH and F measurements were significantly correlated, with purebred dogs exhibiting more runs of homozygosity and higher inbreeding coefficients compared to mixed-breed dogs. Our findings demonstrate that ulcWGS can provide reliable and economical estimations of inbreeding, expanding accessibility to genetic monitoring.
Abstract Chronological age estimation can provide supporting information in forensic casework when traditional identification methods are limited. DNA methylation, a stable epigenetic mark, has emerged as a promising tool for predicting chronological age from trace samples. However, many existing age estimation models rely on linear regression approaches, which often yield biased prediction errors across the age distribution (i.e. model residuals show a significant age dependence). In this study, we compared three approaches for age estimation modeling: multivariable linear regression, random forest regression and maximum likelihood estimation. While the first two approaches are well established, for the third one we constructed and validated a DNA methylation-based LOESS regression maximum likelihood model for age estimation utilizing forensic-relevant CpG markers. In all cases, model performance was evaluated through Leave-One-Out Cross-Validation (LOOCV). We utilized three independent publicly accessible methylation datasets collected using droplet digital PCR (ddPCR) to evaluate the most effective method for accuracy and bias in age estimation. Notably, when we compare the results of the maximum likelihood approach to the other approaches, multivariable linear regression and random forest regression, we find less bias in the age associated residuals compared to the other methods. These findings highlight the utility of non-linear modeling techniques in reducing the biases of epigenetic age estimation for forensic applications.
Aging is a complex biological process marked by a gradual decline in phys-iological function that contributes to increased vulnerability to disease and mortality. Numerous studies have investigated the cellular and molecular aspects of aging at single-cell resolution, yet the heterogeneity of cellular aging in an individual remains poorly understood. To enhance our ability to study aging at the single cell level, we developed a statistical framework to predict the age of individual cells based on their transcriptomic profiles. Our Bayesian approach estimates the most likely age of a cell given its read counts. We applied the model to data from Tabula Muris Senis and examined organ- and cell-type-specific transcriptomic signatures of aging. Compared with standard regression-based methods, our framework achieved higher pre-dictive accuracy. We show that scBayesAge is a powerful tool for dissecting the cellular heterogeneity of aging and age-related functional decline.
Transcriptional pause-release critically regulates cellular RNA biogenesis, yet how dysregulation of this process impacts embryonic development is not fully understood. Rtf1 is a multifunctional transcription regulatory protein involved in modulating promoter-proximal pausing of RNA Polymerase II (RNA Pol II). Using zebrafish and mouse as model systems, we show that Rtf1 activity is essential for the differentiation of the myocardial lineage from mesoderm. Ablation of rtf1 impairs the formation of nkx2.5+/tbx5a+ cardiac progenitor cells, resulting in the development of embryos without cardiomyocytes. Structure-function analysis demonstrates that Rtf1's cardiogenic activity requires its Plus3 domain, which confers interaction with the pausing/elongation factor Spt5. In Rtf1-deficient embryos, the occupancy of RNA Pol II at transcription start sites was reduced relative to downstream occupancy, suggesting a reduction in transcriptional pausing. Intriguingly, attenuating pause release by pharmacological inhibition or morpholino targeting of CDK9 improved RNA Pol II occupancy at the transcription start sites of key cardiac genes and restored cardiomyocytes in Rtf1-deficient embryos. Thus, our findings demonstrate the crucial role that Rtf1-mediated transcriptional pausing plays in controlling the precise spatiotemporal transcription programs that govern early heart development.
The current liquid biopsy paradigm rests mostly on plasma cell-free DNA (cfDNA), and only to a limited extent on other biofluids. Saliva is molecularly rich but difficult to use, as tumour signal must be separated from abundant background noise. Salivary cfDNA is dominated by shorter fragments, contains substantial microbial DNA, and shows a cancer-associated shift toward longer fragments opposite to plasma. Here we develop a bench workflow and analytical framework to detect gastric cancer from salivary cfDNA. We applied Broad-Range cfDNA Sequencing (BRcfDNA-Seq) to capture fragments lost in conventional library preparation and profiled 281 participants under a PRoBE design, with model fitting restricted to a development cohort (n = 132) and performance evaluated in a separate held-out cohort (n = 149). Applied genome-wide, fragmentomics detected gastric cancer but could not separate signal from background, reaching 95.5% sensitivity at only 37.8% specificity. Anchoring the same coverage to stomach-specific CTCF insulator architecture and refining to discriminative loci raised specificity from 38% to 84.1% (AUROC 0.877, 82% sensitivity); mismatched immune references produced smaller gains, supporting tissue matching. Adding demographic covariates increased AUROC to 0.907 but reduced specificity. Tissue-anchored chromatin analysis converts saliva, a molecularly rich compartment, into a usable substrate for cancer detection.
Abstract Background Epigenetic aging bridges the gap between biological and chronological age by exploiting DNA methylation (DNAm) patterns. Over the past decade, successive DNAm-based clocks have been introduced, beginning with the first-generation Horvath and Hannum models and extending to second-generation PhenoAge and the GrimAge family; complementary measures include DNAm-estimated telomere length and mitotic indices such as epiTOC/pcgtAge. We previously conducted a side-by-side evaluation of these metrics in colorectal cancer using publicly available data from The Cancer Genome Atlas (TCGA) COAD and READ cohorts, but an equally systematic assessment in breast cancer has been lacking. Result Here, using TCGA-BRCA tumor methylomes linked to clinical data (analytic n = 781), we compared seven metrics (Horvath, Hannum, PhenoAge, GrimAge1, GrimAge2, epiTOC/pcgtAge, DNAmTL) via Kaplan–Meier grouping (median and tertiles) and Cox models adjusted for menopausal status, age at diagnosis, receptor subtype, stage, race, and ethnicity, with overall survival truncated at 4000 days. Our analysis reproduced expected benchmark patterns: Triple Negative Breast Cancer (TNBC) had the worst outcomes, Luminal A the best, and higher stage and older age predicted poorer survival, supporting analytic validity. We found first-generation clocks did not separate survival, whereas PhenoAge and GrimAge2 stratified outcomes; in multivariable analyses, only GrimAge1 provided independent prognostic information. DNAmTL was inversely associated with mortality in univariate models, and epiTOC stratified tertiles but showed wide, nonsignificant Cox estimates. Conclusions Second-generation clocks demonstrated stronger prognostic signal than first-generation models in unadjusted analyses. Among them, GrimAge1 retained independent prognostic value beyond established clinicopathologic factors in breast cancer, supporting further external validation with richer covariates to refine clinical utility.
Single-cell RNA sequencing technologies provide insights into gene expression at the cellular level, enabling detailed analysis of cellular heterogeneity. In this study, we systematically compared two scRNA-seq platforms,10X Genomics and Parse Biosciences, using human peripheral blood mononuclear cells (PBMCs) and terminally differentiated effector memory CD8+ T cells (TEMRAs). We identified significant differences in gene expression variability and platform-specific biases, such as ribosomal and mitochondrial gene capture. 10X has a bias for shorter genes, while Parse exhibited enhanced detection of longer transcripts. In CD8+ TEMRAs, the expression of key genes related to antimicrobial immune responses were underrepresented in Parse cells compared to 10X cells (e.g. GNLY, PRF1 and GZMB). These findings underscore the need for careful selection of scRNA-seq platforms based on specific research objectives, as platform-specific biases can influence cell type identification and as well as mechanistic insights derived from gene expression data. Our results provide insights for selection of scRNA-seq experimental platforms in immunological studies.
Abstract Background Sex differences have been described in several corneal diseases such as Fuchs endothelial corneal dystrophy and keratoconus, with estrogens implicated in the induction of these differences. Here, we report the identification of sex differences in a cohort of 177 individuals with Corneal Hereditary Endothelial Dystrophy (CHED), a rare corneal endothelial dystrophy associated with biallelic SLC4A11 gene mutations, and in a Slc4a11 −/− mouse model of CHED. Methods Central corneal thickness (CCT) was measured in individuals with CHED and in Slc4a11 −/− and Slc4a11 +/+ mice to identify a correlation between sex and the degree of corneal edema. To investigate potential causes of such a correlation, RNAseq analysis and mitochondrial superoxide measurement were performed on corneal endothelium from male and female Slc4a11 −/− and Slc4a11 +/+ mice, for which body composition analysis was also performed. Gonadectomy or sham surgery was performed in Slc4a11 −/− and Slc4a11 +/+ mice at 4 weeks of age with subsequent longitudinal CCT and body weight monitoring, followed by an analysis of the interaction effect of surgery type, sex and genotype on CCT. Results Male sex is associated with increased CCT, and thus more severe corneal edema, the characteristic clinical feature of CHED, in affected individuals and Slc4a11 −/− mice. The corneal endothelium in male Slc4a11 −/− mice demonstrates increased levels of oxidative stress compared to Slc4a11 −/− female mice, as evidenced by higher levels of glucose- and glutamine-derived mitochondrial superoxide, controlling for age. Removal of gonadal hormones in Slc4a11 −/− mice increases corneal edema in female mice, suggesting a protective role for ovarian hormones. Transcriptomic analysis of corneal endothelium and body composition analysis in Slc4a11 +/+ and Slc4a11 −/− mice suggest that estrogens play a role in promoting corneal endothelial utilization of lipids via β-oxidation as an alternative energy source in the absence of SLC4A11-mediated NH3:H transport function, thereby reducing oxidative stress from glucose and glutamine metabolism. Conclusions Male sex is associated with a more severe corneal phenotype in individuals with CHED and a Slc4a11 −/− mouse model of the disease. Increased corneal edema in female Slc4a11 −/− mice following gonadectomy suggests ovarian hormones play a protective role in maintaining corneal deturgescence in the setting of loss of SLC4A11 function.
Abstract Background SARS-CoV-2 infection leads to a wide range of clinical manifestations ranging from asymptomatic to fatal cases. DNA methylation plays a crucial role in modulating host responses to viral infections; however, its potential for forecasting COVID-19 disease progression has not been fully explored. Results We aimed to explore the connection between DNA methylation and COVID-19 phenotypic trajectories by examining a subset of the IMPACC cohort (n = 75), which includes longitudinal samples from hospitalized COVID-19 patients grouped into five disease trajectory groups (TGs) based on respiratory severity during the acute phase of infection. Our findings reveal that DNA methylation is associated with respiratory status, hospitalization duration, and immune cell composition, including T cell depletion and band neutrophil accumulation during acute infection. DNA methylation profiles at hospital admission differ among TGs and are associated with subsequent disease progression. Furthermore, DNA de-methylation of specific distal enhancers correlated with TG, and DNA methylation patterns generally normalize post-resolution. Comparative analyses revealed that DNA methylation changes are more strongly associated with TGs than transcriptome profiles across time. Conclusions This study highlights the association between DNA methylation and COVID-19 phenotypic trajectories and identifies CpG loci that could serve as potential biomarkers for identifying patients at heightened risk of severe COVID.
Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals, including sex differences in killifish liver, and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.
Background Understanding the genetic architecture of domestic dogs provides unique insights into the processes of domestication, breed formation, and the genetic basis of complex traits and diseases. Dog populations, characterized by their diverse morphologies and behaviors, also exhibit extensive evidence of historical and ongoing admixture. This widespread mixing, driven by both natural migration and selective breeding practices, has profoundly shaped the genomic landscape of modern dog breeds. Though global admixture has been extensively estimated in human population studies, where the number of subgroups is typically limited, there has been more limited analysis in canines, where there may be dozens of ancestral groups, or breeds. Results Here we present a procedure for estimating global admixture in dogs from whole genome sequence data using SCOPE. We created a reference population of 65 dog breeds that included 349 individuals, from which we determined breed-informative SNPs. We demonstrate that SCOPE can accurately infer breed composition in both simulated and real admixed samples, even at low sequencing depths. We also characterized the genetic similarity between our reference dog breeds and recovered previously reported relationships. Conclusion This approach allows us to identify the strength of the genetic signature of breeds and place error bounds on admixture estimates. It also provides evidence that admixture can be accurately inferred in subjects that may originate from multiple ancestral populations.